Novel power system edge side power self-discipline distribution method under guidance of limited knowledge

By deploying agents at the edge of the power system and training the target policy network using local observation information and a power grid security rule base, the problems of low learning efficiency and high security risks of traditional multi-agent deep reinforcement learning in power systems are solved, and efficient and safe active and reactive power coordinated control is achieved.

CN121863451APending Publication Date: 2026-04-14STATE GRID LIAONING ELECTRIC POWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID LIAONING ELECTRIC POWER CO LTD
Filing Date
2025-11-28
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Traditional multi-agent deep reinforcement learning in power system applications suffers from several problems: difficulty in acquiring "complete knowledge" leading to low learning efficiency; pure data-driven learning lacks physical laws, which can easily lead to safety risks; and centralized training coupled with decentralized execution makes it difficult to break free from dependence on data collection.

Method used

By deploying intelligent agents at the edge nodes of the power system, generating local observation states through local observation information, and combining them with the power grid safety operation rule base and reward function, the target policy network is trained using a multi-agent reinforcement learning algorithm under a human-machine hybrid enhancement mechanism to achieve coordinated control of active and reactive power.

Benefits of technology

It reduces data transmission and processing costs, improves learning efficiency, enhances the system's autonomy and security, adapts to the real-time requirements of power systems, and reduces reliance on centralized architecture and generalized data acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121863451A_ABST
    Figure CN121863451A_ABST
Patent Text Reader

Abstract

The invention discloses a novel power system edge side power self-discipline distribution method under the guidance of limited knowledge, and relates to the technical field of power system operation control, intelligent agents are deployed at a plurality of edge side nodes of a power system, and each intelligent agent generates a local observation state based on local limited observation information; on the basis of a preset power grid safe operation rule base, action constraints are set, a mixed reward function containing penalty terms and man-machine interaction rewards is constructed, each agent inputs a local observation state into a strategy network subjected to multi-agent reinforcement learning training, and an active power and reactive power combined control instruction is output; and then, the intelligent agent executes an instruction to perform cooperative control on the power regulation equipment in the jurisdiction range, so that edge-side power self-discipline distribution is realized, only local observation and limited boundary information are depended on, a global model is not needed, and the self-governance capability, the safety level and the operation efficiency of the power grid are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power system operation and control technology, and in particular to a novel power self-regulating allocation method for the edge side of a power system guided by limited knowledge. Background Technology

[0002] With the rapid advancement of new power system construction, a large proportion of distributed energy sources (such as photovoltaic and wind power) are being connected to distribution networks and user sides. The intermittent and fluctuating output of these distributed power sources causes the power flow in the distribution network to change from the traditional unidirectional flow to bidirectional flow, leading to a series of problems such as node voltage exceeding limits and power flow reversal, posing a severe challenge to the safe and stable operation of the power grid. Optimization control methods for grid voltage and power are mainly divided into centralized and distributed types. Centralized control methods rely on accurate global grid models and powerful central computing capabilities, making it difficult to adapt to the real-time fluctuations of distributed power sources. They also suffer from high communication bandwidth pressure and high risks of data privacy leakage. Distributed control methods, such as distributed optimization based on the Alternating Direction Multiplier Method (ADMM), reduce the dependence on a central controller, but still require multiple iterative communications between agents during the solution process, placing extremely high demands on the real-time performance and reliability of the communication system. Furthermore, their optimization effect still largely depends on an accurate physical model of the power grid.

[0003] In related technologies, data-driven deep reinforcement learning (DRL) methods offer a new approach to solving model-free optimization decision-making. Multi-agent deep reinforcement learning (MADRL) learns near-optimal cooperative control strategies through the interaction of multiple agents with the environment. However, the applicant recognizes that in practical applications of traditional multi-agent deep reinforcement learning algorithms in the specific field of power systems, the training and decision-making processes of agents still rely to some extent on the difficult-to-obtain "complete knowledge" (i.e., accurate global models, real-time global states, and long-term accurate prediction information). In the highly physically coupled cooperative environment of distribution networks, global rewards are difficult to quantify fairly and accurately, leading to low learning efficiency. Purely data-driven models lack the guidance of the inherent physical laws of power systems, resulting in poor interpretability of the decision-making process. In "unseen" operating scenarios not covered by training data, they are prone to making decisions that violate safety rules, posing a high risk. Many advanced algorithms adopt a "centralized training, distributed execution" framework. Although distributed decision-making can be achieved in the execution phase, their training process usually still needs to gather local observation information from each agent to construct a global or near-global perspective, failing to completely get rid of the dependence on extensive data collection. Summary of the Invention

[0004] In view of this, this application provides a novel power autonomous allocation method for the edge side of a power system guided by limited knowledge. The main purpose is to solve the problems of low learning efficiency due to the difficulty in obtaining "complete knowledge" in traditional multi-agent deep reinforcement learning in power system applications, the lack of physical laws in pure data-driven learning which easily leads to safety risks, and the difficulty in getting rid of data acquisition dependence due to centralized training and decentralized execution.

[0005] According to the first aspect of this application, a novel power self-regulation allocation method for the edge side of a power system guided by limited knowledge is provided, the method comprising: Each agent collects limited local observation information within its jurisdiction and generates a local observation state based on the limited local observation information. The agents are deployed on the edge nodes of the power system. Each agent inputs its corresponding local observation state into the corresponding target policy network to obtain a joint action consisting of active power control instructions and reactive power control instructions. The agent includes action constraints and a reward function determined based on the power grid safe operation rule base. The reward function includes a penalty term and a human-machine interaction reward term. The target policy network is trained using the action constraints and the reward function through a multi-agent reinforcement learning algorithm under a human-machine hybrid reinforcement mechanism. Each intelligent agent coordinates the control of multiple local execution devices within its jurisdiction based on the corresponding target joint action.

[0006] By employing the above technical solutions, the technical solutions provided in the embodiments of this application have at least the following advantages: This application provides a novel edge-side power self-regulation allocation method for power systems guided by limited knowledge. Each agent collects limited local observation information within its jurisdiction and generates a local observation state based on this information. The corresponding local observation state is then input into a corresponding target policy network to obtain a target joint action consisting of active power control commands and reactive power control commands. Based on this target joint action, multiple local execution devices within the jurisdiction are coordinated for control. The agents are deployed on edge-side nodes of the power system and include action constraints and reward functions determined based on a power grid safety operation rule base. The reward function includes a penalty term and a human-machine interaction reward term. The target policy network is trained using a multi-agent reinforcement learning algorithm under a human-machine hybrid reinforcement mechanism, utilizing the action constraints and reward function. Edge-side deployment and limited local observation eliminate the need for centralized collection of massive global data, reducing data transmission and processing costs while enabling low-latency response through edge-side decision-making, thus meeting the real-time requirements of power systems. Furthermore, the policy network training incorporates action constraints based on power grid safety rules, defining the boundaries of safety decisions from the source. Combined with a reward function containing penalties and human-machine interaction, and a hybrid human-machine enhancement mechanism, this improves policy learning efficiency while incorporating expert experience to reduce safety risks. The distributed multi-agent cooperative control mode of this application allows each agent to make autonomous decisions and collaboratively manage equipment based on local observations, freeing it from dependence on centralized architecture and generalized data acquisition, while simultaneously improving the accuracy of power allocation at the edge.

[0007] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0008] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A schematic diagram of a novel power autonomous allocation method for the edge side of a power system guided by limited knowledge is shown. Figure 2 This paper presents a schematic diagram of a novel method for autonomous power allocation at the edge of a power system guided by limited knowledge. Figure 3 A schematic diagram of a human-machine collaborative reinforcement learning method is shown. Figure 4A schematic diagram of a human-machine hybrid enhancement mechanism is shown; Figure 5 This diagram illustrates an edge-side power self-regulating allocation system architecture guided by limited knowledge. Figure 6 A flowchart of a novel power self-regulating allocation method for the edge side of a power system guided by limited knowledge is shown. Detailed Implementation

[0009] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.

[0010] With the rapid development of new power systems, a high proportion of distributed energy sources (such as photovoltaic and wind power) are widely connected to distribution networks and user sides. The intermittent and fluctuating output of these distributed sources causes the power flow in the distribution network to change from the traditional unidirectional flow to bidirectional flow, leading to increasingly prominent problems such as voltage exceeding limits and power flow reversal, posing a severe challenge to the safe and stable operation of the power grid. Existing voltage and power optimization methods are mainly divided into centralized and distributed types. Centralized control methods rely on accurate global grid models and powerful central computing capabilities, making it difficult to cope with the real-time fluctuations of distributed sources. They also suffer from high communication pressure and privacy risks. Distributed control methods, such as optimization based on the Alternating Direction Multiplier Method (ADMM), reduce the dependence on a central controller, but still require multiple iterative communication iterations, placing high demands on the real-time performance and reliability of the communication system, and also relying on accurate physical models.

[0011] Multi-agent deep reinforcement learning (MADRL) can learn near-optimal cooperative control strategies through the interaction of multiple agents with the environment. However, traditional MADRL algorithms still face the following challenges in practical applications of power systems. The root cause lies in the fact that the training and decision-making processes of agents still rely to some extent on the difficult-to-obtain "complete knowledge", namely, accurate global models, real-time global states, and long-term accurate prediction information: (1) Reliability allocation problem: In a cooperative environment with strong physical coupling, such as a distribution network, it is difficult to fairly allocate global rewards to each agent, resulting in low learning efficiency. Essentially, this is because it is difficult to accurately assess the impact of a single agent's actions on the global system without a global model. (2) Model generalization and security: Purely data-driven models lack the guidance of physical knowledge of power systems (such as Kirchhoff's laws and thermal stability limits), and the decision-making process cannot be explained. They may make decisions that violate safety rules in scenarios that have not been experienced before, which exposes the inherent risks of pure data-driven methods when "knowledge is limited" (not all scenarios can be foreseen). (3) Dependence on global information: Many advanced MADRL algorithms (such as MADDPG) adopt a "centralized training, distributed execution" framework. During the training phase, local information of each agent still needs to be collected, which cannot fully protect the data privacy of power grid operators and users. (4) Difficulty in integrating human experience: Existing methods lack an effective human-machine interaction mechanism for power grids, and cannot integrate the dispatcher's valuable prior knowledge (such as monitoring of key sections and experience judgment of equipment status) and intervention instructions in emergency situations into the learning and decision-making closed loop of the agent. To address the aforementioned issues, this application proposes a novel power autonomous allocation method for the edge side of a power system guided by limited knowledge. This method deeply integrates pre-set security rules and constraints of the power system with multi-agent reinforcement learning based on limited knowledge, and innovatively introduces a confidence allocation and a human-machine hybrid enhancement mechanism for the power grid. This achieves collaborative autonomous allocation of active and reactive power relying solely on local observations and limited boundary information, effectively improving the autonomy, security level, and operational efficiency of the regional power grid. The implementing entity of this application can be a power autonomous allocation system for the edge side guided by limited knowledge. This system provides services to users based on the computing power of a server. The server can be an independent server or a server providing basic cloud computing services such as cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. This enables rapid, secure, and autonomous allocation of active and reactive power at the edge side of the distribution network and microgrid under limited knowledge conditions.

[0012] This application provides a novel power self-regulating allocation method for the edge side of a power system guided by limited knowledge, such as... Figure 1As shown, the method includes: 101. Each agent collects limited local observation information within its jurisdiction and generates a local observation status based on the limited local observation information.

[0013] In this embodiment, each intelligent agent is specifically deployed at an edge node of the power system. The edge refers to nodes in the power system close to distributed energy sources (such as photovoltaics and wind power) and loads, typically located at the end of the distribution network or on the user side, encompassing distribution substations, microgrids, and feeder segments. These locations are the main areas for new energy access and power fluctuations, making rapid response difficult with traditional centralized control. Deploying intelligent agents at the edge, leveraging local computing power to conduct distributed decision-making, effectively addresses the intermittency and volatility of high-proportion new energy sources, enhancing the system's autonomy and reducing communication latency. An intelligent agent is an autonomous decision-making unit deployed at the edge of the power system, its core being a software agent embedded in an edge computing device. Each intelligent agent corresponds to an "edge autonomous unit," which can be a microgrid, distribution substation, feeder, or a tightly coupled cluster of controllable resources (such as a photovoltaic inverter cluster or energy storage system cluster). The intelligent agent possesses the following core functions: real-time acquisition of electrical quantity data within its jurisdiction via local sensors or communication interfaces to achieve state awareness; generation of power control commands based on a pre-trained policy network and a pre-set power grid safety operation rule base to complete decision-making; issuance of adjustment commands to active power regulation equipment and / or reactive power compensation equipment within its jurisdiction via control interfaces to complete execution; and limited communication with neighboring agents to exchange boundary coupling information and local reward values ​​to achieve distributed collaborative decision-making. Each agent only collects local operational data (such as node voltage, load power, equipment output, etc.) within its own jurisdiction and organizes this data into corresponding local observation states. This eliminates the need for cross-regional global data collection, reducing data transmission bandwidth pressure and communication costs. Furthermore, it enables rapid processing of local data through edge-side deployment, significantly shortening subsequent decision-making delays. Localized information collection also avoids privacy leaks and security risks associated with global data transmission.

[0014] 102. Each agent inputs its corresponding local observation state into the corresponding target policy network to obtain a joint target action consisting of active power control instructions and reactive power control instructions.

[0015] In this embodiment, each agent inputs its own organized local observation state into the trained target policy network, which then outputs a joint target action containing active power control instructions and reactive power control instructions. The training of this target policy network is guided by action constraints defined by a power grid safety operation rule base, a reward function containing violation penalties and human-machine interaction rewards, and is completed through a multi-agent reinforcement learning algorithm combined with a human-machine hybrid enhancement mechanism. Action constraints limit the agent's safety decision-making scope from the source, reducing the risk of violations; the reward function both penalizes erroneous actions and incorporates expert experience to improve policy rationality; the combination of multi-agent reinforcement learning and human-machine enhancement allows agents to efficiently learn collaborative control strategies, reducing blind exploration driven purely by data.

[0016] 103. Each intelligent agent coordinates the control of multiple local execution devices within its jurisdiction based on the corresponding target joint action.

[0017] In this embodiment, each agent coordinates and controls multiple local execution devices (such as distributed power sources, energy storage devices, capacitor banks, etc.) within its jurisdiction based on the joint actions output by the target policy network. This enables these devices to execute power commands in coordination, ultimately achieving autonomous power allocation within the region. Devices within the same region are coordinated by a single agent, avoiding action conflicts between devices and improving control accuracy. Local coordination at the edge is faster, allowing for timely responses to sudden situations such as power fluctuations and voltage deviations within the region. Simultaneously, the distributed control mode does not rely on a centralized control center, and control failures in a single region have a smaller impact on the global system, significantly improving the operational robustness of the power system.

[0018] This application provides a novel edge-side power self-regulation allocation method for power systems guided by limited knowledge. Compared with existing technologies, this application method involves each agent collecting limited local observation information within its jurisdiction and generating local observation states based on this information. The corresponding local observation states are then input into the corresponding target policy network to obtain a target joint action consisting of active power control commands and reactive power control commands. Based on this target joint action, multiple local execution devices within the jurisdiction are then coordinated for control. The agents are deployed on edge-side nodes of the power system and include action constraints and reward functions determined based on a power grid safety operation rule base. The reward function includes penalty terms and human-machine interaction reward terms. The target policy network is trained using a multi-agent reinforcement learning algorithm under a human-machine hybrid reinforcement mechanism, utilizing the action constraints and reward functions. Edge-side deployment and limited local observation eliminate the need for centralized collection of massive global data, reducing data transmission and processing costs while enabling low-latency response through edge-side decision-making, thus meeting the real-time requirements of the power system. Furthermore, the policy network training incorporates action constraints based on power grid safety rules, defining the boundaries of safety decisions from the source. Combined with a reward function containing penalties and human-machine interaction, and a hybrid human-machine enhancement mechanism, this improves policy learning efficiency while incorporating expert experience to reduce safety risks. The distributed multi-agent cooperative control mode of this application allows each agent to make autonomous decisions and collaboratively manage equipment based on local observations, freeing it from dependence on centralized architecture and generalized data acquisition, while simultaneously improving the accuracy of power allocation at the edge.

[0019] Furthermore, as a refinement and extension of the specific implementation methods of the above embodiments, and in order to fully illustrate the specific implementation process of this embodiment, this application provides another novel power self-regulation allocation method for the edge side of a power system guided by limited knowledge, such as... Figure 2 As shown, the method includes: 201. The target policy network is obtained by training the multi-agent reinforcement learning algorithm under the human-machine hybrid reinforcement mechanism, using action constraints and reward functions.

[0020] In this embodiment, each agent is responsible for managing multiple active power regulation devices (such as distributed generation equipment and energy storage systems) and reactive power compensation devices (such as photovoltaic inverters, static var compensators, and capacitor banks) within its edge autonomous unit. The core of the agent is a policy network embedded in the edge computing device. This network makes decisions based on the local information of its managed unit and engages in limited communication with neighboring agents to jointly achieve power self-regulation across a wide area. For each agent, the agent acquires a sample of local limited observation information within its managed area and generates a current observation state based on this sample, which is used to store local observation data. Specifically, the agent identifies multiple nodes within its managed area and extracts the node voltage, load active power, load reactive power, photovoltaic active power output, energy storage device state of charge, CB tap position, and CB action count for each node from the agent's local limited observation information to generate the agent's local observation state, as shown in Formula 1 below: Formula 1:

[0021] in, Let be the local observation state vector of node i in the agent at time t. Let be the node voltage of node i in the agent at time t. Let be the active power of node i in the agent at time t. Let be the load reactive power of node i in the agent at time t. The active power output of photovoltaic power generation at time t for node i in the intelligent agent. Let represent the state of charge of the energy storage device at time t for node i in the intelligent agent. Let be the CB gear position of node i in the agent at time t. Let be the number of CB actions performed by node i in the agent at time t.

[0022] The agent determines action constraints based on the power grid safety operation rule base. The action constraints are used to limit the agent's action space. The power grid safety operation rule base includes at least one of voltage safety operation constraints, equipment power capacity constraints, branch transmission power constraints, and energy storage charge and discharge times constraints. The action constraints are used to limit the agent's action space, and the penalty term of the reward function is used to negatively incentivize behaviors that violate the rules in the reward function.

[0023] Specifically, constraints are set for the actions of intelligent agents and the safe operation rules of the power grid. To ensure the safe, stable, and reliable operation of the power grid and to achieve economical and efficient power transmission and distribution, power flow constraints are set as follows: Formula 2: Formula 2:

[0024]

[0025] in, The active power injected at node f on busbar. The reactive power injected at node f on busbar is... The active load at busbar f is... The reactive load at busbar f is... Let f be the voltage across busbar f. Let g be the voltage across busbar g. The electrical conductance between busbar f and busbar g is... The susceptance between busbar f and busbar g. Let f be the phase difference between bus f and bus g, and F be the number of buses. To ensure the safe and stable operation of each node, safety operation constraints are set for each node. The branch transmission power constraint is as follows: Formula 3: Formula 3:

[0026] in, Let be the active power transmitted by branch ij at time t. This represents the maximum limit of active power transmitted by branch ij. This is the branch set. The voltage amplitude constraint is as follows: Formula 4: Formula 4:

[0027] in, Let be the lower limit of the voltage amplitude at node j at time t. Let be the upper limit of the voltage amplitude at node j at time t. Let be the voltage magnitude of node j at time t, and M be the number of nodes. The SVC output constraint is given by the following formula 5: Formula 5:

[0028] in, For the reactive power output of the nth SVC, This represents the upper limit of the reactive power output of the nth SVC. This represents the lower limit of the reactive power output of the nth SVC. The reactive power output range of the photovoltaic inverter is shown in Formula 6 below: Formula 6:

[0029]

[0030]

[0031] in, Let be the active power generated by the c-th photovoltaic inverter at time t. Let be the reactive power generated by the c-th photovoltaic inverter at time t. Let be the active power of the c-th photovoltaic inverter operating in maximum power point tracking mode at time t. A collection of photovoltaic inverters, This represents the upper limit of reactive power for the c-th photovoltaic inverter. This is the lower limit of the reactive power of the c-th photovoltaic inverter. Let be the apparent power of the c-th photovoltaic inverter. The output constraint of the gas turbine unit is as follows: Formula 7: Formula 7:

[0032]

[0033] in, For the active power output of the l-th unit, This represents the upper limit of the active power output of the l-th unit. This is the lower limit of the active power output of the l-th unit. For the reactive power output of the l-th unit, This represents the upper limit of the reactive power output of the l-th unit. This is the lower limit of the reactive power output of the l-th unit. The energy storage charging and discharging power limit is as follows: Formula 8: Formula 8:

[0034]

[0035] in, This represents the upper limit of the charging and discharging power of the a-th energy storage unit, with discharging being positive and charging being negative. This represents the lower limit of energy storage capacity. This represents the upper limit of energy storage capacity. A collection of energy storage systems, This is the index for the energy storage system. CB (switched capacitor bank) is used for reactive power control and regulation; its operation is discrete, as shown in Formula 9 below: Formula 9:

[0036]

[0037]

[0038] in, Let be the number of actions of the k-th CB at time t. This represents the maximum number of actions that can be performed by the k-th CB. For the CB set, This is the gear position of the k-th CB.

[0039] Then, a reward function is set, with the negative value of network loss as the core reward. At the same time, a mixed reward signal of voltage and power limit violation penalties and human-computer interaction feedback is introduced to guide the agent in the exploration process. It not only learns to run in the optimal direction, but also spontaneously abides by all safe operation constraints. This ensures that the final learned scheduling strategy achieves a fusion of engineering feasibility and safety while pursuing economy, and achieves a unity of goal optimization and safety penalty, as shown in the following formula 10: Formula 10:

[0040] in, Let be the reward function of agent n at time t. As the first weighting coefficient, This is the second weighting coefficient. The third weighting coefficient, Let n be the global network loss of agent n. This is the penalty coefficient for exceeding the limit. The total number of intelligent agents. Let n be the set of nodes within the jurisdiction of agent n. Let be the voltage amplitude at node j at time t. Let be the lower limit of the voltage amplitude at node j at time t. Let be the upper limit of the voltage amplitude at node j at time t. Let n be the set of branches within the jurisdiction of agent n. This represents the maximum limit of active power transmitted by branch ij. Let be the active power transmitted by branch ij at time t. For agent n, the human-computer interaction reward is determined through a human-computer hybrid enhancement mechanism. This mechanism integrates the experience, judgment, or feedback of human experts into the training process of the agent, guiding the agent to learn safer, more efficient, or more operationally preferred strategies.

[0041] Next, the reinforcement learning model of the intelligent agent and the human-machine hybrid reinforcement mechanism were established, such as... Figure 3 The diagram illustrates the internal architecture of a single autonomous agent and its integration with a human-machine hybrid augmentation mechanism. Each agent's core consists of a policy network (Actor), a value network (Critic), and an estimation network (Estimate). The agent's decision-making process is as follows: through model building, parameter setting, model learning, and evaluation, an initial reward signal r(t) is output; during model learning, continuous iterative loops are performed, expert policies are stored in a hybrid experience pool, and the agent's subsequent interaction "state-action experience trajectories" are accumulated. The agent generates a decision o(t) based on experience from the pool. The decision is then evaluated by the judgment module based on confidence level. If the confidence level is low, human intervention is triggered to obtain a human score. Conversely, it outputs action a(t); after the action is input into the state environment, a new state o(t+1) is generated and fed back to the agent. The "state-action" experience is fed back into the experience pool. The human score and the initial reward are combined as optimization signals. Finally, through the mode of "expert experience fusion + agent autonomous interaction + dynamic human intervention", the agent's strategy is efficiently and safely iteratively optimized. Among them, local observation Boundary information from neighboring agents is fed into an estimation network, which outputs a boundary estimate. Local observations and boundary estimates are combined through an information fusion module to construct a local joint observation view using a value network. The policy network is based on pure local observation. Output decision action Meanwhile, the value network is based on the fused observation view. The system evaluates actions based on Q-scores to generate training signals for policy updates. The entire decision-making and training process is deeply embedded in a human-machine hybrid augmentation mechanism, specifically manifested in three stages: active learning, imitation learning, and safe takeover.

[0042] To achieve autonomous power allocation on the edge side under limited knowledge, a distributed multi-agent reinforcement learning model based on the Actor-Critic architecture is adopted, and customized improvements are made for the characteristics of the hybrid action space of the power system.

[0043] The Actor network serves as the policy network for each agent, and its input is the agent's local observation state. The state vector includes, but is not limited to, local node voltage, load active and reactive power, distributed power source active power output, and energy storage device state of charge. To simultaneously generate continuous control commands and discrete switching commands, this application makes a significant improvement to the standard Actor network structure: an output layer with dimension D (i.e., the maximum action level of the local discrete device) is added to the Actor network. The output of this output layer is processed by the Softmax function to obtain the probability distribution corresponding to each discrete level (such as capacitor bank switching). Simultaneously, the original output layer used for outputting continuous actions (such as reactive power output adjustment) adopts the Tanh activation function. The outputs of these two output layers together constitute the joint action of the agent. To evaluate the long-term value of the joint actions, a Critic network is provided. In a distributed implementation, each agent's Critic network input includes its local observations. One's own actions And boundary coupling information obtained from neighboring agents through communication links. Similarly, to accurately evaluate the action value in a hybrid action space, this application makes a significant improvement to the standard Critic network structure: by adding an output layer to the Critic network, with a dimension of... (in The number of hidden layer neurons (representing the number of neurons in the hidden layer) is used to output the Q-value corresponding to each discrete action level, thereby achieving a comprehensive value assessment of discrete-continuous mixed actions. To reduce reliance on highly reliable real-time communication and enhance the system's robustness in real industrial environments, this architecture also introduces an Estimate network. The Estimate network takes the agent's local observation state as input and outputs an estimate of the voltage amplitude of the boundary coupling nodes. This estimate is provided to the Critic network as a substitute for actual boundary information, thus significantly reducing the system's communication burden while ensuring training effectiveness.

[0044] Based on the above improvements, the agent initializes the policy network, value network, and estimation network, and inputs the current observation state into the policy network to obtain the current joint action. The policy network contains a discrete action output layer using the Softmax function and a continuous action output layer using the Tanh function. The discrete action output layer is used to output the current discrete action, and the continuous action output layer is used to output the current continuous action. The current joint action consists of the current discrete action and the current continuous action.

[0045] Then, the agent corrects the current joint action based on the action constraints and the human-machine hybrid enhancement mechanism. Specifically, the agent uses the action constraints to perform a safety verification operation on the current joint action; if the safety verification operation passes, the agent uses the human-machine hybrid enhancement mechanism to correct the current joint action. The human-machine hybrid enhancement mechanism includes at least one of the following: sample labeling processing, imitation learning processing, active learning processing, and safety takeover processing. The human-machine hybrid enhancement mechanism refers to the introduction of human experts into the training process. The intervention methods include one or more of the following: (1) labeling key samples of power grid operation (such as voltage critical limits and equipment action boundaries); (2) integrating the decision-making experience of scheduling experts through imitation learning; (3) adopting an active learning strategy, the agent selects decision-making scenario samples with high uncertainty in power grid operation (such as violent fluctuations in new energy power and sudden load changes) and requests expert intervention; (4) performing manual takeover when the agent's decision confidence is lower than a preset threshold.

[0046] By deeply embedding human-machine hybrid enhancement mechanisms into the training and decision-making processes of intelligent agents, its core architecture is as follows: Figure 4 As shown, this is specifically reflected in the following three stages: sample optimization based on active learning, policy guidance based on imitation learning, and safe takeover based on confidence assessment. Human experts access this mechanism through a human-computer interaction interface. The mechanism includes three modules: safe takeover, expert demonstration, and active learning. The safe takeover module first conducts a confidence assessment, followed by a safety judgment (if the current confidence level is...). If an alarm is triggered and a takeover is initiated, a human expert will provide the appropriate action; otherwise, the normal procedure will be executed. The expert demonstration module uses the expert demonstration dataset and calculates the imitation loss using KL divergence, which is then used to train the policy network to generate agent policies. The active learning module first assesses uncertainty and then makes a key sample judgment (if uncertainty is high...). (If a threshold is reached, expert annotation of samples is requested; otherwise, regular training continues.) Finally, expert actions, annotated samples, and agent policy-related data are all fed into the experience replay pool, continuously supporting the iterative iteration of each module and achieving safe and efficient optimization of the agent policy. These three stages are tightly coupled through the experience replay pool and the policy network training process, jointly ensuring that the final trained agent policy is safe and reliable. Specifically: This paper addresses the sample optimization stage based on active learning. During reinforcement learning training, the agent faces a massive state space. Uniform learning across all samples leads to low training efficiency, and the number of samples requiring annotation by human experts is enormous, resulting in high costs. Therefore, this application designs an active learning mechanism based on uncertainty quantification. The core idea of ​​this mechanism is to enable the agent to automatically identify "difficult scenarios" with high decision-making uncertainty and actively request human expert annotation. The technical problem it aims to solve is: how to accurately and automatically select the most valuable key samples for policy improvement from a massive number of states, thereby significantly improving the efficiency of expert annotation and the agent's learning efficiency. Uncertainty measurement is performed on both discrete and continuous actions in the mixed action space. For discrete action uncertainty measurement, the degree of hesitation in decision-making among multiple discrete options (e.g., "putting in," "cutting out," "holding out" of a capacitor bank) is quantified, and the information entropy of the discrete action probability distribution output by its policy network is calculated, as shown in Formula 11 below. Formula 11:

[0047] in, For information entropy, For discrete action probability distributions, Given a discrete set of actions, the entropy value is... The higher the value, the more hesitant the agent is among multiple discrete choices. Regarding the measurement of uncertainty in continuous actions, the predictive stability of the agent's output for continuous actions (such as reactive power setpoints) is quantified. Estimation is performed using multiple stochastic forward propagation techniques. During the inference phase of the policy network, either by randomly activating Dropout layers in the network (i.e., the Monte Carlo Dropout method) or by using a set of identically structured, independently initialized ensemble networks, K forward computations are performed on the same observed state to obtain K continuous action outputs. Then, the statistical variance of these output values ​​is calculated and used as a measure of uncertainty, as shown in Formula 12 below: Formula 12:

[0048] in, The variance represents the continuous action; a larger variance indicates that the network's prediction of continuous actions (such as specific numerical values ​​of reactive power output) is more unstable and uncertain. Triggering conditions for key samples and an expert intervention mechanism are set. Given the difference in the uncertainty dimensions of discrete and continuous actions, they cannot be directly compared; therefore, a unified standard is needed to trigger active learning. Here, the uncertainty measures of the two parts mentioned above are normalized and weighted into a unified uncertainty score, as shown in Formula 13 below: Formula 13:

[0049]

[0050]

[0051] in, Score the uncertainty. The information entropy of discrete actions, The variance of continuous actions, For a discrete set of actions, The upper limit threshold for variance can be set based on the maximum variance observed in historical operating data, or determined through offline simulation. Let n be the current observation state of agent n. The first weighting coefficient, This is the second weighting coefficient. These coefficients can be determined through hyperparameter search (such as grid search, Bayesian optimization) or empirically preset according to the emphasis on power grid security and economy. It is the theoretical maximum entropy value of the discrete action set. The function measures and scales different dimensions to the [0,1] interval. For discrete action uncertainty, maximum entropy normalization is used, and for continuous action uncertainty, maximum variance normalization is used.

[0052] Agent acquires uncertainty threshold The agent compares the uncertainty score with the uncertainty threshold. If the uncertainty score is greater than the uncertainty threshold, the agent marks the current observation state and the current joint action as key samples and triggers an expert intervention request. In response to the expert intervention request, the agent pushes the current observation state and the current joint action to the terminal held by the expert. The agent receives the corrected current joint action uploaded by the expert based on the terminal held by the expert.

[0053] This application addresses the policy guidance process based on imitation learning. Pure reinforcement learning agents require extensive random exploration in the early stages of training. However, in high-risk environments like power grids, this can lead to numerous unsafe or ineffective exploration behaviors, reducing learning efficiency and even inducing system risks. To solve this problem, this application introduces imitation learning as a policy guidance mechanism. This transforms the prior knowledge and operational experience of human experts into computable constraints, guiding the agent to learn efficiently within a safe decision space. This avoids blind exploration in the early stages of training while ensuring that the learning process never deviates from basic safety principles.

[0054] First, expert knowledge is represented and preprocessed by collecting operation records of human experts from historical power grid operation data, i.e., dispatching expert decision records, and then an expert demonstration dataset is constructed. The expert policy function was obtained by fitting an expert demonstration dataset. This dataset consists of a large number of state-action pairs. Composition, in which It is the expert in the corresponding state The actual operation is then executed. This method transforms expert experience, which is difficult to formalize, into structured data that can be directly used by machine learning models, laying the data foundation for policy guidance. Expert data can come from historical operation logs or be specifically generated by experts in the simulation system. To deeply integrate expert knowledge into the reinforcement learning process, the training loss function of the Actor network is reconstructed. The total loss function consists of two parts: the standard reinforcement learning objective... A newly introduced imitation learning regularization term As shown in Formula 14 below: Formula 14:

[0055]

[0056]

[0057] in, For the total loss function, To enhance the gradient loss of the learning strategy, it is responsible for guiding the agent to pursue the maximization of long-term cumulative rewards. To mimic the learning regularization term, the core of which is to minimize the difference between the agent's policy and the expert's policy, this application uses KL divergence (Kullback-Leibler Divergence) to quantify this difference. The regularization weight coefficient can be determined through hyperparameter tuning or by using a dynamic decay strategy: setting a large value in the early stages of training to strongly constrain the agent to imitate the expert; and gradually decreasing it as training progresses. This allows agents to explore more autonomously and surpass expert performance. The index is the expected value of random sampling. This indicates that the expectation is directed at the experience replay buffer. A set of states randomly sampled from the middle The computational experience replay buffer It stores historical experience data (states, actions, rewards, etc.) accumulated by the agent as it explores the environment. The advantage function evaluates the actions taken by agent n. The key signals of relative superiority or inferiority, within the training framework of the method described in this application, The computation is derived from the Critic network, and its process can access extended information available during training (such as local joint observation views). With joint actions This allows for the generation of more accurate value assessments, thereby guiding the optimization of strategies. The KL divergence is a nonnegativity measure in information theory that measures the difference between two probability distributions. Here, it is specifically calculated for the expert policy distribution. With agent policy distribution The differences between them. For expert strategies, i.e., in the sampling state Under these circumstances, human experts or the best historical decision-makers will take action. The conditional probability distribution can be obtained by fitting it to expert demonstration datasets such as pre-collected scheduling operation records and historical optimal power flow calculation results using statistical methods. The policy of an agent is represented by the state observed locally. At that time, the agent selects an action based on the current policy network. The probability density (for continuous actions) or probability (for discrete actions). Let n be the current observation state of agent n. For the current joint action of agent n, The parameters of the policy network, i.e., the set of trainable parameters of the agent policy network (Actor network), are optimized... To minimize the loss, We provide expert demonstration datasets. By incorporating imitation learning regularization into the loss function, we creatively unify and optimize the goals of "learning from experts" and "learning from the environment." This ensures that the agent's strategy can learn reward signals from environmental feedback without deviating excessively from proven safe and effective expert behavior during exploration, thus achieving efficient exploration under safety constraints.

[0058] The agent queries the expert policy library based on the current observation state to obtain expert-recommended actions. These actions are then input into the expert policy function for calculation, resulting in the expert policy. The agent calculates the KL divergence using both its own policy and the expert policy. Specifically, for the current observation state... Expert policies can be characterized as an empirical conditional distribution based on an expert dataset. In practical computation, this can be achieved by applying the dataset... The distribution is fitted by finding neighboring states similar to the current state and statistically analyzing their corresponding expert actions. Different calculation methods are used for KL divergence calculation depending on the action space type. For discrete actions, the KL divergence between two discrete probability distributions is directly calculated, as shown in Formula 15 below: Formula 15:

[0059] in, Denotes KL divergence, Representing the discrete action space One of the specific actions, It represents the set of all possible discrete actions. Indicating expert strategies In a given state Select action The probability, Indicates the agent's current policy In the same state Select action The probability of the true probability distribution is given by the KL divergence. When using expert strategies, distribution is employed. Encoding the agent's policy incurs additional information loss. Minimizing this divergence means driving the agent's policy... The probability distribution approximates the expert strategy as closely as possible across all action dimensions. ,when and When all actions are identical, the KL divergence is zero. For continuous actions, assuming the agent's policy follows a parameterized distribution (such as a Gaussian distribution), the expert policy can also be represented by a fitted distribution (such as a Gaussian distribution). The KL divergence between the two distributions is then calculated as shown in Equation 16 below: Formula 16:

[0060] in, Let represent a random variable in the space of continuous actions. Indicating expert strategies In a given state Next, regarding continuous actions The probability density function, Indicates the agent's current policy In the same state Next, regarding continuous actions The probability density function.

[0061] In the actual training process, during each training iteration, for each agent n, the experience replay pool is first used. A batch of states are sampled, and then each state is processed... Calculate the policy distribution of the current Actor network. Then, by querying the expert demonstration dataset or using a pre-trained expert policy model, the distribution of expert policies for the corresponding state is obtained. Next, the KL divergence regularization term is calculated, followed by the reinforcement learning loss term to calculate the total loss. Finally, the Actor network parameters are updated using gradient descent. The agent obtains a divergence threshold and compares the KL divergence with the divergence threshold. When the KL divergence is greater than or equal to the divergence threshold, the agent uses an expert-recommended action to correct the current joint action.

[0062] This paper addresses the safety takeover process based on confidence assessment. A confidence-based safety takeover mechanism addresses the issue that after an agent completes training and is deployed online, it may encounter unfamiliar scenarios not fully covered by training data due to the complex and ever-changing operation of the power system. In such scenarios, the agent's decisions may be unreliable. The safety takeover mechanism designed in this application aims to provide ultimate safety assurance in this situation. This mechanism can automatically and in real-time assess the reliability of each decision instruction from the agent within a millisecond to second timescale during online operation. Once the confidence of a decision instruction is insufficient, protective measures are immediately triggered, seamlessly transferring control to human experts, thereby completely avoiding power grid operation risks caused by AI model decision errors. To comprehensively assess the reliability of hybrid action decisions, a real-time quantitative assessment of comprehensive confidence is performed. This confidence is composed of the confidence of discrete actions and continuous actions. For the confidence of discrete actions, the normalized information entropy is calculated based on the discrete action probability distribution output by the policy network, as shown in Formula 17 below. Formula 17:

[0063] in, For discrete actions, the lower the entropy value, the more certain the decision is, and the higher the confidence level. It is the maximum entropy when all action probabilities are equal, used to normalize the entropy value to the [0,1] interval. The closer a value is to 1, the more certain the agent is about discrete decisions.

[0064] For continuous action confidence, a multiple random forward propagation technique (such as Monte Carlo Dropout) consistent with the training phase is employed. During online inference, the same state is subjected to k forward computations to obtain k consecutive action outputs. Its confidence level is evaluated by the reciprocal of the normalized variance of the output set, as shown in Formula 18 below: Formula 18:

[0065]

[0066] in, For continuous action confidence, To determine the variance based on a preset upper limit threshold of historical operating data or offline simulation, The smaller the value, the more consistent the prediction results, the more stable the decision-making, and the higher the confidence level of consecutive actions. The closer the value is to 1, the larger the variance, and the closer the confidence level is to 0. To arrive at the overall confidence level, the two confidence levels mentioned above are weighted and combined to obtain the overall confidence level for this joint decision-making process, as shown in Formula 19 below: Formula 19:

[0067] in, The overall confidence level is a value between [0,1]. A higher value indicates a higher overall confidence level in this joint decision. The first confidence level weighting coefficient is used. This is the second confidence level weighting coefficient. , For discrete action confidence, For continuous action confidence, the agent obtains a safety threshold and compares the overall confidence level with the safety threshold. When the overall confidence level is less than the safety threshold, the agent triggers an alarm so that experts can correct the agent's current joint action.

[0068] When the confidence level is too low, the system will trigger a safety takeover and closed-loop learning mechanism. During this phase, a preset confidence level safety threshold will be established. This threshold can be set to 0.8 or 0.85 depending on the system risk level. During system online operation, for each observed state... The intelligent agent will output corresponding actions. And simultaneously calculate the overall confidence level. .when When the overall confidence level falls below this safety threshold, the system determines that it is currently in a low-confidence, high-risk decision-making state. At this point, an alarm will be triggered and the system will take over, allowing human experts to conduct the final analysis and operations. The system will record complete data for this event, covering the observation status. Agent recommends actions Low confidence value And the final disposal actions of human experts These data will be used as low-confidence samples. Ultimately, this sample... It will be fed into the active learning process and used as high-quality expert-annotated data for the iterative update of the policy network.

[0069] After receiving the corrected current joint action, the agent executes the corrected joint action and calculates the next observation state and current reward value using the reward function. The current observation state, current joint action, next observation state, and current reward value are then stored in the local experience replay pool. The agent then updates the value network, policy network, and estimation network based on a multi-agent reinforcement learning algorithm until the updated policy network reaches the convergence criterion. The updated policy network is then used as the agent's target policy network. The multi-agent reinforcement learning algorithm is a distributed multi-agent soft actor-critic algorithm improved for the continuous-discrete hybrid control characteristics of power systems. This process achieves global optimization in distributed scenarios by combining boundary information and reward / penalty information interaction between agents based on local experience learning. Its core training loop is as follows: each autonomous agent interacts independently with its local environment, storing the generated experience tuples (observation, action, reward, next observation) in its local experience replay pool. The reward is a hybrid reward signal that integrates the optimization objective (minimizing network loss), rule-based penalties (voltage exceeding limits, branch power exceeding limits), and human-machine interaction feedback.

[0070] Specifically, the agent updates the value network by minimizing the soft Bellman error, as shown in Equation 20 below: Formula 20:

[0071] in, For the objective value function, This is the current reward value. As a discount factor, For the k-th target value network, For the policy entropy term, This is the next local joint observation view for agent n. For the next set of joint actions of multiple agents, Let n be the next observed state. For the next joint action of agent n, The temperature coefficient is used. This equation is introduced into the target network. To stabilize the training process and include a policy entropy term. By incentivizing exploration, this fully embodies the core idea of ​​maximum entropy reinforcement learning. In the soft Bellman equation, the objective value function... A local joint observation view of intelligent agents was adopted. The policy entropy term Based on its purely local observations This design aims to balance training effectiveness with the reliability of online execution: the Critic network leverages expanded boundary information to obtain more accurate value assessments, thereby generating higher-quality policy gradients; while the optimization objective of the Actor network (including its entropy regularization term) is strictly limited to the purely local observation space available during online execution. This ensures that the trained policy does not depend on potentially unreliable boundary communication during execution, strictly adhering to the principle of autonomous decision-making under "limited knowledge" conditions.

[0072] The agent updates the policy network using policy gradients based on the counterfactual baseline, as shown in Equation 21 below: Formula 21:

[0073]

[0074] in, Let be the gradient of the policy performance objective function with respect to the policy network parameters. Let n be the policy network parameters of agent n. Let be the policy performance objective function of agent n, which aims to maximize the expected cumulative reward. Let be the empirical average expectation of agent n, representing the expectation from the experience replay buffer. Mid-sampled state-action pairs The expectation. Let n be the local experience replay pool for agent n. Let n be the policy function of agent n, representing the locally observed state. Take action below The probability, Let be the individual advantage function of agent n, which measures the degree of advantage of its actions relative to the average level. The current set of joint actions of multiple agents. The temperature coefficient is a hyperparameter greater than 0, used to adjust the importance of the entropy term relative to the dominance function. Let be the policy entropy of agent n, representing the entropy of the policy under given observations. At that time, strategy The information entropy of the actions taken Let be the action value function of the local joint observation view of agent n, representing the joint action performed under the local observation view of agent n. The expected cumulative reward after that, Let be the expected distribution of the policy of agent n, representing the distribution of agent n according to the current policy. New sampling actions Expectations This is the current local joint observation view of agent n. Let n be the current observation state of agent n. For the current joint action of agent n, Let n be the set of current joint actions of multiple agents other than agent n. For the alternative joint action of agent n, For the new joint action, only the action of agent n is replaced with , For agent n, the current local joint observation view is constructed using the boundary coupling information estimate from the estimation network output, the current observation state, the current joint action, and limited action information obtained from neighboring agents. The individual advantage function is calculated based on this local joint observation view, aiming to achieve effective confidence allocation under limited knowledge conditions. Its rationale lies in the fact that the physical coupling of the power system has local characteristics, and the agent's actions primarily affect the neighboring region. By fusing local observations, boundary information, and neighboring actions, the key coupling effects caused by the agent's actions can be accurately captured. By calculating the value difference between a specific action and the average action under the same local view, the advantage function can accurately quantify the contribution of individual actions to the local region, and ultimately achieve the emergence of global collaborative effectiveness through interactions between agents. Counterfactual reasoning accurately quantifies the additional value of the agent's own actions, thus clearly allocating confidence in a distributed architecture and guiding each agent to learn the globally optimal local policy.

[0075] The agent employs supervised learning to update the estimation network. To reduce reliance on communication, the Estimate network fits the boundary coupling node voltages in real time based on local observations, providing the Critic with stable neighborhood state estimates. Simultaneously, low-confidence samples generated during training or key samples actively labeled through learning are continuously fed back into the local experience pool, driving the policy to continuously evolve and optimize within a closed loop of "decision-evaluation-intervention-learning," ensuring the system's continuous self-improvement and reliable operation under limited knowledge conditions.

[0076] Through the improvement of the above strategy, a distributed multi-agent soft actor-critic (DMASAC) algorithm suitable for this application is obtained. This algorithm follows the maximum entropy reinforcement learning framework, and its fundamental training objective is to find an optimal policy for each agent n that combines high reward with strong exploratory power. This objective is formally defined as follows: Equation 22: Formula 22:

[0077] in, This represents the optimal strategy. This indicates the search for the expected value function. This is the comprehensive reward signal of agent n at time t, which includes the basic reward item and the human-computer interaction reward item. Indicates by strategy induced state action trajectory distribution For agent n, under the observations, according to the policy The entropy of taking an action This is a temperature coefficient used to adjust the importance of action entropy relative to reward.

[0078] The Actor network and Critic network work together through an iterative optimization process, the core of which is a loop of "policy selection - value evaluation - parameter update": the training process starts from a randomly initialized policy. To begin, in each step, firstly, each Actor network, based on its local observation state... Actions are selected in a decentralized manner. Then, each Critic network evaluates the long-term value of the joint actions based on a local joint view comprised of the Actor's local observations, actions, and boundary information. Finally, the system uses this value assessment to calculate the policy gradient and update the parameters of the Actor network, thereby generating a new policy with improved performance. This cycle repeats continuously, driving the strategy to evolve gradually to its optimal state under strictly limited knowledge conditions.

[0079] 202. Each agent collects limited local observation information within its jurisdiction and generates a local observation status based on the limited local observation information.

[0080] In this embodiment, each agent responsible for a specific area of ​​the power system (such as a distribution network area or a distributed power generation cluster area) only collects limited local observation information within its own jurisdiction (specifically including data such as node voltage, real-time load power, distributed power generation output, and local equipment operating status within the area), and integrates these scattered local data into a local observation status reflecting the current operating status of the area. At the same time, these agents are deployed at the edge nodes of the power system (i.e., physical end locations close to local equipment, such as distribution terminals of distribution areas or grid connection points of distributed power generation), and the observation status they generate is strictly limited to the area under their jurisdiction, without needing to obtain global operating information of the entire power system across regions. This approach offers several advantages: First, it significantly reduces data transmission costs and communication bandwidth pressure, eliminating the need to collect and transmit massive amounts of global data. Second, edge-side deployment and local information processing can significantly shorten the latency between data collection and status generation, meeting the real-time requirements of power system allocation. Third, it avoids the risks of privacy leaks and data security associated with global information collection, while simplifying system architecture complexity and avoiding the computational burden of centralized global data processing. Fourth, the strong correlation between local observation status and equipment in the jurisdiction makes subsequent decisions more aligned with the actual operating scenarios of the region, improving the accuracy of power allocation.

[0081] 203. Each agent inputs its corresponding local observation state into the corresponding target policy network to obtain a joint target action consisting of active power control instructions and reactive power control instructions.

[0082] In this embodiment, each agent reflects the local observation status of its jurisdiction's operating status and inputs it into a target policy network trained by a multi-agent reinforcement learning algorithm. This network outputs a joint action that simultaneously includes active power control commands (such as distributed power generation output adjustment values ​​and energy storage charging and discharging power commands) and reactive power control commands (such as capacitor bank switching commands and reactive power compensation device adjustment amounts). Simultaneously, the agent incorporates two types of mechanisms to support safe and rational decision-making: one is action constraints determined based on the power grid safety operation rule base, which correspond to the actual safety of the power grid operation. The first is a set of specifications (such as node voltage amplitude range, equipment power capacity limit, branch transmission power limit, etc.), which directly restrict the action space of the agent and prevent it from outputting illegal commands from the source. The second is a dual-guided reward function, which includes both a penalty for violating safety rules (such as triggering negative rewards when voltage exceeds the limit or power exceeds the capacity), which encourages the agent to actively avoid illegal operations during training and operation through negative incentives, and an interactive feedback from grid dispatchers (i.e., human-machine interaction reward items), which transforms the expert's evaluation of the rationality of the control action into a reward signal, guiding the agent to learn strategies that are more in line with the actual scenario. Action constraints lock in safety boundaries from the decision-making source, directly reducing the risks to power grid operation caused by agent misoperation; violation penalties strengthen the agent's compliance with safety rules, enabling the strategy to be optimized more accurately towards compliance during training; human-computer interaction rewards incorporate the practical experience of scheduling experts, making up for the limitations of purely data-driven strategies and improving the actual adaptability of control actions; target joint actions cover both active and reactive power control, meeting the coordinated needs of power grid power balance and voltage regulation, and can more efficiently achieve the goal of autonomous power allocation on the edge side.

[0083] 204. Each intelligent agent coordinates the control of multiple local execution devices within its jurisdiction based on the corresponding target joint action.

[0084] In this embodiment, each agent coordinates the scheduling of multiple local execution devices within its jurisdiction based on the target joint action output by the target policy network. These devices include active power regulation devices (such as distributed photovoltaics, energy storage charging and discharging devices, micro gas turbines, etc.) and reactive power compensation devices (such as parallel capacitor banks, static var compensators, reactive power regulation modules of photovoltaic inverters, etc.). The entire control process is completed under the constraint of "limited knowledge" (relying only on the agent's own local observation state and the edge coupling information transmitted by neighboring agents, such as boundary node voltage and power interaction between adjacent regions), ultimately achieving autonomous and coordinated allocation of active and reactive power on the edge side, ensuring the safe and stable operation of the power grid in this area. At the same time, the decision-making of each agent does not require the acquisition of the global model of the power system (such as the entire power grid topology) and the global operating status (such as the equipment output and load data of the remote area). By coordinating the control of active and reactive power equipment, the system can simultaneously optimize power balance and voltage regulation within the region, making it more suitable for the multi-objective requirements of power grid operation. Relying solely on local observation and information coupled with adjacent edge devices for decision-making avoids the communication costs, bandwidth pressure, and data security risks associated with global information collection, while also reducing reliance on centralized computing power. Moreover, the edge-side autonomous allocation mode has a fast response speed and can promptly address sudden situations such as power fluctuations and voltage deviations within the region. Furthermore, the distributed decision-making architecture enhances system robustness, ensuring that local failures of individual agents do not affect the overall operation of the global system.

[0085] 205. When the power system topology changes, each agent updates the target policy network through knowledge distillation.

[0086] In this embodiment, each agent performs joint actions to coordinate the control of active power regulation and reactive power compensation devices within its jurisdiction, achieving autonomous allocation of active and reactive power at the edge. When the power system topology changes, each agent utilizes transfer learning to adapt to the topology change. Specifically, the agent policy network parameters trained under the original topology are used as teacher models to initialize the agent policy network under the new topology. Knowledge distillation or policy guidance accelerates the training and convergence process of the policy network under the new topology. After training, each agent outputs control commands in real time based on local observations and boundary coupling information (such as adjacent node voltages). The active power regulation devices include distributed generation equipment and energy storage systems, balancing power output by adjusting output. The reactive power compensation devices include photovoltaic inverters, static var compensators (SVCs), and capacitor banks, supporting voltage by adjusting reactive power. Control commands can respond in milliseconds, achieving autonomous allocation of active / reactive power at the edge without requiring an upper-level centralized controller.

[0087] This application proposes a knowledge distillation-based transfer learning mechanism to address the technical problems of "policy inaccuracy caused by topology changes" and "high retraining costs." The core of this mechanism is to treat the well-trained agent in the original topology as the teacher model, transferring its knowledge to the student model in the new topology through distillation loss. This guides the student model to converge quickly, achieving efficient transfer of policy knowledge between different topologies. The teacher model refers to the agent trained in the original topology... In this environment, an agent policy network that has been fully trained until it converges. During transfer learning, its network parameters It will be frozen, serving only as a stable source of knowledge, and will not participate in gradient updates. The student model, however, needs to be updated in the new topology. The policy network is retrained. Its parameters These are parameters to be optimized. Despite the topology changes, the definition of the agent's local observation space remains consistent (e.g., both include node voltages and load power). Therefore, the teacher and student models have the same input dimension, ensuring the feasibility of knowledge transfer.

[0088] The agent uses the target policy network as the teacher model and initializes an agent policy network as the student model for the new power grid topology. Knowledge distillation loss is calculated using the teacher and student models, and the target loss function is constructed using the knowledge distillation loss and the multi-agent reinforcement learning loss for the new power grid topology. Specifically, the training objective of the student model is a multi-objective optimization problem, which aims to both mimic the decision-making style of the teacher model to inherit its robustness and adapt to the new topology environment to complete the optimization task. To this end, a composite loss function is designed, as shown in Equation 23 below: Formula 23:

[0089] in, Let be the target loss function. The knowledge distillation loss is used to force the student model to mimic the decision probability distribution of the teacher model. For the multi-agent reinforcement learning loss of the new power grid topology, These are dynamic weighting coefficients used to balance the importance of imitating the teacher and adapting to the new environment. To learn the model parameters.

[0090] Knowledge distillation loss The core objective is to minimize the difference between the output action probability distributions of the teacher model and the student model. This application uses KL (Kullback-Leibler) divergence as a metric for this difference. For the same state observation sampled from the empirical replay pool, the loss is calculated as shown in Equation 24 below: Formula 24:

[0091]

[0092] in, For knowledge distillation loss, Let KL divergence be the KL divergence. For the policy probability distribution of the teacher model, This represents the policy probability distribution for the student model. Due to the characteristics of the mixed action space, the KL divergence needs to be calculated separately for the discrete and continuous parts. For discrete actions, the KL divergence between the discrete action probability distributions output by the two models can be directly calculated. For continuous actions, if the policy output is a parameterized continuous distribution (e.g., a Gaussian distribution with output mean and variance), then the KL divergence between these two continuous probability distributions should be calculated.

[0093] The emphasis on imitation and exploration should differ at different stages of training. Therefore, a dynamic weight decay strategy is designed, where the weight coefficient decreases as the number of training rounds increases. The specific calculation formula is shown in Formula 25 below: Formula 25:

[0094] in, As the initial weights, The attenuation rate, This represents the current training round number.

[0095] The complete training process is as follows: First, the student model parameters are initialized to the teacher model parameters. Then, the student model explores the new topology environment, and the interaction data is stored in the experience pool. Next, data is sampled from the experience pool, the reinforcement learning loss and knowledge distillation loss of the student model are calculated, and the total loss is calculated based on the dynamic weights. The student model is then updated by minimizing the objective total loss function using gradient descent. Finally, the above interaction, sampling calculation, and loss update steps are repeated until the student model reaches the convergence criterion under the new power grid topology, that is, the performance converges under the new topology. The updated student model is then used as the target policy network under the new power grid topology.

[0096] like Figure 5 The diagram illustrates an edge-side power autonomous allocation system architecture guided by limited knowledge. This system primarily comprises a power system physical layer, an edge agent layer (distributed decision-making), and a local execution device layer, and optionally includes a central system. Through multiple autonomous agents distributed at the edge, collaborative autonomous allocation of active and reactive power is achieved under conditions of limited communication and limited knowledge. The system architecture includes the following three core layers: 1. Local Execution Equipment Layer: This layer consists of clusters of devices such as photovoltaic inverters, energy storage systems, and static var compensators (SVCs) within each edge autonomous unit (e.g., distribution substation, microgrid). It is responsible for executing power regulation commands issued by the intelligent agents. These devices are directly connected to the edge nodes of the power system, enabling rapid power regulation. As the execution end for edge-side power regulation, the local execution equipment layer includes equipment clusters 1, 2, and m, each corresponding to a different autonomous intelligent agent. Equipment cluster 1 is equipped with photovoltaic inverters and energy storage systems; equipment cluster 2 is equipped with photovoltaic inverters and SVCs; and equipment cluster m includes photovoltaic inverters, capacitor banks, and energy storage systems. These devices respectively receive control commands output by their respective autonomous intelligent agents in the edge intelligent agent layer. Through photovoltaic inverters regulating active power output, energy storage systems performing charging and discharging, and SVCs and capacitor banks completing reactive power compensation, the autonomous allocation of active and reactive power on the edge side is ultimately achieved.

[0097] 2. Edge Agent Layer: Consists of multiple autonomous agents deployed in a distributed manner. Each agent obtains its state through a local perception module. A local joint observation view is formed through the information fusion module. ), and the policy network outputs joint actions ( Each agent is embedded in an edge computing device, achieving distributed collaborative decision-making through local perception and limited communication. The edge agent layer contains autonomous agents 1, 2 to m, each following a unified decision-making chain. Taking autonomous agent 1 as an example, it first obtains operational data of its jurisdiction through local perception. This includes voltage (v), active power (p), reactive power (q), and equipment state of charge (SOC), etc.; then, through information fusion, Boundary estimation information obtained from autonomous agent 2 via limited communication ,award Integrating into a merged state ; then The input policy network is Actor-Critic, and the output includes active power control instructions. Reactive power control commands joint action The processes of the other autonomous agents are consistent with those of autonomous agent 1, exchanging boundary estimation information through limited communication. In coordination with the reward r, the combined actions output by all agents are ultimately transmitted as control commands to the local execution device layer.

[0098] 3. Power System Physical Layer: This layer represents the objective power grid environment controlled and operated by the described method, including the distribution network, distributed energy resources, and loads on the edge side. The main power grid is connected to the next-level edge autonomous unit 1 via feeders, enabling power interaction between the main grid and the edge side. Edge autonomous unit 1 contains nodes 1 and 2, and edge autonomous unit 2 contains nodes 3 and 4. They are interconnected via tie lines to complete power dispatch. The power system physical layer establishes a connection with the edge intelligent agent layer through physical coupling links, transmitting physical layer operational data for intelligent decision-making on the edge side.

[0099] In addition, the system includes a lightweight, optional central system responsible for providing expert interaction interfaces, maintaining the transfer learning engine, and performing overall AI and hybrid management. This central system, as an auxiliary module, comprises two core functional units: an expert interaction interface responsible for updating the rule base, human intervention in agent decision-making, and storing experience data; and a transfer learning library, whose activation condition is limited to power system topology changes, used to achieve policy knowledge transfer and adaptation. The central system establishes a connection with the underlying edge agent layer (distributed decision-making) through a human-machine hybrid enhanced link, providing expert experience and policy transfer support for the edge agents' distributed decision-making.

[0100] like Figure 6 The flowchart shown is for a novel edge-side power autonomous allocation method for power systems guided by limited knowledge. After the process starts, step S1 is executed first: deploying an edge-side autonomous agent and generating a local observation state based on limited local observation information within the agent's jurisdiction. Next, step S2 is executed: constructing rules and rewards, simultaneously completing two operations: first, pre-setting a grid safety operation rule base; second, constructing a hybrid reward function. Then, step S3 is executed: human-machine hybrid augmented training and decision-making. At this point, a system topology change is assessed. If the change is confirmed, transfer learning knowledge distillation is performed before returning to step S3; if not, the training and decision-making in step S3 are directly advanced. After completing step S3, step S4 is executed: execution and coordinated control, controlling the active power regulation equipment and reactive power regulation equipment respectively to achieve edge-side power autonomous allocation.

[0101] On the one hand, this application conducts deployment and state awareness work on distributed intelligent agents, and constructs a guidance and constraint mechanism based on pre-defined rules to restrict the actions of the agents. Specifically, autonomous intelligent agents are deployed at edge nodes in multiple electrically coupled areas of the power system (e.g., distribution substations, microgrids, feeder segments). Each autonomous intelligent agent generates a local observation state based on corresponding limited local observation information (e.g., node voltage, injected power, etc.), and sets action constraints and penalty terms in the reward function for each autonomous intelligent agent according to a pre-defined power grid safety operation rule library (including voltage safety constraints, equipment capacity constraints, branch power constraints, etc.), while incorporating human-computer interaction reward terms into the reward function.

[0102] On the other hand, this application employs a distributed multi-agent soft actor-critic (DMASAC) framework based on improved hybrid control characteristics of power systems for collaborative decision-making, and incorporates a confidence allocation and human-machine hybrid reinforcement training mechanism for the power grid. Guided by this mechanism, the policy networks of each agent are trained to generate power control commands. The human-machine hybrid reinforcement training mechanism for the power grid introduces the intervention of power grid dispatching experts during the training process. Specifically, this includes labeling key samples (such as voltage critical states and equipment action boundaries), incorporating expert dispatching decision-making experience through imitation learning, and human intervention when the agent's decision confidence falls below a preset threshold. This injects human expert knowledge into the agent's learning process, ensuring the reliability and security of the final decision-making strategy.

[0103] Finally, each autonomous agent executes joint actions output by the policy network to perform distributed collaborative control of active power regulation equipment and / or reactive power compensation equipment within its jurisdiction, thereby completing the autonomous allocation of active and reactive power at the edge. During the execution of joint actions, a pre-set power grid safety operation rule base is first used to perform power system safety verification on the actions to ensure that they meet basic safety constraints.

[0104] When the power system topology changes (such as planned maintenance or network reconfiguration after fault isolation), the knowledge distillation-based transfer learning process designed in this application is initiated. The parameters of the agent policy network that have been trained under the original topology are used as the teacher model to initialize the agent policy network under the new topology. Through knowledge distillation, the training and convergence process of the policy network under the new topology is accelerated to maintain the system's continuous autonomous operation capability.

[0105] Compared with existing research, this application achieves the following beneficial effects: (1) By leveraging the improved distributed multi-agent architecture and limited knowledge decision-making mechanism, a truly localized autonomous distributed collaborative control that fits the physical characteristics of the distribution network is achieved, fundamentally reducing the dependence on global information, protecting the data privacy of the power grid and users, and improving the system response speed. (2) Through the dual guarantee of pre-set rule constraints that comply with power system safety regulations and human-machine hybrid enhancement, the safety and reliability of the new power system are significantly enhanced even when model knowledge and data knowledge are incomplete. (3) By utilizing the confidence allocation mechanism and the transfer learning method that adapts to changes in power grid topology, the learning efficiency of the agent in a knowledge-constrained environment and its adaptability to system changes are improved. (4) Through systematic guidance under a limited knowledge framework, rapid autonomous allocation of active and reactive power on the edge side is realized, effectively improving the autonomous capability and operating efficiency of the power grid.

[0106] This application provides a novel edge-side power self-regulation allocation method for power systems guided by limited knowledge. Compared with existing technologies, this application method involves each agent collecting limited local observation information within its jurisdiction and generating local observation states based on this information. The corresponding local observation states are then input into the corresponding target policy network to obtain a target joint action consisting of active power control commands and reactive power control commands. Based on this target joint action, multiple local execution devices within the jurisdiction are then coordinated for control. The agents are deployed on edge-side nodes of the power system and include action constraints and reward functions determined based on a power grid safety operation rule base. The reward function includes penalty terms and human-machine interaction reward terms. The target policy network is trained using a multi-agent reinforcement learning algorithm under a human-machine hybrid reinforcement mechanism, utilizing the action constraints and reward functions. Edge-side deployment and limited local observation eliminate the need for centralized collection of massive global data, reducing data transmission and processing costs while enabling low-latency response through edge-side decision-making, thus meeting the real-time requirements of the power system. Furthermore, the policy network training incorporates action constraints based on power grid safety rules, defining the boundaries of safety decisions from the source. Combined with a reward function containing penalties and human-machine interaction, and a hybrid human-machine enhancement mechanism, this improves policy learning efficiency while incorporating expert experience to reduce safety risks. The distributed multi-agent cooperative control mode of this application allows each agent to make autonomous decisions and collaboratively manage equipment based on local observations, freeing it from dependence on centralized architecture and generalized data acquisition, while simultaneously improving the accuracy of power allocation at the edge.

[0107] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification. The above embodiments only illustrate several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of this application's patent. It should be pointed out that for those skilled in the art, several modifications and improvements can be made without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims. Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred implementation scenario, and the modules or processes in the drawings are not necessarily necessary for implementing this application. The above application serial numbers are merely for description and do not represent the superiority or inferiority of the implementation scenario. The above disclosures are only a few specific implementation scenarios of this application. However, this application is not limited to these. Any variations that can be conceived by those skilled in the art should fall within the protection scope of this application.

Claims

1. A novel power self-regulating allocation method for the edge side of a power system guided by limited knowledge, characterized in that, include: Each agent collects limited local observation information within its jurisdiction and generates a local observation state based on the limited local observation information. The agents are deployed on the edge nodes of the power system. Each agent inputs its corresponding local observation state into the corresponding target policy network to obtain a joint action consisting of active power control instructions and reactive power control instructions. The agent includes action constraints and a reward function determined based on the power grid safe operation rule base. The reward function includes a penalty term and a human-machine interaction reward term. The target policy network is trained using the action constraints and the reward function through a multi-agent reinforcement learning algorithm under a human-machine hybrid reinforcement mechanism. Each intelligent agent coordinates the control of multiple local execution devices within its jurisdiction based on the corresponding target joint action.

2. The method according to claim 1, characterized in that, Each agent collects limited local observation information within its jurisdiction and generates a local observation state based on the limited local observation information, including: For each agent, the agent identifies multiple nodes within its jurisdiction and extracts the node voltage, load active power, load reactive power, photovoltaic active power output, energy storage device state of charge, CB tap position, and CB activation count for each node from the agent's limited local observation information to generate the agent's local observation state. in, Let be the local observation state vector of node i in the agent at time t. Let be the node voltage of node i in the intelligent agent at time t. Let be the load active power of node i in the intelligent agent at time t. Let be the load reactive power of node i in the intelligent agent at time t. Let t be the active power output of photovoltaic power generation of node i in the intelligent agent at time t. Let be the state of charge of the energy storage device of node i in the intelligent agent at time t. Let be the CB gear position of node i in the agent at time t. Let be the number of CB actions performed by node i in the agent at time t.

3. The method according to claim 1, characterized in that, The method further includes: For each agent, the agent initializes the policy network, value network, and estimation network; The intelligent agent acquires local limited observation information samples within its jurisdiction and generates the current observation state based on the local limited observation information samples; The agent inputs the current observation state into the policy network to obtain the current joint action. The policy network includes a discrete action output layer using the Softmax function and a continuous action output layer using the Tanh function. The discrete action output layer is used to output the current discrete action, and the continuous action output layer is used to output the current continuous action. The current joint action is composed of the current discrete action and the current continuous action. The intelligent agent modifies the current joint action based on the action constraints and the human-machine hybrid enhancement mechanism; The agent executes the corrected current joint action, calculates the next observation state and the current reward value using the reward function, and stores the current observation state, the current joint action, the next observation state, and the current reward value in the local experience replay pool. The agent updates the value network, the policy network, and the estimation network based on the multi-agent reinforcement learning algorithm until the updated policy network reaches the convergence criterion. The updated policy network is then used as the target policy network for the agent. The multi-agent reinforcement learning algorithm is a distributed multi-agent soft actor-commentator algorithm improved for the continuous-discrete hybrid control characteristics of power systems.

4. The method according to claim 3, characterized in that, The method further includes: The intelligent agent determines the action constraints based on the power grid safety operation rule base. The action constraints are used to limit the action space of the intelligent agent. The power grid safety operation rule base includes at least one of voltage safety operation constraints, equipment power capacity constraints, branch transmission power constraints, and energy storage charge and discharge times constraints. The action constraints are used to limit the action space of the intelligent agent. The agent determines the reward function based on the power grid safety operation rule base. in, Let be the reward function of agent n at time t. As the first weighting coefficient, This is the second weighting coefficient. The third weighting coefficient, Let n be the global network loss of agent n. This is the penalty coefficient for exceeding the limit. The total number of intelligent agents. Let n be the set of nodes within the jurisdiction of agent n. Let be the voltage amplitude at node j at time t. Let be the lower limit of the voltage amplitude at node j at time t. Let be the upper limit of the voltage amplitude at node j at time t. Let n be the set of branches within the jurisdiction of agent n. This represents the maximum limit of active power transmitted by branch ij. Let be the active power transmitted by branch ij at time t. The human-computer interaction reward item is defined for agent n, and the human-computer interaction reward item is determined through the human-computer hybrid enhancement mechanism.

5. The method according to claim 3, characterized in that, The agent updates the value network, the policy network, and the estimation network based on the multi-agent reinforcement learning algorithm, including: The agent updates the value network by minimizing the soft Bellman error. in, For the objective value function, This is the current reward value. As a discount factor, For the k-th target value network, For the policy entropy term, This is the next local joint observation view for agent n. For the next set of joint actions of multiple agents, Let n be the next observed state. For the next joint action of agent n, Temperature coefficient; The agent updates the policy network using policy gradients based on a counterfactual baseline. in, Let be the gradient of the policy performance objective function with respect to the policy network parameters. Let n be the policy network parameters of agent n. Let be the policy performance objective function for agent n. Let n be the empirical average expectation of agent n. Let n be the local experience replay pool for agent n. Let n be the policy function of agent n. Let be the individual advantage function of agent n. The current set of joint actions of multiple agents. For temperature coefficient, Let n be the policy entropy of agent n. Let n be the action value function of the local joint observation view of agent n. Let n be the expected policy distribution of agent n. This is the current local joint observation view of agent n. Let n be the current observation state of agent n. For the current joint action of agent n, Let n be the set of current joint actions of multiple agents other than agent n. For the alternative joint action of agent n, For the current policy of agent n, the current local joint observation view is constructed using the boundary coupling information estimate output by the estimation network, the current observation state, and the current joint action; The agent updates the estimation network using supervised learning.

6. The method according to claim 3, characterized in that, The agent modifies the current joint action based on the action constraints and the human-machine hybrid enhancement mechanism, including: The intelligent agent uses the action constraints to perform a security verification operation on the current joint action. If the security verification operation passes, the intelligent agent will use the human-machine hybrid enhancement mechanism to correct the current joint action. The human-machine hybrid enhancement mechanism includes at least one of sample labeling processing, imitation learning processing, active learning processing, and security takeover processing.

7. The method according to claim 6, characterized in that, The method further includes: The agent calculates an uncertainty score using the current observation state. in, Score the uncertainty. The information entropy of discrete actions, The variance of continuous actions, For a discrete set of actions, This is the upper limit threshold of variance. Let n be the current observation state of agent n. The first weighting coefficient, This is the second weighting coefficient; The agent acquires an uncertainty threshold and compares the uncertainty score with the uncertainty threshold. If the uncertainty score is greater than the uncertainty threshold, the agent marks the current observation state and the current joint action as key samples and triggers an expert intervention request. In response to the expert intervention request, the intelligent agent pushes the current observation state and the current joint action to the terminal held by the expert. The agent receives the corrected current joint action uploaded by the expert based on the terminal it holds.

8. The method according to claim 6, characterized in that, The method further includes: The intelligent agent extracts scheduling expert decision records from historical power grid operation data, establishes an expert demonstration dataset using the scheduling expert decision records, and obtains an expert policy function by fitting the expert demonstration dataset. The agent queries the expert policy library based on the current observation state to obtain expert recommended actions, and inputs the expert recommended actions into the expert policy function for calculation to obtain the expert policy; The agent calculates the KL divergence using its own policy and the expert policy. The agent obtains a divergence threshold and compares the KL divergence with the divergence threshold. When the KL divergence is greater than or equal to the divergence threshold, the agent uses the expert-recommended action to correct the current joint action; The agent updates its policy network based on a total loss function that includes a regularization term for imitation learning. in, Let the total loss function be... To enhance the gradient loss of the learning strategy, To learn regular expressions by imitation, These are the regularization weight coefficients. For random sampling expectation, For the dominant function, Let KL divergence be the KL divergence. For the aforementioned expert strategy, The strategy for the intelligent agent, This represents the current observation state of the agent n. For the current joint action of the agent n, The parameters of the policy network are... This is the expert demonstration dataset.

9. The method according to claim 6, characterized in that, The method further includes: The agent calculates the overall confidence level. in, To assess the overall confidence level, The first confidence level weighting coefficient is used. This is the second confidence level weighting coefficient. For discrete action confidence, For continuous action confidence; The agent obtains a security threshold and compares the overall confidence level with the security threshold. When the overall confidence level is less than the safety threshold, the agent triggers an alarm so that experts can correct the agent's current joint action.

10. The method according to claim 1, characterized in that, The method further includes: When the power system topology changes, the agent uses the target policy network as the teacher model and initializes an agent policy network as the student model for the new power grid topology. The agent calculates the knowledge distillation loss using the teacher model and the student model, and constructs a target loss function using the knowledge distillation loss and the multi-agent reinforcement learning loss of the new power grid topology. in, Let the target loss function be... For the knowledge distillation loss, The multi-agent reinforcement learning loss for the new power grid topology is... For dynamic weighting coefficients, To learn model parameters; The agent updates the student model by minimizing the target total loss function using gradient descent until the updated student model reaches the convergence criterion under the new power grid topology. The agent uses the updated student model as the target policy network under the new power grid topology.