Power grid regulation method, system, device and medium based on agent reinforcement learning

By employing a power grid control method based on agent reinforcement learning, the power grid dispatching mode and resource allocation are dynamically adjusted, solving the response lag and reliability problems of traditional power grid dispatching systems in complex scenarios, and achieving efficient and reliable adaptive power grid dispatching.

CN122118975AInactive Publication Date: 2026-05-29国网浙江省电力有限公司嵊州市供电公司

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
国网浙江省电力有限公司嵊州市供电公司
Filing Date
2026-04-29
Publication Date
2026-05-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional power grid dispatching systems cannot respond effectively to complex power grid scenarios in real time, resulting in low dispatching efficiency and poor reliability. They are unable to make dynamic optimization decisions, especially when there are random fluctuations in new energy output and sudden load changes, making it difficult to detect fluctuation trends in time, missing the best adjustment window, and affecting the reliable and stable operation of the power grid.

Method used

A power grid control method based on agent reinforcement learning is adopted. By calculating the fluctuation characteristics of power grid operation indicators, abnormal states are identified, the target scheduling mode is dynamically adjusted, and operation and control strategies are generated. Furthermore, the operation strategies of the power grid agent are optimized through deep reinforcement learning to achieve adaptive scheduling control.

Benefits of technology

It improves the power grid's sensitivity to abnormal situations and response speed, enables precise hierarchical processing and rational allocation of resources, enhances the efficiency and reliability of power grid dispatching, avoids excessive equipment operation, ensures the power grid remains economical during normal fluctuations, and guarantees stability during severe anomalies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122118975A_ABST
    Figure CN122118975A_ABST
Patent Text Reader

Abstract

The application discloses a power grid regulation method, system, device and medium based on agent reinforcement learning, relates to the field of power grid automatic dispatching, and comprises the following steps: calculating fluctuation characteristic values of operation indexes based on power grid power operation data; judging whether abnormal fluctuation occurs in the power grid based on the fluctuation characteristic values, and if yes, determining a target dispatching mode according to the abnormal degree; controlling corresponding power grid agents to generate an operation regulation strategy and execute a strategy action according to the target dispatching mode; and performing reinforcement learning training on the power grid agents based on operation data of the power grid agents and change data of the operation indexes, so as to optimize the operation regulation strategy and perform dispatching control on the power grid. The application can dynamically perceive the power grid operation state, match a targeted and highly adaptive dispatching strategy, realizes self-adaptive dispatching control on the power grid, and improves the power grid dispatching efficiency and reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power grid automated dispatching, specifically to power grid control methods, systems, equipment, and media based on agent reinforcement learning. Background Technology

[0002] In the power system field, the power grid dispatch automation control system undertakes the core tasks of real-time monitoring of power grid operation status, coordinating resource allocation, and ensuring stable power supply. With the large-scale integration of new energy sources, the increasing complexity of load characteristics, and the widespread deployment of distributed power sources, the power grid exhibits high-dimensionality, strong coupling, and time-varying characteristics, placing higher demands on the dispatch system's rapid response, global optimization, and risk prevention capabilities. Traditional power grid dispatch systems rely on SCADA (Supervisory Control and Data Acquisition) and EMS (Energy Management System) for data acquisition, use static threshold analysis to classify power grid operation status, and rely on manually written rule bases for dispatching, which can only meet steady-state requirements. When facing dynamic scenarios such as random fluctuations in new energy output and sudden load changes, fixed thresholds cannot detect fluctuation trends in a timely manner, easily missing the optimal adjustment window, resulting in delayed power grid response. Furthermore, preset rules are difficult to cover complex coupled scenarios and cannot effectively integrate real-time data with historical experience, thus failing to dynamically optimize decision-making schemes based on new power grid conditions. The cross-regional resource optimization capability is also poor, further affecting the reliable and stable operation of the power grid. Summary of the Invention

[0003] The purpose of this application is to address the problems of low grid dispatch efficiency and poor reliability caused by the inability of existing power grid automated dispatch methods to respond effectively to complex power grid scenarios in real time and to dynamically optimize decision-making. This application proposes a power grid control method, system, equipment, and medium based on agent-based reinforcement learning. By calculating the fluctuation characteristics of power grid operating indicators, abnormal states of the power grid are determined. Based on the degree of abnormality, a target dispatch mode is determined to generate the operation and control strategy of the power grid agent. This allows for dynamic perception of the power grid operating state to match targeted and highly adaptable dispatch strategies. Furthermore, based on the operating data of the power grid agent, reinforcement learning is used to optimize the operating strategy, enabling adaptive dispatch control of the power grid and improving grid dispatch efficiency and reliability.

[0004] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows: In a first aspect, embodiments of this application provide a power grid control method based on agent reinforcement learning, the method comprising: The fluctuation characteristic value of the operation index is calculated based on the power operation data of the power grid; based on the fluctuation characteristic value, it is determined whether the power grid has abnormal fluctuations. If so, the target scheduling mode is determined according to the degree of abnormality; based on the target scheduling mode, deep reinforcement learning is used to control the corresponding power grid agent to generate operation and control strategies and execute strategy actions; based on the operation data of the power grid agent and the change data of the operation index, the power grid agent is trained by reinforcement learning to optimize the operation and control strategies and to perform scheduling and control of the power grid.

[0005] In this scheme, fluctuation characteristic values ​​are calculated in real time using the power grid's operating indicators to identify abnormal fluctuations in the power grid, thereby improving the sensitivity and response speed to power grid anomalies. Target scheduling modes are determined based on the degree of anomaly, thus obtaining targeted operation and control strategies. This enables precise hierarchical processing of anomalies of different degrees, helping to improve the rational allocation of resources by the power grid agent. Furthermore, by conducting reinforcement learning training on the power grid agent, operation and control strategies are optimized, achieving the agent's adaptive adjustment capability to operation and control strategies, improving the agent's resource optimization and allocation capabilities, and further enhancing power grid dispatch efficiency and reliability.

[0006] Preferably, the calculation of fluctuation characteristic values ​​of operating indicators based on power grid operation data includes: acquiring real-time power grid operation data; calculating the indicator values ​​of power grid operation indicators, wherein the operating indicators include at least: voltage deviation indicator, frequency deviation indicator, line load indicator, and output fluctuation indicator of distributed power sources; calculating the average indicator values ​​of the voltage deviation indicator, frequency deviation indicator, line load indicator, and output fluctuation indicator within a period; and calculating the rolling variance of each operating indicator within a period based on the average indicator values ​​of the voltage deviation indicator, frequency deviation indicator, line load indicator, and output fluctuation indicator to obtain the fluctuation characteristic values ​​of each operating indicator.

[0007] Preferably, the step of determining whether the power grid has experienced abnormal fluctuations based on the fluctuation characteristic value, and if so, determining the target scheduling mode according to the degree of abnormality, includes: determining that the power grid has experienced abnormal fluctuations when the fluctuation characteristic value is greater than the historical normal fluctuation value; The trigger intensity level of the power grid agent is determined based on the scheduling decision triggering mechanism, and it is used as the degree of anomaly of the power grid to match the corresponding target scheduling mode.

[0008] Preferably, determining the trigger intensity level of the power grid agent based on the scheduling decision triggering mechanism, and using it as the degree of anomaly to match the corresponding target scheduling mode, includes: determining an initial trigger threshold matching each of the operating indicators based on historical normal operation data of the power grid; calculating a target trigger threshold based on the fluctuation characteristic value, the historical normal fluctuation value, and the initial trigger threshold; comparing the indicator value of each operating indicator with the corresponding target trigger threshold, and determining the trigger intensity level based on the degree of exceeding the target trigger threshold when the indicator value of an operating indicator exceeds the target trigger threshold, wherein the trigger intensity level is used to characterize the degree of anomaly of the power grid; and determining the scheduling mode with the higher trigger intensity level as the target scheduling mode.

[0009] Preferably, the target scheduling mode includes at least a fine-tuning mode and a global mode. The step of using deep reinforcement learning to control the corresponding power grid agent to generate operation control strategies and execute strategy actions based on the target scheduling mode includes: when the target scheduling mode is a fine-tuning mode, activating the power grid agent in the corresponding scheduling area, triggering the power grid agent to generate corresponding operation control strategies according to scheduling requirements, adjusting the operation parameters according to the operation control strategies, and working according to the adjusted operation parameters. When the target scheduling mode is global mode, the power grid agents in multiple scheduling regions are activated, and the decision weights of each power grid agent are dynamically adjusted to generate a collaborative operation and control strategy. The system then operates according to this collaborative operation and control strategy. Specifically, this includes: each power grid agent generating a local adjustment strategy based on a deep reinforcement learning algorithm, and simultaneously broadcasting scheduling trigger events to neighboring agents via pheromone communication links; obtaining the anomaly level based on the scheduling trigger events, and dynamically adjusting the decision weights of the neighboring agents to generate a target local adjustment strategy; and constructing a collaborative operation and control strategy for multiple scheduling regions based on the target local adjustment strategies of each power grid agent, and operating according to this collaborative operation and control strategy.

[0010] Preferably, the step of training the power grid agent using reinforcement learning based on its operational data and the change data of its operational indicators to optimize the operational control strategy and perform grid scheduling control includes: obtaining reinforcement learning training samples based on the degree of grid anomaly, the operational control strategy, the operational data of the power grid agent, and the real-time indicator values ​​of the operational indicators; constructing the state space of the power grid agent based on the reinforcement learning training samples, and mapping observation vectors based on the action parameters of the power grid agent to obtain a target observation vector; using a multi-agent algorithm, taking the state space and the target observation vector as input parameters, to optimize and train the strategy models of each power grid agent under power grid operational constraints to obtain an optimized collaborative operational control strategy for each power grid agent; and performing adaptive regulation of grid scheduling based on the optimized collaborative operational control strategy.

[0011] Preferably, the step of constructing the state space of the power grid agent based on the reinforcement learning training samples and obtaining the target observation vector by mapping the observation vector based on the action parameters of the power grid agent includes: constructing the state space of the power grid agent based on the anomaly level, the domain agent code corresponding to the power grid agent, and the real-time index value and fluctuation characteristic value of the operation index; generating the action space of the power grid agent based on the operation data of the power grid agent and the corresponding operation control strategy; and mapping the action parameters of the power grid agent to a numerical range based on the action space to obtain the target observation vector.

[0012] Secondly, embodiments of this application provide a power grid control system based on agent reinforcement learning, comprising: a data acquisition module for calculating fluctuation characteristic values ​​of operating indicators based on power grid operation data; a trigger decision module for determining whether abnormal fluctuations have occurred in the power grid based on the fluctuation characteristic values, and if so, determining a target scheduling mode according to the degree of abnormality, and controlling the corresponding power grid agent to generate an operation control strategy according to the target scheduling mode; a human-machine interaction module for receiving the operation control strategy and displaying information details on a visual interface, generating a control execution command by having an operator correct control parameters or confirm the control plan; an execution module for executing strategy actions according to the control execution command; and a feedback optimization module for performing reinforcement learning training on the power grid agent based on the operation data of the power grid agent and the change data of the operating indicators, so as to optimize the operation control strategy and perform scheduling control of the power grid.

[0013] Thirdly, embodiments of this application provide a computer device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of the power grid control method based on agent reinforcement learning as described in the first aspect above.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the power grid control method based on agent reinforcement learning as described in the first aspect above.

[0015] The beneficial effects of this application are: 1. By introducing a scheduling mode triggering mechanism, the target triggering threshold is dynamically calculated, enabling the system to automatically adjust the threshold based on real-time grid operation data. This improves the sensitivity and response speed to abnormal situations, overcoming the problem that traditional systems lack adaptive threshold adjustment capabilities, leading to a high misjudgment rate of grid scheduling triggering in high-proportion renewable energy grids. 2. Determine the target scheduling mode based on the trigger intensity level to achieve precise graded processing of anomalies of different degrees, including the fine-tuning mode of the power grid agent in the corresponding region and the collaborative working mode between power grid agents across regions, avoiding unnecessary frequent equipment operation; at the same time, it enables the power grid to maintain economy during normal fluctuations and ensure stability during severe anomalies, thus balancing the sensitivity and cost of power grid anomaly detection. 3. By combining deep reinforcement learning with pheromone communication links, the decision weights of agents in various domains are dynamically adjusted, overcoming the limitations of local optima traps and low collaborative efficiency in traditional multi-agent systems. This achieves a distributed control mode of "local anomaly resolution and regional crisis collaborative handling." Furthermore, by integrating power grid operation constraints, a multi-dimensional state space is constructed, including real-time operation indicators, fluctuation characteristics, and agent collaborative operation data. This enables agents to accurately perceive the degree and scope of power grid anomalies. Based on the physical dimension mapping and data standardization of the action space, reinforcement learning is used to generate refined strategies that conform to equipment adjustment capabilities and safety constraints, effectively solving the problem of low collaborative efficiency among multi-agents in complex power grid scenarios. Attached Figure Description

[0016] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings.

[0017] Figure 1 A flowchart of a power grid control method based on agent reinforcement learning provided in this application embodiment; Figure 2 This application provides a schematic flowchart of a scheduling mode selection and execution method according to an embodiment of the present application. Figure 3 A schematic diagram of a power grid control system module based on agent reinforcement learning provided in an embodiment of this application; Figure 4 This application provides an embodiment of a multi-agent unit reinforcement learning training and triggering mechanism interaction diagram. Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of this application is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely one preferred embodiment of this application and are only used to explain this application. They do not limit the scope of protection of this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] Example 1: As Figure 1 As shown, a power grid control method based on agent reinforcement learning includes steps S101-S104, wherein: S101. Calculate the fluctuation characteristic value of the operation index based on the power operation data of the power grid.

[0020] As an optional implementation, step S101 specifically includes: S1011. Obtain real-time power operation data of the power grid and calculate the index values ​​of the power grid operation indicators, wherein the operation indicators include at least: voltage deviation index, frequency deviation index, line load index and output fluctuation index of distributed power sources. S1012. Calculate the average values ​​of the voltage deviation index, frequency deviation index, line load index, and output fluctuation index within the cycle, respectively. S1013. Calculate the rolling variance of each operating index within a period based on the average values ​​of the voltage deviation index, frequency deviation index, line load index, and output fluctuation index, respectively, to obtain the fluctuation characteristic value of each operating index.

[0021] In this embodiment, the voltage deviation index is used to quantify the degree to which the bus voltage deviates from the rated value, and the index value is calculated based on the actual voltage of the power grid and the rated voltage; the frequency deviation index is used to monitor the frequency stability of the power grid, and the index value is calculated based on the actual frequency of the power grid and the rated frequency; the line load index reflects the ratio of the current load of the transmission line to the rated capacity, and the index value is calculated based on the actual line load and the rated line load; the output fluctuation index is the output fluctuation of distributed new energy sources, which is used to measure the deviation between the actual output of distributed power sources, such as photovoltaic or wind power, and the predicted value.

[0022] In some possible embodiments, the formulas for calculating the values ​​of voltage deviation, frequency deviation, line load, and output fluctuation of distributed power sources can be expressed as follows: =(Uactual - Urated) / Urated × 100% (1); =(factual - frated) / frated × 100% (2); =(Pactual / Prated)×100%(3; =(P current - P predicted) / P predicted × 100% (4); in, Voltage deviation rate, representing the index value of the voltage deviation index; Frequency deviation rate represents the index value of the frequency deviation index. Line load rate represents the index value of the line load indicator; The power output volatility of new energy sources represents the index value of the power output volatility index.

[0023] Furthermore, using the sliding window algorithm, the average value of each indicator within the time window is calculated using the indicator values ​​of each indicator with a window length equal to a preset data point. Then, the fluctuation characteristic value of the corresponding indicator is calculated based on the average indicator value. The formula for calculating the fluctuation characteristic value of the indicator is shown below: (5); in, This represents the fluctuation characteristic value of a certain indicator; n represents the preset data points for the window length, such as 30 data points, which is 30 indicator values ​​for a certain indicator. This represents the average value of the indicator within the time window. This represents the mean variance of historical normal operating data.

[0024] In this embodiment, abnormal fluctuations are quantified by real-time acquisition of key power grid operation indicators such as voltage deviation rate, frequency deviation rate, line load rate, and renewable energy output fluctuation rate. This overcomes the limitations of traditional static data acquisition and enables real-time perception of dynamic changes in the power grid (such as sudden drops in renewable energy output and load abrupt changes), providing accurate and standardized data input for subsequent decision-making. The fluctuation characteristic value of the indicators also represents the rolling variance of the indicators. Calculating the rolling variance improves the quantification accuracy of abnormal fluctuations, providing a reliable basis for the dynamic adjustment of subsequent target trigger thresholds and avoiding misjudgments or omissions due to data lag.

[0025] S102. Based on the fluctuation characteristic value, determine whether the power grid has abnormal fluctuations. If so, determine the target scheduling mode according to the degree of abnormality.

[0026] As an optional implementation, step S102 specifically includes: S1021. When the fluctuation characteristic value is greater than the historical normal fluctuation value, it is determined that the power grid has abnormal fluctuations; S1022. Determine the trigger intensity level of the power grid agent based on the scheduling decision triggering mechanism, and use it as the degree of anomaly of the power grid to match the corresponding target scheduling mode.

[0027] In this embodiment, variance can accurately quantify the degree of fluctuation of various power grid indicators. By comparing it with the historical normal variance, it can dynamically reflect the risk of abnormal indicators, so that the subsequent trigger threshold adjustment can be adapted to the actual fluctuation characteristics of the power grid.

[0028] As an optional implementation, step S1022 specifically includes: The initial trigger threshold matching each of the aforementioned operating indicators is determined based on historical normal operation data of the power grid; The target trigger threshold is calculated based on the fluctuation characteristic value, the historical normal fluctuation value, and the initial trigger threshold. The values ​​of each of the aforementioned operational indicators are compared with the corresponding target trigger thresholds. If the value of an operational indicator exceeds the target trigger threshold, the trigger intensity level is determined based on the degree of exceeding the threshold. The trigger intensity level is used to characterize the degree of anomaly in the power grid. The scheduling mode with the highest trigger intensity level is determined as the target scheduling mode.

[0029] In some possible embodiments, such as Figure 2 As shown, the fluctuation characteristic value of each indicator is compared with the mean variance of the fluctuation data of historical normal operation (here, the historical normal fluctuation value). If the fluctuation characteristic value is greater than the historical normal fluctuation value, that is... If abnormal fluctuations occur, the power grid is identified as abnormal. In the event of abnormal power grid fluctuations, the trigger intensity of the scheduling mode is divided based on the scheduling decision triggering mechanism. The scheduling mode of the power grid agent is determined by comparing the target trigger threshold calculated based on the real-time fluctuation characteristic value with the initial trigger threshold.

[0030] Specifically, the target trigger threshold is a dynamic threshold that adaptively adjusts based on real-time fluctuation characteristics. Its calculation formula is as follows: (6); Where T represents the target trigger threshold; This indicates the initial trigger threshold for each operational indicator; An adjustment coefficient (0.2 ≤ k ≤ 0.5) is used to simulate the adaptive behavior of a flea shortening its jump distance when encountering an obstacle; where the initial trigger threshold is... This can be set based on the 95th percentile of the historical normal operating data of the power grid corresponding to each operating indicator, for example, of ±5%, of It is ±0.2%.

[0031] It should be noted that the target trigger threshold is adaptively adjusted based on real-time fluctuation characteristics, drawing inspiration from the survival strategies of fleas in complex environments. When encountering obstacles, fleas instinctively shorten their jump distance to avoid collisions and better adapt to environmental changes. In the power grid dispatching system, when abnormal fluctuations occur in power grid operating indicators, the target trigger threshold is dynamically adjusted. For example, the threshold is reduced. The initial threshold is set based on the 95th percentile of historical data. When fluctuations exceed limits, the threshold is tightened by adjusting the coefficient k (0.2-0.5), thereby increasing the difference between the actual operating indicator value and the target trigger threshold. This improves the system's ability to perceive and respond to abnormal power grid conditions, enhances its sensitivity to dynamic fluctuations, provides more accurate and reliable trigger signals for power grid intelligent agent collaboration, and generates more reasonable and effective power grid dispatching strategies.

[0032] Specifically, the trigger intensity of the scheduling mode is divided into a first level and a second level. The scheduling mode corresponding to the first trigger intensity level is the fine-tuning mode, and the scheduling mode corresponding to the second trigger intensity level is the emergency mode. Among them, the difference between the target trigger threshold and the actual value of the indicator corresponding to the first trigger intensity level is greater than 10% and less than or equal to 20%, and the difference between the target trigger threshold and the actual value of the indicator corresponding to the second trigger intensity level is greater than 20%. That is, when the actual value of each operating indicator exceeds the target trigger threshold by 10%-20% (inclusive), the fine-tuning mode is triggered, and when the actual value of each operating indicator exceeds the target trigger threshold by 20%, the emergency mode is triggered.

[0033] In this embodiment, the dynamic threshold adjustment for target triggering can improve the timeliness of power grid anomaly response and reduce the missed trigger rate. By dividing the response into fine-tuning mode and emergency mode according to the degree of exceeding the threshold, and matching the response of power grid agents within different ranges, this hierarchical triggering method can effectively avoid the problems of over-reaction or under-reaction of power grid agents, reduce the invalid actions of agent devices, thereby shortening the recovery time of various operating indicators and improving the efficiency of power grid dispatching. In addition, in this embodiment, when the actual value of a power grid operating indicator exceeds the target triggering threshold, the corresponding dispatching strategy is executed according to the power grid dispatching mode matched with the above-mentioned triggering intensity level, ensuring that power grid anomalies are handled in a timely manner, thereby ensuring the safe and stable operation of the power grid.

[0034] S103. Based on the target scheduling mode, deep reinforcement learning is used to control the corresponding power grid agent to generate operation and control strategies and execute strategy actions.

[0035] As an optional implementation, the target scheduling mode includes at least a fine-tuning mode and a global mode, and step S103 specifically includes: S1031. When the target scheduling mode is fine-tuning mode, the power grid agent in the corresponding scheduling area is activated, the power grid agent is triggered to generate a corresponding operation control strategy according to the scheduling requirements, and the operation parameters are adjusted according to the operation control strategy, and the operation is performed according to the adjusted operation parameters. S1032. When the target scheduling mode is global mode, activate the power grid agents in multiple scheduling regions, dynamically adjust the decision weights of each power grid agent to generate a collaborative operation and control strategy, and operate according to the collaborative operation and control strategy, specifically including: Each of the aforementioned power grid agents generates a local adjustment strategy based on a deep reinforcement learning algorithm, and simultaneously broadcasts a scheduling trigger event to neighboring agents via a pheromone communication link. Based on the scheduling trigger event, the degree of anomaly is obtained, and the decision weights of the neighborhood agents are dynamically adjusted to generate a target local adjustment strategy. Based on the target local adjustment strategy of each of the power grid intelligent agents, a collaborative operation and control strategy for multi-schedule regions is constructed, and the operation is carried out in accordance with the collaborative operation and control strategy.

[0036] In some possible embodiments, combined Figure 2 As shown, when the scheduling mode is determined to be fine-tuning mode, the corresponding regional agent is activated to generate the corresponding operation adjustment strategy, such as voltage deviation triggering the substation agent; when the scheduling mode is determined to be emergency mode, the cross-regional collaborative agent is activated, and the final operation adjustment strategy is generated through the collaboration between the agents based on the multi-agent collaborative decision-making mechanism, such as frequency deviation triggering both the generation agent and the load agent.

[0037] Furthermore, the multi-agent collaborative decision-making mechanism specifically includes: activated agents generating local adjustment strategies based on deep reinforcement learning (DRL) algorithms, such as reactive power compensation switching schemes for substation agents and output adjustment step sizes for generator agents; simultaneously, they broadcast trigger events to neighboring agents via pheromone communication links, with trigger events including at least trigger strength, anomaly type, and grid impact range; neighboring agents adjust their own decision weights based on the received trigger strength to generate target local adjustment strategies; the target local adjustment strategies of various agents across regions are combined to obtain a final collaborative operation control strategy, trigger strength, and risk level multi-dimensional grid dispatch decision-making suggestions; then, the grid dispatch decision-making suggestions are confirmed through human-machine interaction, and dispatch execution instructions are generated to execute grid dispatch control.

[0038] Among them, the decision weight of the neighborhood agent is positively correlated with the trigger intensity. The higher the trigger intensity, the greater the increase in the decision weight of the neighborhood agent. The trigger intensity value is related to the actual value of the operating indicator exceeding the target trigger threshold, and the value range is 0-100%. For example, if the actual value of the operating indicator exceeds the target trigger threshold by 25%, the corresponding trigger intensity is 25%.

[0039] In some embodiments, the risk level includes levels 1-3, with lower values ​​indicating a higher degree of grid anomaly. When the trigger intensity level is level 1 and the trigger intensity is less than or equal to 10%, the risk level is 3, indicating that the intelligent agent does not need to be triggered for grid scheduling. When the trigger intensity level is level 1 and the trigger intensity is greater than 10% and less than or equal to 20%, the risk level is 2, indicating abnormal grid fluctuations and requiring the triggering of fine-tuning mode. When the trigger intensity level is level 2 and the trigger intensity is greater than 20%, the risk level is 1, indicating severe abnormal grid fluctuations and requiring the triggering of emergency mode.

[0040] In this embodiment, a local adjustment strategy is generated through a deep reinforcement learning algorithm. Based on the shared trigger event information through the pheromone communication link, each domain agent can dynamically adjust its own decision weights for power grid scheduling according to the trigger event. This enables multi-agent collaborative operation and avoids the local optimum trap in multi-agent collaboration, which can lead to waste or shortage of schedulable resources. At the same time, cross-regional agent collaborative response realizes the mechanism of local anomaly resolution and regional crisis collaborative handling, improving response efficiency and collaboration efficiency.

[0041] S104. Based on the operation data of the power grid agent and the change data of the operation indicators, the power grid agent is trained by reinforcement learning to optimize the operation and control strategy and to perform scheduling and control of the power grid.

[0042] As an optional implementation, step S104 specifically includes: S1041. Obtain reinforcement learning training samples based on the degree of anomaly in the power grid, the operation control strategy, the operation data of the power grid agent, and the real-time index values ​​of the operation indicators. S1042. Construct the state space of the power grid agent based on the reinforcement learning training samples, and perform observation vector mapping based on the action parameters of the power grid agent to obtain the target observation vector. S1043. Based on the multi-agent algorithm, the state space and the target observation vector are used as input parameters to optimize and train the strategy model of each power grid agent under the power grid operation constraints, and obtain the optimized collaborative operation and control strategy of each power grid agent. S1044. Adaptive control of power grid dispatch is performed based on the optimized collaborative operation and control strategy.

[0043] As an optional implementation, step S1042 specifically includes: The state space of the power grid agent is constructed based on the degree of anomaly, the domain agent code corresponding to the power grid agent, and the real-time index value and fluctuation characteristic value of the operation index. The action space of the power grid intelligent agent is generated based on its operational data and corresponding operational control strategies. Based on the action space, the action parameters of the power grid agent are mapped to numerical intervals to obtain the target observation vector.

[0044] In some possible embodiments, the adjustment coefficient of the target trigger threshold or the pheromone communication link is optimized based on the trigger effectiveness index of the above scheduling mode; the trigger effectiveness index includes trigger accuracy and trigger latency, and the calculation formula is as follows: Trigger accuracy = Number of correct triggers (indicators return to normal after triggering) / Total number of triggers × 100%; Trigger delay = the time it takes for the real-time value of the running indicator to exceed the target trigger threshold. Time until receiving the execution instruction of the operation control strategy Time difference ( ).

[0045] In some examples, if the trigger accuracy is less than 80%, the adjustment factor is increased. (Increase by 0.1 each time) to improve trigger sensitivity; if the trigger delay is >20s, optimize the transmission priority of the pheromone communication link, for example, shorten the delay of high trigger intensity events from 10ms to 5ms.

[0046] In some embodiments, after the operation and control strategy is executed, the anomaly type, trigger intensity, operation and control strategy of the agent, correction record of the control strategy, and execution effect of the control strategy (trigger accuracy and trigger delay) are stored in the database in a quintuple data format for optimizing the model of the operation and control strategy generated by each agent.

[0047] Furthermore, each agent is trained using a deep reinforcement learning algorithm. The state space includes real-time operating metrics, historical case features, and neighborhood cooperation signals, while the action space covers device adjustment parameters and cooperation requests.

[0048] Specifically, an agent state space is established based on the power grid's voltage deviation index, frequency deviation index, line load index, and output fluctuation index, as well as the corresponding fluctuation characteristic values ​​of each index and the domain agent ID (i.e., the domain agent code, which is a unique identifier for each agent). It is expressed as follows: (7); Here, i represents an agent, and the domain agent ID identifies the power grid agents in the adjacent areas that need to be coordinated in this scheduling.

[0049] Furthermore, based on the operational control strategies and operational data of each agent, the action space of the agent is generated. It is expressed as follows: (8); , ; ; ; in, Indicates the action parameters of the power generation intelligent agent. Let t be the active power output of the generator. This represents the action parameters of the substation intelligent agent. Indicates the action parameters of the load agent; This is a scheduling power constraint used to ensure that the power change between two adjacent scheduling operations is not less than this value; Mapping the action parameters of each agent to the [0,1] interval according to physical dimensions yields the target observation vector: (9); in, This represents the target observation vector of the i-th agent. This indicates the generator's maximum adjustable power. This indicates the total number of power compensation devices. Indicates the total interrupted load; , and Both represent the maximum adjustment capability of intelligent agent devices.

[0050] Furthermore, based on the execution effect of the control strategy, i.e., the trigger accuracy and trigger delay combined with the power grid operation target, a reward function is designed. Under the constraints of power grid operation, a multi-agent (Proximal Policy Optimization, PPO) algorithm is used to train the policy network model of each agent, and the trigger intensity signal is shared through a pheromone communication link, wherein: The input layer of the policy network model takes the agent's state as input. With target observation vector The output layer outputs the probability distribution of actions (such as the probability of a power-generating agent adjusting its output). The power grid operation constraints include at least voltage amplitude constraints, frequency constraints, and power constraints. A constraint vector is established based on each constraint. , The values ​​of each constraint must be between the minimum and maximum values ​​under normal grid operation. For example, during the training of the strategy network model, the ramp power of the generator set is limited, such as not exceeding the generator ramp rate, to ensure that the output generator agent adjustment strategy meets the ramp power constraint.

[0051] In some embodiments, after the policy network model of each agent is trained, the contribution of the current agent's operation and control strategy to the overall optimization of the power grid (such as reducing the loss rate and improving the voltage qualification rate) is evaluated, so as to adjust the model parameters according to the reward function. The update of the model parameters is represented as follows: (10); in, This represents the updated model parameters; Indicates the parameters of the initial policy network model; This represents the learning rate (dynamically adjusted within the range of 0.001-0.01). This represents the action probability output by the policy network model; Represents the gradient operator; Here, Expectation represents the average value of a random variable; R represents the reward function.

[0052] Specifically, the agent is based on the current state. and observation vector Through policy network Calculate the probability distribution of discrete actions that characterize the content of the control strategy to obtain the optimized control strategy; for example, when the observed frequency is below the threshold (Δf < −0.2Hz), the probability of increasing power generation increases, and when the voltage exceeds the upper limit (Vviolation > 0), the probability of reducing reactive power increases.

[0053] In this embodiment, the multi-agent reinforcement learning training method integrates the physical constraints and trigger strength of power grid operation to construct a multi-dimensional state space containing real-time operating indicators, fluctuation characteristics, and neighborhood cooperation information. This enables agents to accurately perceive the degree and scope of power grid anomalies, improving the accuracy of state perception. Based on the physical dimension mapping of the action space combined with the operation control strategy update (optimization) mechanism of the multi-agent algorithm, agents can generate refined strategies that conform to the equipment adjustment capabilities and safety constraints. For example, the output adjustment step size error of the power generation agent is controlled within ±2% of the rated power. Compared with traditional heuristic algorithms, this significantly improves the economy and safety of the strategy. By sharing trigger strength signals through pheromone communication links, the decision weights of neighborhood agents can be dynamically adjusted according to the degree of anomaly, thereby shortening the cross-regional collaborative response time and effectively solving the problem of low multi-agent collaboration efficiency in complex power grid scenarios.

[0054] Example 2, based on the same inventive concept, also provides a power grid control system based on agent reinforcement learning, corresponding to the power grid control method based on agent reinforcement learning, such as... Figure 3 As shown, the system includes: The data acquisition module is used to calculate the fluctuation characteristics of operating indicators based on the power operation data of the power grid; The trigger decision module is used to determine whether the power grid has abnormal fluctuations based on the fluctuation characteristic value. If so, the target scheduling mode is determined according to the degree of abnormality, and the corresponding power grid agent is controlled to generate an operation and control strategy according to the target scheduling mode. The human-computer interaction module is used to receive the operation control strategy and display the information details on the visual interface, and generate control execution instructions by allowing operators to correct control parameters or confirm control plans. The execution module is used to execute strategy actions according to the control execution instructions; The feedback optimization module is used to perform reinforcement learning training on the power grid agent based on the operation data of the power grid agent and the change data of the operation indicators, so as to optimize the operation and control strategy and perform scheduling and control of the power grid.

[0055] As an optional implementation, the data acquisition module is specifically used to: acquire real-time power operation data of the power grid, calculate the index values ​​of the power grid's operation indicators, the operation indicators including at least: voltage deviation index, frequency deviation index, line load index, and output fluctuation index of distributed power sources; calculate the average index values ​​of the voltage deviation index, frequency deviation index, line load index, and output fluctuation index within a period; and calculate the rolling variance of each operation indicator within a period based on the average index values ​​of the voltage deviation index, frequency deviation index, line load index, and output fluctuation index, so as to obtain the fluctuation characteristic value of each operation indicator.

[0056] As an optional implementation, the triggering decision module specifically includes: a first triggering decision unit, used to determine that the power grid has abnormal fluctuations when the fluctuation characteristic value is greater than the historical normal fluctuation value; and a second triggering decision unit, used to determine the triggering intensity level of the power grid agent based on the scheduling decision triggering mechanism, and use it as the degree of abnormality of the power grid to match the corresponding target scheduling mode.

[0057] As an optional implementation, the second triggering decision unit is specifically used for: determining an initial triggering threshold matching each of the operating indicators based on historical normal power grid data; calculating a target triggering threshold based on the fluctuation characteristic value, the historical normal fluctuation value, and the initial triggering threshold; comparing the indicator value of each operating indicator with the corresponding target triggering threshold, and determining the triggering intensity level based on the degree of exceeding the threshold when the indicator value of an operating indicator exceeds the target triggering threshold, wherein the triggering intensity level is used to characterize the degree of abnormality of the power grid; and determining the scheduling mode with the higher triggering intensity level as the target scheduling mode.

[0058] As an optional implementation, the execution module specifically includes: a first execution unit, configured to, when the target scheduling mode is fine-tuning mode, activate the power grid agent in the corresponding scheduling area, trigger the power grid agent to generate a corresponding operation control strategy according to scheduling requirements, adjust the operation parameters according to the operation control strategy, and perform work according to the adjusted operation parameters; and a second execution unit, configured to, when the target scheduling mode is global mode, activate the power grid agents in multiple scheduling areas, dynamically adjust the decision weights of each power grid agent to generate a collaborative operation control strategy, and perform work according to the collaborative operation control strategy.

[0059] As an optional implementation, the second execution unit is specifically used for: each of the power grid agents generating local adjustment strategies based on a deep reinforcement learning algorithm, and synchronously broadcasting scheduling trigger events to neighboring agents via pheromone communication links; obtaining the degree of anomaly based on the scheduling trigger events, dynamically adjusting the decision weights of the neighboring agents to generate a target local adjustment strategy; constructing a collaborative operation control strategy for multiple scheduling regions based on the target local adjustment strategies of each of the power grid agents, and operating according to the collaborative operation control strategy.

[0060] As an optional implementation, the feedback optimization module specifically includes: a first feedback optimization unit, used to obtain reinforcement learning training samples based on the anomaly level of the power grid, the operation control strategy, the operation data of the power grid agent, and the real-time index values ​​of the operation indicators; a second feedback optimization unit, used to construct the state space of the power grid agent based on the reinforcement learning training samples, and to perform observation vector mapping based on the action parameters of the power grid agent to obtain a target observation vector; a third feedback optimization unit, used to optimize and train the strategy model of each power grid agent under power grid operation constraints using the state space and the target observation vector as input parameters based on a multi-agent algorithm, to obtain an optimized collaborative operation control strategy for each power grid agent; and a fourth feedback optimization unit, used to perform adaptive control of power grid scheduling based on the optimized collaborative operation control strategy.

[0061] In some possible embodiments, combined Figure 4 As shown in the figure, this is a schematic diagram of the interaction between a multi-agent unit reinforcement learning training and triggering mechanism. The data acquisition module collects the voltage deviation rate, frequency deviation rate, line load rate, and new energy output fluctuation rate of the power grid in real time. It calculates the rolling variance of each indicator over a time period using a sliding window algorithm. When the fluctuation of each indicator far exceeds the historical normal variance level, the system determines that an anomaly has occurred and transmits the standardized data and fluctuation information to the triggering decision module.

[0062] Furthermore, after receiving the data, the trigger decision module determines the trigger strength of the power grid agent based on the current fluctuation level. After matching the scheduling mode according to the trigger strength, it activates the corresponding agent to generate a corresponding operation adjustment strategy, i.e., the power grid's scheduling strategy. Following the execution of the decision plan, the feedback optimization module obtains the real-time values ​​of various power grid operation indicators. Combining this with the agent's operation data, including the domain agent code (domain agent ID), it constructs the agent's state space and generates an action space. After standardizing the action parameters and combining them with the power grid operation goals and the trigger mechanism's effect, it designs a reward function to update the scheduling strategy. When the optimal scheduling strategy is obtained, it is output to the trigger decision module.

[0063] In some examples, the output of a regional photovoltaic power station drops sharply from 80% to 20% of its rated power within 10 minutes. In this case, the power grid control system based on agent reinforcement learning operates according to the following process: A1. The data acquisition module collects data in real time, showing that the voltage deviation rate reaches -8%, the frequency deviation rate is -0.3Hz, the line load rate soars to 95%, and the power output fluctuation rate of new energy reaches -60%. At the same time, the 5-minute rolling variance calculated by the sliding window algorithm shows that the fluctuation of various indicators far exceeds the historical normal variance level. The system judges that an anomaly has occurred and transmits the standardized data and fluctuation information to the trigger decision module. A2. After receiving the data, the trigger decision module tightens the target trigger thresholds based on the current fluctuation level, reducing the voltage deviation trigger threshold from the initial ±5% to ±3%, and the frequency deviation threshold to ±0.15Hz. Since both voltage and frequency deviations exceed the tightened thresholds, and the fluctuation rate of new energy output exceeds the threshold by more than 20%, the emergency mode is triggered. Specific details include: In emergency mode, the power generation agents, load agents, and collaborative agents in adjacent areas are activated. The activated power generation agents, based on deep reinforcement learning algorithms, quickly generate local adjustment strategies to increase the output of thermal power units and start standby diesel generators. The load agents, based on an interruptible load database, formulate plans to cut off some non-critical commercial loads. Simultaneously, trigger events are broadcast to neighboring agents via pheromone communication links, including trigger strength (set to 90%), anomaly type (sudden drop in renewable energy output), and impact range (the entire regional power grid). Upon receiving the information, the neighboring agents increase their decision weight from 30% to 70% and provide cross-regional power support solutions, such as transmitting 20MW of power from adjacent areas. Finally, the trigger decision module integrates the strategies of each agent to generate multi-dimensional decision recommendations including adjustment plans, trigger strength (90%), and risk level (level 3). A3. The human-computer interaction module presents decision suggestions in a visual interface, including dynamic threshold curves, over-threshold heatmaps, and detailed adjustment plans. After the operator confirms the feasibility of the plan, the system generates an execution instruction with trigger logs and manual intervention markers, and sends it to the execution module. A4. The execution module, based on the emergency mode, starts the backup diesel generator within one minute, simultaneously sends load shedding commands to non-critical commercial users, and coordinates power support from adjacent areas. During execution, it records real-time changes in indicators such as voltage, frequency, and load rate, and transmits this data to the feedback optimization module. A5. The feedback optimization module calculates that the trigger accuracy of this trigger is 100% and the trigger delay is only 8 seconds. Based on the reinforcement learning reward function, the dynamic threshold parameters of the trigger decision module are optimized, and this typical case is stored in the historical database for reverse optimization of the indicator selection strategy of the data acquisition module and the agent training samples of the trigger decision module.

[0064] This application also provides a computer device, such as... Figure 5 The diagram shown is a schematic representation of the structure of a computer device provided in an embodiment of this application, including: a processor 51, a memory 52, and a bus 53. The memory 52 stores machine-readable instructions executable by the processor 51 (e.g., ...). Figure 3 In the system, the data acquisition module 301, trigger decision module 302, human-computer interaction module 303, execution module 304 and feedback optimization module 305 (and their corresponding execution instructions, etc.) are executed. When the computer device is running, the processor 51 and the memory 52 communicate through the bus 53. When the machine-readable instructions are executed by the processor 51, the steps of the power grid control method based on intelligent agent reinforcement learning in the above embodiment are executed.

[0065] The processor 51 mentioned above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0066] The aforementioned memory 52 may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0067] The aforementioned bus 53 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus 53 can be divided into an address bus, a data bus, a control bus, etc.

[0068] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, performs the steps of the power grid control method based on agent reinforcement learning in the above embodiments.

[0069] Optionally, in embodiments of this application, the computer-readable medium is configured to store program code for the processor to execute the steps of the power grid control method based on agent reinforcement learning described in the above embodiments.

[0070] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated here. Furthermore, in the specific implementation of this application embodiment, the above embodiments can be consulted, and corresponding technical effects can be achieved.

[0071] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or combinations thereof.

[0072] For software implementation, the techniques described herein can be implemented by units that perform the functions described herein. The software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.

[0073] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0074] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0075] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0076] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, the functional units in the various embodiments of this application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0077] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0078] It should be noted that, in this document, relational terms such as "first," "second," etc., are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprise," "include," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprises a…" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0079] The above-described embodiments are preferred embodiments of this application and are not intended to limit the specific scope of this application. The scope of this application includes but is not limited to the specific embodiments described above. All equivalent changes made in accordance with the shape, structure, and method of this application are within the protection scope of this application.

Claims

1. A power grid control method based on agent reinforcement learning, characterized in that: Includes the following steps: The fluctuation characteristic values ​​of operating indicators are calculated based on the power operation data of the power grid; Based on the fluctuation characteristic value, it is determined whether the power grid has abnormal fluctuations. If so, the target scheduling mode is determined according to the degree of abnormality. Based on the target scheduling mode, deep reinforcement learning is used to control the corresponding power grid agent to generate operation and control strategies and execute strategy actions. The power grid agent is trained using reinforcement learning based on its operational data and the changes in its operational indicators, in order to optimize the operational control strategy and perform scheduling and control of the power grid.

2. The power grid control method based on agent reinforcement learning according to claim 1, characterized in that: The fluctuation characteristic values ​​of the operating indicators calculated based on the power grid operation data include: The system acquires real-time power operation data of the power grid and calculates the index values ​​of the power grid's operation indicators, which include at least: voltage deviation index, frequency deviation index, line load index, and output fluctuation index of distributed power sources. Calculate the average values ​​of the voltage deviation index, frequency deviation index, line load index, and output fluctuation index within the cycle, respectively. The rolling variance of each operating index within a period is calculated based on the average values ​​of the voltage deviation index, frequency deviation index, line load index, and output fluctuation index, in order to obtain the fluctuation characteristic value of each operating index.

3. The power grid control method based on agent reinforcement learning according to claim 1, characterized in that: The step of determining whether the power grid has experienced abnormal fluctuations based on the fluctuation characteristic value, and if so, determining the target scheduling mode according to the degree of abnormality, includes: When the fluctuation characteristic value is greater than the historical normal fluctuation value, it is determined that the power grid has abnormal fluctuations; The trigger intensity level of the power grid agent is determined based on the scheduling decision triggering mechanism, and it is used as the degree of anomaly of the power grid to match the corresponding target scheduling mode.

4. The power grid control method based on agent reinforcement learning according to claim 3, characterized in that: The step of determining the trigger strength level of the power grid agent based on the scheduling decision triggering mechanism, and using it as the anomaly level to match the corresponding target scheduling mode, includes: The initial trigger threshold matching each of the aforementioned operating indicators is determined based on historical normal operation data of the power grid; The target trigger threshold is calculated based on the fluctuation characteristic value, the historical normal fluctuation value, and the initial trigger threshold. The values ​​of each of the aforementioned operational indicators are compared with the corresponding target trigger thresholds. If the value of an operational indicator exceeds the target trigger threshold, the trigger intensity level is determined based on the degree of exceeding the threshold. The trigger intensity level is used to characterize the degree of anomaly in the power grid. The scheduling mode with the highest trigger intensity level is determined as the target scheduling mode.

5. The power grid control method based on agent reinforcement learning according to claim 1, characterized in that: The target scheduling mode includes at least a fine-tuning mode and a global mode. The step of generating and executing operation and control strategies for the corresponding power grid agent based on the target scheduling mode using deep reinforcement learning includes: When the target scheduling mode is fine-tuning mode, the power grid agent in the corresponding scheduling area is activated, the power grid agent is triggered to generate the corresponding operation control strategy according to the scheduling requirements, adjust the operation parameters according to the operation control strategy, and work according to the adjusted operation parameters. When the target scheduling mode is global mode, the power grid agents in multiple scheduling regions are activated, the decision weights of each power grid agent are dynamically adjusted to generate a collaborative operation and control strategy, and the system operates according to the collaborative operation and control strategy, specifically including: Each of the aforementioned power grid agents generates a local adjustment strategy based on a deep reinforcement learning algorithm, and simultaneously broadcasts a scheduling trigger event to neighboring agents via a pheromone communication link. Based on the scheduling trigger event, the degree of anomaly is obtained, and the decision weights of the neighborhood agents are dynamically adjusted to generate a target local adjustment strategy. Based on the target local adjustment strategy of each of the power grid intelligent agents, a collaborative operation and control strategy for multi-schedule regions is constructed, and the operation is carried out in accordance with the collaborative operation and control strategy.

6. The power grid control method based on agent reinforcement learning according to claim 1, characterized in that: The process of training the power grid agent using reinforcement learning based on its operational data and changes in operational indicators to optimize the operational control strategy and perform grid dispatch control includes: Reinforcement learning training samples are obtained based on the degree of anomaly in the power grid, the operation and control strategy, the operation data of the power grid agent, and the real-time index values ​​of the operation indicators. The state space of the power grid agent is constructed based on the reinforcement learning training samples, and the observation vector is mapped based on the action parameters of the power grid agent to obtain the target observation vector. Based on the multi-agent algorithm, the state space and the target observation vector are used as input parameters to optimize and train the strategy model of each power grid agent under the power grid operation constraints, so as to obtain the optimized collaborative operation and control strategy of each power grid agent. The optimized collaborative operation and control strategy is used to adaptively control the power grid dispatch.

7. The power grid control method based on agent reinforcement learning according to claim 6, characterized in that: The process of constructing the state space of the power grid agent based on the reinforcement learning training samples, and obtaining the target observation vector by mapping the observation vector based on the action parameters of the power grid agent, includes: The state space of the power grid agent is constructed based on the degree of anomaly, the domain agent code corresponding to the power grid agent, and the real-time index value and fluctuation characteristic value of the operation index. The action space of the power grid intelligent agent is generated based on its operational data and corresponding operational control strategies. Based on the action space, the action parameters of the power grid agent are mapped to numerical intervals to obtain the target observation vector.

8. A power grid control system based on agent reinforcement learning, characterized in that: The system applicable to the power grid control method based on agent reinforcement learning as described in any one of claims 1-7, the system comprising: The data acquisition module is used to calculate the fluctuation characteristics of operating indicators based on the power operation data of the power grid; The trigger decision module is used to determine whether the power grid has abnormal fluctuations based on the fluctuation characteristic value. If so, the target scheduling mode is determined according to the degree of abnormality, and the corresponding power grid agent is controlled to generate an operation and control strategy according to the target scheduling mode. The human-computer interaction module is used to receive the operation control strategy and display the information details on the visual interface, and generate control execution instructions by allowing operators to correct control parameters or confirm control plans. The execution module is used to execute strategy actions according to the control execution instructions; The feedback optimization module is used to perform reinforcement learning training on the power grid agent based on the operation data of the power grid agent and the change data of the operation indicators, so as to optimize the operation and control strategy and perform scheduling and control of the power grid.

9. A computer device, characterized in that: include: The computer device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of the power grid control method based on agent reinforcement learning as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, performs the steps of the power grid control method based on agent reinforcement learning as described in any one of claims 1-7.