A power distribution network fault scene setting value checking method based on reinforcement learning
By constructing a hierarchical multi-agent verification system and utilizing reinforcement learning methods for fault situation perception and dynamic task allocation, the complexity of setting verification under distribution network fault scenarios is solved, enabling efficient and accurate protection device actions and improving the safety and stability of the distribution network.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- POWER RES INST OF STATE GRID SHAANXI ELECTRIC POWER CO LTD
- Filing Date
- 2025-07-11
- Publication Date
- 2026-04-14
AI Technical Summary
Existing methods for verifying setpoints in distribution network fault scenarios are inadequate in terms of complex fault situation perception, distributed protection unit coordination, and dynamic task allocation. They are unable to respond quickly and accurately to cascading faults, leading to maloperation or failure of protection devices, which affects power supply reliability.
A hierarchical multi-agent verification system is constructed, which adopts a reinforcement learning-based approach for fault situation awareness and dynamic task allocation. The system decomposes the global fixed-value verification target through a high-level agent and uses low-level agents to execute the fixed-value verification sub-tasks. The verification strategy is optimized by combining centralized training and distributed execution modes.
It improves the efficiency of handling complex faults, enhances the collaborative capability of distributed protection units, realizes the adaptability and robustness of setting value verification, optimizes operation and maintenance efficiency, and improves the safe and stable operation level of the distribution network.
Smart Images

Figure CN120930983B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of distribution network protection technology, and more specifically, to a method for verifying setpoints in distribution network fault scenarios based on reinforcement learning. Background Technology
[0002] The power distribution network is a crucial component of the power system, and its safe and stable operation is essential for ensuring reliable power supply to users. Protection settings are key parameters that ensure the correct operation of relay protection devices during faults.
[0003] With the increasing complexity of distribution network structures, the large-scale integration of new energy sources, and the variability of load characteristics, distribution network fault scenarios are exhibiting increased complexity and uncertainty. Existing setting verification technologies mainly face the following key technical challenges:
[0004] In terms of complex fault situation awareness, traditional methods often rely on static rules or experience-based judgments, making it difficult to accurately predict the evolution path and development trend of faults. Especially during cascading faults or large-scale disturbances, existing methods lack dynamic fault situation awareness and cannot quickly identify potential fault propagation risks, resulting in insufficient targeting of setting value verification. Regarding distributed protection unit collaboration, existing technologies mainly employ independent verification by each protection unit, lacking an effective collaboration mechanism. Insufficient information sharing and poor coordination among distributed protection units easily lead to improper protection coordination in complex fault scenarios, affecting overall protection performance. In terms of dynamic task allocation and adaptive adjustment, traditional setting value verification methods are mainly based on preset processes and fixed strategies, lacking the ability to dynamically allocate tasks according to real-time fault situations and system states. This static approach is ill-suited to complex and ever-changing fault scenarios and cannot achieve optimized allocation of verification resources.
[0005] However, with the increasing complexity of distribution network structures, the large-scale integration of new energy sources, and the variability of load characteristics, traditional setting verification methods that rely on manual experience or offline setting calculations face significant challenges. Especially in the event of cascading faults or large-scale disturbances, existing methods often struggle to quickly and accurately assess and adjust settings, potentially leading to maloperation or failure of protection devices, expanding the scope of the accident, and impacting power supply reliability.
[0006] While some research has begun to apply machine learning and other techniques to setting verification, most studies focus on single-point or local optimization, lacking in-depth consideration of multi-protection unit collaboration, dynamic task adaptation, and hierarchical decision-making mechanisms under complex fault scenarios. Existing methods still have significant shortcomings in intelligent fault state perception, effective collaboration of distributed protection units, and dynamic allocation of verification tasks.
[0007] Therefore, existing technologies suffer from low learning efficiency, insufficient generalization ability, and poor collaborative effects in addressing complex faults with long time-series dependencies, achieving efficient collaboration among distributed protection units, and flexibly allocating verification tasks based on dynamic fault conditions. There is an urgent need for a novel setpoint verification technology that combines artificial intelligence, particularly reinforcement learning methods, to improve the intelligence level and fault response capabilities of distribution network protection systems. Summary of the Invention
[0008] This invention provides a setpoint verification method for distribution network fault scenarios based on reinforcement learning, which solves the technical problems of low setpoint verification efficiency under complex cascading faults, difficulty in coordination of distributed protection units, and insufficient adaptability to dynamic fault conditions in related technologies.
[0009] This invention provides a method for setting verification in distribution network fault scenarios based on reinforcement learning, comprising:
[0010] Construct a hierarchical multi-agent verification system, which includes at least one high-level agent and multiple low-level agents;
[0011] Based on real-time operation data of the power distribution network, the high-level intelligent agent uses the first reinforcement learning model to perceive the fault situation and decomposes the global setpoint verification target into multiple setpoint verification sub-tasks.
[0012] The high-level intelligent agent assigns multiple fixed-value verification subtasks to the target low-level intelligent agent for execution based on the dynamic task allocation model. The dynamic task allocation model comprehensively evaluates the low-level intelligent agent's capability parameters, task-related distance parameters, and task urgency parameters.
[0013] The target low-level agent uses the second reinforcement learning model to execute the assigned value verification subtask and obtain the value verification result;
[0014] Based on the set value verification results and real-time operation data of the distribution network, a centralized training and distributed execution mode is adopted to jointly optimize the first reinforcement learning model and the second reinforcement learning model.
[0015] Furthermore, the fault situation awareness includes:
[0016] The first reinforcement learning model is used to analyze the real-time operation data and topology information of the distribution network, and output the potential fault evolution path sequence and the probability of occurrence of each path.
[0017] Furthermore, the dynamic task allocation model is constructed based on the ContractNet protocol;
[0018] The high-level intelligent agent acts as the task publisher, publishing the multiple fixed-value verification subtasks.
[0019] Multiple low-level intelligent agents act as contractors, bidding for the multiple fixed-value verification subtasks based on their respective task utility values;
[0020] The high-level intelligent agent selects the low-level intelligent agent with the highest task utility value as the target low-level intelligent agent.
[0021] Furthermore, the task utility value is calculated as follows:
[0022] U ij =w1·capability j (s k ) + w2·(1 / distance ij )+w3·urgency i ;
[0023] U ij represents the utility value of low-level agent j for subtask i, used to measure the overall suitability of the low-level agent for the task, where i is the subtask number and j is the low-level agent number; w1 represents the weight of the capability score.
[0024] w2 represents the weight of the distance factor, and w3 represents the weight of the urgency level; capability j (s k ) indicates the lower level
[0025] Agent j targets the k-th feature s of subtask i. k The verification capability score, s k The k-th feature of subtask i, where k is the feature index, and capability j (·) is the capability assessment function, with a real number output reflecting the professional capability of agent j under this feature; distance ij Distance represents the distance between the low-level agent j and the region involved in subtask i. ij >0; 1 / distance ij This represents the reciprocal of the distance; the smaller the distance, the more frequent the 1 / distance. ij A higher value indicates that the agent is more likely to be assigned to the nearest agent; urgency i This represents the urgency level of subtask i, usually a positive real number. i A higher value indicates a more urgent task.
[0026] Furthermore, the multiple low-level intelligent agents negotiate horizontally with each other through a predefined communication protocol regarding the received value verification subtasks, in order to determine a task collaborative execution scheme or resource sharing protocol.
[0027] Furthermore, the target low-level agent employs a second reinforcement learning model to execute the assigned value verification subtask, including:
[0028] Based on local real-time electrical quantity data, protection setting parameters to be verified, relevant power grid topology fragment information, and verification targets or constraints issued by higher authorities, the protection performance of the current settings under simulated specific fault sub-scenarios is evaluated, and the rationality of the settings is determined.
[0029] Furthermore, the centralized training and distributed execution mode includes:
[0030] During the centralized training phase, all interaction data of high-level agents and low-level agents are aggregated, and the first reinforcement learning model of the high-level agents and the second reinforcement learning model of the low-level agents are trained using the interaction data.
[0031] During the distributed execution phase, each agent uses the trained model parameters to make online decisions and perform verification.
[0032] Furthermore, the first reinforcement learning model is a reinforcement learning model based on long short-term memory networks combined with attention mechanisms.
[0033] Furthermore, the second reinforcement learning model is a deep Q-network model.
[0034] The present invention provides a storage medium, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the above-described method for setting verification in a distribution network fault scenario based on reinforcement learning.
[0035] The beneficial effects of this invention are as follows:
[0036] By constructing a first reinforcement learning model based on a long short-term memory network (LSTM) combined with an attention mechanism, the efficiency of handling complex faults is improved, enabling high-level intelligent agents to accurately predict fault development trends, intelligently decompose verification tasks, shorten the average verification time in complex fault scenarios, and effectively prevent fault propagation.
[0037] By using dynamic task allocation based on the contract network protocol and a horizontal negotiation mechanism among low-level agents, the collaborative capability of distributed protection units is enhanced, enabling efficient information sharing and intelligent collaboration among protection units deployed in different geographical locations and system levels. This solves the problems of distributed units operating independently and insufficient information sharing in traditional methods.
[0038] By using a second reinforcement learning model based on a deep Q-network and a centralized training and distributed execution mode, the adaptability and robustness of the setpoint verification are improved, enabling the entire setpoint verification process to be flexibly adjusted according to real-time fault conditions, power grid topology changes and agent states, ensuring continuous optimization of the global verification strategy.
[0039] This invention can effectively address complex cascading faults with long time-series dependencies in distribution networks, achieve efficient execution of protection setting verification tasks, enhance the autonomous response capability and adaptability to local changes of regional protection units, thereby optimizing operation and maintenance efficiency and fundamentally improving the safe and stable operation level of distribution networks. Attached Figure Description
[0040] Figure 1 This is a flowchart of a setpoint verification method for distribution network fault scenarios based on reinforcement learning, as described in this invention. Detailed Implementation
[0041] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.
[0042] At least one embodiment of the present invention discloses a method for setting verification in a distribution network fault scenario based on reinforcement learning, such as... Figure 1 As shown, it includes:
[0043] Step 1: Construct a hierarchical multi-agent verification system, which includes at least one high-level agent and multiple low-level agents.
[0044] This step aims to establish a hierarchical agent structure to achieve the decomposition and collaborative processing of complex setpoint verification tasks. This structure provides fundamental support and an information exchange platform for subsequent steps such as fault situation awareness, task decomposition and allocation. Specifically, a hierarchical multi-agent verification system is constructed, comprising at least one high-level agent (Meta-Controller) and multiple low-level agents (Controllers).
[0045] A high-level intelligent agent, such as a central server or a group of computing nodes with advanced decision-making capabilities, has the core functions of global or regional fault situation awareness, formulating macro-level verification strategies, and decomposing and allocating verification tasks.
[0046] Lower-level intelligent agents, such as intelligent electronic devices (IEDs) or protective relays deployed at key nodes in the power distribution network, are responsible for receiving instructions from higher-level intelligent agents or specific verification tasks determined through negotiation, performing local setting verification calculations, and reporting the results. Higher-level and lower-level intelligent agents interact with each other, as well as with each other, through a pre-defined communication network.
[0047] Step 2: The high-level intelligent agent uses the first reinforcement learning model to perceive the fault situation based on the real-time operation data of the distribution network, and decomposes the global setpoint verification target into multiple setpoint verification sub-tasks.
[0048] This step builds upon the hierarchical multi-agent verification system constructed in step 101. It utilizes the global perspective of a higher-level agent to analyze potential complex faults and refines the verification tasks, providing input for subsequent task allocation and execution. Specifically, the higher-level agent monitors the operational data of the distribution network in real time, including, for example, the current, voltage, switch status, and power output of each line and renewable energy source.
[0049] The high-level intelligent agent integrates a fault development trend analysis model, which is preferably a model trained based on deep reinforcement learning, using a Long Short-Term Memory (LSTM) network to process time-series data, and combined with an attention model.
[0050] Input data types: real-time operation time-series data of the distribution network and power grid topology information.
[0051] Specific output results include: potential failure evolution path sequences, the probability of occurrence of each path, and identified key influencing nodes or regions.
[0052] Based on the fault development trend obtained from the analysis, the high-level intelligent agent decomposes the complex global setting verification objective (e.g., ensuring that the protection settings of the entire system can respond correctly under a certain cascading fault) into a series of specific, executable sub-tasks.
[0053] Input data types: global setpoint verification target, fault development trend analysis results.
[0054] Specific output results: a list of setpoint verification subtasks for a specific line, a specific area, or a specific fault stage. Each subtask includes the verification object, verification requirements (such as speed and selectivity indicators), and suggested priority.
[0055] Step 3: The high-level intelligent agent assigns multiple fixed-value verification subtasks to the target low-level intelligent agent for execution based on the dynamic task allocation model. The dynamic task allocation model comprehensively evaluates the low-level intelligent agent's capability parameters, task-related distance parameters, and task urgency parameters.
[0056] This step takes the verification subtasks obtained from the decomposition of the high-level agent in step 2 as input, aiming to achieve intelligent and dynamic allocation of verification subtasks to low-level agents and promote collaboration among low-level agents. According to an embodiment of this application, the high-level agent employs a dynamic task allocation model to distribute the verification subtasks obtained from the decomposition in step 2 to the most suitable low-level agents or groups of low-level agents. Optionally, the dynamic task allocation model is built based on the ContractNet Protocol (CNP). The high-level agent, acting as the task publisher (Manager), broadcasts or targets the verification subtasks as contracts. The low-level agents, acting as potential contractors, calculate the task utility value based on their own status (such as computational load, existing verification skills, and historical task completion status) and task requirements (such as the urgency of the task, distance from or relevance to their jurisdiction), and bid to the high-level agent.
[0057] The task utility function is expressed as:
[0058] U ij =w1·capability j (s k ) + w2·(1 / distance ij )+w3·urgency i ;
[0059] U ij represents the utility value of low-level agent j for subtask i, used to measure the overall suitability of the low-level agent for the task, where i is the subtask number and j is the low-level agent number; w1 represents the weight of the capability score, w2 represents the weight of the distance factor, and w3 represents the weight of the urgency level; capability j (s k ) represents the low-level agent's response to the k-th feature s of subtask i. k The verification capability score, s k The k-th feature of subtask i, where k is the feature index, and capability j (·) is the capability assessment function, with a real number output reflecting the professional capability of agent j under this feature; distance ij Distance represents the distance between the low-level agent j and the region involved in subtask i. ij >0; 1 / distanceij This represents the reciprocal of the distance; the smaller the distance, the more frequent the 1 / distance. ij A higher value indicates that the agent is more likely to be assigned to the nearest agent; urgency i This represents the urgency level of subtask i, usually a positive real number. i A higher value indicates a more urgent task.
[0060] Input data types: Validation subtask list, real-time status information (capabilities, load, location, etc.) of each low-level agent.
[0061] Specific output results: the optimal allocation scheme of execution agents (or agent groups) for each verification subtask.
[0062] Based on the bidding results, the high-level intelligent agent selects the low-level intelligent agent (or a group of intelligent agents that can work together to complete the task) with the highest utility value and awards the contract, that is, assigns the sub-task.
[0063] Furthermore, this application also provides a horizontal negotiation model among low-level agents. In some implementations, when a low-level agent receives a task that is too complex or requires cooperation from other agents, it can initiate a negotiation request to other low-level agents through a predefined communication protocol (e.g., a message queue based on a shared task board, or a direct point-to-point request-response model). The negotiation content may include further subdivision of the task, the scope of required auxiliary verification, and the sharing of resources (such as specific historical data or computing power).
[0064] Input data types: tasks received by the lower-level agent, its own processing capacity assessment, and state information of other surrounding agents.
[0065] Specific output results: Task collaboration or resource sharing agreements reached with other intelligent agents.
[0066] Step 4: The target low-level agent uses the second reinforcement learning model to execute the assigned value verification subtask and obtain the value verification result.
[0067] This step takes the specific verification task allocated and negotiated in step 3 as input. Each lower-level agent performs the specific setting value verification calculation and feeds back the results, providing data support for subsequent global collaborative learning and optimization. Specifically, after receiving instructions allocated by higher-level agents or verification tasks determined through negotiation, each lower-level agent, within its own local perception data range (e.g., electrical quantities of its directly connected lines, voltage information of adjacent nodes), invokes the locally deployed setting value verification reinforcement learning model to perform specific protection setting value verification operations. The setting value verification reinforcement learning model is a reinforcement learning model with a deep Q-network (DQN) structure.
[0068] Input data types: local real-time electrical quantity data, protection setting parameters to be verified, relevant power grid topology segment information, and verification targets or constraints issued by higher authorities.
[0069] Specific output results include: a performance evaluation report of the current setting under simulated specific fault sub-scenarios (e.g., whether the action time meets the speed requirement, whether it will incorrectly switch non-faulty areas to evaluate selectivity, whether it can reliably detect the minimum fault current to evaluate sensitivity, and whether it will fail to operate under disturbances to evaluate reliability), and a judgment on whether the setting is reasonable (qualified, too large, too small, or a specific direction / amplitude for adjustment). In some implementations, when performing setting verification, the lower-level agent also simulates the execution of protection actions and predicts the subsequent impact of the actions on the local power grid state, including this prediction information as part of the verification results. After verification, the lower-level agent feeds back the verification results, key state information (such as key electrical quantity characteristics that lead to unreasonable settings), and optional execution confidence levels to the higher-level agent through the communication network, and shares this information with other lower-level agents that need it according to the requirements of the negotiation protocol.
[0070] Step 5: Based on the set value verification results and real-time operation data of the distribution network, the first reinforcement learning model and the second reinforcement learning model are jointly optimized using a centralized training and distributed execution mode.
[0071] This step takes the verification results fed back by the low-level agents in step 4 and the system operation data as input, aiming to continuously optimize the setpoint verification performance of the entire hierarchical multi-agent system through continuous learning, and achieve system-level adaptive and collaborative improvement. According to an embodiment of this application, a centralized training with decentralized execution (CTDE) mode is used to train the entire hierarchical multi-agent verification system. In the centralized training phase:
[0072] Step 5.1: Collect runtime data;
[0073] It summarizes all interaction data generated by high-level and low-level agents during actual operation or simulation, including state, action, reward, next state transition samples, task allocation records, negotiation interaction logs, etc.
[0074] Step 5.2, High-level agent strategy optimization;
[0075] The reinforcement learning model of the high-level agent (based on a Long Short-Term Memory (LSTM) network combined with an attention mechanism, responsible for fault situation awareness, task decomposition, and task allocation) is trained using collected global data. Its reward function comprehensively considers factors such as the accuracy of the final value verification, the efficiency of task completion, and the overall stability of the system. For example, if a task allocation leads to a cascading failure that is not correctly verified and isolated in a timely manner, the high-level allocation strategy will receive a negative reward.
[0076] Step 5.3, Low-level agent strategy optimization;
[0077] The setpoint verification reinforcement learning models of each low-level agent are trained using their local experience and experience shared from other agents (under the CTDE framework, global information or broader experience can be accessed during training). Their reward functions are primarily based on the accuracy of their local setpoint verification and their adherence to higher-level instructions. During the distributed execution phase, each agent uses the model parameters trained in the centralized phase for online decision-making and verification. Through the aforementioned global collaborative learning and policy optimization, higher-level agents can continuously improve the effectiveness of their task decomposition and collaborative allocation, while lower-level agents can continuously improve the accuracy of their local setpoint verification, thereby jointly promoting the global optimization of the entire system's setpoint verification performance in complex power distribution network fault scenarios.
[0078] A storage medium includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the above-described reinforcement learning-based setpoint verification method for power distribution network fault scenarios.
[0079] Here, the present invention provides an implementation example:
[0080] Suppose a city's power distribution network, area A, contains multiple 10kV feeder lines. Lines L1 and L2 are connected via tie switch K1, and lines L2 and L3 are connected via tie switch K2. The area has recently experienced rapid load growth, and the penetration rate of distributed photovoltaic power is gradually increasing, leading to more complex grid operation and a greater risk of cascading faults. It is necessary to verify the protection settings for this area in the event of different types of faults (such as a single-phase ground fault at the end of line L1, or a two-phase short-circuit fault on line L2 that may trigger a cascading fault due to overload tripping of line L3).
[0081] When deploying this method, a high-level intelligent agent (Meta-Controller) is set up in the area, and low-level intelligent agents (Controller-L1, Controller-L2, Controller-L3) are deployed on the critical protection devices (such as IEDs configured for circuit breakers) of the L1, L2, and L3 lines.
[0082] The high-level intelligent agent monitors real-time operating data of the power grid in area A. For example, the current of line L1 is 150A, the current of line L2 is 120A, and the current of line L3 is 100A. Both tie switches K1 and K2 are closed. Based on its internal first reinforcement learning model (a pre-trained LSTM network combined with an attention model), it analyzes that under the current operating conditions, if a ground fault occurs on line L1, there is a high probability (e.g., a predicted probability of 0.7) that improper protection coordination or fluctuations in renewable energy sources will cause the protection of line L2 to malfunction, which may then lead to overload of line L3 (predicted probability 0.4).
[0083] Therefore, the high-level intelligent agent decomposes the global objective of "ensuring the correct setpoints in region A under complex fault conditions" into the following sub-tasks:
[0084] Subtask 1 (T1): Verify the speed and selectivity of the L1 line grounding protection setting under the current load and new energy output. This has a high priority.
[0085] Subtask 2 (T2): Verify the phase-to-phase short-circuit protection settings of L2 line and evaluate its coordination with L1 protection. This has a high priority.
[0086] Subtask 3 (T3): Verify the overload protection settings of line L3, considering the potential power flow shift caused by faults in L1 and L2. (Medium priority)
[0087] High-level intelligent agents T1, T2, and T3 were released;
[0088] The lower-level intelligent agent Controller-L1 calculates the highest utility value for T1 (e.g., Y). L1,T1 =0.85), Controller-L2 calculates the highest utility value for T2 (e.g., U L2,T2 =0.9), Controller-L3 calculates the highest utility value for T3 (e.g., U L3,T3 =0.7). Therefore, T1 is assigned to Controller-L1, T2 is assigned to Controller-L2, and T3 is assigned to Controller-L3.
[0089] When executing T2, Controller-L2 finds that it needs detailed setting parameters of the L1 line to accurately assess the coordination. Therefore, it sends a request to Controller-L1 through the negotiation model. Controller-L1 responds to the request and shares its relevant setting parameters.
[0090] After receiving T1, Controller-L1 uses its local second reinforcement learning model (DQN model) to input the current real-time electrical quantities of the L1 line and the grounding protection setting to be verified (e.g., current setting I).set1 =50A, time constant t set1 =0.2s) and related topology information. The model output evaluates this set value: in the simulated L1 end-to-ground fault scenario, the action time is 0.18s (satisfying the speed requirement), and it will not cause L2 or K1 to malfunction (satisfying the selectivity requirement), thus it is judged as qualified. This result is fed back to the higher-level intelligent agent.
[0091] After obtaining the setting parameters of L1, Controller-L2 similarly sets the phase-to-phase short-circuit protection settings for the L2 line it is responsible for (e.g., I...). set2 =200A,t set2 The value was verified at 0.1s. Model evaluation revealed that under certain boundary conditions, this setpoint might act before the L1 protection, indicating a risk of selective coordination. This result and key boundary condition information were fed back to the higher-level agent.
[0092] After receiving feedback from Controller-L2 indicating a risk of selective cooperation, the higher-level agent, combined with the qualified information from Controller-L1, will adjust its task decomposition and allocation strategy based on a Long Short-Term Memory (LSTM) network and attention mechanism during its centralized training process. For example, in scenarios involving multi-path cooperation, a joint verification subtask might be directly generated and assigned to Controller-L1 and Controller-L2 for joint completion, or the evaluation weights for collaborative capabilities in the task utility function might be adjusted.
[0093] Meanwhile, the Controller-L2's local DQN model will also adjust its Q-value network parameters in subsequent training based on the cooperation risks discovered in this validation (e.g., obtaining a negative reward signal indicating "poor cooperation"), in order to more accurately identify such risks in the future.
[0094] In simulating the complex fault scenario of the distribution network in area A mentioned above, the method of this application demonstrates technical advantages in the following two aspects compared with traditional verification methods based on fixed rules or single agents:
[0095] Improved efficiency in handling complex faults: Through layering and collaboration, potential cascading fault risks can be identified earlier, and verification tasks can be decomposed in a targeted manner, effectively shortening the overall value verification time for complex fault scenarios.
[0096] Enhanced distributed protection coordination capabilities: Through dynamic task allocation and negotiation, distributed protection units can share information and coordinate actions more effectively, improving the accuracy of setting verification and the overall coordination level of protection.
[0097] The following are comparative data of key performance indicators in the simulation experiment:
[0098] The comparison results of the setpoint verification time in complex fault scenarios are shown in Table 1:
[0099] Table 1: Comparison of Setpoint Verification Time in Complex Fault Scenarios
[0100] Method type Average setpoint verification time (seconds) cascading failure misjudgment rate Traditional fixed rule validation method 125 15% Single agent reinforcement learning verification 80 8% The method described in this application 45 3%
[0101] The accuracy comparison of distributed protection collaborative verification is shown in Table 2:
[0102] Table 2: Comparison of Accuracy of Collaborative Verification in Distributed Protection
[0103] Method type Accuracy of critical protection coordination verification Fixed-value defect identification rate Traditional fixed rule validation method 70% 65% Uncooperative multi-agent verification 82% 78% The method described in this application 95% 92%
[0104] As can be seen from the data in Tables 1 and 2, the method described in this application shortens the setting value verification time in complex fault scenarios and significantly reduces the false judgment rate of cascading faults. At the same time, it significantly improves the collaborative verification accuracy and setting value defect identification rate of distributed protection, proving the effectiveness and superiority of this method.
[0105] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments based on the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.
Claims
1. A method for verifying setpoints in a distribution network fault scenario based on reinforcement learning, characterized in that, include: Construct a hierarchical multi-agent verification system, which includes at least one high-level agent and multiple low-level agents; Based on real-time operation data of the power distribution network, the high-level intelligent agent uses the first reinforcement learning model to perceive the fault situation and decomposes the global setpoint verification target into multiple setpoint verification sub-tasks. The high-level intelligent agent assigns multiple fixed-value verification subtasks to the target low-level intelligent agent for execution based on the dynamic task allocation model. The dynamic task allocation model comprehensively evaluates the low-level intelligent agent's capability parameters, task-related distance parameters, and task urgency parameters. The target low-level agent uses the second reinforcement learning model to execute the assigned value verification subtask and obtain the value verification result; Based on the set value verification results and real-time operation data of the distribution network, a centralized training and distributed execution mode is adopted to jointly optimize the first reinforcement learning model and the second reinforcement learning model. The dynamic task allocation model is built based on the ContractNet protocol; The high-level intelligent agent acts as the task publisher, publishing multiple fixed-value verification subtasks; Multiple low-level intelligent agents act as contractors, bidding for multiple fixed-value verification subtasks based on their respective task utility values. The higher-level intelligent agent selects the lower-level intelligent agent with the highest task utility value as the target lower-level intelligent agent; The task utility value is calculated as follows: ; in Represents low-level intelligent agents Pair Task The utility value is used to measure the overall suitability of a low-level agent for the task. Number the subtasks. Assign numbers to lower-level intelligent agents; The weights representing the ability scores Indicates the weight of the distance factor. Weights indicating the degree of urgency; Represents low-level intelligent agents For subtasks The Features Verification capability score For subtasks The One characteristic, For feature index, This is a capability evaluation function, with a real number output reflecting the agent's capabilities. Professional competence under this characteristic; Represents low-level intelligent agents sub-tasks The distance of the area involved ; This represents the reciprocal of the distance; the smaller the distance, the lower the reciprocal. The larger the value, the more likely it is to be allocated to agents that are closer in distance; Subtasks The degree of urgency is usually a positive real number. A higher value indicates a more urgent task; The target low-level agent employs a second reinforcement learning model to execute the assigned value verification subtask, including: Based on local real-time electrical quantity data, protection setting parameters to be verified, relevant power grid topology segment information, and verification targets or constraints issued by higher authorities, the protection performance of the current setting under a specific fault sub-scenario is evaluated, and the rationality of the setting is determined. The first reinforcement learning model is a reinforcement learning model based on long short-term memory networks combined with attention mechanisms; The second reinforcement learning model is a deep Q-network model.
2. The method for verifying setpoints in a distribution network fault scenario based on reinforcement learning according to claim 1, characterized in that, The fault situation awareness includes: The first reinforcement learning model is used to analyze the real-time operation data and topology information of the distribution network, and output the potential fault evolution path sequence and the probability of occurrence of each path.
3. The method for verifying setpoints in a distribution network fault scenario based on reinforcement learning according to claim 1, characterized in that, The multiple low-level intelligent agents negotiate horizontally with each other through a predefined communication protocol on the received value verification subtasks in order to determine a task collaborative execution plan or resource sharing protocol.
4. The method for verifying setpoints in a distribution network fault scenario based on reinforcement learning according to claim 1, characterized in that, The centralized training and distributed execution mode includes: During the centralized training phase, all interaction data between high-level and low-level agents are aggregated, and the first reinforcement learning model of the high-level agents and the second reinforcement learning model of the low-level agents are trained using the interaction data. During the distributed execution phase, each agent uses the trained model parameters to make online decisions and perform verification.
5. A storage medium, characterized in that, The device includes a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the setpoint verification method for distribution network fault scenarios based on reinforcement learning as described in any one of claims 1-4.
Citation Information
Patent Citations
Substation data collaborative identification method, device and equipment and storage medium
CN118820909A
Multi-agent exploration path planning system based on deep learning
CN119270866A