A high-voltage power distribution network operation topology optimization model and a construction method and system of a distributed reinforcement learning framework thereof

By constructing a topology optimization model for high-voltage distribution network operation and its distributed reinforcement learning framework, the problem of optimal operation mode during major maintenance under complex power grid background is solved. This achieves the stability, reliability and security of the power grid during major maintenance, solves the problem of reliance on human experience, and ensures the optimality and security of power grid operation.

CN122413628APending Publication Date: 2026-07-17QUANZHOU POWER SUPPLY COMPANY OF STATE GRID FUJIAN ELECTRIC POWER +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QUANZHOU POWER SUPPLY COMPANY OF STATE GRID FUJIAN ELECTRIC POWER
Filing Date
2026-03-27
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies struggle to determine the optimal operating mode during major maintenance periods in complex power grid contexts, making it difficult to guarantee the stability, reliability, and security of power grid operation. In particular, with the addition of new energy power and new loads, human experience is insufficient to find the optimal topology adjustment strategy.

Method used

A topology optimization model for high-voltage distribution networks and its distributed reinforcement learning framework are constructed. By building a weighted comprehensive objective function and constraints, a reinforcement learning framework with centralized training and distributed execution is established. The state space and action space of the agent are defined, and instantaneous and global reward mechanisms are designed. The model is trained by combining deep reinforcement learning algorithms to achieve optimal topology adjustment decisions.

Benefits of technology

It ensures the stability, reliability, and safety of the power grid during major maintenance periods, solves the problem of reliance on human experience, and ensures the optimal operation and safety of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122413628A_ABST
    Figure CN122413628A_ABST
Patent Text Reader

Abstract

This invention proposes a method and system for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework, comprising: Step S1: Constructing a high-voltage distribution network operation topology optimization model for the target high-voltage distribution network; including constructing the high-voltage distribution network operation topology optimization model and proposing its optimization objective and constraints; Step S2: Building and training the distributed reinforcement learning framework; including building a centralized training and distributed execution reinforcement learning computation framework for the high-voltage distribution network operation topology optimization problem of the target high-voltage distribution network, then determining the agents and the state space and action space of each agent, clarifying the reward function and training algorithm, and carrying out model training; Step S3: Formulating auxiliary decision-making for high-voltage distribution network topology optimization during the maintenance period of the target high-voltage distribution network, and finally obtaining the optimal operation topology of the high-voltage distribution network as a whole, and outputting performance evaluation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention proposes a method and system for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework, which relates to the field of power systems. Background Technology

[0002] Currently, to ensure reliable equipment operation, power grid equipment requires regular maintenance. Major maintenance work on substations in central areas and important interconnection channels can significantly impact the robustness of the power grid topology and increase operational safety risks. Therefore, in practical engineering, dedicated personnel are needed to assess the operational risks of power grids in areas undergoing major maintenance and determine operational topology adjustment strategies, such as changing the power supply path of a main transformer or a section of busbar. However, existing models often use node-load descriptions, oversimplifying the main electrical wiring of the power system and failing to reflect the risks of switching operations during operational mode adjustments. Algorithmically, the current common practice is to manually identify several candidate topology schemes and then perform calculations and comparisons. However, with increasingly complex power grid topologies and the introduction of randomly fluctuating renewable energy and diverse new load types, finding the optimal topology adjustment strategy based on human experience is becoming increasingly difficult. Summary of the Invention

[0003] In view of this, and to fill the gaps and deficiencies in existing technologies, this invention proposes a method and system for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework. This invention aims to solve the problem of determining the optimal operation mode of a high-voltage distribution network during major maintenance in complex power grid contexts, thereby ensuring the stability, reliability, and security of power grid operation during major maintenance periods.

[0004] This invention proposes a method and system for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework, including the following:

[0005] According to a first aspect of the present invention, the present invention proposes a method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework, characterized by comprising the following:

[0006] Step S1: Construct a high-voltage distribution network operation topology optimization model for the target high-voltage distribution network; including constructing a high-voltage distribution network operation topology optimization model based on the main electrical wiring of the target high-voltage distribution network and the power grid operation requirements of the region where the target high-voltage distribution network is located, and proposing its optimization objectives and constraints;

[0007] Step S2: Build and train the distributed reinforcement learning framework; this includes building a centralized training and distributed execution reinforcement learning computation framework for the topology optimization problem of the target high-voltage distribution network, then determining the agents and the state space and action space of each agent, clarifying the reward function and training algorithm, and carrying out model training.

[0008] Step S3: Formulate auxiliary decision-making for the topology optimization of the high-voltage distribution network during the maintenance period of the target high-voltage distribution network. This includes obtaining the topology adjustment decision of each agent based on the trained distributed reinforcement learning model and the original operating topology of the power grid for specific maintenance work. Finally, the optimal operating topology of the high-voltage distribution network as a whole is obtained by combining the results and outputting the performance evaluation results.

[0009] Further, step S1 includes the following:

[0010] Step S11: Construct the weighted comprehensive objective function of the target high-voltage distribution network, including minimizing the weighted comprehensive objective and converting it into a function expression, and constructing a function expression that includes the N-1 risk objective and the switching operation risk objective;

[0011] The functional expression for the weighted synthesis objective minimization transformation includes the following:

[0012] minF=ω L F L +ω0F0;

[0013] Where F represents the weighted composite objective, F L F0 and F0 represent the N-1 risk target and the switching operation risk target, respectively, ω L ω and ω0 correspond to the weights of the two sub-objectives, respectively;

[0014] The function expression, which includes N-1 risk control, includes the following:

[0015] ;

[0016] in, This indicates the direct load loss on the medium-voltage busbar of the substation caused by equipment failure tripping and voltage loss, coupled with the failure of the automatic transfer switch.

[0017] in, This refers to the indirect load loss caused by emergency load control due to equipment overload resulting from the automatic transfer switch operation of the medium-voltage busbar in the substation.

[0018] in: ;

[0019] ;

[0020] ;

[0021] Where s represents the substation of the target high-voltage distribution network;

[0022] In substations, electrical equipment involving bus tie functions is indicated by superscript M, electrical equipment involving main transformer functions is indicated by superscript A, and electrical equipment involving incoming and outgoing lines is indicated by superscript E.

[0023] Among them, the medium-voltage side bus of the substation is represented by b, the main transformer of the substation is represented by a, the outgoing line number is represented by e, and the tie line number is represented by c;

[0024] Among them, the load substation on the opposite side connected via outgoing line e uses s h This indicates that the high-voltage side incoming line of the load substation connected to the opposite side uses e h It indicates that, via the opposite load substation s h Connected to the opposite power substation using s d express;

[0025] Among them, the opposite power supply substation corresponding to the load substation uses s(s) h e h They jointly stated that the incoming line of the power supply substation on the opposite side uses e(s) h e h This indicates that the load substation s is connected to the opposite power supply substation via tie line c. c (s, e) is connected to the high-voltage side incoming line of the opposite power substation using e c (s, e) represents;

[0026] in This indicates the current open / closed status of the main transformer circuit breaker in substation s; This indicates the open / closed state of the main transformer circuit breaker at substation s at the previous moment; if the main transformer circuit breaker is closed at the current moment, then... , If the main transformer circuit breaker trips at the current moment, there is , ;

[0027] in This indicates the current open / closed status of the outgoing circuit breaker in substation s; This indicates the open / closed state of the outgoing circuit breaker of substation s at the previous moment; if the outgoing circuit breaker of substation s is closed at the current moment, then... , If the outgoing circuit breaker of substation s trips at the current moment, , ;

[0028] in This indicates the current open / closed status of the tie line circuit breaker in substation s; This indicates the open / closed state of the tie-line circuit breaker at substation s at the previous moment; if the tie-line circuit breaker at substation s is closed at the current moment, then... , If the tie-line circuit breaker of substation s trips at the current time, , ;

[0029] in This indicates the current open / closed status of the bus tie circuit breaker in substation s; This indicates the open / closed state of the bus tie circuit breaker at substation s at the previous moment; if the bus tie circuit breaker at substation s is closed at the current moment, then... , If the bus tie circuit breaker of substation s trips at the current moment, , ;

[0030] in This indicates the current open / closed status of the outgoing circuit breaker at the opposite power substation; This indicates the open / closed state of the outgoing circuit breaker of the opposite power substation at the previous moment; if the outgoing circuit breaker of the opposite power substation is closed at the current moment, then... , If the outgoing circuit breaker of the opposite power substation trips at the current moment, , ;

[0031] in This indicates the current open / closed status of the outgoing circuit breaker at the opposite power substation; This indicates the open / closed state of the outgoing circuit breaker of the opposite power substation at the previous moment; if the outgoing circuit breaker of the opposite power substation is closed at the current moment, then... , If the outgoing circuit breaker of the opposite power substation trips at the current moment, , ;

[0032] in This indicates the current open / closed status of the main transformer disconnector in substation s; This indicates the open / closed state of the main transformer disconnector at substation s at the previous moment; if the main transformer disconnector is closed at the current moment, then... , If the main transformer disconnect switch is tripped at the current moment, there is , ;

[0033] in This indicates the open / closed status of the outgoing disconnect switch of substation s; This indicates the open / closed state of the outgoing disconnector at substation s at the previous moment; if the outgoing disconnector is closed at the current moment, then... , If the outgoing disconnect switch is tripped at the current moment, there is , ;

[0034] in This indicates the current open / closed status of the tie-line disconnector in substation s; This indicates the open / closed state of the tie-line disconnector at substation s at the previous time; if the tie-line disconnector is closed at the current time, then... , If the disconnecting switch on the tie line is tripped at the current moment, there is , ;

[0035] in, and These respectively represent the safety risks of switching operations of circuit breakers and disconnectors. and These represent the safety risk coefficients for switching operations of circuit breakers and disconnectors, respectively.

[0036] Furthermore, step S1 also includes the following:

[0037] Step S12: Establish the constraints for the operation topology optimization of the target high-voltage distribution network, including high-voltage distribution network topology constraints, power flow constraints of maintenance methods, and operation constraints under N-1 faults.

[0038] The topology constraints of the high-voltage distribution network include radial operation constraints; the configuration conditions for the radial operation constraints are: the high-voltage distribution network maintains a radial structure, there is no circulating current, the open loop point of the high-voltage distribution network cannot be placed on the power substation side, and the high-voltage distribution network has fast protection for its energized lines.

[0039] The power flow constraints for maintenance methods include the following: Under optimized operating topology conditions, the power flow limit constraints for each line and power substation must meet the following requirements:

[0040] The load of the main transformer in the target high-voltage distribution network is equal to the total load of the bus it supplies and the remaining load capacity after sharing with other main transformers that do not belong to the target high-voltage distribution network.

[0041] The operational constraints under N-1 fault include: under N-1 fault, the backup automatic transfer devices of the target high-voltage distribution network can operate correctly, and the lines and main transformers of the target high-voltage distribution network do not experience conditions exceeding the rated overload capacity.

[0042] Step S13: The optimization objectives for the target high-voltage distribution network include the position state variables of circuit breakers and disconnectors that do not meet the special conditions; the position state variables of circuit breakers and disconnectors that meet the special conditions are not included in the optimization objectives; the circuit breakers and disconnectors that meet the special conditions include the following:

[0043] The circuit breakers and disconnect switches corresponding to the equipment under maintenance; the status of disconnect switches in single busbar connections; outside the equipment under maintenance, the outgoing circuit breakers corresponding to the voltage level of the high-voltage distribution network of the power supply substation are in the closed position; outside the equipment under maintenance, the disconnect switches of each switch bay of the power supply substation on the side away from the busbar are in the closed position.

[0044] Further, step S2 includes the following:

[0045] Step S21: Construction of the distributed reinforcement learning framework, including the adoption of a centralized training and distributed execution architecture. The distributed reinforcement learning framework includes two parts: a centralized training module and a distributed execution module. The centralized training module and the distributed execution module communicate in real time through a data interaction interface to ensure the consistency between the training process and the actual operation scenario of the power grid.

[0046] The centralized training module includes: responsible for global data aggregation, model parameter training and updating; integrating power grid operation data, action execution feedback and reward values ​​collected by various distributed agents; adjusting model parameters based on global optimization objectives to avoid global topology inconsistencies caused by local optima of a single agent; and embedding the constructed high-voltage distribution network topology optimization model of the target high-voltage distribution network as a constraint for model training to ensure that the trained model output conforms to power grid operation procedures.

[0047] The distributed execution module includes: partitioning the high-voltage distribution network topology; deploying multiple agents, each agent corresponding to a partition, responsible for collecting real-time power grid operation data within its partition, executing parameter commands issued by the centralized training module, outputting topology adjustment actions for its partition, and uploading the action execution results, power grid status feedback and reward value to the centralized training module, thereby realizing partitioned execution and global coordination.

[0048] Furthermore, step S2 also includes the following:

[0049] Step S22: Define the agent, state space, and action space, including the following:

[0050] The definition of intelligent agents includes defining their deployment and division of labor, which includes the following:

[0051] The intelligent agents are deployed according to power grid zones and are divided into two categories: power substation intelligent agents and load substation intelligent agents. The two types work together to complete topology optimization decisions.

[0052] Power Substation Intelligent Agent: Each power substation is equipped with one dedicated intelligent agent, which is responsible for the status decision of the medium voltage side bus tie switch, main transformer switch, outgoing line switch, tie line switch and corresponding disconnect switch of the substation, while taking into account the selection of the bus power supply mode of the substation and responding to load distribution requirements.

[0053] Load substation intelligent agent: Each load substation is equipped with one dedicated intelligent agent, which is responsible for the status decision of the high-voltage side incoming switch and bus section switch of the substation, and works with the power supply substation intelligent agent to complete the connection and optimization of the power supply path to ensure the continuity of load power supply;

[0054] The definition of the state space includes the following:

[0055] The state space describes the current state of the power grid operation of the agent. It covers the constraint-related parameters in the optimization model and is represented in vector form. The state space of each agent is defined differently based on its role, and its core includes the following parameters:

[0056] Common parameters: Current status of all circuit breakers in this zone, current status of all disconnect switches (0-1 variables), energized status of each busbar, current load rate of each line, current load rate of each main transformer, and specific information on maintenance work;

[0057] Intelligent parameters of power substation: rated capacity of each main transformer, rated load of each outgoing line and tie line, total load of each bus, availability of external power supply channels, and predicted potential load loss under the N-1 fault scenario.

[0058] Individual parameters of the intelligent agent of the load substation: total load of each section of the busbar of this substation, power supply reliability of each incoming line corresponding to the power supply substation, current status of the backup automatic transfer device, and load distribution of each incoming line;

[0059] The action space definition includes the following:

[0060] The action space describes the topology adjustment operations that the agent can perform. The actions must comply with the power grid switching operation procedures. A discrete action space is used, and each action corresponds to a specific operation instruction. The core actions are as follows:

[0061] The operation space of the power substation intelligent body includes: opening or closing of bus tie switch, opening or closing of main transformer switch, opening or closing of outgoing line switch, opening or closing of tie line switch, opening or closing of corresponding isolating switch, and switching between local power supply and external power supply mode.

[0062] The action space of the intelligent agent in the load substation includes: opening and closing operations of incoming line switches, opening and closing operations of bus section switches, and coordinated actions with the intelligent agent in the power supply substation to complete the switching of power supply paths.

[0063] Action constraints: Each agent can only execute one action at a time, and the execution of the action must satisfy the topology constraints in step one. Actions that cause grid circulating current or bus voltage loss are prohibited.

[0064] Furthermore, step S2 also includes the following:

[0065] Step S23: Design the reward function, including a composite reward mechanism of immediate reward and global reward, to guide the agent to learn the optimal topology adjustment strategy. The reward value is calculated as follows:

[0066] Total reward value R = α × immediate reward R1 + β × global reward R2;

[0067] Where α and β are weighting coefficients, and α+β=1;

[0068] Instant reward R1: Evaluates the local effect of a single action performed by an agent, with the core associated risk objective of switching operations. The calculation formula is as follows:

[0069] R1 = K1 - K2 × (Number of circuit breaker switching operations × β_U + Number of disconnector switching operations × β_V);

[0070] Where K1 is the basic reward constant, K2 is the penalty coefficient, β_U and β_V are the safety risk coefficients for switching operations of circuit breakers and disconnectors defined in step one; if the violation of the action leads to the destruction of topological constraints, then R1 = -K3; where K3 is the violation penalty constant, and K3 > K1;

[0071] Global Reward R2: Evaluates the global effect of coordinated actions by all agents, with a core associated risk objective of N-1. This reward is calculated by the centralized training module and then distributed to each agent. The calculation formula is as follows:

[0072] R2 = K4 - K5 × (General user load loss × λ1 + Important user load loss × λ2);

[0073] Where K4 is the global basic reward constant, K5 is the load loss penalty coefficient, λ1 and λ2 are the load loss coefficients for general users and important users, respectively, and λ2>λ1;

[0074] Step S24: Training algorithm selection and model training, including the following:

[0075] The distributed deep reinforcement learning algorithm is selected to adapt to the centralized training and distributed execution framework, which balances the fitting ability of the continuous state space and the execution efficiency of the discrete action space. The training process combines actual operation data of the high-voltage distribution network and simulated maintenance scenarios to ensure the generalization ability of the model.

[0076] Step S241: Data initialization: Collect historical operation data and maintenance case data of high-voltage distribution network, and build a training dataset that covers the power grid operation status under different maintenance scenarios, different load levels, and different fault scenarios;

[0077] Step S242: Model initialization: Initialize the global model parameters of the centralized training module and the local model parameters of each distributed agent, set the training hyperparameters, and embed topology optimization constraints as hard constraints for model training.

[0078] Step S243: Iterative Training: Each distributed agent collects the current power grid state, selects and executes actions based on the current model parameters, collects the power grid state and reward value after the action is executed, and stores the experience in the experience replay pool; the centralized training module randomly samples batch experiences from the experience replay pool of each agent, updates the global model parameters based on the total reward value, and distributes the updated parameters to each agent to complete one iteration;

[0079] Step S244: Constraint verification: After each iteration, verify whether the power grid topology corresponding to the topology adjustment action output by the agent satisfies all constraints. If not, adjust the penalty coefficient of the reward function and re-iterate the training.

[0080] Step S245: Convergence judgment: When the number of iterations reaches the set value, and the fluctuation range of the total reward value of multiple consecutive iterations is ≤5%, and the weighted comprehensive objective function value corresponding to the topology adjustment strategy output by the model reaches the minimum value, training is stopped and the trained distributed reinforcement learning model is saved.

[0081] Further, step S3 includes the following:

[0082] Step S31: Import the maintenance scenario and set the initial state, including the following:

[0083] The import of maintenance scenarios includes: obtaining detailed parameters of specific maintenance work and importing them into the centralized training module of the distributed reinforcement learning framework, including maintenance equipment, maintenance scope, maintenance duration, and load prediction data during maintenance.

[0084] Initial state setting: The normal operating topology of the power grid before maintenance is taken as the initial state. The power grid operating parameters of each partition are collected at this time and assigned to the corresponding distributed agents as the initial input for agent decision-making. At the same time, according to the maintenance requirements, the constraint threshold during the maintenance period is set and the weight coefficient of the reward function is adjusted.

[0085] For a specific maintenance task, the topology adjustment decision of each agent is obtained from the trained distributed reinforcement learning model and the original operating topology of the power grid. The optimal operating topology of the high-voltage distribution network as a whole is obtained by combining the results and outputting the performance evaluation results.

[0086] Step S32: Distributed agent decision execution, including the following:

[0087] Local decision output: Based on the initial state, maintenance scenario parameters and trained model parameters, each distributed agent autonomously selects the optimal topology adjustment action and outputs the local topology adjustment decision for its partition, including the opening / closing operation instructions of switches and disconnect switches, and the power supply mode switching instructions, and uploads the local decision to the centralized training module;

[0088] Global Decision Collaboration: The centralized training module receives local topology adjustment decisions from all agents and, combined with the topology optimization model from step one, performs collaborative verification of the global topology to determine whether the overall topology after combining local decisions satisfies all constraints and whether the weighted comprehensive objective is minimized.

[0089] If the global topology satisfies the constraints and the objective function value is optimal, it is directly determined as the final topology adjustment decision.

[0090] If the global topology does not meet the constraints, or the objective function value does not reach the optimal, the centralized training module issues an adjustment instruction to the corresponding agent. The agent re-outputs the local decision based on the adjustment instruction until the optimal global topology that meets all constraints is formed.

[0091] Feasibility verification of decisions: The final global topology adjustment decision is compared with the high-voltage distribution network switching operation procedures and major maintenance safety specifications to verify the feasibility of the decision. If there are infeasible operations, the decision of the corresponding intelligent agent is fine-tuned to ensure that the decision can be directly implemented.

[0092] Furthermore, step S3 also includes the following:

[0093] Step S33: Optimal running topology output, including the following:

[0094] After the final global topology adjustment decision is verified, the optimal operating topology of the high-voltage distribution network during this major overhaul is output. The output includes:

[0095] Global topology diagram: Marks the location of all power substations and load substations, the final status of each line, switch, and disconnector, the power supply mode of each bus, and the direction of the power supply path;

[0096] Detailed operation instructions: List all operation instructions to be executed in the order of switching operations, including the operating equipment, operation type, operation time node, and clarify the division of labor of each intelligent agent.

[0097] According to a second aspect of the present invention, the present invention proposes a system for constructing and training a distributed reinforcement learning framework based on a high-voltage distribution network operation topology optimization model, comprising an electronic device, wherein the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements a method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework as described in any one of the present invention.

[0098] According to a third aspect of the present invention, the present invention proposes a system for constructing and training a distributed reinforcement learning framework based on a high-voltage distribution network operation topology optimization model, comprising a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements a method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework as described in any one of the present invention.

[0099] The present invention has the following advantages:

[0100] This invention aims to address the challenge of determining the optimal operating mode of a high-voltage distribution network during major maintenance in complex power grid environments, ensuring the stability, reliability, and security of the power grid during such maintenance periods. Furthermore, this invention incorporates distributed reinforcement learning technology to achieve intelligent optimization of the high-voltage distribution network's operating topology, thus eliminating reliance on manual experience. Attached Figure Description

[0101] Figure 1 This is a schematic diagram of the steps of the present invention. Detailed Implementation

[0102] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.

[0103] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0104] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the present invention; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0105] like Figure 1 As shown, this invention proposes a method and system for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework, characterized by including the following:

[0106] According to a first aspect of the present invention, the present invention proposes a method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework, characterized by comprising the following:

[0107] Step S1: Construct a high-voltage distribution network operation topology optimization model for the target high-voltage distribution network; including constructing a high-voltage distribution network operation topology optimization model based on the main electrical wiring of the target high-voltage distribution network and the power grid operation requirements of the region where the target high-voltage distribution network is located, and proposing its optimization objectives and constraints;

[0108] Step S2: Build and train the distributed reinforcement learning framework; this includes building a centralized training and distributed execution reinforcement learning computation framework for the topology optimization problem of the target high-voltage distribution network, then determining the agents and the state space and action space of each agent, clarifying the reward function and training algorithm, and carrying out model training.

[0109] Step S3: Formulate auxiliary decision-making for the topology optimization of the high-voltage distribution network during the maintenance period of the target high-voltage distribution network. This includes obtaining the topology adjustment decision of each agent based on the trained distributed reinforcement learning model and the original operating topology of the power grid for specific maintenance work. Finally, the optimal operating topology of the high-voltage distribution network as a whole is obtained by combining the results and outputting the performance evaluation results.

[0110] Further, step S1 includes the following:

[0111] Step S1 constructs a weighted comprehensive objective function based on the main electrical wiring and actual operation procedures of the high-voltage distribution network, clarifies five types of constraints, and defines the variables to be optimized, laying a mathematical foundation for subsequent distributed reinforcement learning training. Among them, the load loss under N-1 fault only considers the failure scenario of the backup automatic transfer device, and is further subdivided into four types of fault scenarios for quantitative calculation.

[0112] In one embodiment of the present invention, step S1 includes the following:

[0113] The topology optimization of high-voltage distribution network operation takes the minimization of weighted comprehensive objective as its core, while taking into account the two sub-objectives of N-1 risk control and switching operation risk control. It adapts to the safe operation requirements during major power grid maintenance, and the weighting weight can be dynamically adjusted according to the power grid safety priority.

[0114] The target of this invention is high-voltage distribution networks, specifically the status of circuit breakers (commonly known as switches) and outgoing / tie lines / main transformer disconnecting switches connected to the medium-voltage side busbars of power supply substations, as well as the status of circuit breakers connected to the high-voltage side busbars of load substations.

[0115] The switches connected to the medium-voltage side busbars of power substations are divided into four categories: bus tie switches, used to connect different sections of the medium-voltage side busbars of the power substation; main transformer switches, used to connect the medium-voltage side switches of each main transformer in the power substation; outgoing line switches, connected to the load substations via transmission lines; and tie line switches, connected to another power substation via tie lines. In the following formulas, the superscript M is used for bus tie functions, the superscript A is used for main transformer functions, the superscript E is used for outgoing lines, and the superscript C is used for tie line functions. For a given power substation S, its medium-voltage side busbar number is represented by b, its main transformer number by a, its outgoing line number by e, its tie line number by c, and the load substation connected to the opposite side via outgoing line e is represented by s. h (s, e) indicates that the high-voltage side incoming line connected to the opposite load substation is represented by e. k (s, e) represents the load substation s k Connected to the opposite power substation using s d (s, e) indicates that the high-voltage side switches of the load substation are divided into two categories: one is the incoming line switch, which is represented by the superscript. It indicates that the power supply substation on the opposite side uses s(s) k e k ) indicates that the incoming line of the power supply substation on the opposite side uses e(s) k e k The superscript indicates that one type is a busbar sectionalizing switch. This indicates that the power substation S is connected to the opposite power substation via tie line C. c (s, e), connected to the high-voltage side incoming line of the opposite power substation using e c (s, e) represents the relationship between the high-voltage side incoming switch and the connected high-voltage side busbar in the load substation. The relationship between the two busbar sections connected to the sectionalizing switch is also fixed. The circuit breaker status is represented by the 0-1 variable u, and the disconnecting switch status is represented by the 0-1 variable v.

[0116] Further, step S1 includes the following:

[0117] Step S11: Construct the weighted comprehensive objective function of the target high-voltage distribution network, including minimizing the weighted comprehensive objective and converting it into a function expression, and constructing a function expression that includes the N-1 risk objective and the switching operation risk objective;

[0118] The functional expression for the weighted synthesis objective minimization transformation includes the following:

[0119] minF=ω L F L +ω0F0;

[0120] Where F represents the weighted composite objective, F L F0 and F0 represent the N-1 risk target and the switching operation risk target, respectively, ω L ω and ω0 correspond to the weights of the two sub-objectives, respectively;

[0121] The function expression, which includes N-1 risk control, includes the following:

[0122] ;

[0123] in, This indicates the direct load loss on the medium-voltage busbar of the substation caused by equipment failure tripping and voltage loss, coupled with the failure of the automatic transfer switch.

[0124] in, This refers to the indirect load loss caused by emergency load control due to equipment overload resulting from the automatic transfer switch operation of the medium-voltage busbar in the substation.

[0125] in: ;

[0126] ;

[0127] ;

[0128] Where s represents the substation of the target high-voltage distribution network;

[0129] In substations, electrical equipment involving bus tie functions is indicated by superscript M, electrical equipment involving main transformer functions is indicated by superscript A, and electrical equipment involving incoming and outgoing lines is indicated by superscript E.

[0130] Among them, the medium-voltage side bus of the substation is represented by b, the main transformer of the substation is represented by a, the outgoing line number is represented by e, and the tie line number is represented by c;

[0131] Among them, the load substation on the opposite side connected via outgoing line e uses s h This indicates that the high-voltage side incoming line of the load substation connected to the opposite side uses e h It indicates that, via the opposite load substation sh Connected to the opposite power substation using s d express;

[0132] Among them, the opposite power supply substation corresponding to the load substation uses s(s) h e h They jointly stated that the incoming line of the power supply substation on the opposite side uses e(s) h e h This indicates that the load substation s is connected to the opposite power supply substation via tie line c. c (s, e) is connected to the high-voltage side incoming line of the opposite power substation using e c (s, e) represents;

[0133] in This indicates the current open / closed status of the main transformer circuit breaker in substation s; This indicates the open / closed state of the main transformer circuit breaker at substation s at the previous moment; if the main transformer circuit breaker is closed at the current moment, then... , If the main transformer circuit breaker trips at the current moment, there is , ;

[0134] in This indicates the current open / closed status of the outgoing circuit breaker in substation s; This indicates the open / closed state of the outgoing circuit breaker of substation s at the previous moment; if the outgoing circuit breaker of substation s is closed at the current moment, then... , If the outgoing circuit breaker of substation s trips at the current moment, , ;

[0135] in This indicates the current open / closed status of the tie line circuit breaker in substation s; This indicates the open / closed state of the tie-line circuit breaker at substation s at the previous moment; if the tie-line circuit breaker at substation s is closed at the current moment, then... , If the tie-line circuit breaker of substation s trips at the current time, , ;

[0136] in This indicates the current open / closed status of the bus tie circuit breaker in substation s; This indicates the open / closed state of the bus tie circuit breaker at substation s at the previous moment; if the bus tie circuit breaker at substation s is closed at the current moment, then... , If the bus tie circuit breaker of substation s trips at the current moment, , ;

[0137] in This indicates the current open / closed status of the outgoing circuit breaker at the opposite power substation; This indicates the open / closed state of the outgoing circuit breaker of the opposite power substation at the previous moment; if the outgoing circuit breaker of the opposite power substation is closed at the current moment, then... , If the outgoing circuit breaker of the opposite power substation trips at the current moment, , ;

[0138] in This indicates the current open / closed status of the outgoing circuit breaker at the opposite power substation; This indicates the open / closed state of the outgoing circuit breaker of the opposite power substation at the previous moment; if the outgoing circuit breaker of the opposite power substation is closed at the current moment, then... , If the outgoing circuit breaker of the opposite power substation trips at the current moment, , ;

[0139] in This indicates the current open / closed status of the main transformer disconnector in substation s; This indicates the open / closed state of the main transformer disconnector at substation s at the previous moment; if the main transformer disconnector is closed at the current moment, then... , If the main transformer disconnect switch is tripped at the current moment, there is , ;

[0140] in This indicates the open / closed status of the outgoing disconnect switch of substation s; This indicates the open / closed state of the outgoing disconnector at substation s at the previous moment; if the outgoing disconnector is closed at the current moment, then... , If the outgoing disconnect switch is tripped at the current moment, there is , ;

[0141] in This indicates the current open / closed status of the tie-line disconnector in substation s; This indicates the open / closed state of the tie-line disconnector at substation s at the previous time; if the tie-line disconnector is closed at the current time, then... , If the disconnecting switch on the tie line is tripped at the current moment, there is , ;

[0142] in, and These respectively represent the safety risks of switching operations of circuit breakers and disconnectors. and These represent the safety risk coefficients for switching operations of circuit breakers and disconnectors, respectively.

[0143] Further, step S1 includes the following:

[0144] Step S12: Establish the constraints for the operation topology optimization of the target high-voltage distribution network, including high-voltage distribution network topology constraints, power flow constraints of maintenance methods, and operation constraints under N-1 faults.

[0145] The topology constraints of the high-voltage distribution network include radial operation constraints; the configuration conditions for the radial operation constraints are: the high-voltage distribution network maintains a radial structure, there is no circulating current, the open loop point of the high-voltage distribution network cannot be placed on the power substation side, and the high-voltage distribution network has fast protection for its energized lines.

[0146] The power flow constraints for maintenance methods include the following: Under optimized operating topology conditions, the power flow limit constraints for each line and power substation must meet the following requirements:

[0147] The load of the main transformer in the target high-voltage distribution network is equal to the total load of the bus it supplies and the remaining load capacity after sharing with other main transformers that do not belong to the target high-voltage distribution network.

[0148] The operational constraints under N-1 fault include: under N-1 fault, the backup automatic transfer devices of the target high-voltage distribution network can operate correctly, and the lines and main transformers of the target high-voltage distribution network do not experience conditions exceeding the rated overload capacity.

[0149] Furthermore, in one embodiment of the present invention, step S1 further includes the following:

[0150] In one embodiment of the present invention, the high-voltage distribution network topology constraints further include:

[0151] (A) Radial Operation Constraints: High-voltage distribution networks must maintain a radial structure to avoid circulating currents, and the open-loop point cannot be placed on the power substation side to ensure that live lines have rapid protection.

[0152] Power supply side substation:

[0153] When the bus tie switch corresponding to the voltage level of the high-voltage distribution network is in the open position, it is prohibited to place it in the closed position.

[0154] ;

[0155] Among them, U s This indicates the power substation, and s represents the status of the high-voltage side bus tie switch, which is a given parameter.

[0156] Power supply path tracing: Each busbar can only be powered by this station or by one external power source; the station's power supply can be multiple transformers connected to the same busbar section, or multiple transformers connected to different busbar sections, with bus couplers in parallel.

[0157] When this station is powered:

[0158] (Connected to the same section as the transformer, and powered by the corresponding transformer);

[0159] (Double busbar connection, with only one main transformer disconnector in the closed position);

[0160] (On the busbar, power is supplied by all transformers).

[0161] (When the bus tie is disconnected, power is not supplied by transformers connected to different bus sections).

[0162] (The logical relationship between the power supply to the busbar and the specific main transformer supplying the power).

[0163] (The logical relationship between the power supply to the busbar and the specific main transformer supplying the power).

[0164] Where I is the energized indicator variable, which is a 0-1 variable; the superscript L indicates that the power supply is from this station; and LA is used to indicate which transformer in this station supplies the power.

[0165] When powered by an external power source:

[0166] (The busbar must be energized; power supply can be either supplied by the station itself or transferred from an external power source.)

[0167] (If a busbar has an external power supply point, it will be supplied with power.)

[0168] (If a busbar has a connecting line transfer point, it is transferred to another line.)

[0169] (There can only be one busbar transfer line).

[0170] (The outgoing line e of the power substation s is actually the line being transferred to another power source);

[0171] (The outgoing line ed of the power substation sd on the opposite side of the transferred line is the transferred line);

[0172] (The line bay e of the power substation s can only be one of three states: being transferred to another, transferring to another, or a normal power supply line).

[0173] In this context, the superscript Z of the energized indicator variable I indicates that the power is being transferred; ZE indicates which incoming line is supplying the power; and ZC indicates which tie line is supplying the power. The 0-1 variable X distinguishes the state of the power substation's outgoing line, indicating whether it is in one of three states: being transferred, transferring power, or a normal power supply line.

[0174] The following explains what constitutes external power supply transfer: When external power supply is transferred, all switches tracing back to the busbar of the previous power station are in the closed position, and the busbar of the previous power station supplies power to this station. This is divided into two situations: tie lines and lines passing through load substations.

[0175] (When switching from tie line C to the busbar, the tie line switch at this station needs to be in the closed position.)

[0176] (When switching from tie line C to the busbar, the tie line switch of the power supply substation on the other side needs to be in the closed position.)

[0177] (When power is transferred from tie line C, the busbar BC connected to the line switch EC of the opposite substation SC is powered by this station).

[0178] Power is transferred via load substation lines:

[0179] (When external power supply e is being transferred, the external power supply e switch must be in the closed position.)

[0180] (When the external power supply is transferred, the switch ed of the line connecting the sd of the substation on the other side must be in the closed position.)

[0181] (When the external power source e is transferred to the busbar of the power supply substation s, all high-voltage side switches k of the intermediate load substation sh must be in the closed position; otherwise, at least one of the high-voltage side switches k of the intermediate load substation sh must be in the open loop.)

[0182] (When the external power supply e is transferred, the busbar bd connected to the line switch ed of the opposite substation sd is powered by this station).

[0183] (When external power supply e is transferred, the external power supply e disconnect switch must be connected to bus b, or connected to another bus but the bus tie is in the closed position.)

[0184] (When external power supply e is transferred, the external power supply e disconnect switch must be connected to bus b, or connected to another bus but the bus tie is in the closed position.)

[0185] (If the external power source e is connected to another busbar and the bus tie is broken, then the power will not be supplied by the external power source e).

[0186] (Double busbar connection, only one disconnector on the same line is in the closed position);

[0187] ,when (Power supply is not allowed to be transferred through the false double-circuit substation lines on the busbar.) It is a set used to indicate which outgoing lines of substation s are connected to pseudo-double-circuit substations.

[0188] Power flow constraints for maintenance methods: Ensure that power flow limits for each line and power source substation are met under optimized operating topology conditions.

[0189] Calculation of main transformer load in power supply side substation:

[0190] The load of the main transformer is the sum of the loads of the busbars it supplies (the sum of the loads of the power supply lines and the transfer lines) and the load allocated to other main transformers.

[0191] ;

[0192] ;

[0193] ;

[0194] Load calculations for each outgoing line (calculated uniformly from the power supply side substation to avoid duplication):

[0195] ;

[0196] ;

[0197] Power supply line load = total load of the substation busbars supplying power to it:

[0198] ;

[0199] ;

[0200] ;

[0201] Among them, the 0-1 variable IX is used to indicate whether a certain section of the busbar in the load substation is powered by a certain incoming line. This is used to indicate the load substation sh connected to the outgoing line e of the power substation, with the corresponding incoming line eh, and the set of section switch numbers that the k-th busbar passes through when it is powered by the incoming line.

[0202] The load of the transferred line = the total load of the substation busbar of the transferred power source:

[0203] ;

[0204] Load of the transfer line = Total load of the substation busbar of the transferred power source + Total load of the entire station on the transfer channel:

[0205] ;

[0206] Tie line load = Total load of the busbar on the supply side:

[0207] .

[0208] Furthermore, step S1 also includes the following:

[0209] Step S13: The optimization objectives for the target high-voltage distribution network include the position state variables of circuit breakers and disconnectors that do not meet the special conditions; the position state variables of circuit breakers and disconnectors that meet the special conditions are not included in the optimization objectives; the circuit breakers and disconnectors that meet the special conditions include the following:

[0210] The circuit breakers and disconnect switches corresponding to the equipment under maintenance; the status of disconnect switches in single busbar connections; outside the equipment under maintenance, the outgoing circuit breakers corresponding to the voltage level of the high-voltage distribution network of the power supply substation are in the closed position; outside the equipment under maintenance, the disconnect switches of each switch bay of the power supply substation on the side away from the busbar are in the closed position.

[0211] Furthermore, in one embodiment of the present invention, step S1 further includes the following:

[0212] Operating constraints under N-1 fault (considering backup automatic transfer logic): Under the anticipated N-1 fault of the bus / main transformer / line, the dummy equipment automatic transfer devices should all operate correctly, and the line and main transformer should not be severely overloaded; the load loss due to equipment voltage loss when there are no other backup channels.

[0213] Load losses include:

[0214] Because the objective of this invention is to optimize the network topology of high-voltage distribution networks, the N-1 fault scenario only needs to consider situations that cause the power supply busbar of the substation to trip / lose voltage. Specific scenarios could include main transformer tripping during single main transformer maintenance; busbar tripping during single busbar maintenance; or incoming line tripping during single incoming line maintenance. Therefore, this analysis does not delve into the specific causes of tripping / lose voltage on the substation's power supply busbar, but only analyzes the consequences of load loss.

[0215] The causes of load loss in high-voltage distribution networks are: 1) equipment failure, tripping, and loss of voltage combined with the failure of automatic transfer switch; 2) equipment overload and emergency load control.

[0216] Situations where automatic transfer switch (ATS) fails include: equipment maintenance causing the ATS mode to be unavailable; and simultaneous loss of backup power. Here, a 0-1 variable B is used to determine whether the ATS at the load substation is functioning correctly.

[0217] The reason for equipment overload after a fault is that after the automatic transfer switch is activated, the load is superimposed in the same direction, specifically including line overload and main transformer overload.

[0218] Calculation of direct load loss due to equipment failure tripping and voltage loss combined with automatic transfer switch failure: Load supplied by busbar - Load successfully transferred to standby:

[0219] ;

[0220] in, Indicates direct load loss; B indicates whether busbar b of substation s has experienced a fault / loss of voltage; B is used to indicate whether the automatic transfer switch of each substation is in operation.

[0221] Calculation of indirect load loss due to equipment overload and emergency load control after automatic switching action: Overload load - maximum load capacity of equipment.

[0222] ;

[0223] Outgoing line e directly supplies load to substation backup:

[0224] ;

[0225] The power supply substation connected to the same busbar as the outgoing line e is a backup power supply substation:

[0226] ;

[0227] The power supply substation connected to the same busbar section as the power supply substation via tie line C is available for backup operation.

[0228] ;

[0229] The above is an embodiment of step S13.

[0230] The variables to be optimized include:

[0231] The variables to be optimized in the topology optimization problem are the position and state variables of circuit breakers and disconnectors. However, to reduce the problem scale, based on the actual operation of the power grid, not all circuit breakers and disconnectors in the distribution network need to be included in the optimization. Scenarios that do not need to be included in the optimization include: 1) Circuit breakers and disconnectors corresponding to equipment under maintenance; 2) The state of disconnectors in single busbar connections; 3) Other than equipment under maintenance, the outgoing circuit breakers corresponding to the voltage level of the high-voltage distribution network of the power supply substation are in the closed position; 4. Other than equipment under maintenance, the disconnectors in the switch bays of the power supply substation far from the busbar are in the closed position.

[0232] Further, step S2 includes the following:

[0233] Step S21: Construction of the distributed reinforcement learning framework, including the adoption of a centralized training and distributed execution architecture. The distributed reinforcement learning framework includes two parts: a centralized training module and a distributed execution module. The centralized training module and the distributed execution module communicate in real time through a data interaction interface to ensure the consistency between the training process and the actual operation scenario of the power grid.

[0234] In one embodiment of the present invention, step S2 includes the following:

[0235] The architecture adopts a "centralized training-distributed execution" approach, which is adapted to the distributed topology characteristics of multiple substations and lines in high-voltage distribution networks. It balances training efficiency and execution flexibility. The core of the framework is divided into two parts: a centralized training module and a distributed execution module. The two communicate in real time through a data interaction interface to ensure the consistency between the training process and the actual operation scenario of the power grid.

[0236] The centralized training module is responsible for global data aggregation, model parameter training and updating. It integrates power grid operation data, action execution feedback and reward values ​​collected by various distributed agents, adjusts model parameters based on global optimization objectives, and avoids the problem of unreasonable global topology caused by local optima of a single agent. At the same time, it embeds the high-voltage distribution network topology optimization model (including objective function and constraints) constructed in step one as the constraint basis for model training, ensuring that the output of the trained model conforms to the power grid operation procedures.

[0237] Distributed execution module: Multiple agents are deployed according to the high-voltage distribution network topology partition (divided according to the power supply range of power substations). Each agent corresponds to a partition and is responsible for collecting real-time power grid operation data (such as bus voltage, line load, switch status, etc.) within its partition, executing parameter commands issued by the centralized training module, outputting topology adjustment actions for its partition, and uploading the action execution results, power grid status feedback and reward value to the centralized training module to achieve "partition execution and global coordination".

[0238] Further, step S2 includes the following:

[0239] The deployment and division of labor among intelligent agents include the following:

[0240] The intelligent agents are deployed according to power grid zones and are divided into two categories: power substation intelligent agents and load substation intelligent agents. The two types work together to complete topology optimization decisions.

[0241] Power Substation Intelligent Agent: Each power substation is equipped with one dedicated intelligent agent, which is responsible for the status decision of the medium voltage side bus tie switch, main transformer switch, outgoing line switch, tie line switch and corresponding disconnect switch of the substation, and takes into account the selection of the power supply mode of the bus (local power supply / external power transfer) and responds to load distribution requirements.

[0242] Load substation intelligent agent: Each load substation is equipped with one dedicated intelligent agent, which is responsible for the status decision of the high-voltage side incoming switch and bus section switch of the substation, and works with the power supply substation intelligent agent to complete the connection and optimization of the power supply path to ensure the continuity of power supply to the load.

[0243] The state space definition includes the following:

[0244] The state space describes the current state of the power grid operation of the agent and needs to fully cover the constraint-related parameters in the optimization model of step one. It is represented in vector form, and the state space of each agent is defined differently according to its division of labor. The core includes the following common and unique parameters:

[0245] Common parameters: Current status of all circuit breakers (switches) in this zone (0-1 variables, 0 for open position, 1 for closed position), current status of all disconnecting switches (0-1 variables), energized status of each busbar (0-1 variables, 0 for undervoltage, 1 for energized position), current load rate of each line, current load rate of each main transformer, and specific information on major maintenance work (maintenance equipment, maintenance scope, maintenance duration).

[0246] Intelligent parameters of power substation: rated capacity of each main transformer, rated load of each outgoing line and tie line, total load of each bus, availability of external power supply channels (0-1 variable, 0 for unavailable, 1 for available), and predicted potential load loss under N-1 fault scenarios.

[0247] Individual parameters of the intelligent agent of the load substation: total load of each section of the busbar in this substation, power supply reliability of each incoming line corresponding to the power supply substation, current status of the standby automatic transfer device (0-1 variable, 0 for failure, 1 for normal), and load distribution of each incoming line.

[0248] The action space definition includes the following:

[0249] The action space describes the topology adjustment operations that the agent can perform. Actions must conform to the power grid switching operation procedures and correspond one-to-one with the variables to be optimized (switches and disconnectors) in the optimization model of step one. A discrete action space (0-1 actions) is used, with each action corresponding to a specific operation instruction. The core actions are as follows:

[0250] The operating space of the power substation intelligent agent includes: bus tie switch opening / closing operation, main transformer switch opening / closing operation, outgoing line switch opening / closing operation, tie line switch opening / closing operation, corresponding isolating switch opening / closing operation, and selection of local power supply / external power transfer mode switching (only when the transfer channel is available).

[0251] The action space of the intelligent agent in the load substation includes: opening / closing of incoming line switches, opening / closing of bus section switches, and coordinated actions with the intelligent agent in the power supply substation to complete the switching of power supply paths (such as adjusting the distribution of incoming line loads).

[0252] Action constraints: Each agent can only execute one action at a time, and the execution of the action must meet the topological constraints of step one (such as radial operation, power supply path tracing, etc.). Actions that cause grid circulating current, bus voltage loss, or other violations are prohibited.

[0253] Furthermore, step S2 also includes the following:

[0254] Step S23: Design a reward function, including a composite reward mechanism of immediate reward plus global reward, to guide the agent to learn the optimal topology adjustment strategy.

[0255] The reward function is used to evaluate the quality of the agent's actions. It is highly consistent with the weighted comprehensive objective function constructed in step one (minimizing N-1 risk and switching operation risk). A composite reward mechanism of "immediate reward + global reward" is adopted to guide the agent to learn the optimal topology adjustment strategy. The reward value is calculated as follows:

[0256] The total reward value R = α × immediate reward R1 + β × global reward R2, where α and β are weighting coefficients (α + β = 1), which can be dynamically adjusted according to the power grid security priority to ensure the power grid security during major maintenance periods.

[0257] Instant reward R1: Evaluates the local effect of a single action performed by an agent, with the core associated risk objective of switching operations. The calculation formula is as follows:

[0258] R1 = K1 - K2 × (Number of circuit breaker switching operations × β_U + Number of disconnector switching operations × β_V)

[0259] Wherein, K1 is the basic reward constant (assigned when the action is compliant), K2 is the penalty coefficient, and β_U and β_V are the safety risk coefficients for switching operations of circuit breakers and disconnectors defined in step one, respectively; if the action is non-compliant (such as causing the destruction of topological constraints), then R1=-K3 (K3 is the non-compliance penalty constant, and K3>K1).

[0260] Global Reward R2: Evaluates the global effect of coordinated actions by all agents, with a core associated risk objective of N-1. This reward is calculated by the centralized training module and then distributed to each agent. The calculation formula is as follows:

[0261] R2 = K4 - K5 × (General user load loss × λ1 + Important user load loss × λ2)

[0262] Wherein, K4 is the global basic reward constant, K5 is the load loss penalty coefficient, λ1 and λ2 are the load loss coefficients for general users and important users respectively (λ2>λ1), and the load loss value adopts the load loss calculation method of N-1 risk target in step one (only considering the failure scenario of the backup automatic transfer device).

[0263] Furthermore, step S2 also includes the following:

[0264] Step S24: Training algorithm selection and model training, including the following:

[0265] The Distributed Deep Reinforcement Learning (DDPG) algorithm was selected to adapt to the "centralized training-distributed execution" framework, balancing the fitting ability of the continuous state space with the execution efficiency of the discrete action space. The training process combines actual operating data of the high-voltage distribution network with simulated maintenance scenarios to ensure the generalization ability of the model.

[0266] In one embodiment of the present invention, step S2 further includes the following:

[0267] Step S241: Data initialization: Collect historical operating data of high-voltage distribution network (switch status, load data, fault data, etc.) and major maintenance case data to build a training dataset that covers the power grid operating status under different maintenance scenarios, different load levels, and different fault scenarios.

[0268] Step S242: Model initialization: Initialize the global model parameters of the centralized training module and the local model parameters of each distributed agent, set the training hyperparameters (learning rate, number of iterations, size of the experience replay pool, weight coefficients α, β, etc.), and embed the topology optimization constraints from step one as hard constraints for model training.

[0269] Step S243: Iterative Training: Each distributed agent collects the current power grid state (the initial state is the historical normal operation state or the simulated maintenance initial state), selects and executes actions based on the current model parameters, collects the power grid state and reward value after the action is executed, and stores the experience (state, action, reward, next state) in the experience replay pool; the centralized training module randomly samples batch experiences from the experience replay pool of each agent, updates the global model parameters based on the total reward value, and distributes the updated parameters to each agent to complete one iteration.

[0270] Step S244: Constraint Verification: After each iteration, verify whether the power grid topology corresponding to the topology adjustment action output by the agent satisfies all the constraints in Step 1 (topology constraints, power flow constraints, N-1 fault constraints). If not, adjust the penalty coefficient of the reward function and iterate the training again.

[0271] Step S245: Convergence judgment: When the number of iterations reaches the set value, and the total reward value of multiple consecutive iterations tends to stabilize (fluctuation range ≤ 5%), and the weighted comprehensive objective function value corresponding to the topology adjustment strategy output by the model reaches the minimum value, training is stopped and the trained distributed reinforcement learning model is saved.

[0272] Further, step S3 includes the following:

[0273] The maintenance scenario import and initial state setting include the following:

[0274] Maintenance scenario import: Obtain detailed parameters for specific major maintenance tasks and import them into the centralized training module of the distributed reinforcement learning framework, including maintenance equipment (such as main transformers in central substations, important interconnection lines, etc.), maintenance scope (involved switches, lines, busbars), maintenance duration, and load forecast data during the maintenance period (load forecast values ​​for general users and important users).

[0275] Initial state setting: The normal operating topology of the power grid before maintenance is used as the initial state. At this time, the power grid operation parameters of each zone (switching status, disconnecting switch status, bus energized status, line and main transformer load rate, etc.) are collected and assigned to the corresponding distributed agents as the initial input for agent decision-making. At the same time, according to the maintenance requirements, the constraint thresholds during the maintenance period (such as the upper limit of line load rate, the upper limit of main transformer load rate, etc.) are set, and the weight coefficients of the reward function are adjusted (prioritizing to increase the weight β of the N-1 risk target).

[0276] For specific major maintenance tasks, the topology adjustment decisions of each agent are obtained from the trained distributed reinforcement learning model and the original operating topology of the power grid. The optimal operating topology of the high-voltage distribution network as a whole is obtained by combining the results and outputting the performance evaluation results.

[0277] The distributed agent decision-making execution includes the following:

[0278] Local decision output: Based on the initial state, maintenance scenario parameters, and trained model parameters, each distributed agent autonomously selects the optimal topology adjustment action (action selection follows the "maximum reward value" principle), outputs the local topology adjustment decision for its partition, including the opening / closing operation instructions of switches and disconnect switches, and the power supply mode switching instructions (if any), and uploads the local decision to the centralized training module.

[0279] Global Decision Collaboration: The centralized training module receives local topology adjustment decisions from all agents and, combined with the topology optimization model from step one, performs collaborative verification of the global topology to determine whether the overall topology after combining local decisions satisfies all constraints and whether the weighted comprehensive objective is minimized.

[0280] If the global topology satisfies the constraints and the objective function value is optimal, it is directly determined as the final topology adjustment decision.

[0281] If the global topology does not meet the constraints (such as circulating current, bus undervoltage, equipment overload, etc.), or the objective function value does not reach the optimum, the centralized training module issues adjustment instructions to the corresponding agent. The agent re-outputs local decisions based on the adjustment instructions until an optimal global topology that meets all constraints is formed.

[0282] Feasibility verification of decisions: The final global topology adjustment decision is compared with the high-voltage distribution network switching operation procedures and major maintenance safety specifications to verify the feasibility of the decision (such as whether the operation sequence is compliant and whether there are any safety hazards). If there are infeasible operations, the decision of the corresponding intelligent agent is fine-tuned to ensure that the decision can be directly implemented.

[0283] The optimal running topology output includes the following:

[0284] After the final global topology adjustment decision is verified, the optimal operating topology of the high-voltage distribution network during this major overhaul is output. The output includes:

[0285] Global topology diagram: Marks the location of all power substations and load substations, the final status (open / closed) of each line, switch, and disconnector, the power supply mode of each bus (local power supply / external power transfer), and the power supply path.

[0286] Detailed operation instructions: List all operation instructions to be executed in the order of switching operations, including operating equipment (switches, disconnecting switches), operation type (open / close), operation time nodes, and clearly define the division of labor of each intelligent agent.

[0287] Topology parameter details: total load of each bus, load rate of each line, load rate of each main transformer, and load status of external power supply channels. Ensure that all parameters meet the constraints of step one.

[0288] The performance evaluation results output includes the following:

[0289] Based on the optimal operating topology and combined with the weighted comprehensive objective function from step one, the topology performance evaluation results are output from three dimensions: safety, economy, and reliability. This provides auxiliary basis for power grid operation and dispatch during maintenance. The evaluation indicators are as follows:

[0290] N-1 Risk Assessment: Calculate the load loss value of the optimal topology under the N-1 fault scenario (failure of the backup automatic transfer device) (distinguishing between general users and important users), compare it with the N-1 load loss value of the normal topology before maintenance, and evaluate the fault resilience of the topology; if the load loss value decreases, it indicates that the N-1 risk control effect of the topology has improved.

[0291] Risk assessment of switching operations: Calculate the number of circuit breakers and disconnectors required for optimal topology adjustment, calculate the target value of switching operation risk, compare it with the switching operation risk value of manually proposed schemes, and evaluate the convenience and safety of topology adjustment.

[0292] Weighted comprehensive objective evaluation: Calculate the weighted comprehensive objective function value of the optimal topology, clarify the specific contribution ratio of the N-1 risk objective and the switching operation risk objective, and verify the minimization effect of the objective function.

[0293] Reliability assessment: Calculate the power supply reliability indicators (average power availability, power supply guarantee rate for important users) of the optimal topology during the maintenance period, and assess whether the topology can meet the power supply reliability requirements during major maintenance.

[0294] The final output is a performance evaluation report, which clarifies the advantages and disadvantages of the optimal operating topology and provides decision-making reference for power grid dispatchers. If the evaluation results do not meet the expected requirements, the maintenance scenario parameters can be re-imported, the model hyperparameters can be adjusted, and the topology adjustment decision can be regenerated.

[0295] The above are preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention shall fall within the protection scope of the present invention.

Claims

1. A method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework, characterized in that, Includes the following: Step S1: Construct a high-voltage distribution network operation topology optimization model for the target high-voltage distribution network; This includes constructing a high-voltage distribution network operation topology optimization model based on the main electrical wiring of the target high-voltage distribution network and the power grid operation requirements of the area where the target high-voltage distribution network is located, and proposing its optimization objectives and constraints. Step S2: Build and train the distributed reinforcement learning framework; This includes building a reinforcement learning computation framework for centralized training and distributed execution of the topology optimization problem of the target high-voltage distribution network, then determining the agents and the state space and action space of each agent, clarifying the reward function and training algorithm, and carrying out model training; Step S3: Formulate auxiliary decision-making for the topology optimization of the high-voltage distribution network during the maintenance period of the target high-voltage distribution network. This includes obtaining the topology adjustment decision of each agent based on the trained distributed reinforcement learning model and the original operating topology of the power grid for specific maintenance work. Finally, the optimal operating topology of the high-voltage distribution network as a whole is obtained by combining the results and outputting the performance evaluation results.

2. The method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework according to claim 1, characterized in that, Step S1 includes the following: Step S11: Construct the weighted comprehensive objective function of the target high-voltage distribution network, including minimizing the weighted comprehensive objective and converting it into a function expression, and constructing a function expression that includes the N-1 risk objective and the switching operation risk objective; The functional expression for the weighted synthesis objective minimization transformation includes the following: minF=ω L F L +ω0F0; Where F represents the weighted composite objective, F L F0 and F0 represent the N-1 risk target and the switching operation risk target, respectively, ω L ω and ω0 correspond to the weights of the two sub-objectives, respectively; The function expression, which includes N-1 risk control, includes the following: ; in, This indicates the direct load loss on the medium-voltage busbar of the substation caused by equipment failure tripping and voltage loss, coupled with the failure of the automatic transfer switch. in, This refers to the indirect load loss caused by emergency load control due to equipment overload resulting from the automatic transfer switch operation of the medium-voltage busbar in the substation. in: ; ; ; Where s represents the substation of the target high-voltage distribution network; In substations, electrical equipment involving bus tie functions is indicated by superscript M, electrical equipment involving main transformer functions is indicated by superscript A, and electrical equipment involving incoming and outgoing lines is indicated by superscript E. Among them, the medium-voltage side bus of the substation is represented by b, the main transformer of the substation is represented by a, the outgoing line number is represented by e, and the tie line number is represented by c; Among them, the load substation on the opposite side connected via outgoing line e uses s h This indicates that the high-voltage side incoming line of the load substation connected to the opposite side uses e h It indicates that, via the opposite load substation s h Connected to the opposite power substation using s d express; Among them, the opposite power supply substation corresponding to the load substation uses s(s) h e h They jointly stated that the incoming line of the power supply substation on the opposite side uses e(s) h e h This indicates that the load substation s is connected to the opposite power supply substation via tie line c. c (s, e) is connected to the high-voltage side incoming line of the opposite power substation using e c (s, e) represents; in This indicates the current open / closed status of the main transformer circuit breaker in substation s; This indicates the open / closed state of the main transformer circuit breaker at substation s at the previous moment; if the main transformer circuit breaker is closed at the current moment, then... , If the main transformer circuit breaker trips at the current moment, there is , ; in This indicates the current open / closed status of the outgoing circuit breaker in substation s; This indicates the open / closed state of the outgoing circuit breaker of substation s at the previous moment; if the outgoing circuit breaker of substation s is closed at the current moment, then... , If the outgoing circuit breaker of substation s trips at the current moment, , ; in This indicates the current open / closed status of the tie line circuit breaker in substation s; This indicates the open / closed state of the tie-line circuit breaker at substation s at the previous moment; if the tie-line circuit breaker at substation s is closed at the current moment, then... , If the tie-line circuit breaker of substation s trips at the current time, , ; in This indicates the current open / closed status of the bus tie circuit breaker in substation s; This indicates the open / closed state of the bus tie circuit breaker at substation s at the previous moment; if the bus tie circuit breaker at substation s is closed at the current moment, then... , If the bus tie circuit breaker of substation s trips at the current time, , ; in This indicates the current open / closed status of the outgoing circuit breaker at the opposite power substation; This indicates the open / closed state of the outgoing circuit breaker of the opposite power substation at the previous moment; if the outgoing circuit breaker of the opposite power substation is closed at the current moment, then... , If the outgoing circuit breaker of the opposite power substation trips at the current moment, , ; in This indicates the current open / closed status of the outgoing circuit breaker at the opposite power substation; This indicates the open / closed state of the outgoing circuit breaker of the opposite power substation at the previous moment; if the outgoing circuit breaker of the opposite power substation is closed at the current moment, then... , If the outgoing circuit breaker of the opposite power substation trips at the current moment, , ; in This indicates the current open / closed status of the main transformer disconnector in substation s; This indicates the open / closed state of the main transformer disconnector at substation s at the previous moment; if the main transformer disconnector is closed at the current moment, then... , If the main transformer disconnect switch is tripped at the current moment, there is , ; in This indicates the open / closed status of the outgoing disconnect switch of substation s; This indicates the open / closed state of the outgoing disconnector at substation s at the previous moment; if the outgoing disconnector is closed at the current moment, then... , If the outgoing disconnect switch is tripped at the current moment, there is , ; in This indicates the current open / closed status of the tie-line disconnector in substation s; This indicates the open / closed state of the tie-line disconnector at substation s at the previous time; if the tie-line disconnector is closed at the current time, then... , If the disconnecting switch on the tie line is tripped at the current moment, there is , ; in, and These respectively represent the safety risks of switching operations of circuit breakers and disconnectors. and These represent the safety risk coefficients for switching operations of circuit breakers and disconnectors, respectively.

3. The method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework according to claim 2, characterized in that, Step S1 also includes the following: Step S12: Establish the constraints for the operation topology optimization of the target high-voltage distribution network, including high-voltage distribution network topology constraints, power flow constraints of maintenance methods, and operation constraints under N-1 faults. The topology constraints of the high-voltage distribution network include radial operation constraints; the configuration conditions for the radial operation constraints are: the high-voltage distribution network maintains a radial structure, there is no circulating current, the open loop point of the high-voltage distribution network cannot be placed on the power substation side, and the high-voltage distribution network has fast protection for its energized lines. The power flow constraints for maintenance methods include the following: Under optimized operating topology conditions, the power flow limit constraints for each line and power substation must meet the following requirements: The load of the main transformer in the target high-voltage distribution network is equal to the total load of the bus it supplies and the remaining load capacity after sharing with other main transformers that do not belong to the target high-voltage distribution network. The operational constraints under N-1 fault include: under N-1 fault, the backup automatic transfer devices of the target high-voltage distribution network can operate correctly, and the lines and main transformers of the target high-voltage distribution network do not experience conditions exceeding the rated overload capacity. Step S13: The optimization objectives for the target high-voltage distribution network include the position state variables of circuit breakers and disconnectors that do not meet the special conditions; the position state variables of circuit breakers and disconnectors that meet the special conditions are not included in the optimization objectives; the circuit breakers and disconnectors that meet the special conditions include the following: The circuit breakers and disconnect switches corresponding to the equipment under maintenance; the status of disconnect switches in single busbar connections; outside the equipment under maintenance, the outgoing circuit breakers corresponding to the voltage level of the high-voltage distribution network of the power supply substation are in the closed position; outside the equipment under maintenance, the disconnect switches of each switch bay of the power supply substation on the side away from the busbar are in the closed position.

4. The method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework according to claim 1, characterized in that, Step S2 includes the following: Step S21: Construction of the distributed reinforcement learning framework, including the adoption of a centralized training and distributed execution architecture. The distributed reinforcement learning framework includes two parts: a centralized training module and a distributed execution module. The centralized training module and the distributed execution module communicate in real time through a data interaction interface to ensure the consistency between the training process and the actual operation scenario of the power grid. The centralized training module includes: responsible for global data aggregation, model parameter training and updating; integrating power grid operation data, action execution feedback and reward values ​​collected by various distributed agents; adjusting model parameters based on global optimization objectives to avoid global topology inconsistencies caused by local optima of a single agent; and embedding the constructed high-voltage distribution network topology optimization model of the target high-voltage distribution network as a constraint for model training to ensure that the trained model output conforms to power grid operation procedures. The distributed execution module includes: partitioning the high-voltage distribution network topology; deploying multiple agents, each agent corresponding to a partition, responsible for collecting real-time power grid operation data within its partition, executing parameter commands issued by the centralized training module, outputting topology adjustment actions for its partition, and uploading the action execution results, power grid status feedback and reward value to the centralized training module, thereby realizing partitioned execution and global coordination.

5. The method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework according to claim 4, characterized in that, Step S2 also includes the following: Step S22: Define the agent, state space, and action space, including the following: The definition of intelligent agents includes defining their deployment and division of labor, which includes the following: The intelligent agents are deployed according to power grid zones and are divided into two categories: power substation intelligent agents and load substation intelligent agents. The two types work together to complete topology optimization decisions. Power Substation Intelligent Agent: Each power substation is equipped with one dedicated intelligent agent, which is responsible for the status decision of the medium voltage side bus tie switch, main transformer switch, outgoing line switch, tie line switch and corresponding disconnect switch of the substation, while taking into account the selection of the bus power supply mode of the substation and responding to load distribution requirements. Load substation intelligent agent: Each load substation is equipped with one dedicated intelligent agent, which is responsible for the status decision of the high-voltage side incoming switch and bus section switch of the substation, and works with the power supply substation intelligent agent to complete the connection and optimization of the power supply path to ensure the continuity of load power supply; The definition of the state space includes the following: The state space describes the current state of the power grid operation of the agent. It covers the constraint-related parameters in the optimization model and is represented in vector form. The state space of each agent is defined differently based on its role, and its core includes the following parameters: Common parameters: Current status of all circuit breakers in this zone, current status of all disconnect switches (0-1 variables), energized status of each busbar, current load rate of each line, current load rate of each main transformer, and specific information on maintenance work; Intelligent parameters of power substation: rated capacity of each main transformer, rated load of each outgoing line and tie line, total load of each bus, availability of external power supply channels, and predicted potential load loss under the N-1 fault scenario. Individual parameters of the intelligent agent of the load substation: total load of each section of the busbar of this substation, power supply reliability of each incoming line corresponding to the power supply substation, current status of the backup automatic transfer device, and load distribution of each incoming line; The action space definition includes the following: The action space describes the topology adjustment operations that the agent can perform. The actions must comply with the power grid switching operation procedures. A discrete action space is used, and each action corresponds to a specific operation instruction. The core actions are as follows: The operation space of the power substation intelligent body includes: opening or closing of bus tie switch, opening or closing of main transformer switch, opening or closing of outgoing line switch, opening or closing of tie line switch, opening or closing of corresponding isolating switch, and switching between local power supply and external power supply mode. The action space of the intelligent agent in the load substation includes: opening and closing operations of incoming line switches, opening and closing operations of bus section switches, and coordinated actions with the intelligent agent in the power supply substation to complete the switching of power supply paths. Action constraints: Each agent can only execute one action at a time, and the execution of the action must satisfy the topology constraints in step one. Actions that cause grid circulating current or bus voltage loss are prohibited.

6. The method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework according to claim 5, characterized in that, Step S2 also includes the following: Step S23: Design the reward function, including a composite reward mechanism of immediate reward and global reward, to guide the agent to learn the optimal topology adjustment strategy. The reward value is calculated as follows: Total reward value R = α × immediate reward R1 + β × global reward R2; Where α and β are weighting coefficients, and α+β=1; Instant reward R1: Evaluates the local effect of a single action performed by an agent, with the core associated risk objective of switching operations. The calculation formula is as follows: R1 = K1 - K2 × (Number of circuit breaker switching operations × β_U + Number of disconnector switching operations × β_V); Where K1 is the basic reward constant, K2 is the penalty coefficient, β_U and β_V are the safety risk coefficients for switching operations of circuit breakers and disconnectors defined in step one; if the violation of the action leads to the destruction of topological constraints, then R1 = -K3; where K3 is the violation penalty constant, and K3 > K1; Global Reward R2: Evaluates the global effect of coordinated actions by all agents, with a core associated risk objective of N-1. This reward is calculated by the centralized training module and then distributed to each agent. The calculation formula is as follows: R2 = K4 - K5 × (General user load loss × λ1 + Important user load loss × λ2); Where K4 is the global basic reward constant, K5 is the load loss penalty coefficient, λ1 and λ2 are the load loss coefficients for general users and important users, respectively, and λ2>λ1; Step S24: Training algorithm selection and model training, including the following: The distributed deep reinforcement learning algorithm is selected to adapt to the centralized training and distributed execution framework, which balances the fitting ability of the continuous state space and the execution efficiency of the discrete action space. The training process combines actual operation data of the high-voltage distribution network and simulated maintenance scenarios to ensure the generalization ability of the model. Step S241: Data initialization: Collect historical operation data and maintenance case data of high-voltage distribution network, and build a training dataset that covers the power grid operation status under different maintenance scenarios, different load levels, and different fault scenarios; Step S242: Model initialization: Initialize the global model parameters of the centralized training module and the local model parameters of each distributed agent, set the training hyperparameters, and embed topology optimization constraints as hard constraints for model training. Step S243: Iterative Training: Each distributed agent collects the current power grid state, selects and executes actions based on the current model parameters, collects the power grid state and reward value after the action is executed, and stores the experience in the experience replay pool; the centralized training module randomly samples batch experiences from the experience replay pool of each agent, updates the global model parameters based on the total reward value, and distributes the updated parameters to each agent to complete one iteration; Step S244: Constraint verification: After each iteration, verify whether the power grid topology corresponding to the topology adjustment action output by the agent satisfies all constraints. If not, adjust the penalty coefficient of the reward function and re-iterate the training. Step S245: Convergence judgment: When the number of iterations reaches the set value, and the fluctuation range of the total reward value of multiple consecutive iterations is ≤5%, and the weighted comprehensive objective function value corresponding to the topology adjustment strategy output by the model reaches the minimum value, training is stopped and the trained distributed reinforcement learning model is saved.

7. The method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework according to claim 1, characterized in that, Step S3 includes the following: Step S31: Import the maintenance scenario and set the initial state, including the following: The import of maintenance scenarios includes: obtaining detailed parameters of specific maintenance work and importing them into the centralized training module of the distributed reinforcement learning framework, including maintenance equipment, maintenance scope, maintenance duration, and load prediction data during maintenance. Initial state setting: The normal operating topology of the power grid before maintenance is taken as the initial state. The power grid operating parameters of each partition are collected at this time and assigned to the corresponding distributed agents as the initial input for agent decision-making. At the same time, according to the maintenance requirements, the constraint threshold during the maintenance period is set and the weight coefficient of the reward function is adjusted. For a specific maintenance task, the topology adjustment decision of each agent is obtained from the trained distributed reinforcement learning model and the original operating topology of the power grid. The optimal operating topology of the high-voltage distribution network as a whole is obtained by combining the results and outputting the performance evaluation results. Step S32: Distributed agent decision execution, including the following: Local decision output: Based on the initial state, maintenance scenario parameters and trained model parameters, each distributed agent autonomously selects the optimal topology adjustment action and outputs the local topology adjustment decision for its partition, including the opening / closing operation instructions of switches and disconnect switches, and the power supply mode switching instructions, and uploads the local decision to the centralized training module; Global Decision Collaboration: The centralized training module receives local topology adjustment decisions from all agents and, combined with the topology optimization model from step one, performs collaborative verification of the global topology to determine whether the overall topology after combining local decisions satisfies all constraints and whether the weighted comprehensive objective is minimized. If the global topology satisfies the constraints and the objective function value is optimal, it is directly determined as the final topology adjustment decision. If the global topology does not meet the constraints, or the objective function value does not reach the optimal, the centralized training module issues an adjustment instruction to the corresponding agent. The agent re-outputs the local decision based on the adjustment instruction until the optimal global topology that meets all constraints is formed. Feasibility verification of decisions: The final global topology adjustment decision is compared with the high-voltage distribution network switching operation procedures and major maintenance safety specifications to verify the feasibility of the decision. If there are infeasible operations, the decision of the corresponding intelligent agent is fine-tuned to ensure that the decision can be directly implemented.

8. The method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework according to claim 7, characterized in that, Step S3 also includes the following: Step S33: Optimal running topology output, including the following: After the final global topology adjustment decision is verified, the optimal operating topology of the high-voltage distribution network during this major overhaul is output. The output includes: Global topology diagram: Marks the location of all power substations and load substations, the final status of each line, switch, and disconnector, the power supply mode of each bus, and the direction of the power supply path; Detailed operation instructions: List all operation instructions to be executed in the order of switching operations, including the operating equipment, operation type, operation time node, and clarify the division of labor of each intelligent agent.

9. A system for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework, comprising an electronic device, wherein the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework as described in any one of claims 1 to 8.

10. A system for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework, comprising a computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a method for constructing a high-voltage distribution network operation topology optimization model and its distributed reinforcement learning framework as described in any one of claims 1 to 8.