A Substation Operation and Maintenance Decision-making Method Based on Multi-Agent Reinforcement Learning

By adopting parallel PER two-layer deep Q network and partitioning method in the substation system, combining community detection and semi-Markov decision-making processes, the problem of system-level random failures in large-scale substation systems is solved, and efficient operation and maintenance decision-making and cost reduction are achieved.

CN118644225BActive Publication Date: 2025-05-27NANJING QIZHENG INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410679783.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-29
Publication Date
2025-05-27
Estimated Expiration
2044-05-29

AI Technical Summary

Technical Problem

When dealing with large-scale substation systems, existing multi-agent reinforcement learning methods are difficult to effectively deal with system-level random failures, and the model training time is long, making it difficult to optimize the hyperparameters.

Method used

The parallel PER two-layer deep Q network is adopted, combining partitioning and governance methods and community detection to build a semi-Markov decision-making process, optimize the substation operation and maintenance strategy, reduce operation and maintenance costs, and improve system stability and operation efficiency.

Benefits of technology

Effectively dealing with random system-level failures improves the operation and maintenance efficiency and decision-making accuracy of the substation system, and reduces the operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118644225B_ABST
    Figure CN118644225B_ABST
Patent Text Reader

Abstract

The present invention provides a substation system operation and maintenance decision-making method based on multi-agent reinforcement learning. This method belongs to the field of intelligent decision-making technology, and includes establishing a substation system model containing a semi-Markov decision process, and accordingly establishing a parallel PER double-layer deep Q-network algorithm. Further, the system is segmented by using the Newman-fast community detection algorithm, and finally, the operation and maintenance decision-making of the substation system is realized by using a multi-agent divide-and-conquer method. The method provided by the present invention optimizes the operation and maintenance strategy of the substation, improves the stability and efficiency of the system and reduces costs, and effectively improves the intelligent decision-making level based on the substation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a substation operation and maintenance decision-making method based on improved multi-agent reinforcement learning, and belongs to the field of intelligent decision-making technology. Background Art

[0002] Substations are of crucial significance in the power system, and their existence plays an important role in aspects such as power transportation, distribution, and supply. The substation equipment system will age as the service life increases, thereby increasing the risk of failure or degradation. Component failures caused by degradation may lead to system-level failure events, such as power outages, service interruptions, network loss, or reduced system availability and maintainability. In particular, its failures or degradations often occur randomly, which also increases the difficulty of substation operation and maintenance work. For the operation and maintenance strategy planning of such a system composed of components prone to uncertain failures, classical system reliability theory often introduces a Markov model to handle the characteristics of components and systems. However, in the classical Markov model, the sojourn time follows a geometric distribution, and physically, it means that the probability of system components failing remains unchanged, and a constant probability of failure occurs at any time. This characteristic is different from the characteristics of the actual physical system. In reality, component failures in physical systems are often caused by component aging. In other words, the probability of a component failing will change over time.

[0003] On the other hand, when applying the Markov decision process, for large-scale systems, the state and action spaces will grow exponentially with the increase in the number of components, and it is easy to fall into the curse of dimensionality. Therefore, researchers have introduced deep reinforcement learning algorithms to overcome these difficulties. In addition, multi-agent reinforcement learning is considered an effective means to improve the scalability of deep reinforcement learning because it introduces a multi-agent learning strategy to achieve the minimum cost. However, although the multi-agent reinforcement learning algorithm has successfully reduced the computational complexity, even in a relaxed environment, model training or hyperparameter tuning still requires a relatively long time. Therefore, in existing multi-agent reinforcement learning methods, the system topologies processed are simple, and the number of components is also limited.

[0004] As disclosed in Chinese Patent No. CN117252252A, a multi-agent reinforcement learning intelligent decision-making method is disclosed, including: determining the state vectors of multiple agents in the target problem at the current time step; inputting the state vectors of adjacent agents into the graph attention network included in the algorithm model of the target agent to obtain corresponding influence weights, and performing weighted average processing on the state vectors of adjacent agents based on the influence weights to obtain corresponding mean field vectors; inputting the state vector of the target agent and the mean field vector into the actor network included in the algorithm model of the target agent to obtain the processing decision corresponding to the target agent, so as to control the target agent to execute corresponding actions according to the processing decision at the current time step. The method provided by the present invention can be applied to large-scale agent intelligent decision-making, greatly improving the efficiency and accuracy of multi-agent reinforcement learning intelligent decision-making, and effectively improving the decision-making level of agents.

[0005] It considers the problems of many variables and large dimensions when facing large-scale systems, and does not take into account the possible random events in system components.

[0006] As a key node in the power system, the substation is an important part of civil infrastructure and has a profound impact on modern economic society. Therefore, there is an urgent need to propose a substation operation and maintenance method that can take into account system-level random failures, so that it can have good performance in the environment of large-scale system variables and large-scale system parameters. Summary of the Invention

[0007] The present invention aims to provide a substation operation and maintenance decision-making method based on multi-agent reinforcement learning. By using a parallel PER double-layer deep Q network and combining the divide-and-conquer method and community detection, this method can efficiently handle the operation and maintenance problems of the substation system under system-level random failures.

[0008] To achieve the above object, the present invention is implemented by the following technical solutions.

[0009] On the one hand, the present invention provides a substation system operation and maintenance decision-making method based on multi-agent reinforcement learning, which is characterized by including:

[0010] S1: Using sensors or detectors deployed on substation system equipment or devices to obtain substation system state information;

[0011] S2: According to the target substation state information, combined with the system-level fault and operation and maintenance model, construct a substation system fault and operation and maintenance model;

[0012] S3: According to the substation system fault and operation and maintenance model, combined with multi-agent reinforcement learning, use the divide-and-conquer method to train the parallel PER double-layer deep Q network to obtain a decision set for the substation system fault and operation and maintenance model;

[0013] Among them, the substation system fault and operation and maintenance model introduces a semi - Markov decision process to describe the fault or degradation process of the substation system equipment or devices; the divide - and - conquer method refers to when training the parallel PER double - layer deep Q - network, using the Newman - fast algorithm to perform community detection and division on the substation system fault and operation and maintenance model, so as to divide the substation system fault and operation and maintenance model into m subsystems, and through the multi - agent reinforcement learning method, configure an agent j for each of the subsystems to respectively generate the optimal decision for each of the subsystems

[0014] Furthermore, step S2 further includes the following steps:

[0015] S2.1: Build the substation system model;

[0016] S2.2: Optimize the system - level sequential maintenance;

[0017] S2.3: Introduce the PER double - layer deep Q - network;

[0018] S2.4: Decompose the total cost function;

[0019] The process of building the substation system model in step S2.1 further includes the following steps:

[0020] S2.1.1: Introduce a finite - time semi - Markov decision process to describe the fault, aging or failure of the substation system equipment; once all the substation system state information is determined by being monitored by the sensors or detectors, the QoS loss of the substation system can be calculated by the maximum - flow algorithm, which is the difference between the current maximum - flow capacity and the maximum - flow capacity in the initial system state; S2.1.2: Based on the system state information, the agent selects one of the two operation and maintenance operations of "no action (N)" and "restore to a brand - new state (R)" for the component at time t to implement; determine the system - level operation from the system state, thereby defining the system - level state transition probability;

[0021] S2.1.3: Introduce the weighted - summation formula method to transform the multi - objective function describing the substation system operation and maintenance cost into a single - objective function;

[0022] ​​When optimizing the system-level sequential maintenance, the agent determines the action decision according to the system state, so as to minimize the Q value, which is used to characterize the total discounted cost of the substation system in its life cycle; the Q value is updated by continuously selecting the optimal decision that minimizes the discounted cost;

[0023] Further, step S3 further includes the following steps:

[0024] S3.1: Introduce the Newman-fast algorithm to decompose the substation system into m subsystems, and assign an agent to each subsystem to learn the operation and maintenance strategy of the subsystem;

[0025] S3.2: Introduce a predefined function, decompose the total cost function into the decentralized cost functions of the subsystems through multi-agent credit assignment, and allocate the total cost to the subsystems according to the predefined function;

[0026] S3.3: Introduce a parallel processing method, divide the processing units into five groups, and the hyperparameters in each group take different values and each uses - The greedy algorithm explores the optimal decentralized strategy simultaneously;

[0027] The Newman-fast algorithm sets each vertex in the system topology graph as an independent cluster, and selects two clusters with the largest modularity for merging in each iteration until the entire network is merged into one cluster; the entire merging process obtains a bottom-up tree graph, and the partition with the largest modularity is selected from all the hierarchical partitions of the tree graph as the result of the community detection;

[0028] When tuning the hyperparameters of the parallel processing, after one training cycle, the superiority of the decision based on the current hyperparameter value is judged by the expected life cycle cost; before the start of the next cycle, the online network parameters and the target network parameters of all the processing units will be synchronized with the parameters with the optimal performance in the previous cycle. Through synchronization, the strategies that are not explored or rarely explored due to being close to the optimal strategy will be propagated to other processing units.

[0029] Further, the PER double-layer deep Q network is a model-free reinforcement learning algorithm, which introduces two parameterized deep Q networks to approximate the Q value. The two parameterized deep Q networks are an online network and a target network respectively; it also includes introducing the PER method and evaluating the importance of each experience through the temporal difference error.

[0030] Further, the online network and the target network are respectively used for action selection and Q value evaluation; the parameters of the target network Every The step size will be updated to the parameters of the online network ; The PER method is used to solve the correlation and data utilization; by combining the calculation of the loss function value and updating the online network parameters by gradient descent.

[0031] Preferably, when using the PER double-layer deep Q-network algorithm for Q-value calculation, a community detection algorithm is introduced, multiple agents are assigned to the decentralized cost functions decomposed from the total cost function, and the agents synchronously explore the strategies for minimizing the decentralized costs under different hyperparameter values in multiple processing units.

[0032] Compared with the prior art, the technical method adopted by the present invention introduces a semi-Markov decision process, establishes a system model that is more in line with the actual substation industrial scenario, and accordingly constructs a parallel PER double-layer deep Q-network algorithm. Further, the Newman-fast community detection algorithm is used, combined with the divide-and-conquer method based on multiple agents, to optimize the operation and maintenance strategy of the substation, reduce the operation and maintenance cost, and improve the stability and operation efficiency of the substation system. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 The figure shows a schematic diagram of the substation operation and maintenance management method of improved multi-agent reinforcement learning;

[0034] Figure 2 The figure shows a schematic diagram of the steps for determining the system-level fault and operation and maintenance model;

[0035] Figure 3 The figure shows a schematic diagram of the steps for building the substation system model;

[0036] Figure 4 The figure shows a schematic diagram of the improved parallel multi-agent double-layer deep Q-network algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] It should be noted that:

[0038] The technical solution of the present invention will be described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.

[0039] The term "and / or" only describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " generally represents an "or" relationship between the associated objects before and after.

[0040] Embodiment 1

[0041] Currently, traditional substation system operation and maintenance management methods include deep network learning, DQN, DRL, etc. These methods have problems such as high computational cost, low accuracy, and low serial processing efficiency. In addition, typical multi-agent reinforcement learning also has problems such as a single topological structure, slow model training speed, and long hyperparameter adjustment time.

[0042] This embodiment introduces a substation operation and maintenance management method based on improved multi-agent reinforcement learning. Please refer to Figure 1 , and the specific steps are as follows:

[0043] S1: Monitor the operation of substation equipment through sensors, collect and sample the voltage information V of the equipment with faults or degradation, and provide scenarios and data support for the system-level fault model. It should be noted that the equipment information collected according to this method can be various monitorable information such as voltage, current, or magnetic flux, and in this embodiment, the equipment voltage value V is selected as a reference for consideration.

[0044] S2: Determine the system-level fault and operation and maintenance model. As Figure 2 shown, it is divided into the following steps:

[0045] S2.1: Build the substation system model. As Figure 3 shown, it is divided into the following steps:

[0046] S2.1.1: Describe the substation system. First, considering that the faults, aging, or failures of the actual substation system equipment are highly random, a finite-time semi-Markov decision process (S, A, T, G, ) is introduced. Different from the classical Markov process, the sojourn time of the semi-Markov process changes with time, so its state transition probability no longer follows a geometric distribution and is more practical. Consider a multi-modal flow system model with n components, and assume that the initial state of each component in the system is brand new and complete, and at the same time, the observation of the system state is accurate. The system degradation state is a vector of component states, where is the state of component at time t, defined as the flow rate of the i-th component. Since this embodiment selects the equipment voltage V as the reference equipment information, the system state S actually represents the voltage value V of the equipment at this time. The total flow between two predetermined terminals is a performance index for quantifying the system QoS. Once the flow rates of all components are determined, that is, is observed, the loss of the system QoS It can be calculated by the maximum flow algorithm, which is the difference between the current maximum flow capacity and the maximum flow capacity under the initial system state. According to the topology of the system, the flow losses caused by the aging, failure or malfunction of each component have different impacts on the system QoS.

[0047] S2.1.2: Describe the maintenance and degradation processes. Based on the state information of the system, the agent selects one of the two operation and maintenance operations, "No action (N)" and "Restore to new state (R)", for the component at time t. Select one to implement from the two operation and maintenance operations of "No action (N)" and "Restore to new state (R)". The system-level operation is defined as the set vector of all component actions. When operation R is executed, the state of the target component is deterministically changed to a new state; otherwise, it will continue to deteriorate with a time-varying probability, called the state transition probability where is the sojourn time. In this way, the stochastic degradation process of a single component is modeled as an independent discrete-time semi-Markov process. Once the system-level action formed by all components is determined by the system state , the system-level state transition probability can be defined by the state transition probabilities of all components .

[0048] S2.1.3: Describe the operation and maintenance costs. During the system operation and maintenance process, multiple variables need to be considered simultaneously in each link, such as maintenance costs, system QoS, and component failures. Due to the trade-off relationship between system costs and system risks, there is usually no single solution that can optimize all variables simultaneously. Therefore, it is necessary to appropriately select a compromise solution among these conflicting optimization goals. To solve this problem, the multi-objective function is transformed into a single-objective function through a scalarization method, that is, the weighted summation formula method is introduced for processing. The total cost at time t is defined as the sum of the costs of multiple objective functions:

[0049] ,

[0050] where is the maintenance cost of component i; is the cost brought by the downtime of component i; is the system damage cost caused by the system QoS loss. The system QoS loss and the penalty factor jointly determine the system damage cost :

[0051] ,

[0052] where represents the system state a function that represents the maximum traffic capacity between two predetermined terminals under a system state; represents the initial state vector of the system, i.e., each component is in a brand-new and complete state.

[0053] S2.2: Optimize the system-level sequential maintenance. In system operation and maintenance optimization, the goal of the agent is to determine a decision that maps from the system state to the action to minimize the Q-value , where is defined as the total discounted cost of the system over its entire lifetime :

[0054] ,

[0055] where is the discount factor that converts future costs to present value. Just as component-level optimality does not guarantee system-level optimality, the optimal decision at each time step does not guarantee long-term optimality in a finite-time environment of the system. Therefore, when and only when an operation and maintenance decision has an expected cost less than or equal to another operation and maintenance decision for all states, is defined as a better operation and maintenance decision, which means that the optimal operation and maintenance decision always makes the Q-value reach the minimum. Its physical meaning is that by obtaining the voltage information V of the target substation equipment (note that V here is the system state S) and making the corresponding decision A, the operation and maintenance cost P of the substation system is minimized. In the Q-value iteration, considering that the goal of the optimal decision is to make the total discounted cost reach the minimum Q-value , so the Q-value is updated by continuously selecting the action with the minimum value:

[0056] .

[0057] After initializing all to zero, these values are updated and iterated according to the following Bellman equation:

[0058] ,

[0059] where is at state and action The k-th iteration estimate. It should be noted that this Q-value iteration algorithm works well in simple environments but is not suitable for environments with large state and action spaces. In particular, it is almost impossible to evaluate the Q-values of every pair of combinable state actions in complex environments. The substation system environment is complex, with a large variety of equipment in large quantities, which is mathematically reflected in a high-dimensional state space. In this case, it is very difficult to calculate the Q-value.

[0060] S2.3: Introduce a double deep Q-network with PER. To overcome the computational limitations, an off-policy reinforcement learning algorithm called the Q-learning algorithm is applied. In the Q-learning algorithm, the following equation is used to iteratively update :

[0061] ,

[0062] where is the learning rate; is the target value. Generally, at this time, a parameterized deep Q-network is introduced in the deep reinforcement learning framework to approximate the Q-value. However, since the operator is used when calculating the target value , it will cause the deep Q-network to tend to underestimate the Q-value in a stochastic environment. By using two different deep Q-networks (online network and target network) for action selection and Q-value evaluation respectively, this problem is avoided. Therefore, the target value is rewritten as:

[0063]

[0064] where is the Q-value estimated by the online network parameterized by ; is the Q-value estimated by the target network parameterized by . The parameter is periodically updated to the parameter every steps. During the process of the deep reinforcement learning network, problems such as correlation issues and low data utilization efficiency may occur. Therefore, considering introducing the ER technique, the agent stores the experience as a tuple in the replay buffer D and updates the online network parameter based on a set of uniformly sampled tuples from the replay buffer, where is the decomposition of the cost . Due to the reduced correlation, the uniformly sampled samples significantly reduce the variance of the update, thus suppressing the oscillation or divergence of the parameters during the training process. By combining, the double-layer deep Q-network calculates the loss function and updates the online network parameters through gradient descent minimize the following as:

[0065] ,

[0066] ,

[0067] where represents a uniform distribution over the replay buffer; is a set of tuples sampled from . It should be noted that this classical ER method does not take into account the importance of experience and simply performs random sampling, which may lead to some less important experiences being frequently selected, thus affecting the learning effect. Therefore, priority experience replay, i.e., the PER method, is introduced. Its advantage lies in determining the selection probability according to the priority of experience. The agent evaluates the importance of each experience through the TD error (i.e., the temporal difference error), which is the difference between the expected Q-values before and after the experience. Therefore, the alternative sampling density of a set of experience tuples with a uniform distribution is expressed as:

[0068] ,

[0069] where represents the TD error. Further, the loss function can be calculated more effectively as follows:

[0070] ,

[0071] where represents the probability mass function of the discrete uniform distribution; is the likelihood ratio to compensate for the bias caused by introducing the alternative sampling density . Further, since the online network and the target network have not been sufficiently trained in the initial stage of learning, this stage is highly oscillatory. Therefore, even if there is a slight bias, stability should be the first consideration in this stage, and the bias can be corrected later. For this purpose, in the above formula should be replaced by , which is defined as:

[0072] ,

[0073] where is the number of experiences in the experience buffer D; is a hyperparameter that controls the degree of compensating bias. When ​Gradually approaching 1 from a low value (the typical range is 0.4 to 0.6), the deviation is completely compensated. When taking the intermediate value, the double-layer deep Q-network and PER will play a substantial synergistic role to make up for the limitations of the current Q-learning algorithm.

[0074] S2.4: Cost decomposition of multi-agent reinforcement learning. With the exponential increase of system states and action decision instructions, the state space and action space will fall into the curse of dimensionality. Therefore, it is considered to divide the original problem into multiple sub-problems to address the challenge. Using the divide-and-conquer strategy of multi-agent reinforcement learning, multiple agents are deployed in each divided action space and state space, and the agents independently or jointly achieve cost minimization through learning strategies. Without a cost function jointly contributed by agents, a given environment can be simplified into several smaller independent environments, that is, the optimal decision set in the sub-problem produces the same solution as the global optimal strategy. Then the size of the action space is reduced from to , where m is the number of agents, and is redefined as the number of available actions of the j-th agent. In many complex environments, the cost function takes various combinations of states and actions as inputs, which leads to the multi-agent credit assignment problem, meaning that the contribution of a single agent to the cost function must be accurately inferred. For this reason, the centralized cost is decomposed into the sum of decentralized costs through a value decomposition network, and the problem can be divided into multiple independent environments.

[0075] S3:: Improving the parallel multi-agent double-layer deep Q-network. The multi-agent reinforcement learning method is combined with the double-layer deep Q-network with PER to improve scalability with minimal accuracy loss, allowing the semi-Markov decision process framework to be applied to complex environments. However, with the increase in the number of components in the system, the accuracy loss gradually accumulates and will eventually lead to the failure of optimal policy search. Therefore, a new improved parallel multi-agent deep Q-network is proposed. This algorithm assigns multiple agents to subsystems based on the Newman-fast community detection method and enables the agents to explore the decentralized cost minimization strategy under different hyperparameter values in multiple processing units. As Figure 4 shown, it is divided into the following steps:

[0076] S3.1: Decompose the system into several subsystems, namely communities composed of different components. There are connection relationships among the devices in the substation, and this connection relationship can be abstracted into a topological graph, that is, the system network, where each independent device is a node in the network, and the connection relationship between devices is an edge in the network. By grouping components that are deeply related functionally or tightly connected into one subsystem, the existing system can be simplified into a system of another subsystem. A classic method is to use the Girvan - Newman algorithm for community division. This algorithm will successively remove the edges with the highest betweenness centrality. The betweenness centrality is the proportion of the number of shortest paths passing through this edge in all shortest paths in the network to the total number of shortest paths. Since the edges with high betweenness centrality are removed, it means that the entire graph of the system is divided into several isolated clusters, and when all edges are removed, the process ends. In the Girvan - Newman algorithm, modularity Stop this process at an appropriate time. Modularity represents the difference between the actual number of edges within the existing clusters in the reconstructed graph and the expected number of edges within the clusters while retaining the degree of each node, and is defined as:

[0077] ,

[0078] where, is the number of isolated clusters; is the number of edges in the topological graph; is the number of edges within cluster y; is the sum of the degrees of nodes in cluster y. Although the Girvan - Newman algorithm can accurately partition the network through modularity, due to its high time complexity ( ), it is only applicable to networks of medium and small scales. For large - scale system networks such as substation systems, the applicability of the Girvan - Newman algorithm significantly decreases. Therefore, the Newman - fast algorithm is introduced. First, each vertex in the system topological graph is set as an independent cluster, and in each iteration, the two clusters that produce the largest are selected for merging until the entire network is merged into one cluster. Initialize the system network:

[0079] ,

[0080] ,

[0081] where, is the ratio of the connecting edges between cluster y and cluster y' to the total number of edges; is the ratio of the number of all edges associated with the points inside cluster y to the total number of edges. Next, successively merge the pairs of clusters with connecting edges in the maximum or minimum direction of and calculate the modularity increment :

[0082] 。

[0083] Repeat this process to merge cluster pairs continuously until the entire system network is merged into one cluster. Further, the entire merging process results in a bottom-up tree graph, where the leaf nodes represent the vertices in the system network, and each layer of partitioning represents a specific partitioning of the system network. Finally, select the partitioning with the largest value as the result of community partitioning. The time complexity of the Newman-fast algorithm is , and for large-scale system networks, compared with the classical Girvan-Newman algorithm, it has higher executability, faster execution speed, and obvious advantages.

[0084] After identifying m subsystems in the target system using the Newman-fast algorithm, assign an agent to each subsystem to learn the operation and maintenance strategies of the subsystem. However, directly using this community detection result to train a parallel multi-agent deep Q-network will have problems: due to the barrel effect, the convergence of the network is controlled by the slowest learning speed among the agents (usually the agent with the largest state space and action space). If the components are concentrated in a specific cluster, it will lead to a situation where the dimension of a certain subsystem in the community detection result is significantly larger than that of other subsystems. At this time, the learning time of the agent will increase exponentially, further resulting in a decrease in the learning efficiency of the parallel multi-agent deep Q-network. Therefore, after the initial community partitioning based on the Newman-fast algorithm, reassign some components in the largest subsystem to other adjacent subsystems. During this process, to prevent the loss of calculation accuracy caused by simplification, the number of edges connecting subsystems should be minimized as much as possible.

[0085] S3.2: Decompose the total cost. After assigning agents to each subsystem through S3.1, the multi-agent system makes effective operation and maintenance decisions by comprehensively considering the states of the components within each subsystem. To enable all subsystems to make decisions independently and effectively, the total cost should be decomposed into the independent costs of each subsystem based on effective multi-agent credit assignment, so a value decomposition function or a predefined function is introduced. In this way, the parallel multi-agent deep Q-network finds the subsystems that are assumed to cause system losses and selectively allocates the total cost to these subsystems according to the predefined function. Specifically, by splitting the total cost and adding the independent costs of the subsystems, for subsystem , the total independent cost is:

[0086] ,

[0087] where is redefined as the action selected by agent j at time t; is the decentralized cost transferred from the total cost to the subsystem ; is a hyperparameter that determines the weight of the decentralized cost. The decentralized cost is predefined as:

[0088] ,

[0089] where is the QoS loss of the subsystem ; is a subset of representing the state vector of the components in the subsystem Since is calculated based on the maximum traffic between two predetermined terminals, will be defined as the total traffic change of the subsystem . Further, the independent Q-value of each agent j is calculated, and more effective and stable Q-learning is achieved by combining step S2.3. Specifically, for , the online network with parameters and the target network with parameters select the system state vector and the current time t as inputs, and then output the Q-values and respectively according to the actions in each vector form. Then, based on the estimated Q-value , the optimal action is transformed into a -dimensional one-hot encoded vector, and it is multiplied by the vector to update the online Q-value. Different from which is updated online at each time step, the target network parameters are updated from to the current time every steps. It should be noted that different from the value decomposition network, the expected total life cycle cost of the system is not equal to the sum of the decentralized costs . Since all agents choose actions that minimize their respective Q-values, the system-level action set

[0090] is redefined as:

[0091] In addition, during the learning process, the performance of the double-layer deep Q-network significantly depends on the selection of the hyperparameter value . Therefore, the hyperparameter value should be appropriately determined according to the specific system environment.

[0092] S3.3: Parallel processing for hyperparameter tuning. To achieve effective hyperparameter tuning, a parallel processing method is introduced. The processing units (GPUs / CPUs) are divided into five groups, with different hyperparameter values in each group. In each group, the agent uses the -greedy algorithm to explore the optimal decentralized policy simultaneously under the hyperparameters. After sufficient training of the agent, i.e., after one training cycle, the superiority of the decision based on the current hyperparameter values is judged by the expected life cycle cost . Before starting the next cycle, the online network parameters and the target network parameters in all processing units are synchronized with the parameters of the best performance in the previous cycle, thus improving the main policy. Through synchronization, some policies that have not been explored or have been rarely explored due to being close to the optimal policy are propagated to other processing units, resulting in a significant performance improvement. Thus, the hyperparameters will be adjusted in the following way:

[0093]

[0094] where is the step size with exponential decay, and . In addition, by comparing the results of parallel processing, when the expected life cycle cost has converged sufficiently, the algorithm will terminate early.

[0095] In summary of the above embodiments, the present invention uses a semi-Markov decision process to determine a risk-informed operation and maintenance strategy for a substation system network to evaluate the risk of the system. Since the number of state and action spaces grows exponentially, it is difficult to find a solution to the semi-Markov decision process. The present invention proposes a multi-agent deep reinforcement learning framework to overcome the dimensionality problem. The method adopts a divide-and-conquer strategy, and through the Newman-fast algorithm applicable to large-scale scenarios similar to the substation system environment for community detection, multiple subsystems are identified, and each agent learns to implement the operation and maintenance strategy of the corresponding subsystem. The agent establishes a strategy to minimize the decentralized cost of the subsystem, including the decentralized cost. This learning process is carried out simultaneously in several parallel processing units, and the trained strategy is periodically synchronized with the best strategy, thus improving the main policy.

[0096] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0097] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general - purpose computer, a special - purpose computer, an embedded processor, or other programmable data - processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data - processing devices produce a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0098] These computer program instructions can also be stored in a computer - readable memory that can direct a computer or other programmable data - processing device to work in a specific manner, such that the instructions stored in the computer - readable memory produce a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0099] These computer program instructions can also be loaded onto a computer or other programmable data - processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer - implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0100] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above - mentioned specific embodiments. The above - mentioned specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention's claims. These all fall within the protection scope of the present invention.

Claims

1. A substation system operation and maintenance decision-making method based on multi-agent reinforcement learning, characterized in that: include: S1: Obtain substation system status information using sensors or detectors deployed on substation system equipment or devices; S2: constructing a substation system fault and operation and maintenance model based on the substation system status information and in combination with a system-level fault and operation and maintenance model; S3: According to the substation system fault and operation and maintenance model, combined with multi-agent reinforcement learning, the parallel PER double-layer deep Q network is trained using a divide-and-conquer method to obtain a decision set for the substation system fault and operation and maintenance model; The substation system failure and operation and maintenance model introduces a semi-Markov decision process to describe the failure or degradation process of the substation system equipment or devices; the divide-and-conquer method refers to using the Newman-fast algorithm to perform community detection and division on the substation system failure and operation and maintenance model when training the parallel PER double-layer deep Q network, so as to divide the substation system failure and operation and maintenance model into m subsystems, and through the multi-agent reinforcement learning, for each of the subsystems Configure agent j to generate a response to each of the subsystems The optimal decision; The step S2 further comprises the following steps: S2.1: Building the substation system model; S2.2: Optimize system-level sequential maintenance; S2.3: Introduce the PER double-layer deep Q network; S2.4: Decompose the total cost function; The process of building the substation system model in step S2.1 further includes the following steps: S2.1.1: A finite-time semi-Markov decision process is introduced to describe the failure, aging or failure of the substation system equipment; once all the substation system status information is determined by monitoring by the sensors or detectors, the QoS loss of the substation system is calculated by the maximum flow algorithm, which is the difference between the current maximum flow capacity and the maximum flow capacity under the initial system state; S2.1.2: Based on the system status information, the agent is a component at time t From the two operation and maintenance operations of "no action" and "restore to a new state" Select one to implement; determine the system level operation according to the system state, thereby defining the system level state transition probability; S2.1.3: Introducing a weighted sum formula method to transform the multi-objective function describing the operation and maintenance cost of the substation system into a single objective function; The step S3 further comprises the following steps: S3.1: Introduce the Newman-fast algorithm to decompose the substation system into m subsystems, and assign one agent to each subsystem to learn the operation and maintenance strategy of the subsystem; S3.2: introducing a predefined function, decomposing the total cost function into decentralized cost functions of each subsystem through multi-agent credit allocation, and allocating the total cost to the subsystems according to the predefined function; S3.3: Introduce a parallel processing method and divide the processing units into five groups. The hyperparameter values ​​in each group are different and are used separately. -The greedy algorithm simultaneously explores the optimal dispersion strategy; The Newman-fast algorithm sets each vertex in the system topology graph as an independent cluster, and selects two clusters with the largest modularity for merging in each iteration until the entire network is merged into one cluster; The whole merging process obtains a bottom-up tree diagram, and the partition with the largest modularity is selected from all hierarchical partitions of the tree diagram as the result of the community detection; The PER double-layer deep Q network is a model-free reinforcement learning algorithm that introduces two parameterized deep Q networks to approximate the Q value, and the two parameterized deep Q networks are an online network and a target network respectively; when the parallel processing hyperparameters are tuned, after one cycle of training, the superiority of the decision based on the current hyperparameter value is judged by the expected life cycle cost; before the start of the next cycle, the online network parameters and target network parameters of all the processing units will be synchronized with the parameters of the optimal performance in the previous cycle. Through synchronization, strategies that have not been explored or rarely explored due to being close to the optimal decentralized strategy will be propagated to other processing units.

2. The substation system operation and maintenance decision-making method based on multi-agent reinforcement learning according to claim 1 is characterized in that: When optimizing the system-level sequential maintenance, the intelligent agent determines the action decision according to the system status to minimize the Q value, which is used to characterize the total discounted cost of the substation system during its life cycle; the Q value is updated by continuously selecting the optimal decision that minimizes the discounted cost.

3. According to claim 2, the substation system operation and maintenance decision-making method based on multi-agent reinforcement learning is characterized in that: It also includes the introduction of the PER method and the evaluation of the importance of each experience through time difference errors.

4. The substation system operation and maintenance decision-making method based on multi-agent reinforcement learning according to claim 3 is characterized in that: The online network and the target network are used for action selection and Q value evaluation respectively; The parameters of the target network Every The step size is then updated to the parameters of the online network ; The PER method is used to solve the problems of relevance and data utilization; The online network parameters are updated by combining the calculated loss function values ​​and gradient descent.

5. The substation system operation and maintenance decision-making method based on multi-agent reinforcement learning according to claim 4 is characterized in that: When using the PER double-layer deep Q network algorithm to calculate the Q value, a community detection algorithm is introduced to assign multiple agents to the decentralized cost function decomposed from the total cost function, and enable the agents to synchronously explore the decentralized cost minimization strategy under different hyperparameter values ​​in multiple processing units.

Citation Information

Patent Citations

  • Multi-agent reinforcement learning intelligent decision-making method and device

    CN117252252A

  • Multi-agent power generation optimal scheduling method based on reinforcement learning

    CN110728406A

  • Intelligent maintenance decision-making method and system for underwater production control system

    CN118096121A