An intelligent operation and maintenance system for cluster systems based on reinforcement learning

By adopting an intelligent operation and maintenance system based on reinforcement learning in the cluster system, combined with a variety of reinforcement learning algorithms, the maintenance timing and strategies are optimized, and the problems of unsatisfactory data utilization, complex environment and unscientific maintenance cycle in the cluster system operation and maintenance are solved, and efficient and economical operation and maintenance management are achieved.

CN119090492BActive Publication Date: 2025-05-13HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411595396.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-05-13
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

The existing cluster system operation and maintenance model has problems such as unsatisfactory data utilization, complex external environment information, and unscientific maintenance cycle, resulting in high maintenance costs, wasted resources and long failure time.

Method used

The intelligent operation and maintenance system of cluster system based on reinforcement learning is adopted, including cluster modeling module, multi-dimensional situational awareness module, intelligent decision planning module and performance evaluation module, combined with reinforcement learning algorithms such as deep Q network (DQN), Actor-Critic Monte Carlo Tree Search (AC-MCTS) and distributed near-end strategy optimization (DPPO), to optimize maintenance timing and strategies.

Benefits of technology

It realizes independent operation and dynamic management of cluster systems in complex environments, reduces failure rate and operation and maintenance costs, quickly restores system functions, and ensures task coordination and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119090492B_ABST
    Figure CN119090492B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent operation and maintenance system for cluster systems based on reinforcement learning. The system includes a cluster modeling module, a multi-dimensional situational awareness module, an intelligent decision-making planning module, and an efficiency evaluation module, and combines a variety of reinforcement learning algorithms to achieve autonomous operation and maintenance and dynamic management of cluster systems in complex environments. The system optimizes maintenance timing and strategies by introducing reinforcement learning algorithms such as deep Q network (DQN), Actor-Critic Monte Carlo Tree Search (AC-MCTS), and distributed proximal policy optimization (DPPO), thereby reducing the failure rate and operation and maintenance costs of the system; quickly restores system functions, reduces system performance losses caused by sudden failures; and ensures task coordination and system stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent operation and maintenance system for a cluster system based on reinforcement learning. Background Art

[0002] With the increasing expansion of human living space and the deepening vision of comprehensive, coordinated and sustainable development, cluster systems have gradually become a very important representative product form, such as various infrastructure networks, various equipment clusters, and various intelligent manufacturing systems. During the operation of cluster systems, the failure of large-scale equipment in them usually leads to very serious consequences, and even causes casualties. Effective operation and maintenance (O&M) is a necessary means to ensure the continuous safe and stable operation of cluster systems. For independently operated equipment, it is usually necessary to observe the status of the equipment and its line replaceable unit (LRU) and then perform corresponding operation and maintenance work. However, unlike the operation and maintenance of a single device, the operation and maintenance at the cluster system level needs to consider the operating status of large-scale heterogeneous equipment. At the same time, all devices share limited operation and maintenance resources. Therefore, the operation and maintenance at the cluster system level needs to consider large-scale decision-making space and has many optimization solutions. The current operation and maintenance of cluster systems still face problems such as poor data utilization, complex external environment information, and unscientific maintenance cycles. Traditional operation and maintenance modes such as "after-the-fact maintenance" and "regular maintenance" are usually not the best operation and maintenance solutions in terms of maintenance costs, security resources, etc. Therefore, it is inevitable to transform to the current representative cutting-edge operation and maintenance modes such as "predictive maintenance", "post-disaster repair", and "dynamic reconstruction". At the same time, with the development of digital, information, and intelligent technologies, seeking more efficient, scientific, and economical new operation and maintenance solutions in the intelligent era has become a common research goal in the field of cluster system operation and maintenance. At present, there is a lack of open and basic intelligent operation and maintenance solutions for cluster systems, and no common framework for intelligent operation and maintenance has been formed.

[0003] Research on intelligent operation and maintenance technology for cluster systems is crucial to provide solutions for safe and reliable intelligent operation and maintenance solutions. In actual operation, cluster systems exhibit core characteristics such as scale, coordination, and uncertainty. Among them, the scale feature is reflected in the cross-level interaction of large-scale cluster systems, as well as the changing state of time-dynamic cluster systems and space-time dynamic cluster systems; coordination requires cluster systems to have multiple core coordination capabilities, including intelligent collaborative decision-making, robust collaborative control, and distributed collaborative perception; uncertainty is reflected in the randomness of cluster system failure behavior, the variability of external environment and operating conditions, and the unpredictability of component degradation.

[0004] Traditional cluster system operation and maintenance models mostly rely on regular maintenance and post-maintenance. However, these models have the following significant shortcomings: 1. Post-maintenance lag: When a device fails, the traditional post-maintenance model may cause system downtime, increase failure time and maintenance costs, and fail to ensure the normal operation of the system in a timely manner. 2. Regular maintenance resource waste: The preset maintenance cycle may not reflect the actual status of the device, resulting in early maintenance when the device has not failed, wasting resources; and when the device needs maintenance, the best maintenance time may be missed due to not reaching the scheduled cycle.

[0005] To solve these problems, intelligent operation and maintenance technology has gradually become a research hotspot in the field of cluster systems. In recent years, with the rapid development of technologies such as artificial intelligence, big data, and deep learning, intelligent operation and maintenance systems based on these emerging technologies can more effectively realize real-time monitoring and management of cluster systems. By collecting the operating data of equipment in the system in real time and combining intelligent algorithms for fault prediction and decision-making, intelligent operation and maintenance systems can achieve more flexible predictive maintenance and dynamic management.

[0006] For example, the applicant previously addressed the problem of "predictive maintenance" of a three-level time-dynamic cluster system under the condition of system performance degradation in the long-term operation stage. By comprehensively weighing the system degradation state and the imbalanced maintenance benefits, a predictive maintenance plan was generated through a method based on DQN deep reinforcement learning. For details, please refer to a DQN-based cluster system predictive maintenance decision-making method disclosed in the Chinese invention patent application (application number: 202411357728X, application date: 20240927); for the problem of "post-disaster repair" of multiple teams in a two-level time-dynamic cluster system facing large-scale local damage, a corrective maintenance plan was generated by coordinating the two-level decisions of maintenance timing and path planning through a method based on Actor-Critic deep reinforcement learning. For details, please refer to the Chinese invention patent application (application number: 2024112822477, application date: 20240913), which discloses a method for post-disaster restorative maintenance of cluster systems based on the AC-MCTS algorithm; for the "dynamic reconstruction" problem of a three-level spatiotemporal dynamic cluster system with multi-formation characteristics, a two-level dynamic reconstruction of functional reconstruction between nodes within the clusters of the three-level cluster system and functional reconstruction between clusters is analyzed by a method based on DPPO deep reinforcement learning to generate a dynamic reconstruction plan. For details, please refer to the Chinese invention patent application (application number: 2024114796053, application date: 20241023), which discloses a dynamic reconstruction decision method for cluster systems based on DPPO deep reinforcement learning. Summary of the invention

[0007] In order to solve the above technical problems, the purpose of the present invention is to provide a cluster system intelligent operation and maintenance system based on reinforcement learning, which includes a cluster modeling module, a multi-dimensional situational awareness module, an intelligent decision-making planning module and an efficiency evaluation module, and combines a variety of reinforcement learning algorithms to achieve autonomous operation and maintenance and dynamic management of cluster systems in complex environments. The system optimizes maintenance timing and strategy by introducing reinforcement learning algorithms such as deep Q network (DQN), Actor-Critic Monte Carlo Tree Search (AC-MCTS) and Distributed Proximal Policy Optimization (DPPO), thereby reducing the failure rate and operation and maintenance costs of the system; quickly restore system functions, reduce system performance losses caused by sudden failures; and ensure task coordination and system stability.

[0008] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions:

[0009] A cluster system intelligent operation and maintenance system based on reinforcement learning, wherein the cluster system is a power system cluster, a drone cluster system or a large-scale battery cluster system in an industrial energy storage system, and the system comprises:

[0010] 1) Cluster modeling module, which is used to establish a multi-level model of the cluster system, including three levels: cluster, cluster and node. The equipment in the cluster system is modeled respectively. Based on the model, the temporal dynamic characteristics and spatiotemporal dynamic characteristics of the cluster system in different operation stages are analyzed, and a cluster system model with multi-dimensional attributes is constructed.

[0011] 2) Multi-dimensional situation awareness module, which is used to collect and process multi-dimensional situation information of cluster systems in complex environments, including node status, link status, environmental changes and task information, and to conduct real-time monitoring and trend prediction of the operation status of cluster systems by combining temporal dynamic characteristics and spatiotemporal dynamic characteristics;

[0012] 3) Intelligent decision-making and planning module, which makes intelligent decisions on cluster system operation and maintenance issues based on reinforcement learning algorithms. This module includes:

[0013] 3.1) Predictive maintenance submodule, including learning the optimal maintenance strategy and generating predictive maintenance solutions based on the Deep Q-Network (DQN) algorithm for cluster system degradation and imbalanced maintenance benefits;

[0014] 3.2) Post-disaster repair submodule, including the Actor-Critic Monte Carlo Tree Search (AC-MCTS) algorithm, which coordinates the maintenance sequence and multi-team path planning of the cluster system in a large-scale distributed environment to generate emergency repair plans;

[0015] 3.3) Dynamic Reconfiguration Submodule, including the Distributed Proximal Policy Optimization (DPPO) algorithm, which considers task coordination and topological configuration for different levels of topological structures in the cluster system and generates dynamic reconstruction solutions;

[0016] 4) Cluster performance evaluation module, which is used to construct and calculate the performance indicators of the cluster system, including network efficiency, balance, reliability and elasticity, to quantify the performance of the cluster system at different operation and maintenance stages, and provide feedback and optimization direction for intelligent decision-making planning.

[0017] Preferably, the cluster modeling module comprises:

[0018] Modeling and analysis tools for different levels in cluster systems, including but not limited to topological modeling tools for describing the two-level structure of cluster-node, and multi-level modeling tools for describing the three-level structure of cluster-cluster-node;

[0019] The modeling and analysis methods for two major types of systems, namely time dynamic clusters and space-time dynamic clusters, take into account the interactive relationships, synergies and uncertainties between each level.

[0020] Preferably, the cluster modeling module is established as follows:

[0021] 1.1) For the cluster system object S under study, first establish a node set N to describe all nodes n in the cluster S, as shown below:

[0022] N={n1, n2,…, n n ,…}(1.1)

[0023] Where n n Represents the nth node in the cluster;

[0024] Before forming a cluster, when some or all nodes in a node set need to be combined into a cluster in a certain order, it is necessary to establish a cluster set C to describe all clusters c in the cluster at the current moment according to factors such as the specific task requirements in the current scenario, as shown below:

[0025] C={c1, c2,…, c m ...} (1.2)

[0026] Among them, c m represents the mth cluster in the cluster at the current moment;

[0027] 1.2) Construct a one-to-many data structure between clusters based on a hash table. The key of the hash table is still the dynamic number of the cluster, and the stored data element is the node set of the cluster. It is constructed using a nested hash table. The key of the nested hash table is the dynamic number of each node in the cluster, and the stored data element is the unique identifier of the node, that is, the number of the node in the node set N. The three-level cluster model SM is constructed by nesting hash tables. III , describing the mapping between the unique identifier of a node and its dynamic number in the cluster, as follows:

[0028] (1.3)

[0029] where n (i,j) Represents cluster c i The jth node in .

[0030] Preferably, the multi-dimensional situation awareness module comprises:

[0031] A multi-dimensional situation information collection and processing unit based on complex network theory and perception technology is used to collect the operation data of the cluster system in real time and extract multi-dimensional situation features through data fusion and pattern recognition technology;

[0032] Based on machine learning and data-driven situation prediction units, a multi-dimensional situation prediction model is constructed according to the temporal and spatial dynamic characteristics of the cluster system to provide an estimate of the future situation.

[0033] Preferably, the data structure establishment method of the multi-dimensional situation awareness module is as follows:

[0034] 2.1) Represent a cluster as an undirected G graph with N nodes and K edges. A G graph is described by two matrices: the adjacency matrix A and the physical distance matrix L. A is an N × N adjacency matrix. If there is a connected node n i and node n j The edge of the adjacency matrix is ij is 1, otherwise it is 0; the element l of the matrix L ij Represents node n i and n j The geographical distance between nodes n i and node n j There are no edge elements between them, l ij It is also known;

[0035] 2.2) For nodes, after creating an empty G graph, the node set N of the cluster system object S can generate the required data. The node set generates different node arrays for addition according to the normal, degraded, and damaged states of the nodes at the current moment; for edges, it is necessary to define the link set L of the cluster system, and then generate edge data and add it to the created G graph; the construction of the cluster system link set L refers to the data structure of the adjacency matrix in the underlying encoding of NetworkX, and on this basis, adds attribute information including operation and maintenance status, and adds the required attribute information according to the actual scenario. The link set L is as follows:

[0036] (1.4)

[0037] Where s(a ij ) represents the connection node n i and node n j The state of the edge, element l ij Represents node n i and n j The geographical distance between them.

[0038] Preferably, the Deep Q-Network (DQN) algorithm comprises the following steps:

[0039] 3.1.1) Constructing a degradation state model of a cluster system, wherein the cluster system includes a plurality of nodes, and the degradation state of the nodes changes over time during the operation of the system;

[0040] 3.1.2) Design and train a DQN model to approximate the optimal value function Q, provide an estimate of the value function Q, and evaluate the current cluster system maintenance status feature X;

[0041] 3.1.3) Estimate the value function Q based on the DQN model, obtain a maintenance state probability π through the ε-greedy algorithm, and then generate the optimal maintenance action a for the current cluster maintenance state characteristics * ;

[0042] 3.1.4) Execute a series of maintenance actions in the maintenance strategy until the performance of the cluster system is restored to the predetermined threshold;

[0043] 3.1.5) By providing feedback on the executed maintenance actions, the obtained maintenance data is added to the experience replay buffer, and the experience replay buffer is used to further train the DQN model to optimize the next maintenance decision.

[0044] Preferably, the Actor-Critic Monte Carlo Tree Search (AC-MCTS) algorithm includes the following steps:

[0045] 3.2.1) Construct the corrective maintenance process of multiple maintenance teams and generate the decision feature tensor X of the cluster system at any time point t t ;

[0046] 3.2.2) Using Actor-Critic Neural Network to t Evaluate and output the prior parameters, which are used as input to the Monte Carlo tree search algorithm;

[0047] 3.2.3) Monte Carlo tree search algorithm in X t The search operation of the maintenance action is performed under the conditions, including:

[0048] i) Select: Change X t Acts as the root node of the Monte Carlo tree search and selects the maintenance action with the best action value based on the upper confidence interval algorithm;

[0049] ii) Expansion and evaluation: Add leaf nodes in the search tree to the queue, evaluate their strategic value using the ResNet neural network, and update the statistics of related nodes in the search tree;

[0050] iii) Backtracking: Based on the evaluation results, backtrack along the search path and update the number of visits and action value of each branch on the search path;

[0051] iv) Execution: By iterating the above steps, the improved maintenance state transition probability π is obtained and the global optimal maintenance action a is selected t * , transfer the cluster system from time point t to t+1;

[0052] 3.2.4) Repeat steps a) to c) until the corrective maintenance task of the cluster system is completed.

[0053] Preferably, the distributed proximal policy optimization (DPPO) algorithm includes the following steps:

[0054] 3.3.1) Construct a dynamic reconstruction decision framework for cluster systems: Analyze the task coordination and topological configuration characteristics of the three-level spatiotemporal dynamic cluster system, construct a three-level dynamic reconstruction decision framework of cluster-cluster-node, and analyze the topological configuration characteristics between clusters within the cluster and between nodes within the cluster;

[0055] 3.3.2) Multi-dimensional situation feature extraction: Design a deep neural network model, use the ResNet module to process the task area features, business coverage features, node location features and damage area features, apply the LSTM module to extract the temporal features of high-dimensional information, and use the Attention mechanism to focus on the internal and cross-level interactions of the cluster to extract multi-dimensional situation features;

[0056] 3.3.3) DPPO-based reinforcement learning algorithm model: Apply the policy gradient algorithm of the dominant actor-critic and combine it with the experience replay technology for asynchronous update. By maximizing the expected policy reward, the neural network is updated using TD (λ), V-trace and UPGO training to achieve dynamic reconstruction strategy optimization.

[0057] 3.3.4) Design of alliance learning model: Design alliance games and virtual self-learning mechanisms, including three types of agents: master agent, master explorer and alliance explorer, conduct distributed learning training, and optimize dynamic reconstruction strategies;

[0058] 3.3.5) Dynamic reconstruction strategy generation: In the dynamic reconstruction process, the dynamic reconstruction agent is used to generate a reconstruction action set based on the multi-dimensional situation information of the cluster system, and redeploy the cluster system topology configuration to restore the task performance.

[0059] Preferably, the cluster effectiveness evaluation module includes:

[0060] a) Construction of performance index system: Construct a performance evaluation index system based on the general and special characteristics of the cluster system. The general characteristics include reliability, maintainability, security, safety and environmental adaptability, and the special characteristics include time dynamic characteristics and space-time dynamic characteristics;

[0061] b) Data collection: The real-time operation data of the cluster system is collected through the multi-dimensional situation awareness module, and the data includes node health status, link status, task execution status and environmental change information;

[0062] c) Evaluation of the network efficiency of the cluster system: The network average efficiency formula E(G) is used to evaluate the efficiency of information transmission between nodes in the cluster system. The formula is:

[0063] ;

[0064] Where N is the number of nodes, w i and w j are the weights of node i and node j respectively, d ij is the shortest path length between node i and node j; d) Task execution balance evaluation: For the three-level spatiotemporal dynamic cluster, the cluster balance ε is used bEvaluate the balance of cluster system task execution capability. The balance calculation formula is:

[0065] ;

[0066] Where I represents the total number of clusters in cluster S, J i Represents cluster c i The total number of nodes, ε ij Represents node n (i,j) The sum of the overlapping areas of the coverage area of ​​​​the node and the coverage areas of other nodes in the same cluster within the task area; is the average redundancy value of all nodes in each cluster.

[0067] Preferably, the system comprises:

[0068] Open interfaces for interacting with external systems, supporting the integration and expansion of different types of reinforcement learning methods in the intelligent operation and maintenance framework;

[0069] The multi-dimensional situation data management module based on big data analysis is used to store, manage and analyze the historical operation data and real-time collection data of the cluster system, providing data support for multi-dimensional situation awareness and intelligent decision-making.

[0070] The present invention adopts the above-mentioned technical solution. The system includes a cluster modeling module, a multi-dimensional situational awareness module, an intelligent decision-making planning module and an efficiency evaluation module, and combines a variety of reinforcement learning algorithms to achieve autonomous operation and dynamic management of the cluster system in a complex environment. The system can achieve the following core technical effects by introducing reinforcement learning algorithms such as deep Q network (DQN), Actor-Critic Monte Carlo Tree Search (AC-MCTS) and distributed proximal policy optimization (DPPO):

[0071] 1. Improve the modeling accuracy and flexibility of cluster systems: Through the cluster modeling module, the present invention can accurately model the multi-level structure in the cluster system, including detailed descriptions of the three levels of cluster, cluster and node. Combining the temporal dynamic characteristics and spatiotemporal dynamic characteristics, the system can comprehensively analyze the state of the cluster at different operating stages, effectively improving the modeling accuracy and flexibility of the cluster system, and providing a reliable data foundation for intelligent operation and maintenance;

[0072] 2. Realize real-time perception and prediction of multi-dimensional situations: The multi-dimensional situation perception module in the present invention, combined with complex network theory and perception technology, can collect multi-dimensional operation data of the cluster system in real time, and extract multi-dimensional situation features through data fusion and pattern recognition technology. By constructing a situation prediction model, the system can not only monitor the current situation, but also predict future situation change trends, providing more forward-looking support for the system's intelligent decision-making, and effectively improving the system's ability to cope with complex environmental changes;

[0073] 3. Intelligent decision-making optimization and autonomous operation and maintenance: The intelligent decision-making planning module in the system realizes the automation of intelligent operation and maintenance decision-making of the cluster system by introducing reinforcement learning algorithms (such as DQN, AC-MCTS, and DPPO). Through learning and optimization, the system can generate optimal decisions in various scenarios such as degradation, failure, and post-disaster recovery, and generate efficient operation and maintenance solutions for different operation and maintenance problems (such as predictive maintenance, post-disaster repair, and dynamic reconstruction). This adaptive decision-making capability based on reinforcement learning can improve the system's autonomous operation and maintenance efficiency, reduce human intervention, and improve the system's stability and operating efficiency;

[0074] 4. Effectively deal with cluster system degradation and failure: This invention uses the deep Q network (DQN) algorithm to predict the node degradation state in the cluster system and generate the optimal maintenance strategy, which improves the system's fault prevention ability. Through the balanced analysis of maintenance benefits and failure risks, the system can reasonably arrange the maintenance sequence, reduce the failure rate, extend the service life of the equipment, and reduce operation and maintenance costs;

[0075] 5. Significantly improved post-disaster emergency repair efficiency: Based on the Actor-Critic Monte Carlo Tree Search (AC-MCTS) algorithm, the present invention can quickly plan the repair path and timing when a large-scale cluster system fails or suffers a disaster, coordinate the collaborative repair work of multiple teams, effectively shorten the system recovery time, and reduce the system downtime and performance degradation caused by failures or disasters;

[0076] 6. Dynamic reconstruction ensures task execution and system stability: The dynamic reconstruction submodule in the system, based on the distributed proximal policy optimization (DPPO) algorithm, can dynamically adjust the topology structure at different levels of the cluster system to ensure the efficiency of task execution and the stability of the cluster system. By analyzing multi-dimensional situation information, the system can generate a reconstruction plan in a timely manner to ensure the continuous and stable operation of the system;

[0077] 7. Improve the performance evaluation and feedback capabilities of the cluster system: The cluster performance evaluation module can comprehensively and accurately analyze the performance of the cluster system at different operation and maintenance stages through quantitative evaluation of indicators such as network efficiency, balance, reliability and elasticity. Through real-time performance evaluation and feedback, the system can continuously optimize its operation and maintenance decisions, provide data support and improvement directions for subsequent maintenance work, thereby improving the performance of the overall system;

[0078] 8. Optimize the resource utilization and operation and maintenance costs of cluster systems: Through the system's intelligent predictive maintenance and decision-making planning, the system can effectively reduce the frequency of sudden failures in cluster systems and reasonably allocate maintenance resources, thereby optimizing resource utilization and reducing operation and maintenance expenses. Especially in large-scale cluster systems, the reduction of operation and maintenance costs and the improvement of system stability have important economic benefits.

[0079] In summary, the present invention realizes efficient and autonomous operation and maintenance of cluster systems by integrating multi-level modeling, situational awareness, intelligent decision-making and performance evaluation modules, combined with reinforcement learning algorithms, and significantly improves the system's predictability, ability to respond to emergencies, and the overall efficiency and reliability of system operation. These technical effects are suitable for intelligent operation and maintenance scenarios of large-scale distributed cluster systems and have broad application prospects and commercial value. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 This is the overall architecture diagram of the intelligent operation and maintenance of the present invention.

[0081] Figure 2 Modeling example diagrams for various types of cluster systems.

[0082] Figure 3 This is a cluster system operation and maintenance status diagram.

[0083] Figure 4 Degraded state analysis for cluster nodes.

[0084] Figure 5 This is the Agent-Environment architecture diagram for the intelligent operation and maintenance of the cluster system.

[0085] Figure 6 This is a diagram of the cluster effectiveness evaluation index system.

[0086] Figure 7 It is a series-parallel hybrid cluster diagram.

[0087] Figure 8 Figure 2 shows the evolution of cluster performance under extreme events and recovery measures. DETAILED DESCRIPTION

[0088] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of the present invention.

[0089] The present invention explains the concept connotation of cluster system, expounds the composition and characteristics of cluster system, and further classifies cluster system in terms of composition and characteristics. In terms of composition, the multi-level composition form of cluster-cluster-node is considered. In terms of characteristics, the temporal dynamic characteristics and spatiotemporal dynamic characteristics are targeted, and the conditions such as scale, synergy, and uncertainty are comprehensively considered. Four common technologies of modeling, perception, decision-making, and evaluation are integrated to build an intelligent operation and maintenance framework of cluster system, which includes: cluster system modeling module, multi-dimensional situation perception module, intelligent decision-making planning module, and cluster effectiveness evaluation module, and the complex interactive relationship between the above four modules is sorted out, such as Figure 1 As shown. The cluster system modeling module needs to conduct modeling analysis from multiple levels for different types of cluster systems; the multi-dimensional situation awareness module needs to deconstruct and classify the multi-dimensional situation information of the cluster system operation stage in a complex environment, and qualitatively and quantitatively describe the key information to support multi-dimensional situation awareness; the intelligent decision-making and planning module needs to sort out the decision-making and planning mathematical models of various operation and maintenance problems for the three main types of operation and maintenance problems, namely preventive maintenance, corrective maintenance, and operation management, and design a reinforcement learning architecture to support decision-making and planning; the cluster effectiveness evaluation module needs to construct its indicator system for the special and general characteristics of the cluster system, and propose an evaluation algorithm for key technical indicators to evaluate the performance of the cluster system in a complex environment and guide the decision-making and planning module to optimize the solution.

[0090] The construction of the intelligent operation and maintenance framework of cluster systems is a pioneering work, which aims to build an open, basic common technology framework to provide a high-precision, high-efficiency, and low-cost intelligent operation and maintenance solution system for various cluster systems in different operation and maintenance scenarios. The present invention builds its framework by integrating the common technologies of intelligent operation and maintenance, and conducts research on some algorithms on this basis. Future researchers can continue to expand the relevant research on intelligent operation and maintenance technologies based on the common technology framework. In addition, the proposed common technology framework can provide an environment for interaction and training for different types of reinforcement learning methods. By unifying the interface specifications in the intelligent operation and maintenance framework and refining the common processes of reinforcement learning algorithms, it can be compatible with many different types of deep reinforcement learning algorithms.

[0091] 1.1 Cluster system modeling module

[0092] The present invention discusses the cluster system modeling module in the intelligent operation and maintenance framework. It first explains the conceptual connotation of the cluster system, then analyzes the multi-level morphology of the cluster system from the perspective of composition, and analyzes the temporal and spatial dynamic characteristics from the perspective of characteristics. On this basis, modeling and analysis are carried out for different categories of cluster systems.

[0093] 1.1.1 Concept of Cluster System

[0094] At this stage, cluster systems have gradually become a very important representative product form, such as various infrastructure systems, various equipment clusters, various intelligent manufacturing systems, etc. In a broad sense, a system refers to a whole formed by a certain internal connection within a certain space and time range. It is a system composed of different systems and has a certain order. From this perspective, the top level of the general system is the system, that is, the system of systems. For the system, what is usually considered is the whole composed of different heterogeneous systems. In contrast, clusters usually focus on the whole composed of homogeneous systems, which is the next level of the system. In addition, compared with network systems, clusters pay more attention to the synergy of large-scale homogeneous systems, but different systems in the cluster still have obvious network characteristics.

[0095] At present, a lot of exploratory research has been carried out on cluster systems at home and abroad. Cluster systems are usually composed of large-scale relatively simple systems, and global collaborative behavior is achieved through local collaboration between individual systems. In order to further study cluster systems, this paper will conduct a multi-level deconstruction analysis on the composition of cluster systems. In addition, cluster systems will be classified according to their different characteristics.

[0096] Regarding the composition of the cluster system, the present invention deconstructs it from the functional forms of three levels: Swarm, Cluster, and Node. The cluster system can construct a cluster-node (Swarm-Node, SN) form, in which the nodes inside the cluster are combined into the functional form of the cluster system according to a certain order and internal connections between nodes, and clusters can interact with each other through node cross-cluster connections; in addition, the cluster system can also construct a cluster-cluster-node (Swarm-Cluster-Node, SCN) form, in which the nodes inside the cluster are combined into clusters according to a certain order and internal connections, and clusters interact with different nodes and different clusters to form the functional form of the cluster system.

[0097] Regarding the characteristics of cluster systems, the present invention divides cluster systems into two categories: time dynamic clusters and space-time dynamic clusters, based on the two core characteristics of time dynamic characteristics and space-time dynamic characteristics. A time dynamic cluster refers to a cluster system whose spatial attributes are mainly static characteristics, while its temporal attributes are mainly dynamic characteristics. For example, various infrastructure systems such as power networks usually show a relatively static state in their spatial position characteristics, but their operation, failure, and degradation properties show dynamic change characteristics over time. A space-time dynamic cluster refers to a cluster system whose spatial and temporal attributes both show dynamic characteristics. For example, various equipment clusters such as cluster drone systems usually show a moving state in their spatial position characteristics during missions, and their operation, failure, and degradation properties still show dynamic change characteristics over time.

[0098] Based on the above explanation of the concept of cluster system, for the cluster system cluster-cluster-node multi-level form, considering the two core characteristics of cluster system time dynamics & spatiotemporal dynamics, further analyze the multi-dimensional situation that needs to be discussed in the operation of cluster system under scale, coordination and uncertainty conditions, focusing on the state information related to cluster system operation and maintenance. For example, for the cluster-node two-level form of time dynamic cluster system, assuming that the spatial position of the nodes in the cluster presents a large-scale distributed layout, in the long-term operation stage, each node and link in the cluster will experience operation and maintenance situations such as degradation and failure. At the same time, when facing large-scale natural disasters such as typhoons and earthquakes, there will be large-scale node destruction in the local area of ​​the cluster. In addition, for the cluster-cluster-node three-level form of spatiotemporal dynamic cluster system, assuming that the spatial position of each node in its operation stage is dynamically adjusted according to real-time changes, and clusters are constructed in real time according to a certain order and internal connections between nodes to complete the collaborative interaction between multiple clusters and multiple nodes. In such a scenario, the operation and maintenance situation inside the cluster system is further described: for the cluster level, it is necessary to consider failure, degradation, local attack and other situations; for the node level, it is necessary to consider the failure and degradation of nodes and links.

[0099] 1.1.2 Cluster System Modeling and Analysis

[0100] Based on the explanation of the concept of cluster system, this paper models and analyzes the cluster system from the perspective of composition, targeting the multi-level forms of clusters, clusters, and nodes, and from the perspective of characteristics, targeting the temporal dynamic characteristics and spatiotemporal dynamic characteristics of the cluster system. This topic mainly studies three typical cluster systems, including two-level temporal dynamic clusters, three-level temporal dynamic clusters, and three-level spatiotemporal dynamic clusters, such as Figure 2As shown in the figure, a simple description is given from the perspectives of composition and dynamic characteristics. For two-level time dynamic clusters, clusters composed of large-scale distributed systems are representative examples. Considering that the distributed systems in the cluster are combined in a certain order and connected to the internal network of the cluster to form a two-level cluster-node form, each system in the cluster exhibits static characteristics in the spatial dimension, such as the spatial position and other attributes remain static, and exhibits dynamic characteristics in the time dimension, such as health, degradation, damage and other operation and maintenance status dynamically change over time. Specific case objects include infrastructure networks represented by large-scale distributed power systems. For three-level time dynamic clusters, clusters composed of large-scale series-parallel-bridge systems are representative examples. Considering that each system in the cluster is combined in a certain series-parallel-bridge order to form a cluster form, clusters are combined based on parallel or n / k redundant structures to form a cluster form. The static and dynamic characteristics of each system in the cluster in the spatial and temporal dimensions are similar to those of two-level time dynamic clusters. Specific case objects include large-scale battery clusters for industrial energy storage systems. For the three-level spatiotemporal dynamic cluster, the cluster composed of multiple formation systems is a representative example. Considering that the systems in the cluster are combined into a formation according to a certain topological configuration and collaborative relationship, the formations are combined into a cluster based on multi-task collaborative relationships. The systems in the cluster show dynamic characteristics in the spatial dimension. For example, their spatial position and other attributes need to be adjusted in real time according to changes in the situation in a complex environment. The dynamic characteristics shown in the time dimension are similar to those of the time dynamic cluster. Specific case objects include cluster systems composed of multiple UAV formations.

[0101] Based on the above analysis of the cluster system, the cluster system model is analyzed from the perspective of multi-level composition. The spatiotemporal dynamic and temporal dynamic characteristics are mainly reflected in the multi-dimensional attributes of the cluster system model, which will be described in detail in the next section on the introduction of the multi-dimensional situation awareness module. For the cluster system object S under study, a node set N is first established to describe all nodes n in the cluster S, as shown below:

[0102] N={n1, n2,…, n n ,…}(1.1)

[0103] The lowercase cursive letter n nRepresents the nth node in the cluster. The number counting variable n will also be used in the subsequent parts of the present invention. This parameter only represents the integer number. The node set N is the number of each node, which is the unique identifier of each node. In the present invention, the integer number is used as the unique identifier of the node. At the same time, the product code of each node in the actual scene can also be used as the unique identifier. In the intelligent operation and maintenance framework, for the cluster system object S, the data structure of its node set N is initialized in the form of a set, and the node set is constructed based on a hash table. The keyword of the hash table is the unique identifier of the node, which is the node number in the present invention. The stored data elements are the relevant attributes of the node, which involve spatiotemporal dynamics and time dynamics. It will be described in detail in the next section about the introduction of the multi-dimensional situation awareness module. In addition, before forming a cluster form, when some or all nodes in the node set need to be combined into a cluster form in a certain order, it is necessary to establish a cluster set C to describe all clusters c in the cluster at the current moment according to factors such as the specific task requirements in the current scene, as shown below:

[0104] C={c1, c2,…, c m …} (1.2)

[0105] The lowercase cursive letter c m Represents the mth cluster in the cluster at the current moment. The number counting variable m will also be used in the subsequent parts of the present invention. This parameter only represents an integer number. Cluster set C The number of each cluster is a dynamic identifier of each cluster, which is related to factors such as tasks, and changes dynamically accordingly. It is generated when a cluster needs to be formed. The cluster set is constructed based on a hash table. The keyword of the hash table is the cluster number, and the stored data elements are the relevant attributes of the cluster, which involve spatiotemporal dynamics and time dynamics. It will be described in detail in the next section on the introduction of the multidimensional situational awareness module. When the cluster set needs to be dynamically adjusted due to changes in factors such as real-time tasks, the corresponding cluster data can be dynamically added, deleted or modified based on the data structure of the hash table.

[0106] On the basis of defining node sets and cluster sets, the multi-level form of clusters is further analyzed. For two-level clusters, all nodes in the node set can form a two-level cluster form by combining them in a certain order and the internal connections between nodes. By establishing its node set N, the one-to-many tree structure relationship between the cluster and all its nodes can be described. However, for a three-level cluster, after establishing its node set, before forming a three-level cluster form, it is necessary to define its three-level model. Specifically, the cluster system object S has a one-to-many tree structure relationship with all clusters in its cluster set, and each cluster has a one-to-many tree structure relationship with all nodes therein, thereby forming a cluster-cluster-node three-level model. The intelligent operation and maintenance framework of the present invention constructs the above three-level model by nesting hash tables in hash tables. First, a one-to-many data structure between clusters is constructed based on the hash table. The keyword of the hash table is still the dynamic number of the cluster, and the stored data element is the node set of the cluster. It is constructed using a nested hash table. The keyword of the nested hash table is the dynamic number of each node in the cluster, and the stored data element is the unique identifier of the node, that is, the number of the node in the node set N. Therefore, the three-level cluster model SM is constructed by nesting hash tables in hash tables. III , the mapping relationship between the unique identifier of a node and its dynamic number in the cluster can be described as follows:

[0107] (1.3)

[0108] where n (i,j) Represents cluster c i The number counting variables i and j will also be used in the subsequent parts of the present invention. This parameter only represents the integer number. n It represents the nth node in the node set N in formula (1.1). Its number n is the unique identifier of each node and is a unique fixed number. Cluster c i The dynamic number i of the cluster is consistent with the dynamic number of each cluster in formula (1.2). When the cluster C changes with factors such as real-time tasks and is dynamically adjusted, the three-level cluster model SM III Then, the corresponding cluster level mapping relationship is dynamically added, deleted or modified based on the data structure of the hash table. i It consists of multiple nodes, where node n (i,j)The subscript number is the dynamic identification of the node in the cluster, which is related to factors such as tasks and changes dynamically. The corresponding node level mapping relationship is dynamically added, deleted or modified based on the data structure of the nested hash table. The keyword of the nested hash table is the counting variable j in the dynamic identification of the node. The data element stored in the nested hash table is the unique identification of the node, that is, the number of the node in the node set N. In short, the number of each node in the cluster in the node set N is the unique identification of the node, which always remains unchanged. When the three-level cluster model SM III When a node in the node set N needs to be called to form a cluster according to factors such as task requirements, the node will be given a corresponding dynamic number. The dynamic numbers of the cluster and the node can be dynamically rebuilt according to factors such as changes in real-time task requirements.

[0109] 1.2 Multi-dimensional Situational Awareness Module

[0110] The present invention discusses the multi-dimensional situation awareness module in the intelligent operation and maintenance framework. Firstly, the conceptual connotation of cluster system operation and maintenance situation and perception is explained. On this basis, the multi-dimensional situation awareness of cluster system intelligent operation and maintenance is analyzed.

[0111] In the situational awareness of cluster system operation and maintenance, the "state" of operation and maintenance refers to the situation, state or environment of the current operation stage of the cluster system, while the "momentum" of operation and maintenance refers to the potential trend or influence of the future operation stage of the cluster system. There is interaction and influence between the "state" and "momentum" of operation and maintenance. When the "state" of operation and maintenance changes, it will also trigger changes in the "momentum" of operation and maintenance. For example, if a large-scale destruction occurs, including natural disasters such as hurricanes or earthquakes, this will cause the current state of the cluster system to change, and this change may trigger a series of subsequent influences and potential forces, such as the destruction of large-scale nodes in the cluster system, the breakage of links, the evacuation of personnel, etc. Therefore, the change of the "state" of operation and maintenance will trigger the change of the "momentum" of operation and maintenance. On the other hand, when the "momentum" of operation and maintenance changes, it will also trigger the change of the "state" of operation and maintenance. In the situational awareness of intelligent operation and maintenance of cluster systems, it is crucial to understand and analyze the interaction between the "momentum" and the "state" of operation and maintenance. By observing and predicting this interaction, we can better respond to and adapt to different scenarios and environments to support better decisions and actions.

[0112] At the same time, in the situational awareness of intelligent operation and maintenance of cluster systems, changes in the "feeling" of operation and maintenance will trigger changes in the "knowing" of operation and maintenance. This is because the perception of operation and maintenance situation is the process of collecting and obtaining information through sensors or sensors. When sensors observe certain changes in complex environments, they will convert them into various forms of signals, which will then be processed, parsed and stored as meaningful operation and maintenance situation information. This is the generation of "knowing" of operation and maintenance. Conversely, changes in the "knowing" of operation and maintenance will also trigger changes in the "feeling" of operation and maintenance. In the process of perceiving the operation and maintenance situation of cluster systems, through the processing, parsing and storage of operation and maintenance situation information, we can obtain understanding and cognition of the complex environment, so that we can respond to changes in the complex environment. These responses may include adjusting the parameters of the cluster system perception devices, reorganizing the functional form of the cluster system, changing the goals or methods of operation and maintenance situation awareness, etc., which will affect the results of the cluster system operation and maintenance situation awareness. In summary, in the situational awareness of intelligent operation and maintenance of cluster systems, the "feeling" and "knowing" of operation and maintenance are interrelated and mutually influential. Changes in the "feeling" of operation and maintenance trigger changes in the "knowing" of operation and maintenance, and changes in the "knowing" of operation and maintenance will further affect changes in the "feeling" of operation and maintenance. This interaction enables the situational awareness of intelligent operation and maintenance of cluster systems to adapt to changes in complex environments and provide accurate and real-time operation and maintenance situation information to support decision-making planning and performance evaluation of intelligent operation and maintenance.

[0113] Based on the above discussion, to build a multi-dimensional situation awareness module for intelligent operation and maintenance of cluster systems, it is necessary to comprehensively analyze the interaction between the "state" and "potential" of cluster system operation and maintenance, as well as the interaction between the "sensing" and "knowledge" of operation and maintenance, and then construct its operation and maintenance situation map for the multi-dimensional situation information of the cluster system operation stage under complex environment, such as Figure 3 As shown. For the "state" of operation and maintenance, based on the definitions of storage attribute information for the node set N and the cluster set C in formulas (1.1) and (1.2), and the definition of storage attribute information for the three-level cluster model SM in formula (1.3), III Based on the definition of , various information in the cluster system operation scenario is further deconstructed, and the environmental information, cluster-cluster-node operation information, degradation information, etc. are classified, and then the similar and highly correlated information is integrated, and then the cluster-cluster-node information is distinguished according to the temporal dynamic characteristics and spatiotemporal dynamic characteristics. On this basis, the environmental object εnv is defined and the environmental information set H is constructed, as shown in Figure 3It is shown that the set is composed of multi-dimensional situation information, including task information, damage information, map information, etc. It also includes the constructed cluster information. For preventive maintenance and corrective maintenance, it also needs to include the guarantee system information. For the "trend" of operation and maintenance, focus on the trend between the current "state" and the future "state", analyze the real-time observation data of the multi-dimensional situation information in the environmental information set H, study the relevant prediction methods, describe the potential trend of the multi-dimensional situation information in the future operation stage, and then predict and update the "state" of operation and maintenance. For the "sense" of operation and maintenance, it is necessary to collect and obtain the operation and maintenance situation information in a complex environment through sensors or sensors, and convert it into various forms of signals. For the "knowledge" of operation and maintenance, it is necessary to qualitatively and quantitatively describe the operation and maintenance situation information in various scenarios during the operation stage of the cluster system. The present invention focuses on the framework of common technologies for intelligent operation and maintenance. In the multidimensional situational awareness module, only the input interface design of multidimensional situation data is done for the "sense" of operation and maintenance, and no in-depth research is done on technologies such as detection and diagnosis. However, the intelligent operation and maintenance framework proposed by the present invention still supports future researchers to integrate related technologies such as detection and diagnosis in the multidimensional situational awareness module; in addition, the multidimensional situational awareness module is designed for the "knowledge" of operation and maintenance. By designing environmental objects and their environmental information set H, and associating the node set N and cluster set C defined by the cluster modeling module, it has been able to process, parse and store the operation and maintenance situation information of the cluster system to support the decision-making planning and performance evaluation of intelligent operation and maintenance. Therefore, in the subsequent chapters of the present invention, the "sense" and "knowledge" of operation and maintenance are no longer distinguished, but the word "perception" is used to uniformly describe the input of environmental observations, as well as the processing, parsing and storage of situation information. The subsequent content of the present invention will be simply discussed with some algorithms integrated in the multidimensional situational awareness module, including data structure examples of operation and maintenance "state" and prediction methods of operation and maintenance "potential".

[0114] (1) Data structure example of “state”

[0115] In the node set N defined in the cluster system modeling module, the node's state, location and other attributes can be stored. However, for some specific cluster systems, it is necessary to consider not only the state of all nodes in the node set, but also the state of all links between nodes.

[0116] Since the intelligent operation and maintenance framework of the present invention is developed in the Python language, starting from the requirement of analyzing the links between all nodes in the cluster system, it is easy to think of the NetworkX tool, which is an algorithm tool for graph theory and complex network modeling. This tool is also developed in the Python language and contains rich algorithm tools for graph theory and complex network modeling and analysis, which can facilitate researchers to carry out complex network simulation modeling and data analysis. Although the NetworkX tool is a quite mature tool in the field of graph theory and complex networks, the implementation of the data structures of nodes, edges, and adjacency matrices in its underlying code does not endow it with the attributes of operation and maintenance situation. Therefore, when applying the NetworkX tool, the present invention comprehensively considers the operation and maintenance situation of the cluster system, and reconstructs the data structures of the nodes and edges of the cluster system based on the data structures of nodes and edges in the NetworkX underlying code, endowing them with relevant attributes of operation and maintenance situation.

[0117] Before introducing the data structures of the nodes and edges of the above cluster system, first, based on the complex network theory, the nodes and edges of the cluster system are simply defined and described. The cluster is represented as an undirected graph G with N nodes and K edges. Assume that the graph G is sparse, that is, K << N (N - 1) / 2. In addition, consider that the graph G is connected, that is, there is at least one path with a finite number of steps connecting any two nodes. Such a graph G needs two matrices to describe: the adjacency matrix A and the physical distance matrix L. A is an N × N adjacency matrix. If there is an edge connecting node n i and node n j , then the element a ij in the adjacency matrix is 1, otherwise it is 0. The element l ij of the matrix L represents the geographical distance between node n i and n j . Even if there is no edge element between node n i and node n j , l ijIt is also known. According to the above definition, the G graph is a collection of nodes and known node pairs (which can be called edges, links, etc.). In the NetworkX tool, a node can be any hashable object, such as text, a string, another graph, a custom node object, etc. In addition, NetworkX can also add nodes containing node attributes at the same time. For nodes, after creating an empty G graph, regardless of whether the added node contains node attributes, the node set N of the cluster system object S of the present invention can generate the required data. The node set generates different node arrays for addition according to the normal, degraded, and damaged states of the nodes at the current moment. For edges, it is necessary to define the link set L of the cluster system, and then generate the edge data and add it to the created G graph. The present invention refers to the data structure of the adjacency matrix in the underlying encoding of NetworkX for the construction of the cluster system link set L, that is, hash nested hash, and adds attribute information such as operation and maintenance status on this basis. In addition, the required attribute information can still be added according to the actual scenario. The link set L is as follows:

[0118] (2.1)

[0119] Where s(a ij ) represents the connection node n i and node n j The state of the edge of the present invention is mainly aimed at the operation and maintenance state. Different operation and maintenance problems may involve different states such as normal, degraded, and damaged. The state parameters can also be customized according to specific problems. In addition, the link set L only stores a ij =1 represents the link and its attributes. It should be noted that when analyzing the links between nodes in the cluster system, the present invention does not consider the spin link for the time being. In addition, most of the problems in the present invention are analyzed for undirected graphs, that is, s(a ij ) and s(a ji ) has the same meaning, but the construction of this data structure still supports the description of the link status in the directed graph required in related problems.

[0120] Based on the above construction of the node set N and the link set L, we can generate node arrays and link arrays describing different states for the normal and damaged states of all nodes and their links at a specific moment in the present or future, and then use the NetworkX tool to simulate and analyze the complex network model of the cluster system.

[0121] (2) Examples of “potential” prediction methods

[0122] Unlike the "state" of operation and maintenance, which focuses on the current state of the cluster, the "potential" of operation and maintenance focuses on the future changing trend of the cluster. In the operation and maintenance of cluster systems, node degradation is a type of time dynamic feature that is of great concern. At this stage, the degradation of nodes can be analyzed using the State of Health (SOH), Remaining Useful Life (RUL), etc., and the corresponding prediction methods are also very rich. This invention introduces a SOH random degradation model as an example of the prediction method of the "potential" of operation and maintenance. The proposed intelligent operation and maintenance framework still supports the integration of richer prediction methods.

[0123] First, the SOH of the actual node object needs to be defined. For example, in the case study of Chapter 3 of this invention, for large-scale battery pack integration in industrial energy storage systems, based on the analysis of the capacity degradation model of each battery pack, its SOH is defined as the percentage of the current capacity to the initial capacity, so 0≤SOH≤100%. Then, in order to use the SOH of all nodes in the cluster to predict the trend of cluster status changes, it is assumed that the SOH of each node at any time follows the normal distribution N (µ,σ 2 ), this assumption means that the defined SOH only provides the average SOH of the nodes in the cluster, that is, µ=SOH, and there are still slight differences between each node due to materials, manufacturing processes and other environmental factors. Assume that the standard deviation σ is (1-µ) / 6, and its initial value is zero. Based on the assumption of discretization of the degradation process, there are multiple discrete degradation states in the node degradation process, and the cluster degradation state can also be constructed by the SOH of the discretized nodes. Based on the normal distribution N (µ, σ 2 ) As a result, by analyzing the probability density function of each node's SOH, we can obtain the probability density function of each node's SOH at each stage after a certain period of operation. For example, a node starts running from the state of SOH = 100%. According to a certain degradation law, when the node runs for 100 hours, 200 hours, 300 hours, 400 hours, 500 hours, 600 hours, and 700 hours, its SOH probability density function is as follows: Figure 4 shown.

[0124] Will Figure 4 The node SOH in is divided into six levels: ≥95%, 90~95%, 85~90%, 80~85% and ≤80%. The probability of each SOH level of the node in different operation stages is shown in Table 1.

[0125] Table 1 Probability of different SOH levels during cluster node operation

[0126]

[0127] The present invention introduces the multidimensional situation awareness module in the intelligent operation and maintenance framework, explains the connotation of operation and maintenance "state" and "trend", and constructs a cluster system operation and maintenance situation diagram. On this basis, the data structure of the operation and maintenance "state" is designed, focusing on the introduction of environmental objects and their environmental information sets H, as well as cluster objects S and their related cluster models. Then, according to the temporal dynamic characteristics and spatiotemporal dynamic characteristics in the operation and maintenance situation diagram, the relevant prediction methods are studied to form a multidimensional situation awareness module to support the decision-making planning and performance evaluation of the cluster system.

[0128] 1.3 Intelligent Decision-making and Planning Module

[0129] The intelligent decision-making planning module of the cluster system intelligent operation and maintenance framework is designed based on reinforcement learning theory. Cluster intelligent operation and maintenance includes three main problems: preventive maintenance, corrective maintenance, and operation management. Each type of problem still includes various representative frontier problems. For various representative operation and maintenance problems, reinforcement learning agents can be designed according to specific needs to support intelligent operation and maintenance decision-making. The various operation and maintenance decision-making agents designed need to maximize the specific operation and maintenance benefits and learn which specific operation and maintenance actions to perform - how to map the operation and maintenance state space to the operation and maintenance action space. The operation and maintenance decision-making agent is not told which operation and maintenance actions to take, but must try to find out which operation and maintenance actions can generate the greatest operation and maintenance benefits. This is the core idea of ​​designing the operation and maintenance decision-making method based on reinforcement learning, that is, the idea of ​​trial and error search. For the most challenging representative frontier operation and maintenance problems, in the trial and error search process, any operation and maintenance action selected will not only affect the immediate benefits in the operation and maintenance scenario, but also affect the next operation and maintenance state, thereby affecting all potential future operation and maintenance delayed benefits in the operation and maintenance scenario. In short, the trial-and-error search of operation and maintenance actions and the delayed benefits of operation and maintenance situations will be the core elements in designing various operation and maintenance decision-making agents.

[0130] In addition, since the reinforcement learning problem is based on reward-driven completion of the agent-environment interaction, thereby realizing the perception-action-learning cycle, the design of the operation and maintenance decision-making agent requires the construction of a simulation environment for the intelligent operation and maintenance of the cluster system, providing the designed operation and maintenance decision-making agent with multi-dimensional situation observation information. In the previous section, the multi-dimensional situation awareness module has completed the construction of the relevant intelligent operation and maintenance framework, which can support the provision of operation and maintenance situation simulation information Obversations (εnv) as state observation input for the operation and maintenance decision-making agent. In addition, to realize the perception-action-learning cycle in the scenario of cluster intelligent operation and maintenance, it is necessary to describe various MDPs in various intelligent operation and maintenance scenarios. The core of the sequential decision-making architecture of intelligent operation and maintenance scenarios is the evaluative feedback of operation and maintenance decisions. That is, the operation and maintenance actions in the sequential decision-making need to consider not only the immediate benefits in the current operation and maintenance scenario, but also the potential development of multi-dimensional situations in the operation and maintenance scenario, thereby leading to fluctuations in benefits in future operation and maintenance scenarios. Therefore, considering the importance of timely benefits and delayed benefits to operation and maintenance decision-making planning, the design of the value function of operation and maintenance actions is particularly important. The value function of operation and maintenance actions needs to be designed based on the benefits of operation and maintenance. It is necessary to build an indicator system to evaluate the effectiveness of the cluster under various operation and maintenance situations. This will be introduced in the next section on the cluster effectiveness evaluation module. In summary, to design an intelligent decision-making and planning module based on reinforcement learning, it is necessary to build an intelligent operation and maintenance simulation environment for the cluster system, design decision-making and planning algorithm agents for various intelligent operation and maintenance scenarios, and consider the operation and maintenance situation observation and operation and maintenance decision benefits in the operation and maintenance simulation environment. The cluster system intelligent operation and maintenance agent-environment interaction interface is sorted out, such as Figure 5 shown.

[0131] In the intelligent decision-making and planning module, the designed operation and maintenance decision agent mainly targets three types of problems: preventive maintenance, corrective maintenance, and operation management. It is necessary to sort out the input-output relationship between various types of operation and maintenance problems and the cluster system modeling module, the multi-dimensional situation awareness module, and the cluster effectiveness evaluation module. Among the three main types of operation and maintenance problems, we can focus on three key representative frontier problems and carry out research on operation and maintenance decision-making methods based on deep reinforcement learning. In view of the "predictive maintenance" problem of the three-level time dynamic cluster system under the condition of system performance degradation in the long-term operation stage, a method based on DQN deep reinforcement learning is used to comprehensively weigh the system degradation state and the imbalance characteristics of maintenance benefits to generate a predictive maintenance plan. For details, please refer to a DQN-based cluster system predictive maintenance decision method disclosed in the Chinese invention patent application (application number: 202411357728X, application date: 20240927), which will not be described in detail in this specification; for the "post-disaster repair" problem of multiple teams in a two-level time dynamic cluster system facing large-scale local damage, a two-level decision of maintenance timing and path planning is coordinated through an Actor-Critic deep reinforcement learning method to generate a corrective maintenance plan. For details, please refer to the Chinese invention patent application (application number: 202411357728X, application date: 20240927). The patent application (application number: 2024112822477, application date: 20240913) discloses a method for post-disaster restorative maintenance of cluster systems based on the AC-MCTS algorithm, which will not be described in detail in this specification; for the "dynamic reconstruction" problem of a three-level spatiotemporal dynamic cluster system with multi-formation characteristics, a two-level dynamic reconstruction of the functional reconstruction between nodes within the cluster of the three-level cluster system and the functional reconstruction between clusters is analyzed by a method based on DPPO deep reinforcement learning to generate a dynamic reconstruction plan, for details, see the Chinese invention patent application (application number: 2024114796053, application date: 20241023) discloses a cluster system dynamic reconstruction decision method based on DPPO deep reinforcement learning, which will not be described in detail in this specification. In addition, the designed intelligent decision-making planning module still supports future researchers to integrate richer operation and maintenance decision-making methods, and ultimately provides a high-precision, high-efficiency, and low-cost intelligent operation and maintenance solution system for various cluster systems in different operation and maintenance scenarios.

[0132] 1.4 Cluster effectiveness evaluation module

[0133] The present invention discusses the cluster effectiveness evaluation module in the intelligent operation and maintenance framework. In the proposed intelligent operation and maintenance framework, the cluster effectiveness evaluation module is intended to evaluate the cluster effectiveness under the multi-dimensional situation of the cluster system operation stage, so as to guide the intelligent decision-making planning module to perform iterative optimization. The effectiveness of the cluster system refers to the ability of the entire cluster system to complete the specified tasks under the specified complex environment and the considered organizational, tactical, survival, and support conditions. It reflects the quality of the cluster system and its ability to complete various tasks. It is comprehensively determined by its general characteristics and special characteristics to make a comprehensive and correct assessment of its quality and ability to complete tasks. The main factors affecting the cluster effectiveness are general characteristics and special characteristics. Among them, the general characteristics include the reliability, maintainability, supportability, testability, safety, durability, human factors and environmental adaptability of the cluster. Special characteristics are divided into two categories: time dynamic characteristics and time-space dynamic characteristics according to the characteristics of the cluster system. Based on the above analysis, we further construct a cluster system performance evaluation index system. The construction of the index system is the basis for studying cluster effectiveness evaluation. By analyzing the task requirements of each task stage of the cluster system, gathering the typical task processes, and comprehensively analyzing the special characteristics and general characteristics of the cluster system, we can summarize the evolution law of typical task functions. On this basis, we select relevant cluster system index items, and analyze and discuss the various performance influencing factors of the cluster system. These factors are not independent of each other, and the complex correlations between them need to be sorted out. Finally, we establish an index system for cluster effectiveness evaluation, such as Figure 6 As shown in the figure, including the general characteristics and special characteristics of cluster systems, the performance indicators involved are quite rich. This paper conducts relevant research on the performance indicators used in the operation and maintenance decision-making problems in subsequent chapters and gives evaluation algorithms. The performance evaluation module developed based on Python language and integrated into the intelligent operation and maintenance framework is completed. The proposed intelligent operation and maintenance framework still supports the expansion of related performance evaluation algorithm research according to the specific operation and maintenance problem requirements.

[0134] 1.4.1 Cluster System-Specific Features

[0135] In the cluster system modeling module, the time dynamic characteristics and the time-space dynamic characteristics are described according to the characteristics of the cluster system. Therefore, the present invention also divides the special characteristics of the cluster system into time dynamic characteristics and time-space dynamic characteristics, and takes the network efficiency of the two-level time dynamic cluster and the balance of the three-level time-space dynamic cluster as examples of the performance evaluation algorithm for detailed introduction.

[0136] 1.4.1.1 Cluster System Network Efficiency

[0137] For two-level time dynamic clusters, when considering the status of all nodes in the node set, it is also necessary to consider the status of all links between nodes. At this time, the network efficiency of the cluster system is a key indicator for evaluating its effectiveness, which can measure the efficiency of information exchange or energy transmission between nodes in the cluster system. Based on the definition of node set N and link set L in the cluster modeling module and the multi-dimensional situational awareness module and the construction of their data structure, the cluster is represented as an undirected G graph with N nodes and K edges, and its adjacency matrix A and physical distance matrix L are defined. In addition, an important parameter of G is the node n i The degree of node n i The number of relevant edges k i , that is, n i The number of neighbors, k i The average value of is k=2K / N. When the adjacency matrix A is known, the distance between any two nodes n can be calculated. i and node n j The shortest path length d between ij Assuming G is connected, for , d ij exists and is positive. To quantify the structural properties of G, two different parameters are usually evaluated: characteristic path length L and clustering coefficient C. L is the average distance between two nodes. , C describes the local characteristics of the network and can be defined as , where C i Represents node n i The adjacent node subgraph G i The number of edges in the system divided by the maximum number of edges allowed is k i (k i -1) / 2. By representing the real network as a weighted G graph, the G graph may even be non-sparse and non-connected. Such a G graph requires two matrices to describe: the adjacency matrix A={a ij}, which is defined the same as for the unweighted graph, and the physical distance matrix L = {l ij}. Matrix element l ij It can be the spatial distance between two nodes or the strength of the interaction between them: even if node n i and node n j There is no edge between them, so we can also assume that l ij is known. For example, ij It can be the geographical distance between nodes in the power system cluster. i and node n j The shortest path length d between ij is the node n in the G graph i and node n jThe minimum physical distance of all possible paths. Therefore, the shortest path length matrix D = {d ij} is obtained by using the matrix A={a ij} and the matrix L={l ij}, it can be defined as:

[0138] (4.1)

[0139] When node n i and node n j When there is an edge between ij ≥l ij ,∀i,j. Node n i and node n j The efficiency between ij It can be defined as:

[0140] (4.2)

[0141] where w i is node n i In this paper, the node degree is used as the weight. In addition, the node weight can be defined according to the specific problem. i and node n j If there is no path between ij =+∞ and∈ ij =0, this situation applies to two scenarios: in the initial state, there is no path between two nodes in the G graph; in the state of large-scale local damage, all paths between two nodes in the G graph fail, resulting in no path. In these two scenarios, the G graph is a non-connected graph. In summary, the average efficiency of G can be defined as:

[0142] (4.3)

[0143] In order to normalize the average efficiency E, consider the ideal complete graph G ideal With all possible edges in G, the total number of edges is N(N-1) / 2. In this ideal case, information or energy is transmitted in the most efficient way, because in this ideal case , and the average efficiency E is at its maximum value at this time The efficiency E(G) considered in the rest of this paper is always divided by E(G ideal), so 0≤E(G)≤1. Although E(G)=1 exists and makes sense when there is an edge between every pair of nodes. However, in reality, due to the large number of nodes in the cluster system, it is almost difficult to satisfy the requirement that there is a link between every pair of nodes. Therefore, it is difficult for the network model of the cluster system to achieve an average efficiency close to 1. In addition, formula (4.3) defines the global efficiency of the G graph, so it can be called E glob The data structures of the node set N and link set L constructed in the multi-dimensional situational awareness module are based on the data structures of the underlying code of the NetworkX tool. The nodes and links are further endowed with operation and maintenance status attributes, which can generate operation and maintenance status information at different operation stages to simulate and analyze the connected graph of the cluster system in a normal state and the disconnected graph in a damaged state. The network average efficiency algorithm of the NetworkX tool is improved to evaluate the network efficiency of the cluster system for specific operation and maintenance status.

[0144] 1.4.1.2 Cluster system balance

[0145] For the three-level spatiotemporal dynamic cluster, due to the different number of nodes in each cluster, there are certain differences in the overall performance of each cluster during the task execution process. In order to ensure that the cluster has the ability to continuously execute tasks and adapt to environmental uncertainties, in the process of cluster configuration and deployment within the cluster, multiple clusters must have relatively balanced task capabilities. For this reason, the present invention defines the balance degree to quantify the balanced task capabilities, and then realizes the evaluation of the performance of the three-level spatiotemporal dynamic cluster. This paper mainly analyzes the detection capability of the cluster system and defines the balance degree of the cluster system. For the task capability of each node, assuming that its coverage area is a circular area, cluster c i The jth node n in (i,j) The coordinates of the node are the center coordinates of the area covered by the node, which can be expressed as (x ij ,y ij ). Define the task area coverage of cluster S Cov S , as shown below:

[0146] Cov S =[cov1,…,cov i ,…] (4.4)

[0147] Where cov i Represents cluster c i The coverage rate of the task area is obtained by dividing the total area covered by all nodes in the cluster by the area of ​​the task area. Define cluster c i The redundancy matrix M i ε ={ε ij} represents the repeated coverage of the task area of ​​the cluster, and the element ε in the matrixij Represents node n (i,j) The sum of the overlapping areas of the coverage area of ​​the node and the coverage areas of other nodes in the same cluster within the task area. Taking into account the node distribution characteristics and the overall performance level of the cluster, in order to avoid the situation where a cluster has insufficient capacity redundancy due to taking on too many tasks, or a cluster has excessive capacity redundancy due to taking on too few tasks, thereby reducing the efficiency of cluster task execution, the cluster balance degree ε is defined. b , which represents the mean square error of the redundancy values ​​of all nodes that can work normally in each cluster during the task execution phase, as shown below:

[0148] (4.5)

[0149] Where I represents the total number of clusters in cluster S, J i Represents cluster c i The total number of nodes, is the average redundancy value of all nodes in each cluster. On the basis of defining the cluster balance, in order to facilitate the selection of the optimal action in the subsequent operation and maintenance decision-making problem, it is necessary to know which nodes in the cluster are damaged and how to take corresponding operation and maintenance actions, which has the greatest impact on the recovery of cluster balance. To this end, the reconstruction process for local damage during the three-level spatiotemporal dynamic cluster task is considered. On the basis of defining the cluster system balance calculation model, the cluster loss degree is introduced to characterize the recovery effect of the overall cluster balance in the operation and maintenance stage, and the loss degree of cluster S in the reconstruction stage is defined as S for:

[0150] (4.6)

[0151] where t init represents the initial time of reconstruction, t init represents the reconstruction recovery time threshold, ε b thr Indicates the cluster balance threshold, loss S (t) represents the real-time loss value of cluster balance at time t in the reconstruction process, ε b (t) represents the real-time balance degree of the cluster at time t in the reconstruction process. In the reconstruction process facing local damage during the cluster task, it is necessary to minimize the loss of balance capacity ζ S To optimize the goal and meet the constraints of reconstruction time threshold.

[0152] 1.4.2 General Features of Cluster Systems

[0153] The common characteristics of cluster systems include reliability, maintainability, security, testability, safety, environmental adaptability and elasticity of the cluster. The reliability and elasticity of the cluster system are introduced in detail as an example of the performance evaluation algorithm.

[0154] 1.4.2.1 Cluster System Reliability

[0155] The present invention applies the reliability block diagram (RBD) modeling method to perform reliability modeling on the cluster system. RBD is the most basic reliability model. Its basic idea is to describe the logical relationship between node functions and cluster functions based on the functional correlation between nodes in the cluster system. The reliability block diagram of the cluster system consists of boxes, lines and logical relationships representing node objects or node functions. The lines can be directed or undirected, which reflects the direction of the cluster function flow. Undirected lines mean bidirectional. The basic RBD models of the cluster system include series cluster model, parallel cluster model, series-parallel mixed cluster model and k / n redundant cluster model.

[0156] (1) Series cluster model

[0157] In a cluster system, nodes are mainly connected to other nodes in two ways: parallel or series structure. In the series structure, all nodes in the cluster system must operate normally for the entire cluster system to operate normally. In the parallel structure (redundant structure), at least one node must operate normally to ensure the normal operation of the cluster system. In the discussion of the present invention, all nodes are indispensable, that is, these nodes must operate normally to ensure the continuous operation of the cluster system. According to this concept, if any of the two serial nodes fails, the cluster system will fail.

[0158] For a series cluster system consisting of n independent series nodes, the reliability of the cluster system can be expressed as

[0159] R S (t)=R1(t)×R2(t)×...×R n (t)≤min{R1(t),R2(t),…,R n (t)} (4.7)

[0160] Where 0 <R i (t)<1,i=1,2,…n represents node n i For the series cluster model, it is very important that all nodes have high reliability, especially for a cluster system containing a huge number of nodes.

[0161] (2) Parallel cluster model

[0162] Two or more nodes are connected in parallel, which is also called a redundant structure. This structure means that the cluster system will fail only when all nodes fail. If one or more nodes are operating normally, the cluster system will continue to operate.

[0163] The reliability of a cluster system with n independent nodes in parallel is equal to 1 minus the probability that all n nodes fail (that is, the probability that at least one node works normally). The reliability of the parallel cluster system can be expressed as

[0164] (4.8)

[0165] Where 0 <R i (t)<1,i=1,2,…n represents node n i reliability.

[0166] (3) Series-parallel hybrid cluster model

[0167] like Figure 7 As shown in Figure 1, a hybrid cluster usually contains nodes that are connected in series and in parallel. In order to calculate the reliability of a hybrid cluster, the cluster can be decomposed into series or parallel clusters. If the reliability of each cluster is known, the cluster reliability can be obtained based on the structural relationship between the clusters. A ,c B ,c C The reliability can be expressed as:

[0168] R cA =1-(1-R1)(1-R2)(4.9)

[0169] R cB =R cA (R3)(4.10)

[0170] R cC =R4R5(4.11)

[0171] And cluster c B With cluster c C First connect in parallel, then connect in series with another node R6, the cluster reliability can be expressed as:

[0172] R S =[1-(1-R B )(1-R C )](R6)(4.12)

[0173] (4) k / n redundant cluster model

[0174] The redundant cluster model is a generalization of n parallel nodes. This model requires that at least k of the n identical and independent nodes in the cluster must work properly during their operation phase. The reliability of a k / n redundant cluster can be calculated using the binomial distribution theorem. Assuming that the reliability of each node is R and is independent of each other, the reliability of a k / n redundant cluster can be expressed as follows based on the probability that x nodes can work properly:

[0175] (4.13)

[0176] 1.4.2.2 Cluster System Elasticity

[0177] Resilience is currently widely used to characterize the ability of a cluster to recover after receiving a local attack. Therefore, the present invention introduces cluster resilience to evaluate the cluster's recovery ability, with the aim of improving cluster performance and shortening recovery time. In order to develop an indicator that quantifies cluster resilience, it is necessary to define a parameter that represents the functional performance of the cluster. This parameter may vary from cluster to cluster and is called the cluster performance parameter (Figure of Merit, FOM) in the proposed intelligent operation and maintenance framework. It should be noted that the defined cluster FOM, as a quantitative performance indicator, can be defined as a performance parameter of the actual cluster according to a specific case. In the event of large-scale extreme events such as earthquakes and typhoons, the typical evolution trend of the cluster FOM is as follows: Figure 8 Before an extreme event, cluster performance usually remains at FOM (-) The cluster runs at a level that may be lower than the expected performance level FOM*. The extreme event that occurs at time t1 causes the cluster performance to deteriorate sharply, and stabilizes at the performance degradation value FOM after time t2. D The cluster then performs reconstruction or maintenance from time t2, and after a period of time, the performance is restored and stabilized at FOM (+) . Figure 8 Two different recovery trajectories of the cluster are shown, which may be the result of different strategies adopted by the cluster or the result of different cluster capabilities. Figure 8 It can be observed that the cluster performance in the first trajectory is (1) Restore to FOM (+) In the second track, cluster performance takes longer time t4 (2) To restore to FOM (+) , suggesting that the former are more adaptable to extreme events than the latter.

[0178] from Figure 8 A simple resilience metric can be derived from , defined as the inverse of the time required to recover from an extreme event, as follows:

[0179] (4.14)

[0180] This definition of resilience is similar to the time-to-return used by clusters represented in state space. In addition, the resilience metric can be defined as resilience loss (RL):

[0181] RL=∫ t1 t4 (FOM * -FOM(t))dt(4.15)

[0182] The RL of a cluster quantifies the degree of performance loss of the cluster due to extreme events. The comprehensive metric GR for the resilience of an interdependent infrastructure cluster system can be defined by comprehensively considering the metrics of robustness ROBU, rapidity RAPI, average performance loss per unit time TAPL, and recovery ability RA, as shown below:

[0183] (4.16)

[0184] RAPI r and RAPI d They represent the rapidity measurement indexes of the recovery and destruction stages respectively. All the indicators in the above formula can be obtained as a function of FOM, as shown below:

[0185] (4.17)

[0186] (4.18)

[0187] (4.19)

[0188] (4.20)

[0189] Where K RP is the number of slopes detected by the slope detection technique during the recovery phase. Another comprehensive measure of resilience considers the fault condition F prof and recovery status R prof Merge into:

[0190] (4.21)

[0191] where F prof and R prof The performance of loss and recovery in each stage is shown below:

[0192] (4.22)

[0193] (4.23)

[0194] Where Q S It represents the performance of the cluster. The above indicators can all be used as indicators to measure the static cluster elasticity. For the measurement indicators of elasticity changing over time, for example, a statement of dynamic elasticity in the field of economics is as follows:

[0195] DR=∑ N i=1 FOM DR (t i )-FOM DU (t i ) (4.24)

[0196] Among them, FOM DR and FOM DU Represents the performance with and without cluster elasticity properties. In addition, the elasticity of the cluster to n events can be quantified as follows:

[0197] (4.25)

[0198] Where R(τ) represents the elasticity function, which is a function of time, FOM, and stress induced by extreme events. The dynamic measure of elasticity can be summarized as incorporating the spatial dependence of elasticity into the dynamic measure:

[0199] (4.26)

[0200] Where θ can be time t, or time t and space s, and ρ represents the total performance loss of the cluster. In addition to static, dynamic, and comprehensive metrics, elasticity can also be quantified as a random variable:

[0201] (4.27)

[0202] The variable h(D i ,λ i ,φ) represents event D i The entropy of the probability distribution of an event based on the parameters λ described by φ i The distribution of ρ i (S p ,F r ,F d , F0) represents the elastic factor, which is used as the speed recovery factor S p 、New steady-state factor F r , damage state factor F d and the original steady-state factor F0. The above entropy and elasticity factors can be expressed as:

[0203] (4.28)

[0204] (4.29)

[0205] (4.30)

[0206] where t δ ,t r * ,t r and a dec They represent the relaxation time, the final recovery time, the time to complete the initial recovery action, and the parameter controlling the elastic attenuation.

[0207] The above is a description of the embodiments of the present invention. Through the above description of the disclosed embodiments, professionals and technicians in the field can implement or use the present invention. Various modifications to these embodiments will be obvious to professionals and technicians in the field. The general principles defined in the present invention can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in the present invention, but will conform to the widest range consistent with the principles and novelties disclosed in the present invention.

Claims

1. A cluster system intelligent operation and maintenance system based on reinforcement learning, wherein the cluster system is a power system cluster, a drone cluster system or a large-scale battery cluster system in an industrial energy storage system, characterized in that the system comprises: 1) Cluster modeling module, which is used to establish a multi-level model of the cluster system, including three levels: cluster, cluster and node. The equipment in the cluster system is modeled respectively. Based on the model, the temporal dynamic characteristics and spatiotemporal dynamic characteristics of the cluster system in different operation stages are analyzed, and a cluster system model with multi-dimensional attributes is constructed. 2) Multi-dimensional situation awareness module, which is used to collect and process multi-dimensional situation information of cluster systems in complex environments, including node status, link status, environmental changes and task information, and to conduct real-time monitoring and trend prediction of the operation status of cluster systems by combining temporal dynamic characteristics and spatiotemporal dynamic characteristics; 3) Intelligent decision-making and planning module, which makes intelligent decisions on cluster system operation and maintenance issues based on reinforcement learning algorithms. This module includes: 3.1) Predictive maintenance submodule, including learning the optimal maintenance strategy and generating predictive maintenance solutions based on the deep Q-network algorithm for cluster system degradation and imbalanced maintenance benefits; 3.2) Post-disaster repair submodule, including the Actor-Critic Monte Carlo tree search algorithm, which coordinates the repair sequence and multi-team path planning of the cluster system in a large-scale distributed environment and generates emergency repair plans; 3.3) Dynamic reconstruction submodule, including generating dynamic reconstruction schemes based on distributed proximal strategy optimization algorithm for topological structures at different levels in the cluster system, taking into account task coordination and topological configuration; 4) Cluster performance evaluation module, which is used to construct and calculate the performance indicators of the cluster system, including network efficiency, balance, reliability and elasticity, to quantify the performance of the cluster system at different operation and maintenance stages, and provide feedback and optimization direction for intelligent decision-making planning; The cluster modeling module is established as follows: 1.1) For the cluster system object S under study, first establish a node set N to describe all nodes n in the cluster S, as shown below: N={n1,n2,…,n n ,…}(1.1) in, n n Represents the nth node in the cluster; Before forming a cluster, when some or all nodes in a node set need to be combined into a cluster in a certain order, it is necessary to establish a cluster set C to describe all clusters c in the cluster at the current moment according to factors such as the specific task requirements in the current scenario, as shown below: C={c1,c2,…,c m …}(1.2) Among them, c m Represents the mth cluster in the cluster at the current moment; 1.2) Construct a one-to-many data structure between clusters based on a hash table. The key of the hash table is still the dynamic number of the cluster, and the stored data element is the node set of the cluster. It is constructed using a nested hash table. The key of the nested hash table is the dynamic number of each node in the cluster, and the stored data element is the unique identifier of the node, that is, the number of the node in the node set N. The three-level cluster model SM is constructed by nesting hash tables. III , describing the mapping between the unique identifier of a node and its dynamic number in the cluster, as follows: (1.3) where n (i,j) Represents cluster c i The jth node in ; The multi-dimensional situation awareness module includes: A multi-dimensional situation information collection and processing unit based on complex network theory and perception technology is used to collect the operation data of the cluster system in real time and extract multi-dimensional situation features through data fusion and pattern recognition technology; Based on machine learning and data-driven situation prediction units, a multi-dimensional situation prediction model is constructed for the temporal and spatial dynamic characteristics of the cluster system to provide an estimate of the future situation; The data structure establishment method of the multi-dimensional situation awareness module is as follows: 2.1) Represent a cluster as an undirected graph G with N nodes and K edges. A G graph is described by two matrices: the adjacency matrix A and the physical distance matrix L. A is an N×N adjacency matrix. If there is a connected node n i and node n j The edge of the adjacency matrix is ij is 1, otherwise it is 0; the element l of the matrix L ij Represents node n i and n j The geographical distance between nodes n i and node n j There are no edge elements between them, l ij It is also known; 2.2) For nodes, after creating an empty G graph, the node set N of the cluster system object S can generate the required data. The node set generates different node arrays for addition according to the normal, degraded, and damaged states of the nodes at the current moment; for edges, it is necessary to define the link set L of the cluster system, and then generate edge data and add it to the created graph G; the construction of the cluster system link set L refers to the data structure of the adjacency matrix in the underlying encoding of NetworkX, and on this basis, adds attribute information including operation and maintenance status, and adds the required attribute information according to the actual scenario. The link set L is shown below: (1.4) Where s(a ij ) represents the connection node n i and node n j The state of the edge, element l ij Represents node n i and n j The geographical distance between them.

2. The cluster system intelligent operation and maintenance system according to claim 1, characterized in that: The deep Q network-based algorithm includes the following steps: 3.1.1) Constructing a degradation state model of a cluster system, wherein the cluster system includes a plurality of nodes, and the degradation state of the nodes changes over time during the operation of the system; 3.1.2) Design and train a DQN model to approximate the optimal value function Q, provide an estimate of the value function Q, and evaluate the current cluster system maintenance status feature X; 3.1.3) Estimate the value function Q based on the DQN model, obtain a maintenance state probability π through the ε-greedy algorithm, and then generate the optimal maintenance action a for the current cluster maintenance state characteristics * ; 3.1.4) Execute a series of maintenance actions in the maintenance strategy until the performance of the cluster system is restored to the predetermined threshold; 3.1.5) By providing feedback on the executed maintenance actions, the obtained maintenance data is added to the experience replay buffer, and the experience replay buffer is used to further train the DQN model to optimize the next maintenance decision.

3. The cluster system intelligent operation and maintenance system according to claim 1, characterized in that: The Actor-Critic Monte Carlo tree search algorithm includes the following steps: 3.2.1) Construct the corrective maintenance process of multiple maintenance teams and generate the decision feature tensor X of the cluster system at any time point t t ; 3.2.2) Using Actor-Critic Neural Network to t Evaluate and output the prior parameters, which are used as input to the Monte Carlo tree search algorithm; 3.2.3) Monte Carlo tree search algorithm in X t The search operation of the maintenance action is performed under the conditions, including: i) Select: Change X t Acts as the root node of the Monte Carlo tree search and selects the maintenance action with the best action value based on the upper confidence interval algorithm; ii) Expansion and evaluation: Add leaf nodes in the search tree to the queue, evaluate their strategic value using the ResNet neural network, and update the statistics of related nodes in the search tree; iii) Backtracking: Based on the evaluation results, backtrack along the search path and update the number of visits and action value of each branch on the search path; iv) Execution: By iterating the above steps, the improved maintenance state transition probability π is obtained and the global optimal maintenance action a is selected t * The group system transfers from time point t to t+1; 3.2.4) Repeat steps 3.2.1) to 3.2.3) until the corrective maintenance task of the cluster system is completed.

4. The cluster system intelligent operation and maintenance system according to claim 1, characterized in that: The distributed proximal strategy optimization algorithm includes the following steps: 3.3.1) Construct a dynamic reconstruction decision framework for cluster systems: Analyze the task coordination and topological configuration characteristics of the three-level spatiotemporal dynamic cluster system, construct a three-level dynamic reconstruction decision framework of cluster-cluster-node, and analyze the topological configuration characteristics between clusters within the cluster and between nodes within the cluster; 3.3.2) Multi-dimensional situation feature extraction: Design a deep neural network model, use the ResNet module to process the task area features, business coverage features, node location features and damage area features, apply the LSTM module to extract the temporal features of high-dimensional information, and use the Attention mechanism to focus on the internal and cross-level interactions of the cluster to extract multi-dimensional situation features; 3.3.3) DPPO-based reinforcement learning algorithm model: Apply the policy gradient algorithm of the dominant actor-critic and combine it with the experience replay technology for asynchronous update. By maximizing the expected policy reward, the neural network is updated using TD (λ), V-trace and UPGO training to achieve dynamic reconstruction strategy optimization. 3.3.4) Design of alliance learning model: Design alliance games and virtual self-learning mechanisms, including three types of agents: master agent, master explorer and alliance explorer, conduct distributed learning training, and optimize dynamic reconstruction strategies; 3.3.5) Dynamic reconstruction strategy generation: In the dynamic reconstruction process, the dynamic reconstruction agent is used to generate a reconstruction action set based on the multi-dimensional situation information of the cluster system, and redeploy the cluster system topology configuration to restore the task performance.

5. The cluster system intelligent operation and maintenance system according to claim 1, characterized in that: The cluster effectiveness evaluation module includes: a) Construction of performance index system: Construct a performance evaluation index system based on the general and special characteristics of the cluster system. The general characteristics include reliability, maintainability, security, safety and environmental adaptability, and the special characteristics include time dynamic characteristics and space-time dynamic characteristics; b) Data collection: The real-time operation data of the cluster system is collected through the multi-dimensional situation awareness module, and the data includes node health status, link status, task execution status and environmental change information; c) Evaluation of the network efficiency of the cluster system: The network average efficiency formula E(G) is used to evaluate the efficiency of information transmission between nodes in the cluster system. The formula is: ; Where N is the number of nodes, w i and w j are the weights of node i and node j respectively, d ij is the shortest path length between node i and node j; d) Task execution balance evaluation: For the three-level spatiotemporal dynamic cluster, the cluster balance ε is used b The balance of the task execution capability of the cluster system is calculated by the following formula: ; Where I represents the total number of clusters in cluster S, J i Represents cluster c i The total number of nodes, ε ij Represents node n (i,j) The sum of the overlapping areas of the coverage area of ​​​​the node and the coverage areas of other nodes in the same cluster within the task area; is the average redundancy value of all nodes in each cluster.

6. The cluster system intelligent operation and maintenance system according to claim 1, characterized in that: The system comprises: Open interfaces for interacting with external systems, supporting the integration and expansion of different types of reinforcement learning methods in the intelligent operation and maintenance framework; The multi-dimensional situation data management module based on big data analysis is used to store, manage and analyze the historical operation data and real-time collection data of the cluster system, providing data support for multi-dimensional situation awareness and intelligent decision-making.

Citation Information

Patent Citations

  • Multi-agent reinforcement learning method and system based on hierarchical consistency learning

    CN114118374A

  • Multi-agent cooperative computing resource scheduling method, device and system

    CN116909742A