A performance evaluation method and device for intelligent operation and maintenance of cluster systems

By building a mathematical model of the cluster system and introducing deep reinforcement learning algorithms, the complexity and insufficient accuracy of cluster system performance evaluation are solved, and efficient, precise management and intelligent operation and maintenance decisions of the cluster system are realized.

CN119127650BActive Publication Date: 2025-05-13HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411595397.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-05-13
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

The existing cluster system performance evaluation methods have the complexity of multi-dimensional situational awareness, insufficient accuracy of performance evaluation, and lack of elastic evaluation, making it difficult to accurately evaluate the operating status and recovery capabilities of the cluster system.

Method used

A method of intelligent operation and maintenance efficiency evaluation of cluster systems based on mathematical models is proposed. By building a reasonable performance index system, network average efficiency calculation model, task execution balance evaluation model and elastic evaluation model, the efficiency of cluster systems is scientifically and accurately evaluated, and a deep reinforcement learning algorithm is introduced for real-time situation information analysis.

Benefits of technology

It improves the overall performance evaluation accuracy of the cluster system, optimizes the efficiency of intelligent operation and maintenance decision-making, enhances the system's flexibility and recovery capabilities, and improves the overall efficiency of task execution capabilities and operation and maintenance management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119127650B_ABST
    Figure CN119127650B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and in particular to a performance evaluation method and device for intelligent operation and maintenance of cluster systems. The present invention can scientifically and accurately evaluate the performance of cluster systems at different operation stages by constructing a reasonable performance indicator system, a network average efficiency calculation model, a task execution balance evaluation model, and a flexibility evaluation model. The present invention can effectively improve the efficiency of intelligent operation and maintenance decision-making of cluster systems by introducing a deep reinforcement learning algorithm to collect, analyze, and predict multi-dimensional situation information in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a performance evaluation method and device for intelligent operation and maintenance of a cluster system. Background Art

[0002] With the rapid development of modern information and intelligent technology, cluster systems, as a complex system composed of large-scale equipment, have been widely used in many fields, such as energy, transportation, communications, and national defense. Cluster systems are usually composed of multiple devices or nodes that cooperate with each other. Their complex structure and wide application environment lead to the high complexity of operation and maintenance. In order to ensure the continuous and efficient operation of cluster systems, intelligent operation and maintenance technology has become an important research direction in the field of cluster systems.

[0003] Traditional cluster system operation and maintenance mainly relies on manual monitoring and scheduled maintenance cycles. Common operation and maintenance modes include regular maintenance and post-event maintenance. However, these modes have obvious shortcomings: 1. Post-event maintenance: When the cluster system fails, it is repaired, which often leads to long system downtime, increasing system maintenance costs and risks. 2. Regular maintenance: The preset maintenance cycle cannot flexibly respond to the operating status of the equipment in the cluster system, which is prone to over-frequent or insufficient maintenance.

[0004] In order to improve the efficiency and reliability of cluster system operation and maintenance, intelligent predictive maintenance and real-time performance evaluation methods have gradually become research hotspots. In recent years, with the development of technologies such as big data, artificial intelligence, and deep learning, intelligent decision-making technologies based on reinforcement learning and deep learning have been gradually applied to the intelligent operation and maintenance of cluster systems. By building a multi-dimensional situational awareness and performance evaluation model for cluster systems, potential problems can be predicted in time before the system fails, and corresponding operation and maintenance decisions can be made to maximize the reliability and efficiency of system operation.

[0005] Despite this, existing cluster system performance evaluation methods still face some key problems and challenges: 1. Complexity of multi-dimensional situational awareness: There is a high degree of coordination and uncertainty between the devices in the cluster system. There are certain technical challenges in real-time monitoring of the multi-dimensional situation information in the system (such as node status, task execution status, environmental changes, etc.) and conducting comprehensive analysis. 2. Insufficient accuracy of performance evaluation: Due to the high complexity of cluster systems, traditional performance evaluation methods are difficult to accurately evaluate the operating status of the system, especially in the case of dynamic changes in large-scale cluster systems. There is a lack of scientific and reasonable performance evaluation models. 3. Lack of resilience assessment: When facing emergencies or failures, how to evaluate the resilience and recovery capabilities of cluster systems in order to quickly formulate repair and reconstruction plans is an important difficulty in existing technologies.

[0006] In the early stage, the applicant aimed at the problem of "predictive maintenance" of three-level time dynamic cluster system under the condition of system performance degradation in the long-term operation stage. By comprehensively weighing the system degradation state and the imbalance characteristics of maintenance benefits, a predictive maintenance plan was generated through a method based on DQN deep reinforcement learning. For details, please refer to the Chinese invention patent application (application number: 202411357728X, application date: 20240927) disclosed a predictive maintenance decision method for cluster system based on DQN; for the problem of "post-disaster repair" of multiple teams in a two-level time dynamic cluster system facing large-scale local damage, a corrective maintenance plan was generated by coordinating the two-level decisions of maintenance timing and path planning through an Actor-Critic deep reinforcement learning method. For details, please refer to the Chinese invention patent application (application number: 2024112822477, application date: 20240913), which discloses a method for post-disaster restorative maintenance of cluster systems based on the AC-MCTS algorithm; for the "dynamic reconstruction" problem of a three-level spatiotemporal dynamic cluster system with multi-formation characteristics, a two-level dynamic reconstruction of functional reconstruction between nodes within the three-level cluster system and functional reconstruction between clusters is analyzed through a method based on DPPO deep reinforcement learning to generate a dynamic reconstruction plan. For details, please refer to the Chinese invention patent application (application number: 2024114796053, application date: 20241023), which discloses a dynamic reconstruction decision method for cluster systems based on DPPO deep reinforcement learning. Summary of the invention

[0007] In order to solve the above problems, the present invention proposes a cluster system intelligent operation and maintenance efficiency evaluation method based on a mathematical model, which can scientifically and accurately evaluate the efficiency of the cluster system at different operation stages by constructing a reasonable efficiency index system, a network average efficiency calculation model, a task execution balance evaluation model and a flexibility evaluation model. The present invention introduces a deep reinforcement learning algorithm to collect, analyze and predict multi-dimensional situation information in real time, which can effectively improve the efficiency of intelligent operation and maintenance decision-making of the cluster system.

[0008] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions:

[0009] A performance evaluation method for intelligent operation and maintenance of a cluster system, wherein the cluster system is a power system cluster, a drone cluster system, or a large-scale battery cluster system in an industrial energy storage system, and the method comprises the following steps:

[0010] a) Construction of performance index system: Construct a performance evaluation index system based on the general and special characteristics of the cluster system. The general characteristics include reliability, maintainability, security, safety and environmental adaptability, and the special characteristics include time dynamic characteristics and space-time dynamic characteristics;

[0011] b) Data collection: The real-time operation data of the cluster system is collected through the multi-dimensional situation awareness module, and the data includes node health status, link status, task execution status and environmental change information;

[0012] c) Evaluation of the network efficiency of the cluster system: The network average efficiency formula E(G) is used to evaluate the efficiency of information transmission between nodes in the cluster system. The formula is:

[0013] ;

[0014] Where N is the number of nodes, w i and w j are the weights of node i and node j respectively, d ij is the shortest path length between node i and node j;

[0015] d) Task execution balance evaluation: For the three-level spatiotemporal dynamic cluster, the cluster balance ε is used b Evaluate the balance of cluster system task execution capability. The balance calculation formula is:

[0016] ;

[0017] Where I represents the total number of clusters in cluster S, J i Represents cluster c i The total number of nodes, is the average redundancy value of all nodes in each cluster; ε ij Represents node n (i,j) The sum of the overlapping areas of the coverage area of ​​the node and the coverage areas of other nodes in the same cluster within the task area.

[0018] Preferably, the network efficiency evaluation adopts normalization processing to normalize the average efficiency E(G) to 0≤E(G)≤10, and when there is an edge between each pair of nodes in the cluster system, E(G)=1.

[0019] As a preference, the balance evaluation of the cluster system is performed by using a redundancy matrix M i ε ={ε ij}Calculate the overlapping area of ​​cluster tasks and calculate the task capacity of each cluster based on the task distribution; the element ε in the matrix ij Represents node n= (i,j) The sum of the overlapping areas of the coverage area of ​​the node and the coverage areas of other nodes in the same cluster within the task area.

[0020] Preferably, the performance evaluation method further includes a network connectivity analysis based on time dynamic characteristics, by constructing an adjacency matrix A={a ij} and the distance matrix moment L = {lij}, analyze the network connectivity status of the cluster system at different stages.

[0021] Preferably, the reliability of the cluster system is calculated by a reliability block diagram model, wherein the model includes a series connection, a parallel connection, a hybrid connection or a k / n redundant cluster model, and the reliability formula is:

[0022] a) Reliability of the series structure R S (t) Formula:

[0023] R S (t)=R1(t)×R2(t)×...×R n (t)≤min{R1(t),R2(t),…,R n (t)};

[0024] Where 0 <R i (t)<1,i=1,2,…n represents node n i reliability;

[0025] b) Reliability of parallel structure R S (t) Formula:

[0026] R S (t)=1-∏ n i=1 [1-R i (t)];

[0027] Where 0 <R i (t)<1,i=1,2,…n represents node n i reliability;

[0028] c) In a hybrid cluster, cluster c A , c B , c C The reliability of is expressed as:

[0029] R cA =1-(1-R1)(1-R2);

[0030] R cB =R cA (R3);

[0031] R cC =R4R5;

[0032] And cluster c B With cluster c C First connect in parallel, then connect in series with another node R6, the cluster reliability is expressed as:

[0033] R S =[1-(1-R B)(1-R C )](R6);

[0034] d) k / n redundant cluster model

[0035] Assuming that the reliability of each node is R and they are independent of each other, based on the probability that x nodes can operate normally, the reliability of the k / n redundant cluster is expressed as:

[0036] .

[0037] Preferably, the resilience assessment of the cluster system includes calculating the inverse of the recovery time to obtain the resilience metric R of the system recovering from a failure. FOM :

[0038] ;

[0039] It is characterized in that t1 is the time when the fault occurs, and t4 is the time when it returns to the normal state;

[0040] The elasticity of the cluster system is evaluated by the elasticity loss formula RL, which is:

[0041] RL=∫ t1 t4 (FOM * -FOM(t))dt;

[0042] It is characterized in that FOM is the target performance and FOM(t) is the performance value at time t.

[0043] As a preferred method, the comprehensive metric GR for the resilience of the interdependent infrastructure cluster system can be defined by comprehensively considering the metrics of robustness ROBU, rapidity RAPI, average performance loss per unit time TAPL, and recovery ability RA, as shown below:

[0044] ;

[0045] RAPI r and RAPI d They represent the rapidity measurement indexes of the recovery and destruction stages respectively. All the indicators in the above formula can be obtained as a function of FOM, as shown below:

[0046] ;

[0047] ;

[0048] ;

[0049] ;

[0050] Its K RP is the number of slopes detected by the slope detection technique during the recovery phase;

[0051] Another comprehensive measure of resilience is to consider the fault condition F prof and recovery status R prof Merge into:

[0052] ;

[0053] where F prof and R prof The performance of loss and recovery in each stage is shown below:

[0054] ;

[0055] ;

[0056] Where Q S It represents the performance of the cluster.

[0057] As a preferred option, the performance evaluation results of the cluster system are used to generate intelligent operation and maintenance decisions and guide operation and maintenance plans for predictive maintenance, post-disaster repair and dynamic reconstruction.

[0058] Preferably, the method performs a comprehensive analysis of data from multiple operation and maintenance stages, outputs performance reports of the cluster system in different stages such as normal, degradation, destruction and recovery, and provides a basis for subsequent optimization.

[0059] Furthermore, the present invention discloses a performance evaluation device for intelligent operation and maintenance of a cluster system, wherein the device executes the method, including:

[0060] a) A data collection module, which is used to collect the operation data of the cluster system in real time, including node health status, link status, task execution status and environmental change information;

[0061] b) an efficiency index construction module, which is used to construct an efficiency evaluation index system for the cluster system. The index is based on the general characteristics and special characteristics of the cluster system to construct an efficiency evaluation index system. The general characteristics include reliability, maintainability, security, safety and environmental adaptability, and the special characteristics include time dynamic characteristics and space-time dynamic characteristics;

[0062] c) an efficiency calculation module, used to calculate and analyze the current efficiency of the cluster system based on the efficiency index;

[0063] d) Result output module, which is used to generate performance reports of the cluster system and provide feedback and optimization suggestions for intelligent operation and maintenance decisions.

[0064] The present invention adopts the above technical solution, combines mathematical models and reinforcement learning algorithms, and effectively solves the shortcomings of traditional cluster system operation and maintenance methods through accurate performance evaluation and intelligent decision-making, and realizes efficient and accurate management of cluster systems. Its specific technical effects include the following aspects:

[0065] 1. Improve the overall performance evaluation accuracy of the cluster system: This invention can accurately quantify the information transmission efficiency and task execution capability of the cluster system by constructing a mathematical model based on the average network efficiency and task execution balance. b , can accurately reflect the performance of the cluster system under different operating conditions, including normal operation, failure occurrence, unbalanced task allocation, etc. Compared with the traditional simple indicator evaluation method, this method is more accurate and comprehensive in the dynamic environment.

[0066] 2. Optimize the efficiency of intelligent operation and maintenance decision-making: By introducing deep reinforcement learning algorithms (such as DQN, DPPO, etc.), the present invention can predict the situation changes of cluster systems in real time and generate the optimal operation and maintenance decision-making plan for different operation and maintenance scenarios (such as predictive maintenance, post-disaster repair, dynamic reconstruction). Through real-time analysis of multi-dimensional situation information such as node health status and link status, the present invention can formulate response plans in advance according to future situation change trends, avoid sudden failures, and improve the system's response speed and decision-making quality.

[0067] 3. Enhance the elasticity and recovery capability of cluster systems: This invention proposes an evaluation model based on elastic loss degree RL, which can effectively evaluate the recovery capability of cluster systems after emergencies. By calculating the recovery time RFOM and recovery performance loss of the cluster system after a failure, the system can quickly quantify the impact of the failure on the overall performance and guide the formulation of a rapid recovery strategy. This evaluation model can significantly improve the self-healing ability and risk resistance of cluster systems in emergencies, and reduce system downtime and losses.

[0068] 4. Improving the task execution capability of large-scale cluster systems: The present invention uses the task execution balance formula ε b The task execution capability of the cluster system is evaluated to ensure the reasonable allocation and balanced execution of tasks among different clusters, avoiding task delays or system performance degradation caused by overload or idleness of some nodes or clusters. This method is particularly suitable for three-level spatiotemporal dynamic cluster systems, which can dynamically adjust task allocation and improve the overall execution efficiency of the system.

[0069] 5. Provide a comprehensive performance evaluation report to support operation and maintenance optimization: The performance evaluation method of the present invention can generate a system performance report through multi-dimensional and comprehensive evaluation, covering the performance of the cluster system in the normal, degradation, destruction and recovery stages. The report can provide reliable data support for the operation and maintenance team, help optimize subsequent operation and maintenance strategies, and improve the long-term stability and operational efficiency of the cluster system.

[0070] 6. Comprehensively consider the general and special characteristics of cluster systems: This method not only evaluates the general characteristics of cluster systems (such as reliability, maintainability, environmental adaptability, etc.), but also combines the special characteristics of cluster systems (such as time dynamic characteristics and space-time dynamic characteristics), providing a more comprehensive perspective for performance evaluation. This multi-level evaluation method ensures that the system can maintain stable and efficient operation in a variety of application scenarios.

[0071] In summary, the present invention has demonstrated significant technical effects in improving the accuracy of cluster system performance evaluation, optimizing intelligent operation and maintenance decisions, and enhancing system elasticity and recovery capabilities. It can be widely used in cluster system operation and maintenance management in multiple fields to improve the overall system performance and intelligence level. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 This is the Agent-Environment architecture diagram for the intelligent operation and maintenance of the cluster system.

[0073] Figure 2 This is a diagram of the cluster effectiveness evaluation index system.

[0074] Figure 3 It is a series-parallel hybrid cluster diagram.

[0075] Figure 4 Figure 2 shows the evolution of cluster performance under extreme events and recovery measures. DETAILED DESCRIPTION

[0076] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0077] The intelligent decision-making planning module of the cluster system intelligent operation and maintenance framework of the present invention is designed based on reinforcement learning theory. Cluster intelligent operation and maintenance includes three main problems: preventive maintenance, corrective maintenance, and operation management. Each type of problem still includes various representative frontier problems. For various representative operation and maintenance problems, reinforcement learning agents can be designed according to specific needs to support the decision-making of intelligent operation and maintenance. The various operation and maintenance decision agents designed all need to maximize the specific operation and maintenance benefits and learn which specific operation and maintenance actions to perform - how to map the operation and maintenance state space to the operation and maintenance action space. The operation and maintenance decision agent is not told which operation and maintenance actions to take, but must try to find out which operation and maintenance actions can generate the greatest operation and maintenance benefits. This is the core idea of ​​designing the operation and maintenance decision method based on reinforcement learning, that is, the idea of ​​trial and error search. For the most challenging representative frontier operation and maintenance problems, any operation and maintenance action selected during the trial and error search process will not only affect the immediate benefits in the operation and maintenance scenario, but also affect the next operation and maintenance state, thereby affecting all potential future operation and maintenance delay benefits in the operation and maintenance scenario. In short, the trial-and-error search of operation and maintenance actions and the delayed benefits of operation and maintenance situations will be the core elements in designing various operation and maintenance decision-making agents.

[0078] In addition, since the reinforcement learning problem is based on reward-driven completion of the agent-environment interaction, thereby realizing the perception-action-learning cycle, the design of the operation and maintenance decision-making agent requires the construction of a simulation environment for the intelligent operation and maintenance of the cluster system, providing the designed operation and maintenance decision-making agent with multi-dimensional situation observation information. In the previous section, the multi-dimensional situation awareness module has completed the construction of the relevant intelligent operation and maintenance framework, which can support the provision of operation and maintenance situation simulation information Obversations (εnv) as state observation input for the operation and maintenance decision-making agent. In addition, to realize the perception-action-learning cycle in the scenario of cluster intelligent operation and maintenance, it is necessary to describe various MDPs in various intelligent operation and maintenance scenarios. The core of the sequential decision-making architecture of intelligent operation and maintenance scenarios is the evaluative feedback of operation and maintenance decisions. That is, the operation and maintenance actions in the sequential decision-making need to consider not only the immediate benefits in the current operation and maintenance scenario, but also the potential development of multi-dimensional situations in the operation and maintenance scenario, thereby leading to fluctuations in benefits in future operation and maintenance scenarios. Therefore, considering the importance of timely benefits and delayed benefits to operation and maintenance decision-making planning, the design of the value function of operation and maintenance actions is particularly important. The value function of operation and maintenance actions needs to be designed based on the benefits of operation and maintenance. It is necessary to build an indicator system to evaluate the effectiveness of the cluster under various operation and maintenance situations. This will be introduced in the next section on the cluster effectiveness evaluation module. In summary, to design an intelligent decision-making and planning module based on reinforcement learning, it is necessary to build an intelligent operation and maintenance simulation environment for the cluster system, design decision-making and planning algorithm agents for various intelligent operation and maintenance scenarios, and consider the operation and maintenance situation observation and operation and maintenance decision benefits in the operation and maintenance simulation environment. The cluster system intelligent operation and maintenance agent-environment interaction interface is sorted out, such as Figure 1 shown.

[0079] In the intelligent decision-making and planning module, the designed operation and maintenance decision agent mainly targets three types of problems: preventive maintenance, corrective maintenance, and operation management. It is necessary to sort out the input-output relationship between various types of operation and maintenance problems and the cluster system modeling module, the multi-dimensional situation awareness module, and the cluster effectiveness evaluation module. Among the three main types of operation and maintenance problems, we can focus on three key representative frontier problems and carry out research on operation and maintenance decision-making methods based on deep reinforcement learning. In view of the "predictive maintenance" problem of the three-level time dynamic cluster system under the condition of system performance degradation in the long-term operation stage, a method based on DQN deep reinforcement learning is used to comprehensively weigh the system degradation state and the imbalance characteristics of maintenance benefits to generate a predictive maintenance plan. For details, please refer to a DQN-based cluster system predictive maintenance decision method disclosed in the Chinese invention patent application (application number: 202411357728X, application date: 20240927), which will not be described in detail in this specification; for the "post-disaster repair" problem of multiple teams in a two-level time dynamic cluster system facing large-scale local damage, a two-level decision of maintenance timing and path planning is coordinated through an Actor-Critic deep reinforcement learning method to generate a corrective maintenance plan. For details, please refer to the Chinese invention patent application (application number: 202411357728X, application date: 20240927). The patent application (application number: 2024112822477, application date: 20240913) discloses a method for post-disaster restorative maintenance of cluster systems based on the AC-MCTS algorithm, which will not be described in detail in this specification; for the "dynamic reconstruction" problem of a three-level spatiotemporal dynamic cluster system with multi-formation characteristics, a two-level dynamic reconstruction of the functional reconstruction between nodes within the cluster of the three-level cluster system and the functional reconstruction between clusters is analyzed by a method based on DPPO deep reinforcement learning to generate a dynamic reconstruction plan, for details, see the Chinese invention patent application (application number: 2024114796053, application date: 20241023) discloses a cluster system dynamic reconstruction decision method based on DPPO deep reinforcement learning, which will not be described in detail in this specification. In addition, the designed intelligent decision-making planning module still supports future researchers to integrate richer operation and maintenance decision-making methods, and ultimately provides a high-precision, high-efficiency, and low-cost intelligent operation and maintenance solution system for various cluster systems in different operation and maintenance scenarios.

[0080] 1.1 Cluster performance evaluation module

[0081] The present invention discusses the cluster effectiveness evaluation module in the intelligent operation and maintenance framework. In the proposed intelligent operation and maintenance framework, the cluster effectiveness evaluation module is intended to evaluate the cluster effectiveness under the multi-dimensional situation of the cluster system operation stage, so as to guide the intelligent decision-making planning module to perform iterative optimization. The effectiveness of the cluster system refers to the ability of the entire cluster system to complete the specified tasks under the specified complex environment and the considered organizational, tactical, survival, and support conditions. It reflects the quality of the cluster system and its ability to complete various tasks. It is comprehensively determined by its general characteristics and special characteristics to make a comprehensive and correct assessment of its quality and ability to complete tasks. The main factors affecting the cluster effectiveness are general characteristics and special characteristics. Among them, the general characteristics include the reliability, maintainability, supportability, testability, safety, durability, human factors and environmental adaptability of the cluster. Special characteristics are divided into two categories: time dynamic characteristics and time-space dynamic characteristics according to the characteristics of the cluster system. Based on the above analysis, we further construct a cluster system performance evaluation index system. The construction of the index system is the basis for studying cluster effectiveness evaluation. By analyzing the task requirements of each task stage of the cluster system, gathering the typical task processes, and comprehensively analyzing the special characteristics and general characteristics of the cluster system, we can summarize the evolution law of typical task functions. On this basis, we select relevant cluster system index items, and analyze and discuss the various performance influencing factors of the cluster system. These factors are not independent of each other, and the complex correlations between them need to be sorted out. Finally, we establish an index system for cluster effectiveness evaluation, such as Figure 2 As shown in the figure, including the general characteristics and special characteristics of cluster systems, the performance indicators involved are quite rich. This paper conducts relevant research on the performance indicators used in the operation and maintenance decision-making problems in subsequent chapters and gives evaluation algorithms. The performance evaluation module developed based on Python language and integrated into the intelligent operation and maintenance framework is completed. The proposed intelligent operation and maintenance framework still supports the expansion of related performance evaluation algorithm research according to the specific operation and maintenance problem requirements.

[0082] 1.1.1 Cluster System-Specific Features

[0083] In the cluster system modeling module, the time dynamic characteristics and the time-space dynamic characteristics are described according to the characteristics of the cluster system. Therefore, the present invention also divides the special characteristics of the cluster system into time dynamic characteristics and time-space dynamic characteristics, and takes the network efficiency of the two-level time dynamic cluster and the balance of the three-level time-space dynamic cluster as examples of the performance evaluation algorithm for detailed introduction.

[0084] 1.1.1.1 Cluster system network efficiency

[0085] For two-level time dynamic clusters, when considering the status of all nodes in the node set, it is also necessary to consider the status of all links between nodes. At this time, the network efficiency of the cluster system is a key indicator for evaluating its effectiveness, which can measure the efficiency of information exchange or energy transmission between nodes in the cluster system. Based on the definition of node set N and link set L in the cluster modeling module and the multi-dimensional situational awareness module and the construction of their data structure, the cluster is represented as an undirected G graph with N nodes and K edges, and its adjacency matrix A and physical distance matrix L are defined. In addition, an important parameter of G is the node n i The degree of node n i The number of relevant edges k i , that is, n i The number of neighbors, k i The average value of is k=2K / N. When the adjacency matrix A is known, the distance between any two nodes n can be calculated. i and node n j The shortest path length d between ij Assuming G is connected, for , d ij exists and is positive. To quantify the structural properties of G, two different parameters are usually evaluated: characteristic path length L and clustering coefficient C. L is the average distance between two nodes. , C describes the local characteristics of the network and can be defined as , where C i Represents node n i The adjacent node subgraph G i The number of edges in the system divided by the maximum number of edges allowed is k i (k i -1) / 2. By representing the real network as a weighted G graph, the G graph may even be non-sparse and non-connected. Such a G graph requires two matrices to describe: the adjacency matrix A={a ij}, which is defined the same as for the unweighted graph, and the physical distance matrix L = {l ij}. Matrix element l ij It can be the spatial distance between two nodes or the strength of their interaction: even if node n i and node n j There is no edge between them, so we can also assume that l ij is known. For example, ij It can be the geographical distance between nodes in the power system cluster. i and node n j The shortest path length d between ij is the node n in the G graph i and node n jThe minimum physical distance of all possible paths. Therefore, the shortest path length matrix D = {d ij} is obtained by using the matrix A={a ij} and the matrix L={l ij}, it can be defined as:

[0086] (1.1)

[0087] When node n i and node n j When there is an edge between ij ≥l ij , Node n i and node n j The efficiency between ij It can be defined as:

[0088] (1.2)

[0089] where w i is node n i The present invention uses the degree of the node as the weight. In addition, the weight of the node can be defined according to the specific problem. i and node n j If there is no path between ij =+∞ and∈ ij =0, this situation applies to two scenarios: in the initial state, there is no path between two nodes in the G graph; in the state of large-scale local damage, all paths between two nodes in the G graph fail, resulting in no path. In these two scenarios, the G graph is a non-connected graph. In summary, the average efficiency of G can be defined as:

[0090] (1.3)

[0091] In order to normalize the average efficiency E, consider the ideal complete graph G ideal With all possible edges in G, the total number of edges is N(N-1) / 2. In this ideal case, information or energy is transmitted in the most efficient way, because in this ideal case d ij =l ij , and the average efficiency E is at its maximum value at this time The efficiency E(G) considered in the rest of this paper is always divided by E(G ideal), so 0≤E(G)≤1. Although E(G)=1 exists and makes sense when there is an edge between every pair of nodes. However, in reality, due to the large number of nodes in the cluster system, it is almost difficult to satisfy the requirement that there is a link between every pair of nodes. Therefore, it is difficult for the network model of the cluster system to achieve an average efficiency close to 1. In addition, formula (1.3) defines the global efficiency of the G graph, so it can be called E glob The data structures of the node set N and link set L constructed in the multi-dimensional situational awareness module are based on the data structures of the underlying code of the NetworkX tool. The nodes and links are further endowed with operation and maintenance status attributes, which can generate operation and maintenance status information at different operation stages to simulate and analyze the connected graph of the cluster system in a normal state and the disconnected graph in a damaged state. The network average efficiency algorithm of the NetworkX tool is improved to evaluate the network efficiency of the cluster system for specific operation and maintenance status.

[0092] 1.1.1.2 Cluster system balance

[0093] For the three-level spatiotemporal dynamic cluster, due to the different number of nodes in each cluster, there are certain differences in the overall performance of each cluster during the task execution process. In order to ensure that the cluster has the ability to continuously execute tasks and adapt to environmental uncertainties, in the process of cluster configuration and deployment within the cluster, multiple clusters must have relatively balanced task capabilities. For this reason, the present invention defines the balance degree to quantify the balanced task capabilities, and then realizes the evaluation of the performance of the three-level spatiotemporal dynamic cluster. This paper mainly analyzes the detection capability of the cluster system and defines the balance degree of the cluster system. For the task capability of each node, assuming that its coverage area is a circular area, cluster c i The jth node n in (i,j) The coordinates of the node are the center coordinates of the area covered by the node, which can be expressed as (x ij ,y ij ). Define the task area coverage of cluster S Cov S , as shown below:

[0094] Cov S =[cov1,…,cov i ,…](1.4)

[0095] Where cov i Represents cluster c i The coverage rate of the task area is obtained by dividing the total area covered by all nodes in the cluster by the area of ​​the task area. Define cluster c i The redundancy matrix M i ε ={ε ij} represents the repeated coverage of the task area of ​​the cluster, and the element ε in the matrixij Represents node n (i,j) The sum of the overlapping areas of the coverage area of ​​the node and the coverage areas of other nodes in the same cluster within the task area. Taking into account the node distribution characteristics and the overall performance level of the cluster, in order to avoid the situation where a cluster has insufficient capacity redundancy due to taking on too many tasks, or a cluster has excessive capacity redundancy due to taking on too few tasks, thereby reducing the efficiency of cluster task execution, the cluster balance degree ε is defined. b , which represents the mean square error of the redundancy values ​​of all nodes that can work normally in each cluster during the task execution phase, as shown below:

[0096] (1.5)

[0097] Where I represents the total number of clusters in cluster S, J i Represents cluster c i The total number of nodes, is the average redundancy value of all nodes in each cluster. On the basis of defining the cluster balance, in order to facilitate the selection of the optimal action in the subsequent operation and maintenance decision-making problem, it is necessary to know which nodes in the cluster are damaged and how to take corresponding operation and maintenance actions, which has the greatest impact on the recovery of cluster balance. To this end, the reconstruction process for local damage during the three-level spatiotemporal dynamic cluster task is considered. On the basis of defining the cluster system balance calculation model, the cluster loss degree is introduced to characterize the recovery effect of the overall cluster balance in the operation and maintenance stage, and the loss degree of cluster S in the reconstruction stage is defined as S for:

[0098] (1.6)

[0099] where t init represents the initial time of reconstruction, t init represents the reconstruction recovery time threshold, ε b thr Indicates the cluster balance threshold, loss S (t) represents the real-time loss value of cluster balance at time t in the reconstruction process, ε b (t) represents the real-time balance degree of the cluster at time t in the reconstruction process. In the reconstruction process facing local damage during the cluster task, it is necessary to minimize the loss of balance capacity ζ S To optimize the goal and meet the constraints of reconstruction time threshold.

[0100] 1.1.2 General features of cluster systems

[0101] The common characteristics of cluster systems include reliability, maintainability, security, testability, safety, environmental adaptability and elasticity of the cluster. The reliability and elasticity of the cluster system are introduced in detail as an example of the performance evaluation algorithm.

[0102] 1.1.2.1 Cluster System Reliability

[0103] The present invention applies the reliability block diagram (RBD) modeling method to perform reliability modeling on the cluster system. RBD is the most basic reliability model. Its basic idea is to describe the logical relationship between node functions and cluster functions based on the functional correlation between nodes in the cluster system. The reliability block diagram of the cluster system consists of boxes, lines and logical relationships representing node objects or node functions. The lines can be directed or undirected, which reflects the direction of the cluster function flow. Undirected lines mean bidirectional. The basic RBD models of the cluster system include series cluster model, parallel cluster model, series-parallel mixed cluster model and k / n redundant cluster model.

[0104] (1) Series cluster model

[0105] In a cluster system, nodes are mainly connected to other nodes in two ways: parallel or series structure. In the series structure, all nodes in the cluster system must operate normally for the entire cluster system to operate normally. In the parallel structure (redundant structure), at least one node must operate normally to ensure the normal operation of the cluster system. In the discussion of the present invention, all nodes are indispensable, that is, these nodes must operate normally to ensure the continuous operation of the cluster system. According to this concept, if any of the two serial nodes fails, the cluster system will fail.

[0106] For a series cluster system consisting of n independent series nodes, the reliability of the cluster system can be expressed as

[0107] R S (t)=R1(t)×R2(t)×...×R n (t)≤min{R1(t),R2(t),…,R n (t)}(1.7)

[0108] Where 0 <R i (t)<1,i=1,2,…n represents node n i For the series cluster model, it is very important that all nodes have high reliability, especially for a cluster system containing a huge number of nodes.

[0109] (2) Parallel cluster model

[0110] Two or more nodes are connected in parallel, which is also called a redundant structure. This structure means that the cluster system will fail only when all nodes fail. If one or more nodes are operating normally, the cluster system will continue to operate.

[0111] The reliability of a cluster system with n independent nodes in parallel is equal to 1 minus the probability that all n nodes fail (that is, the probability that at least one node works normally). The reliability of the parallel cluster system can be expressed as

[0112] (1.8)

[0113] Where 0 <R i (t)<1,i=1,2,…n represents node n i reliability.

[0114] (3) Series-parallel hybrid cluster model

[0115] A hybrid cluster usually contains nodes that are both serial and parallel. To calculate the reliability of a hybrid cluster, the cluster can be decomposed into serial or parallel clusters. If the reliability of each cluster is known, the cluster reliability can be calculated based on the structural relationship between the clusters. Figure 3 In the hybrid cluster shown in the figure, cluster c A ,c B ,c C The reliability can be expressed as:

[0116] R cA =1-(1-R1)(1-R2)(1.9)

[0117] R cB =R cA (R3)(1.10)

[0118] R cC =R4R5(1.11)

[0119] And cluster c B With cluster c C First connect in parallel, then connect in series with another node R6, the cluster reliability can be expressed as:

[0120] R S =[1-(1-R B )(1-R C )](R6)(1.12)

[0121] (4) k / n redundant cluster model

[0122] The redundant cluster model is a generalization of n parallel nodes. This model requires that at least k of the n identical and independent nodes in the cluster must work properly during their operation phase. The reliability of a k / n redundant cluster can be calculated using the binomial distribution theorem. Assuming that the reliability of each node is R and is independent of each other, the reliability of a k / n redundant cluster can be expressed as follows based on the probability that x nodes can work properly:

[0123] (1.13)

[0124] 1.1.2.2 Cluster System Elasticity

[0125] Resilience is currently widely used to characterize the ability of a cluster to recover after receiving a local attack. Therefore, the present invention introduces cluster resilience to evaluate the cluster's recovery ability, with the aim of improving cluster performance and shortening recovery time. In order to develop an indicator that quantifies cluster resilience, it is necessary to define a parameter that represents the functional performance of the cluster. This parameter may vary from cluster to cluster and is called the cluster performance parameter (Figure of Merit, FOM) in the proposed intelligent operation and maintenance framework. It should be noted that the defined cluster FOM, as a quantitative performance indicator, can be defined as a performance parameter of the actual cluster according to a specific case. In the event of large-scale extreme events such as earthquakes and typhoons, the typical evolution trend of the cluster FOM is as follows: Figure 4 Before an extreme event, cluster performance usually remains at FOM (-) The cluster runs at a level that may be lower than the expected performance level FOM*. The extreme event that occurs at time t1 causes the cluster performance to deteriorate sharply, and stabilizes at the performance degradation value FOM after time t2. D The cluster then performs reconstruction or maintenance from time t2, and after a period of time, the performance is restored and stabilized at FOM (+) . Figure 4 Two different recovery trajectories of the cluster are shown, which may be the result of different strategies adopted by the cluster or the result of different cluster capabilities. Figure 4 It can be observed that the cluster performance in the first trajectory is (1) Restore to FOM (+) In the second track, cluster performance takes longer time t4 (2) To restore to FOM (+) , suggesting that the former are more adaptable to extreme events than the latter.

[0126] from Figure 4 A simple resilience metric can be derived from , defined as the inverse of the time required to recover from an extreme event, as follows:

[0127] (1.14)

[0128] This definition of resilience is similar to the time-to-return used by clusters represented in state space. In addition, the resilience metric can be defined as resilience loss (RL):

[0129] RL=∫ t1 t4 (FOM * -FOM(t))dt(1.15)

[0130] The RL of a cluster quantifies the degree of performance loss of the cluster due to extreme events. The comprehensive metric GR for the resilience of an interdependent infrastructure cluster system can be defined by comprehensively considering the metrics of robustness ROBU, rapidity RAPI, average performance loss per unit time TAPL, and recovery ability RA, as shown below:

[0131] (1.16)

[0132] RAPI r and RAPI d They represent the rapidity measurement indexes of the recovery and destruction stages respectively. All the indicators in the above formula can be obtained as a function of FOM, as shown below:

[0133] (1.17)

[0134] (1.18)

[0135] (1.19)

[0136] (1.20)

[0137] Where K RP is the number of slopes detected by the slope detection technique during the recovery phase. Another comprehensive measure of resilience considers the fault condition F prof and recovery status R prof Merge into:

[0138] (1.21)

[0139] where F prof and R prof The performance of loss and recovery in each stage is shown below:

[0140] (1.22)

[0141] (1.23)

[0142] Where Q S It represents the performance of the cluster. The above indicators can all be used as indicators to measure the static cluster elasticity. For the measurement indicators of elasticity changing over time, for example, a statement of dynamic elasticity in the field of economics is as follows:

[0143] DR=∑ N i=1 FOM DR (t i )-FOM DU (t i )(1.24)

[0144] Among them, FOM DR and FOM DU Represents the performance with and without cluster elasticity properties. In addition, the elasticity of the cluster to n events can be quantified as follows:

[0145] (1.25)

[0146] Where R(τ) represents the elasticity function, which is a function of time, FOM, and stress induced by extreme events. The dynamic measure of elasticity can be summarized as incorporating the spatial dependence of elasticity into the dynamic measure:

[0147] (1.26)

[0148] Where θ can be time t, or time t and space s, and ρ represents the total performance loss of the cluster. In addition to static, dynamic, and comprehensive metrics, elasticity can also be quantified as a random variable:

[0149] (1.27)

[0150] The variable h(D i ,λ i ,φ) represents event D i The entropy of the probability distribution of an event based on the parameters λ described by φ i The distribution of ρ i (S p ,F r ,F d , F0) represents the elastic factor, which is used as the speed recovery factor S p , new steady-state factor F r , damage state factor F d and the original steady-state factor F0. The above entropy and elasticity factors can be expressed as:

[0151] (1.28)

[0152] (1.29)

[0153] (1.30)

[0154] where t δ ,t r * ,t r and a dec They represent the relaxation time, the final recovery time, the time to complete the initial recovery action, and the parameter controlling the elastic attenuation.

[0155] The present invention introduces the relevant contents of the cluster performance evaluation module in the intelligent operation and maintenance framework, establishes a cluster performance evaluation index system, including the general characteristics and special characteristics of the cluster system, analyzes the task requirements of each task stage according to the typical task process of the cluster system, selects relevant index items based on the general characteristics, special characteristics and task function evolution laws of the cluster system, and conducts relevant research on the performance indicators used in the operation and maintenance decision-making problems in subsequent chapters and provides an evaluation algorithm. The performance evaluation module is developed based on the Python language and integrated into the intelligent operation and maintenance framework. The proposed intelligent operation and maintenance framework still supports the research on related performance evaluation algorithms based on specific operation and maintenance problem requirements.

Claims

1. A performance evaluation method for intelligent operation and maintenance of a cluster system, wherein the cluster system is a power system cluster, a drone cluster system, or a large-scale battery cluster system in an industrial energy storage system, characterized in that: The method comprises the following steps: a) Construction of performance index system: Construct a performance evaluation index system based on the general and special characteristics of the cluster system. The general characteristics include reliability, maintainability, security, safety and environmental adaptability, and the special characteristics include time dynamic characteristics and space-time dynamic characteristics; b) Data collection: The real-time operation data of the cluster system is collected through the multi-dimensional situation awareness module, and the data includes node health status, link status, task execution status and environmental change information; c) Evaluation of the network efficiency of the cluster system: The network average efficiency formula E(G) is used to evaluate the efficiency of information transmission between nodes in the cluster system. The formula is: ; Where N is the number of nodes, w i and w j are the weights of node i and node j respectively, d ij is the shortest path length between node i and node j; d) Task execution balance evaluation: For the three-level spatiotemporal dynamic cluster, the cluster balance ε is used b Evaluate the balance of cluster system task execution capability. The balance calculation formula is: ; Where I represents the total number of clusters in cluster S, J i Represents cluster c i The total number of nodes, is the average redundancy value of all nodes in each cluster; ε ij Represents node n (i,j) The sum of the overlapping areas of the coverage area of ​​the node and the coverage areas of other nodes in the same cluster within the task area.

2. The performance evaluation method for cluster system intelligent operation and maintenance according to claim 1, characterized in that: The network efficiency evaluation adopts normalization processing to normalize the average efficiency E(G) to 0≤E(G)≤10, and when there is an edge between each pair of nodes in the cluster system, E(G)=1.

3. The performance evaluation method for cluster system intelligent operation and maintenance according to claim 1, characterized in that: The balance evaluation of the cluster system is performed by the redundancy matrix M i ε ={ε ij }Calculate the overlapping area of ​​cluster tasks and calculate the task capacity of each cluster based on the task distribution.

4. The performance evaluation method for cluster system intelligent operation and maintenance according to claim 1, characterized in that: The performance evaluation method also includes a network connectivity analysis based on time dynamic characteristics, by constructing an adjacency matrix A={a ij } and the distance matrix moment L = {l ij }, analyze the network connectivity status of the cluster system at different stages.

5. The performance evaluation method for cluster system intelligent operation and maintenance according to claim 1, characterized in that: The reliability of the cluster system is calculated by a reliability block diagram model, which includes a series, parallel, hybrid or k / n redundant cluster model. The reliability formula is: a) Reliability of the series structure R S (t) Formula: R S (t)=R1 (t)×R2 (t)×...×R n (t)≤min{R1 (t),R2 (t),…,R n (t)}; Where 0<R i (t)<1,i=1,2,…n represents node n i reliability; b) Reliability of parallel structure R S (t) Formula: R S (t)=1-∏ n i=1 [1-R i (t)]; Where 0<R i (t)<1,i=1,2,…n represents node n i reliability; c) In a hybrid cluster, cluster c A , c B , c C The reliability of is expressed as: R cA =1-(1-R1)(1-R2); R cB =R cA (R3); R cC =R4R5; And cluster c B With cluster c C First connect in parallel, then connect in series with another node R6, the cluster reliability is expressed as: R S =[1-(1-R B )(1-R C )](R6); d) k / n redundant cluster model Assuming that the reliability of each node is R and independent of each other, based on the probability that x nodes can operate normally, the reliability of the k / n redundant cluster is expressed as: 。 6. The performance evaluation method for cluster system intelligent operation and maintenance according to claim 1, characterized in that: The resilience evaluation of cluster systems involves calculating the inverse of the recovery time to obtain the resilience metric R of the system to recover from failures. FOM : ; Where: t1 is the time when the fault occurs, and t4 is the time to recover to the normal state; The elasticity of the cluster system is evaluated by the elasticity loss formula RL, which is: Among them, FOM is the target performance, FOM * is the expected performance level, and FOM(t) is the performance value at time t.

7. The performance evaluation method for cluster system intelligent operation and maintenance according to claim 6, characterized in that: The comprehensive metric GR for the resilience of an interdependent infrastructure cluster system is defined by comprehensively considering the metrics of robustness ROBU, rapidity RAPI, average performance loss per unit time TAPL, and recovery ability RA, as shown below: ; RAPI r and RAPI d They represent the rapidity measurement indexes of the recovery and destruction stages respectively. All the indicators in the above formula can be obtained as a function of FOM, as shown below: ; ; ; ; Where K RP is the number of slopes detected by the slope detection technique during the recovery phase; Another comprehensive measure of resilience is to consider the fault condition F prof and recovery status R prof Merge into: ; where F prof and R prof The performance of loss and recovery in each stage is shown below: ; ; Where Q S It indicates the performance of the cluster.

8. The performance evaluation method for cluster system intelligent operation and maintenance according to claim 1, characterized in that: The performance evaluation results of the cluster system are used to generate intelligent operation and maintenance decisions and guide operation and maintenance plans for predictive maintenance, post-disaster repair and dynamic reconstruction.

9. The performance evaluation method for cluster system intelligent operation and maintenance according to claim 1, characterized in that: The method comprehensively analyzes data from multiple operation and maintenance stages, outputs performance reports of cluster systems in different stages such as normal, degradation, destruction and recovery, and provides a basis for subsequent optimization.

10. A performance evaluation device for intelligent operation and maintenance of cluster systems, characterized in that: The device performs the method according to any one of claims 1 to 9, including: a) A data collection module, which is used to collect the operation data of the cluster system in real time, including node health status, link status, task execution status and environmental change information; b) an efficiency index construction module, which is used to construct an efficiency evaluation index system for the cluster system. The index is based on the general characteristics and special characteristics of the cluster system to construct an efficiency evaluation index system. The general characteristics include reliability, maintainability, security, safety and environmental adaptability, and the special characteristics include time dynamic characteristics and space-time dynamic characteristics; c) an efficiency calculation module, used to calculate and analyze the current efficiency of the cluster system based on the efficiency index; d) Result output module, which is used to generate performance reports of the cluster system and provide feedback and optimization suggestions for intelligent operation and maintenance decisions.

Citation Information

Patent Citations

  • Energy consumption balancing and coverage keeping method for underwater wireless sensor network

    CN108055683A

  • Method and system for evaluating operation efficiency of high-performance computing cluster

    CN117130851A