Data Center Operation and Maintenance System Based on Cloud Node Evaluation
Through the data center operation and maintenance system based on cloud node evaluation, the degree-centric and median data are collected, node collections are divided and differentiated monitoring frequency is set, which solves the problem of node monitoring limitations in the existing technology, and efficient operation and maintenance resource allocation and exception handling of the data center are realized.
Patent Information
- Application Number
- CN202510701397.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-28
AI Technical Summary
In the prior art, the data center operation and maintenance system only focuses on the abnormal performance of a single node, and lacks quantitative assessment of the associated risks between nodes, resulting in insufficient protection of core nodes and excessive consumption of operation and maintenance resources by ordinary nodes, and the inability to achieve reasonable allocation of resources.
Through a data center operation and maintenance system based on cloud node evaluation, we collect the degree-centric and median data of each cloud node, calculate the core index, divide the node sets and set differentiated monitoring frequency, adopt differentiated processing strategies for abnormal nodes in different sets, generate evaluation reports and alarm signals, and optimize operation and maintenance resource allocation.
It realizes hierarchical management of nodes of different degrees of importance, optimizes operation and maintenance resource allocation, improves the pertinence and efficiency of exception handling, and avoids the problems of excessive response or insufficient response.
Smart Images

Figure CN120223502B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data center operation and maintenance, and specifically to a data center operation and maintenance system based on cloud node evaluation. Background Art
[0002] In today's digital age, data centers undertake the tasks of storing, processing, and transmitting massive amounts of data. Their stable operation is crucial for the normal operation of enterprises and society. Data centers are usually composed of numerous cloud nodes, and these nodes cooperate with each other to provide various services.
[0003] However, the existing data center operation and maintenance systems based on cloud node evaluation still have the following deficiencies in actual application:
[0004] Node monitoring limitations: Only focusing on the abnormal performance of a single node, regarding each cloud node as an independent individual for monitoring, lacking a quantitative assessment of the associated risks between nodes;
[0005] In addition, for node anomalies, a unified processing logic is usually adopted, without considering the differences in the importance of different nodes in the data center. The nodes in the data center can be divided into core nodes and ordinary nodes, etc. Core nodes, as the hubs and transmission bridges in the data center network, play a key role in the stable operation of the entire system; ordinary nodes are responsible for some relatively minor functions, but the existing operation and maintenance strategies do not optimize for these differences, resulting in insufficient protection of core nodes and possible overconsumption of operation and maintenance resources for ordinary nodes, and unable to achieve reasonable allocation of resources.
[0006] Therefore, a data center operation and maintenance system based on cloud node evaluation is introduced. Summary of the Invention
[0007] The purpose of the present invention is to solve the problems pointed out in the background art, and to propose a data center operation and maintenance system based on cloud node evaluation.
[0008] The purpose of the present invention can be achieved through the following technical solutions: A data center operation and maintenance system based on cloud node evaluation, including:
[0009] Node evaluation module: Collect the dependency data of each cloud node in the data center at a preset frequency, store it in the database, process and analyze the dependency data stored in the database, and determine the core index corresponding to each cloud node; where the dependency data includes degree centrality and betweenness centrality ;
[0010] Operation and maintenance monitoring module: Extract the core indices corresponding to each cloud node, sort them from largest to smallest, and after sorting, divide each cloud node into a core node set, a sub-core node set, and an ordinary node set according to a preset dividing line; set the monitoring frequencies corresponding to different node sets, and evaluate the CPU usage rate and memory usage rate of each cloud node according to the corresponding monitoring frequencies, that is, after normalizing the CPU usage rate and memory usage rate of each cloud node at the current time point, multiply them by the corresponding set weight coefficients respectively, and then sum to obtain the operation status index of each cloud node;
[0011] Abnormal association module: Set the reference thresholds of the operation status indices corresponding to different node sets, receive the operation status indices of each cloud node and compare them with the corresponding reference thresholds, and perform corresponding steps based on the comparison results to push the cloud node evaluation report to the operation and maintenance personnel.
[0012] As a preferred embodiment of the present invention, the determination of the core indices corresponding to each cloud node is specifically as follows:
[0013] For the N nodes included in the data center, the degree centrality of cloud node j is defined as: equals the number of direct connections between node j and other nodes; where j represents the number of the cloud node, j = 1, 2,......, N;
[0014] For node j, its betweenness centrality is defined as: ; where s and t represent the specific numbers of any two groups of nodes among the N cloud nodes, and s ≠ t; represents the total number of shortest paths from node s to t; by setting a reference value for the path length of the shortest path, the paths shorter than the path length reference value are marked as the shortest paths, represents the number of shortest paths from node s to t that pass through node j;
[0015] Extract the degree centrality and betweenness centrality of each cloud node in the data center, and multiply them by the corresponding set weight ratios respectively, so as to obtain the core indices corresponding to each cloud node.
[0016] As a preferred embodiment of the present invention, the setting of the monitoring frequency is specifically as follows: Implement high-frequency real-time monitoring for core nodes, set it to 10 seconds / time, implement medium-frequency monitoring for sub-core nodes, set it to 10 minutes / time, and implement low-frequency monitoring for ordinary nodes, set it to 0.5 hours / time.
[0017] As a preferred embodiment of the present invention, the operating status index of each cloud node is received and compared with the corresponding reference threshold, and corresponding steps are executed based on the comparison result. If the cloud node is in the core node set and higher than the corresponding reference threshold, the following steps are specifically executed:
[0018] S1: If the cloud node higher than the reference threshold is in the core node set, mark the cloud node higher than the reference threshold as an abnormal core node, and extract all directly connected nodes of the abnormal core node from the dependent data of the node evaluation module as peripheral nodes;
[0019] Retrieve the operating status index of the current peripheral node at the current time point, and at the same time extract the total traffic carried by the abnormal core node when it was operating normally before being higher than the reference threshold; Through the formula Calculate the expected load increment of the peripheral node; where the remaining available path number is the number of paths through which the peripheral node can share traffic after the core node fails;
[0020] Based on the expected load increment of the peripheral node, obtain the operating prediction index of the peripheral node through a preset calculation logic;
[0021] Identify the node set where the peripheral node is located, extract the reference threshold of the corresponding node set, compare the calculated operating prediction index of the peripheral node with the corresponding set reference threshold respectively, and mark the peripheral node higher than the corresponding reference threshold as a potential hazard node;
[0022] Extract the operating prediction index corresponding to the potential hazard node, calculate the difference from the corresponding reference threshold, record the calculated difference as the probability difference, and convert the calculated probability difference of the potential hazard node into a probability value using a set conversion rule;
[0023] Extract the number corresponding to the abnormal core node, the number of the potential hazard node, and the probability value corresponding to the potential hazard node, and input them into a pre-constructed report template. Then, extract each group of historical abnormal cases corresponding to the abnormal core node number from the database, and analyze the similarity evaluation between each group of historical abnormal cases and the abnormal core node at the current time point , select the historical abnormal case with the highest similarity evaluation, and obtain its final processing strategy as a reference strategy, and input it into a pre-constructed report template;
[0024] After the input is completed, it is used as a cloud node evaluation report of the abnormal core node and an alarm signal is generated and pushed to the operation and maintenance personnel.
[0025] As a preferred embodiment of the present invention, the operating prediction index of the peripheral node obtained through a preset calculation logic is specifically:
[0026] Based on historical data, calculate the CPU usage rate and memory usage rate of each cloud node under different traffic loads to form a data set;
[0027] Use the fitting linear relationship to determine the relationship formulas (1) between the traffic corresponding to each cloud node and the CPU usage rate, and (2) between the traffic and the memory usage rate. Formulas (1) and (2) are respectively expressed as CPU usage rate = a × traffic + b and memory usage rate = c × traffic + d; where a and c are the sensitivity coefficients of traffic to the CPU usage rate and memory usage rate respectively, and b and d are constant terms, calculated through linear regression;
[0028] After determining the relationship formulas (1) and (2), add the traffic data of the surrounding nodes at the current time point to the expected load increment to obtain the subsequent performance traffic of the surrounding nodes;
[0029] Further, through the subsequent performance traffic of the surrounding nodes, calculate the estimated CPU usage rate and estimated memory usage rate corresponding to the surrounding nodes at the next time point using the relationship formulas (1) and (2). Multiply the estimated CPU usage rate and estimated memory usage rate of the surrounding nodes by the corresponding set weight coefficients respectively, and then sum to obtain the operation estimation index of the surrounding nodes.
[0030] As a preferred implementation mode of the present invention, the probability difference calculated according to the hidden danger node is converted into a probability of occurrence value using a set conversion rule, specifically:
[0031] Set a range set, expressed as ; where represents the interval where each group of differences corresponding to the probability difference is located, n is the total number of intervals where the differences are located, represents the probability of occurrence values of each group; and and are in a matching relationship, and are in a matching relationship, and so on; input the probability difference calculated for the surrounding nodes into the range set for matching to determine the probability of occurrence value.
[0032] As a preferred implementation mode of the present invention, analyze the similarity estimates between each group of historical abnormal cases and the abnormal core nodes at the current time point , specifically:
[0033] Each group of historical abnormal cases includes the abnormal core node number, the operation status index at the time of abnormality, the hidden danger node number, the number of hidden danger nodes, and the final processing strategy;
[0034] Match the hidden danger node numbers of the abnormal core nodes at the current time point with the hidden danger node numbers of each group of historical abnormal cases, and count the number of successfully matched numbers as , calculate the difference between the number of potential hazard nodes of the abnormal core node at the current time point and the number of potential hazard nodes of each group of historical abnormal cases, and record the absolute value as , calculate the difference between the operation status index of the abnormal core node at the current time point and the operation status index of each group of historical abnormal cases, and record the absolute value as ;
[0035] For , and , after normalization, calculate according to the formula to obtain the similarity estimate between each group of historical abnormal cases and the abnormal core node at the current time point ; where are respectively , and corresponding weight coefficients.
[0036] As a preferred embodiment of the present invention, receive the operation status index of each cloud node and compare it with the corresponding reference threshold, and execute the corresponding steps based on the comparison result. If the cloud node is in the secondary core node set and higher than the corresponding reference threshold, specifically execute:
[0037] If the cloud node higher than the reference threshold is in the secondary core node set, then also mark the cloud node higher than the reference threshold as an abnormal core node. Similarly, in step S1, generate a cloud node evaluation report and generate a warning signal to be pushed to the operation and maintenance personnel.
[0038] As a preferred embodiment of the present invention, receive the operation status index of each cloud node and compare it with the corresponding reference threshold, and execute the corresponding steps based on the comparison result. If the cloud node is in the ordinary node set and higher than the corresponding reference threshold, specifically execute:
[0039] If the cloud node higher than the reference threshold is in the ordinary node set, after marking the cloud node higher than the reference threshold as a node to be processed, send the number of the node to be processed to the operation and maintenance personnel to remind the operation and maintenance personnel to process it during idle time. At the same time, temporarily adjust the monitoring frequency of the node to be processed to the monitoring frequency corresponding to the secondary core node set. After the adjustment is completed, if the number of times the node to be processed is higher than the reference threshold continues times, send a warning signaling and the number of the node to be processed to the operation and maintenance personnel, and at the same time push the preset restricted processing time range.
[0040] Compared with the prior art, the beneficial effects of the present invention are:
[0041] The present invention collects the degree centrality and betweenness centrality dependence data of each cloud node, calculates the core index to measure the importance and influence of the node in the network, divides the node set according to the core index, and adopts different monitoring frequencies for different sets. When a core node is abnormal, the information of the surrounding nodes is extracted, the expected load increment is calculated, and the state change of the surrounding nodes is predicted, solving the problem of node monitoring limitations in the prior art, which only focuses on the abnormal performance of a single node, monitors each cloud node as an independent individual, and lacks a quantitative assessment of the associated risks between nodes;
[0042] The present invention divides cloud nodes into a core node set, a sub-core node set, and an ordinary node set according to the core index. For different sets, different monitoring frequencies are set, and different processing strategies are adopted for abnormal nodes in different sets. When the core and sub-core nodes are abnormal, an evaluation report and an alarm signal are generated. When an ordinary node is abnormal, it is first marked for processing, and the monitoring frequency and time-limited processing are upgraded according to the situation, realizing hierarchical management of nodes with different importance levels and optimizing the allocation of operation and maintenance resources;
[0043] When dealing with an anomaly, the present invention converts the difference between the operation prediction index of the potential hazard node and the reference threshold into a probability value according to the set conversion rule. At the same time, information such as the potential hazard node number, quantity, and operation status index of the current abnormal core node is matched with historical abnormal cases, calculates a similar valuation, and selects the final processing strategy of the historical abnormal case with the highest similar valuation as a reference. This data-driven method provides a precise processing strategy for the current anomaly, improving the pertinence and efficiency of anomaly processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] For the convenience of those skilled in the art to understand, the present invention will be further described below with reference to the accompanying drawings.
[0045] Figure 1 It is a principle block diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0047] Please refer to Figure 1 As shown, a data center operation and maintenance system based on cloud node evaluation includes a node evaluation module, an operation and maintenance monitoring module, and an anomaly association module;
[0048] The node evaluation module is used to collect the dependency data of each cloud node in the data center at a preset frequency, store it in the database, process and analyze the dependency data stored in the database, and determine the core index corresponding to each cloud node; the dependency data includes degree centrality and betweenness centrality ;
[0049] Specifically:
[0050] For the N nodes included in the data center, the degree centrality of cloud node j is defined as: equal to the number of direct connections between node j and other nodes; where j represents the number of the cloud node, j = 1, 2,..., N;
[0051] Suppose a data center network contains 5 nodes (A, B, C, D, E), and the connection relationships are as follows:
[0052] A is connected to B, C, D;
[0053] B is connected to A, C;
[0054] C is connected to A, B, E;
[0055] D is connected to A;
[0056] E is connected to C;
[0057] Then the degree centrality of node A is 3;
[0058] It is default to perform normalization processing after determining the degree centrality of a certain node, and obtain it through the formula / (N - 1), where subtracting one from the denominator represents the maximum possible number of connections of the node;
[0059] For node j, its betweenness centrality is defined as: ; where s and t represent the specific numbers of any two groups of nodes among the N cloud nodes, and s ≠ t; represents the total number of shortest paths from node s to t; by setting a reference value for the path length of the shortest path, the paths shorter than the path length reference value are marked as the shortest paths, represents the number of times passing through node j among each shortest path from node s to t;
[0060] It is default to perform normalization processing after determining the betweenness centrality of a certain node, and obtain it through the formula , where the denominator represents the number of all possible node pairs;
[0061] Extract the degree centrality and betweenness centrality , and multiply them by the corresponding set weight ratios respectively, so as to obtain the core indices corresponding to each cloud node; among them, the degree centrality and betweenness centrality The sum of the weight ratios is equal to one;
[0062] Nodes with high degree centrality and betweenness centrality are mostly hubs and transmission bridges in the data center network;
[0063] The operation and maintenance monitoring module is used to extract the core indices corresponding to each cloud node, and perform a descending sort. After sorting, according to a preset dividing line; the dividing line is set according to the value of the total number N of cloud nodes, and the division rules can be formulated according to specific business requirements and experience. For example, nodes ranked in the top certain proportion (such as 20%) can be divided into core nodes, nodes ranked in the middle certain proportion (such as 20% - 40%) can be divided into sub-core nodes, and the remaining nodes can be divided into ordinary nodes. It is also possible to directly specify the ranking range, such as the first 5 groups are core nodes, the 6 - 10 groups are sub-core nodes, and the rest are ordinary nodes; divide each cloud node into a core node set, a sub-core node set, and an ordinary node set; set the monitoring frequencies corresponding to different node sets, and evaluate the CPU usage rate and memory usage rate of each cloud node according to the corresponding monitoring frequencies, that is, after normalizing the CPU usage rate and memory usage rate of each cloud node at the current time point, multiply them by the corresponding set weight coefficients respectively, and then sum to obtain the operation status index of each cloud node;
[0064] Monitoring frequency setting, for example:
[0065] Implement high-frequency real-time monitoring on core nodes (such as collecting indicators per second), medium-frequency monitoring on sub-core nodes (per minute), and low-frequency monitoring on ordinary nodes (per hour), and focus on basic connectivity. Since core nodes are of high importance and play a key role in the stable operation of the entire cloud system, a higher monitoring frequency needs to be set to promptly discover and handle possible problems. The importance of sub-core nodes is secondary, and the monitoring frequency can be appropriately reduced. The importance of ordinary nodes is relatively low, and the monitoring frequency can be further reduced to reduce the consumption of monitoring resources;
[0066] The anomaly correlation module is used to set the reference thresholds for the operation status indices corresponding to different node sets, receive the operation status indices of each cloud node and compare them with the corresponding reference thresholds, and execute corresponding steps based on the comparison results to push the cloud node evaluation report to the operation and maintenance personnel;
[0067] Specifically:
[0068] S1: If the cloud nodes higher than the reference threshold are in the core node set, mark the cloud nodes higher than the reference threshold as abnormal core nodes, and extract all the directly connected nodes of the abnormal core nodes from the dependency data of the node evaluation module as peripheral nodes;
[0069] Retrieve the running status index of the current peripheral nodes at the current time point, and at the same time extract the total traffic carried by the abnormal core nodes when they were running normally before exceeding the reference threshold; Through the formula Calculate the expected load increment of the peripheral nodes; where the remaining available path number is the number of paths through which the peripheral nodes can share traffic after the core node fails; where the weight of the peripheral nodes is set by the technical personnel according to the importance of the peripheral nodes in the path, and the sum is 1;
[0070] Through historical data, count the CPU usage rate and memory usage rate of each cloud node under different traffic loads to form a data set;
[0071] Use the fitting linear relationship to determine the relationship formula (1) between the traffic corresponding to each cloud node and the CPU usage rate, and the relationship formula (2) between the traffic and the memory usage rate. Formula (1) and (2) are respectively expressed as CPU usage rate = a×traffic + b and memory usage rate = c×traffic + d; where a and c are the sensitivity coefficients of traffic to CPU usage rate and memory usage rate respectively, and b and d are constant terms, calculated through linear regression;
[0072] Continuously optimize the relationship formula through the continuous update of historical data;
[0073] After determining relationship formula (1) and (2), add the traffic data of the peripheral nodes at the current time point to the expected load increment to obtain the subsequent performance traffic of the peripheral nodes;
[0074] Further, through the subsequent performance traffic of the peripheral nodes, calculate the estimated CPU usage rate and estimated memory usage rate corresponding to the peripheral nodes at the next time point through relationship formula (1) and (2), multiply the estimated CPU usage rate and estimated memory usage rate of the peripheral nodes by the corresponding set weight coefficients respectively, and then sum to obtain the running estimation index of the peripheral nodes;
[0075] Identify the node set where the peripheral nodes are located, extract the reference threshold of the corresponding node set, compare the running estimation index calculated by the peripheral nodes with the corresponding set reference threshold respectively, and mark the peripheral nodes higher than the corresponding reference threshold as potential hazard nodes;
[0076] Extract the running estimation index corresponding to the potential hazard nodes, and calculate the difference from the corresponding reference threshold, and record the calculated difference as the probability difference;
[0077] Set a range set, expressed as ; where represents the interval in which the differences of each group corresponding to the probability difference are located, and n is the total number of intervals where the differences are located. represents the occurrence probability values of each group; the range is set between 1% - 100%, and the higher the probability difference, the higher the corresponding matching occurrence probability value; and and are in a matching relationship. and are in a matching relationship, and so on; the probability differences calculated by the surrounding nodes are input into the range set for matching to determine the occurrence probability value.
[0078] Extract the numbers corresponding to the abnormal core nodes, the numbers of hidden danger nodes, and the occurrence probability values corresponding to the hidden danger nodes, and input them into the pre-constructed report template. Then, extract each group of historical abnormal cases corresponding to the numbers of the abnormal core nodes from the database. Each group of historical abnormal cases includes the numbers of abnormal core nodes, the operation status indices at the time of abnormality, the numbers of hidden danger nodes, the quantities of hidden danger nodes, and the final handling strategies.
[0079] Match the numbers of hidden danger nodes of the abnormal core node at the current time point with the numbers of hidden danger nodes of each group of historical abnormal cases, and count the number of successfully matched numbers as , calculate the difference between the quantity of hidden danger nodes of the abnormal core node at the current time point and the quantity of hidden danger nodes of each group of historical abnormal cases, and take the absolute value and record it as , calculate the difference between the operation status index of the abnormal core node at the current time point and the operation status indices of each group of historical abnormal cases, and take the absolute value and record it as ;
[0080] For , and , after normalizing them, calculate according to the formula to obtain the similarity estimates between each group of historical abnormal cases and the abnormal core node at the current time point ; Select the historical abnormal case with the highest similarity estimate , and obtain its final handling strategy as the reference strategy, and input it into the pre-constructed report template; where are respectively , and corresponding weight coefficients, and the sum of them is one.
[0081] After the input is completed, it serves as the cloud node evaluation report of the abnormal core node and generates an alarm signal, which is pushed to the operation and maintenance personnel. The operation and maintenance personnel immediately solve the node abnormality according to the cloud node evaluation report, and the measures taken include but are not limited to adjusting resource allocation, load migration, etc.
[0082] S2: If the cloud nodes above the reference threshold are in the secondary core node set, then mark the cloud nodes above the reference threshold as abnormal core nodes as well. Similarly, generate a cloud node evaluation report and generate a warning signal in step S1 and push it to the operation and maintenance personnel;
[0083] S3: If the cloud nodes above the reference threshold are in the ordinary node set, then mark the cloud nodes above the reference threshold as nodes to be processed, and send the numbers of the nodes to be processed to the operation and maintenance personnel to remind the operation and maintenance personnel to process them during idle time. At the same time, temporarily adjust the monitoring frequency of the nodes to be processed to the monitoring frequency corresponding to the secondary core node set. After the adjustment is completed, if the number of times the nodes to be processed are above the reference threshold continues for a certain number of times, send a warning signaling and the numbers of the nodes to be processed to the operation and maintenance personnel, and at the same time push the preset restricted processing time range; for example, resolve the anomaly within 1 hour;
[0084] Abnormal core nodes (S1): Trigger in-depth evaluation, associate with the analysis of the load impact of surrounding nodes, generate a detailed evaluation report containing historical case reference strategies, and achieve accurate positioning and preventive maintenance of key nodes;
[0085] Abnormal secondary core nodes (S2): Similarly trigger the warning mechanism to ensure the anomaly response efficiency of secondary key nodes;
[0086] Abnormal ordinary nodes (S3): Adopt a progressive strategy of "reminder - monitoring upgrade - time-limited processing" to avoid excessive consumption of operation and maintenance resources. At the same time, balance the anomaly handling priorities of ordinary nodes by dynamically adjusting the monitoring frequency (such as increasing it to the monitoring frequency of secondary core nodes);
[0087] Usually, a unified processing logic is adopted for all nodes, lacking hierarchical management of node importance. This mechanism significantly optimizes the allocation of operation and maintenance resources through a differentiated strategy, avoiding "over-response" and "under-response";
[0088] When ordinary nodes are abnormal, temporarily increase the monitoring frequency. When the anomaly persists, trigger time-limited processing, which not only avoids being ignored for ordinary node anomalies but also prevents low-frequency anomalies from interfering with operation and maintenance, achieving dynamic optimization of monitoring resources;
[0089] By extracting the directly connected nodes (surrounding nodes) of abnormal core nodes, combining with the linear fitting model of historical traffic-resource utilization rate (the relationship formula between CPU / memory utilization rate and traffic), quantitatively calculate the expected load increment of surrounding nodes and subsequent resource utilization rate, and predict in advance the cascading impact of abnormal core nodes on surrounding nodes; this mechanism can effectively identify potential hidden danger nodes and avoid the chain reaction caused by the spread of anomalies.
[0090] Through the "probability difference - occurrence probability" mapping model and the calculation of historical case similarity (matching the hidden danger node numbers, quantities, and operating status indices), automatically associate historical processing strategies, provide a reference solution for the current anomaly, and improve the pertinence and efficiency of anomaly handling through data-driven automated strategy recommendation;
[0091] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor limit the invention to the specific embodiments only. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the relevant technical fields can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A data center operation and maintenance system based on cloud node evaluation, characterized in that Including: Node evaluation module: Collect the dependency data of each cloud node in the data center at a preset frequency, store it in the database, process and analyze the dependency data stored in the database, and determine the core index corresponding to each cloud node; Among them, the dependent data includes degree centrality and betweenness centrality ; Operation and maintenance monitoring module: Extract the core index corresponding to each cloud node, sort it from large to small, and after sorting, according to the preset dividing line; Divide each cloud node into a core node set, a sub-core node set, and an ordinary node set; Set the monitoring frequencies corresponding to different node sets, and evaluate the CPU usage rate and memory usage rate of each cloud node according to the corresponding monitoring frequencies, that is, after normalizing the CPU usage rate and memory usage rate of each cloud node at the current time point, multiply them by the corresponding set weight coefficients respectively, and then sum to obtain the operation status index of each cloud node; Abnormal association module: Set the reference thresholds for the operation status indexes corresponding to different node sets, receive the operation status indexes of each cloud node, compare them with the corresponding reference thresholds, and execute the corresponding steps based on the comparison results to push the cloud node evaluation report to the operation and maintenance personnel; If the cloud node is in the core node set and higher than the corresponding reference threshold, the specific execution is: S1: If the cloud node higher than the reference threshold is in the core node set, mark the cloud node higher than the reference threshold as an abnormal core node, and extract all the directly connected nodes of the abnormal core node from the dependency data of the node evaluation module as the surrounding nodes; Retrieve the operation status index of the current surrounding nodes at the current time point, and at the same time extract the total traffic carried by the abnormal core node when it was running normally before exceeding the reference threshold; Calculate the expected load increment of the peripheral nodes through the formula where the remaining available path number is the number of paths through which the peripheral nodes can share traffic after the core node fails; Based on the expected load increment of the surrounding nodes, through historical data, count the CPU usage rate and memory usage rate of each cloud node under different traffic loads to form a data set; Use the fitting linear relationship to determine the relationship formula (1) between the traffic and the CPU usage rate and the relationship formula (2) between the traffic and the memory usage rate corresponding to each cloud node. Formula (1) and (2) are respectively expressed as CPU usage rate = a × traffic + b and memory usage rate = c × traffic + d; where a and c are the sensitivity coefficients of traffic to CPU usage rate and memory usage rate respectively, and b and d are constant terms, calculated through linear regression; After determining the relationship formulas (1) and (2), add the traffic data of the surrounding nodes at the current time point to the expected load increment to obtain the subsequent performance traffic of the surrounding nodes; Furthermore, through the subsequent performance traffic of the surrounding nodes, calculate the estimated CPU usage rate and estimated memory usage rate corresponding to the surrounding nodes at the next time point through the relationship formulas (1) and (2), multiply the estimated CPU usage rate and estimated memory usage rate of the surrounding nodes by the corresponding set weight coefficients respectively, and then sum to obtain the operation prediction index of the surrounding nodes; Identify the node set where the surrounding nodes are located, extract the reference threshold of the corresponding node set, compare the operation prediction index calculated by the surrounding nodes with the corresponding set reference threshold respectively, and mark the surrounding nodes higher than the corresponding reference threshold as potential hazard nodes; Extract the operation prediction index corresponding to the hidden danger node, calculate the difference from the corresponding reference threshold, and record the calculated difference as the probability difference. Based on the probability difference calculated for the hidden danger node, set a range set, expressed as ; where represents the interval where each group of differences corresponding to the probability difference is located, and n is the total number of intervals where the differences are located, represents the occurrence probability value of each group; and and are in a matching relationship, and are in a matching relationship, and so on; Match the probability differences calculated by the peripheral nodes within the input range set to determine the occurrence probability value; Extract the numbers corresponding to the abnormal core nodes, the numbers of potential hazard nodes, and the occurrence probability values corresponding to the potential hazard nodes, and input them into a pre-constructed report template. Then, extract each group of historical abnormal cases corresponding to the abnormal core node numbers from the database. Each group of historical abnormal cases includes the abnormal core node numbers, the operation status indices at the time of abnormality, the numbers of potential hazard nodes, the numbers of potential hazard nodes, and the final handling strategies; Match the hidden danger node numbers of the abnormal core node at the current time point with the hidden danger node numbers of each group of historical abnormal cases, and count the number of successfully matched numbers as , calculate the difference between the number of hidden danger nodes of the abnormal core node at the current time point and the number of hidden danger nodes of each group of historical abnormal cases, and record the absolute value as , calculate the difference between the operation status index of the abnormal core node at the current time point and the operation status index of each group of historical abnormal cases, and record the absolute value as ; Pair , and After normalization processing, calculate according to the formula to obtain the similarity estimates between each group of historical abnormal cases and the abnormal core nodes at the current time point ; among them are respectively , and The corresponding weight coefficients, select the historical abnormal case with the highest similarity estimate and obtain its final processing strategy as a reference strategy, and input it into the pre-constructed report template; After the input is completed, it serves as the cloud node evaluation report for the abnormal core node and generates an alarm signal, which is pushed to the operation and maintenance personnel.
2. The data center operation and maintenance system based on cloud node evaluation according to claim 1, wherein, The determination of the core indices corresponding to each cloud node is specifically as follows: For the N nodes contained in the data center, the degree centrality of cloud node j is defined as: equal to the number of direct connections between node j and other nodes; where j represents the number of the cloud node, j = 1, 2,......, N; For node j, its betweenness centrality is defined as: ; where s and t represent the specific numbers of any two groups of nodes among the N cloud nodes, and s ≠ t; represents the total number of shortest paths from node s to t; by setting a reference value for the path length of the shortest path, the paths shorter than the reference value of the path length are marked as the shortest paths, represents the number of shortest paths from node s to t that pass through node j; Extract the degree centrality and betweenness centrality of each cloud node in the data center and betweenness centrality , and multiply them by the corresponding set weight ratios respectively to obtain the core index corresponding to each cloud node.
3. The data center operation and maintenance system based on cloud node evaluation according to claim 2, characterized in that The monitoring frequency setting is specifically as follows: Implement high-frequency real-time monitoring of the core nodes, set to 10 seconds / time; implement medium-frequency monitoring of the sub-core nodes, set to 10 minutes / time; implement low-frequency monitoring of the ordinary nodes, set to 0.5 hours / time.
4. The data center operation and maintenance system based on cloud node evaluation according to claim 3, wherein Receive the operation status indices of each cloud node and compare them with the corresponding reference thresholds. Based on the comparison results, execute the corresponding steps. If the cloud node is in the sub-core node set and is higher than the corresponding reference threshold, specifically execute: If the cloud node higher than the reference threshold is in the sub-core node set, then also mark the cloud node higher than the reference threshold as an abnormal core node. Similarly, in step S1, generate a cloud node evaluation report and generate a warning signal to be pushed to the operation and maintenance personnel.
5. The data center operation and maintenance system based on cloud node evaluation according to claim 4, characterized in that, Receive the operation status indices of each cloud node and compare them with the corresponding reference thresholds. Based on the comparison results, execute the corresponding steps. If the cloud node is in the ordinary node set and is higher than the corresponding reference threshold, specifically execute: If a cloud node higher than the reference threshold is in the set of ordinary nodes, after marking the cloud node higher than the reference threshold as a node to be processed, the number of the node to be processed is sent to the operation and maintenance personnel to remind the operation and maintenance personnel to process it during idle time. At the same time, the monitoring frequency of the node to be processed is temporarily adjusted to the monitoring frequency corresponding to the sub-core node set. After the adjustment is completed, if the number of times the node to be processed is higher than the reference threshold continues times, a warning signaling and the number of the node to be processed are sent to the operation and maintenance personnel, and at the same time, a preset restricted processing time range is pushed.
Citation Information
Patent Citations
Method and device for determining importance of network nodes
CN106301868A
Host node monitoring method and device based on cloud platform and computer equipment
CN110890977A
Python monitoring task resource use method and system
CN119512883A
Optical cable routing anti-intrusion method and system
CN119628935A