Automatic Selection Method and System for Monitoring Indicators of Large-Scale Computer Systems

By quantifying the dynamic correlation strength and anomaly analysis of in-band and out-of-band monitoring indicators in large-scale computer systems, key out-of-band monitoring indicators are automatically selected, solving the problems of large sensor base, high data pressure and high latency, and realizing efficient automated selection of monitoring indicators and system status detection.

CN121807655BActive Publication Date: 2026-05-26NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NAT UNIV OF DEFENSE TECH
Filing Date
2026-03-10
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In large-scale computer systems, the number of sensors collecting out-of-band monitoring indicators is large, the data transmission pressure is high, and the processing latency is large. Existing technologies mainly rely on manual experience-based configuration, which is static and rigid, and cannot adapt to the needs of system expansion and has low flexibility.

Method used

By quantifying the dynamic correlation strength between in-band and out-of-band monitoring indicators, multiple out-of-band monitoring indicators with strong correlation to in-band monitoring indicators are selected. Combined with in-band anomaly analysis, key out-of-band monitoring indicators are automatically selected to optimize sensor selection.

Benefits of technology

It effectively reduces the number of sensors used for data acquisition, reduces data transmission pressure and processing delay, improves the efficiency of automated and indiscriminate acquisition of monitoring indicators, and enhances the timeliness and accuracy of system status detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121807655B_ABST
    Figure CN121807655B_ABST
Patent Text Reader

Abstract

This invention discloses an automatic selection method and system for monitoring indicators of large-scale computer systems. The automatic selection method for out-of-band monitoring indicators includes: quantifying the dynamic correlation strength between M in-band and N out-of-band monitoring indicators from the same computer node using time-series data sequences; selecting multiple out-of-band monitoring indicators with strong dynamic correlation to the in-band indicators to obtain a set of key out-of-band monitoring indicators; and further optimizing the key out-of-band monitoring indicators by extracting in-band anomalies and performing spatiotemporal correlation analysis between in-band indicator anomalies and out-of-band hardware alarms to obtain the final selected out-of-band indicators. This invention aims to automatically and unsupervisedly select out-of-band monitoring indicators that are highly indicative of system operation by associating the system's in-band operating status, thereby reducing the problems of a large number of data acquisition sensors, high data transmission pressure, and large processing latency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer system monitoring technology, and specifically to a method and system for automatically selecting monitoring indicators for large-scale computer systems. Background Technology

[0002] Large-scale computer systems, exemplified by supercomputers, consist of tens of thousands to millions of heterogeneous computing nodes. Timely detection of hardware anomalies is crucial for ensuring reliable system operation. This relies on a distributed state-aware and centralized feedback-based operation and maintenance management system that dynamically collects the operational status of hardware units distributed across various functional components of the system. This supports maintenance personnel in assessing, analyzing, and quickly locating and troubleshooting faults. Each computing node in a large-scale computer system is equipped with hundreds of sensors to sense the hardware operating status of various circuit components, such as temperature, voltage, power consumption, and humidity. To avoid interference with business traffic from sensor data collection and transmission, a separate network is often used in the design for sensor status transmission and aggregation—the so-called out-of-band monitoring metrics. In contrast, there is data collection focused on the operating system and business operation status—the in-band monitoring metrics. Considering the tens of thousands of computing nodes in the era of exascale computing, the number of sensors used for system state awareness reaches millions. Unifying these at the system operation and maintenance level generates millions of monitoring metrics, with daily out-of-band data exceeding 100GB. This places enormous pressure on transmission and storage, posing a significant challenge to timely analysis and processing. For example, large-scale computer systems at the petabyte level or above (supercomputer systems) can generate over 100GB of software monitoring data (in-band metrics) and sensor status monitoring data (out-of-band metrics) daily. More than half of this monitoring data comes from millions of sensors. This massive amount of out-of-band monitoring data imposes huge transmission and storage overhead on the system monitoring process, and important system statuses are also submerged in the massive amount of monitoring data, making them difficult to detect and identify in a timely and effective manner. Current monitoring and maintenance mainly adopts blacklist and whitelist filtering techniques, which are based on manual experience and configuration. These static and rigid configurations have low flexibility and poor scalability in meeting the needs of system expansion and new systems. This invention addresses the problem of high dimensionality of out-of-band metrics and high overhead of indiscriminate collection in large-scale computer systems. Summary of the Invention

[0003] The technical problem to be solved by this invention is to provide an automatic selection method and system for monitoring indicators of large-scale computer systems, addressing the aforementioned problems of the prior art. This invention aims to select out-of-band monitoring indicators that are highly indicative of system operation in an automated and unsupervised manner by associating the in-band operating status of the system, thereby effectively reducing the problems of large number of data acquisition sensors, high data transmission pressure, and large processing delay.

[0004] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0005] An automatic selection method for monitoring indicators of a large-scale computer system includes the following steps:

[0006] S101, for time-series data sequences of M in-band monitoring indicators and N out-of-band monitoring indicators from the same computer node, quantify the dynamic correlation strength between out-of-band monitoring indicators and in-band monitoring indicators, and select multiple out-of-band monitoring indicators with strong dynamic correlation strength with in-band monitoring indicators to obtain a set of key out-of-band monitoring indicators.

[0007] S102, for the set of key out-of-band monitoring indicators, extract in-band anomalies and perform spatiotemporal correlation analysis between in-band indicator anomalies and out-of-band hardware alarms to further optimize the key out-of-band monitoring indicators to obtain the final selected out-of-band indicators.

[0008] Optionally, step S101 includes:

[0009] S201, for time-series data sequences of M in-band monitoring indicators and N out-of-band monitoring indicators from the same computer node, calculate the arbitrary i-th... Time series data of in-band indicators With the Time series data of out-of-band indicators The dynamic time warping technique is used to calculate the first... The in-band indicators and the first Correlation scores of out-of-band indicators ;

[0010] S202, All related scores Construct an M×N correlation matrix;

[0011] S203, calculate the dynamic correlation strength between each out-of-band indicator and the M in-band monitoring indicators based on the correlation matrix:

[0012] ;

[0013] in, For the first The dynamic correlation strength between one out-of-band indicator and M in-band monitoring indicators;

[0014] S204: Sort the dynamic correlation strength of each out-of-band indicator with M in-band monitoring indicators in descending order, and select the top-ranked indicators by a specified number. Each out-of-band indicator serves as a set of key out-of-band monitoring indicators.

[0015] Optionally, in step S201, the calculation of the first... The in-band indicators and the first Correlation scores of out-of-band indicators The function expression is:

[0016] ;

[0017] in, For the first Time series data of in-band indicators The first in One time series data, For the first Time series data of out-of-band indicators The first in One time series data, This represents the optimal alignment path between two time series. , , ~ For the first Time series data of in-band indicators The first to the second One time series data, ~ For the first Time series data of out-of-band indicators The first to the second One time series data, and These represent the time series data lengths for in-band and out-of-band indicators, respectively.

[0018] Optionally, the functional expression for calculating the dynamic correlation strength of each out-of-band indicator to the M in-band monitoring indicators in step S203 is as follows:

[0019] ;

[0020] in, For the first The dynamic correlation strength between one out-of-band indicator and M in-band monitoring indicators. This represents the total number of in-band monitoring metrics. For the first The in-band indicators and the first The correlation score of each out-of-band indicator.

[0021] Optionally, step S102 includes:

[0022] S301, use the specified abnormal data identification algorithm to identify anomalies in the time series data sequence of in-band monitoring indicators;

[0023] S302, for outliers of the i-th in-band index at any time t. According to the time of its occurrence Using this as a reference point, we statistically analyze its forward adjacent time windows. Within the same node, the alarm trigger frequency of each out-of-band indicator is calculated, and the abnormal correlation frequency between each out-of-band indicator and the in-band monitoring indicator in the key out-of-band monitoring indicator set is calculated.

[0024] S303, sort the out-of-band indicators in the key out-of-band monitoring indicator set in descending order according to the frequency of their abnormal correlation with in-band monitoring indicator anomalies, and select a specified number of the top ones. One out-of-band indicator was selected as the final out-of-band indicator.

[0025] Optionally, the calculation function expression for the frequency of abnormal correlation between each out-of-band indicator and the in-band monitoring indicator in step S302 is as follows:

[0026] ;

[0027] in, For the first The frequency of abnormal correlations between out-of-band indicators and in-band monitoring indicators. A set of outliers The size of the forward adjacent time window. Let t be the time of the next forward adjacent time window. for Time of the first An out-of-band indicator triggered an alert.

[0028] Optionally, before step S101, preprocessing is further included for the time-series data sequences of M in-band monitoring metrics and N out-of-band monitoring metrics from the same computer node, including if any... Time series data of in-band indicators Insufficient length Then, its length is padded to the required value using interpolation. If any number Time series data of out-of-band indicators Insufficient length Then, its length is padded to the required value using interpolation. ,in and These represent the time series data lengths for in-band and out-of-band indicators, respectively.

[0029] The present invention also provides an automatic selection system for monitoring indicators of large-scale computer systems, including a microprocessor and a memory interconnected thereto, wherein the microprocessor is programmed or configured to execute the automatic selection method for monitoring indicators of large-scale computer systems.

[0030] The present invention also provides a computer-readable storage medium storing a computer program or instructions that are programmed or configured to execute, via a processor, a method for automatically selecting monitoring metrics of the large-scale computer system.

[0031] The present invention also provides a computer program product, including a computer program or instructions that are programmed or configured to execute, via a processor, an automatic selection method for monitoring metrics of the large-scale computer system.

[0032] Compared with existing technologies, the present invention can mainly achieve the following beneficial effects: The automatic selection method of monitoring indicators for large-scale computer systems of the present invention, by associating the in-band operating status of the system, explores the semantic correlation between large-scale computer systems and entity units (physical computing nodes running operating systems) and multi-path monitoring (in-band and out-of-band). Through two stages of initial screening of out-of-band monitoring channels and selection of out-of-band sensors that are sensitive to anomalies, the automatic screening of out-of-band indicators with high correlation to in-band status, strong correlation with in-band indicators, and the ability to indicate in-band abnormal events can be achieved in an automated and unsupervised manner, effectively reducing the problems of large number of sensors, high data transmission pressure, and large processing delay. Attached Figure Description

[0033] Figure 1 This is a schematic diagram of the basic process of the method in an embodiment of the present invention.

[0034] Figure 2 This is a schematic diagram illustrating the principle of the method in an embodiment of the present invention. Detailed Implementation

[0035] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings in the embodiments of the present invention.

[0036] like Figure 1 and Figure 2 As shown, the automatic selection method for monitoring indicators of a large-scale computer system in this embodiment includes the following steps: S101, initial screening of out-of-band monitoring channels: for the time-series data sequences of M in-band monitoring indicators and N out-of-band monitoring indicators from the same computer node, quantify the dynamic correlation strength between the out-of-band monitoring indicators and the in-band monitoring indicators. Figure 2 (represented as in-band / out-band correlation analysis), selecting multiple out-of-band monitoring indicators with strong dynamic correlation to in-band monitoring indicators yields a set of key out-of-band monitoring indicators. Figure 2 (This is represented as associated and merged statistics); S102, Selection of out-of-band sensors that are sensitive to anomalies: For the set of key out-of-band monitoring indicators, extract in-band anomalies and perform in-band indicator anomaly analysis ( Figure 2Spatiotemporal correlation analysis of in-band anomaly extraction and out-of-band hardware alarms (represented in the middle) Figure 2 The out-of-band alarm statistics (represented as spatiotemporal correlation) are further optimized to obtain the final selected out-of-band indicators. Among them, computer nodes, in large-scale computer systems (supercomputing systems), refer to the basic units that can be allocated computing tasks. They are generally composed of one or more CPUs running an operating system, along with their associated memory and hardware circuits, and are the basic objects of hardware out-of-band monitoring.

[0037] Step S101 is the initial screening stage of out-of-band monitoring indicators based on in-band and out-of-band correlation analysis. Since in-band indicators directly reflect system tasks, this invention proposes quantifying the dynamic correlation strength between out-of-band sensors and in-band node indicators, and selecting key out-of-band monitoring indicators with strong correlation to in-band states. Specifically, step S101 in this embodiment includes:

[0038] S201, for time-series data sequences of M in-band monitoring indicators and N out-of-band monitoring indicators from the same computer node, performs alignment and correlation of the two indicator time sequences. Dynamic time warping technology is used to measure the similarity of the in-band and out-of-band time sequence patterns of the same monitoring entity (node), that is, calculating the similarity of any given time sequence. Time series data of in-band indicators With the Time series data of out-of-band indicators The dynamic time warping technique is used to calculate the first... The in-band indicators and the first Correlation scores of out-of-band indicators ;

[0039] S202, All related scores Construct an M×N correlation matrix;

[0040] S203, calculate the dynamic correlation strength between each out-of-band indicator and the M in-band monitoring indicators based on the correlation matrix:

[0041] ;

[0042] in, For the first The dynamic correlation strength between one out-of-band indicator and M in-band monitoring indicators;

[0043] S204: Sort the dynamic correlation strength of each out-of-band indicator with M in-band monitoring indicators in descending order, and select the top-ranked indicators by a specified number. Each out-of-band indicator serves as a set of key out-of-band monitoring indicators.

[0044] In step S201 of this embodiment, the calculation of the first... The in-band indicators and the first Correlation scores of out-of-band indicators The function expression is:

[0045] ;

[0046] in, For the first Time series data of in-band indicators The first in One time series data, For the first Time series data of out-of-band indicators The first in One time series data, This represents the optimal alignment path between two time series. , , ~ For the first Time series data of in-band indicators The first to the second One time series data, ~ For the first Time series data of out-of-band indicators The first to the second One time series data, and These represent the time-series data lengths for in-band and out-of-band metrics, respectively. Dynamic time warping addresses the issue of asynchronous data acquisition frequencies between the in-band and out-of-band metric acquisition channels due to proxying. This leads to timing misalignment issues.

[0047] By pairwise correlation calculations of M in-band monitoring indicators and N out-of-band monitoring indicators, an M×N time-series correlation matrix of in-band and out-of-band indicators is obtained. Further, this matrix is ​​transformed into an overall correlation capability assessment of the N out-of-band indicators, integrating the correlation strength of multiple in-band indicators from the same sensor channel. The functional expression for calculating the dynamic correlation strength of each out-of-band indicator to the M in-band monitoring indicators in step S203 is as follows:

[0048] ;

[0049] in, For the first The dynamic correlation strength between one out-of-band indicator and M in-band monitoring indicators. This represents the total number of in-band monitoring metrics. For the first The in-band indicators and the first The correlation scores of each out-of-band indicator are calculated. Based on this, the overall correlation scores of the N out-of-band monitoring indicators are calculated. The metrics are sorted, and the top K1 strongly correlated metrics are selected using a top-K strategy to form a subset of out-of-band monitoring metrics. This subset has good statistical indicative value for in-band operating status.

[0050] Step S102 is the sensitive indicator selection stage based on in-band indicator anomalies. Given that in-band indicator anomalies often foreshadow the triggering of out-of-band hardware alarms, this invention proposes to combine an initial screening of a subset of out-of-band monitoring indicators with spatiotemporal correlation analysis of in-band indicator anomalies and out-of-band hardware alarms to further optimize the monitoring channels of out-of-band sensors. Step S102 includes:

[0051] S301, an anomaly detection algorithm is used to identify anomalies in the time series data of in-band monitoring indicators. The required anomaly detection algorithm can be used as needed. For example, as an optional implementation, an efficient 3-sigma algorithm can be used to perform anomaly detection on the in-band operating status, identifying anomalies in the time series data of in-band monitoring indicators that exceed... Data points within a range are defined as outliers, where The sum of the historical indicator's time series The standard deviation of the indicator within the corresponding statistical range;

[0052] S302, regarding the obtained set of anomalies outliers of the i-th in-band index at any time t. According to the time of its occurrence Using this as a reference point, we statistically analyze its forward adjacent time windows. Within the same node, the alarm trigger frequency of each out-of-band indicator is calculated, and the abnormal correlation frequency between each out-of-band indicator and the in-band monitoring indicator in the key out-of-band monitoring indicator set is calculated.

[0053] S303, sort the out-of-band indicators in the key out-of-band monitoring indicator set in descending order according to the frequency of their abnormal correlation with in-band monitoring indicator anomalies, and select a specified number of the top ones. One out-of-band indicator was selected as the final out-of-band indicator.

[0054] The calculation function expression for the abnormal correlation frequency between each out-of-band indicator and the in-band monitoring indicator in step S302 of this embodiment is as follows:

[0055] ;

[0056] in, For the first The frequency of abnormal correlations between out-of-band indicators and in-band monitoring indicators. A set of outliers The size of the forward adjacent time window. Let t be the time of the next forward adjacent time window. for Time of the first An out-of-band indicator triggered an alarm once (i.e., sensor n triggered a hardware alarm). From the initially screened subset of indicators, the top K2 out-of-band alarm sensors most relevant to the in-band indicator anomaly are selected according to F_n. By prioritizing out-of-band sensors that are spatiotemporally coupled with the abnormal state of the indicator, the retained out-of-band monitoring indicators are ensured to be highly sensitive to system operational anomalies.

[0057] This embodiment's method includes two inputs: multi-dimensional in-band indicator time series and multi-dimensional out-of-band indicator time series of a large-scale computer system. As an optional implementation, before step S101, it further includes preprocessing the time series data sequences of M in-band monitoring indicators and N out-of-band monitoring indicators from the same computer node, including if any... Time series data of in-band indicators Insufficient length Then, its length is padded to the required value using interpolation methods (such as using the average of the sequences). If any number Time series data of out-of-band indicators Insufficient length Then, its length is padded to the required value using interpolation. ,in and The time series data lengths for in-band and out-of-band indicators are specified to ensure data integrity. Then, two stages of filtering can be carried out to filter indicators with strong correlation and indicators with strong anomaly indications.

[0058] To verify the automatic selection method for monitoring indicators of large-scale computer systems in this embodiment, the method was applied to a supercomputing system and compared with existing random selection and expert experience-driven methods. The comparison was conducted on two datasets, A and B, using thousands of collected indicator data (in-band indicators such as CPU utilization, memory usage, disk I / O, and network traffic, as well as sensor readings for temperature, voltage, and fan speed) and millions of alarm data (abnormalities such as sensor reading threshold exceeding limits and hardware failures detected by node-level monitoring units). The retention success rate of the Top-3 key indicators and the retention success rate of the Top-5 key indicators were selected as evaluation indicators. The main results are shown in Table 1.

[0059] Table 1: Key Indicator Retention Success Rate for Different Indicator Selection Methods

[0060]

[0061] As shown in Table 1, the top-5 key indicator retention success rate (the ability of selected monitoring indicators to retain key system status information) of the method in this embodiment can reach 94%, with an indicator selection time of less than 4.5 seconds. Under the same conditions, the key indicator retention success rate of the traditional expert experience-driven static indicator selection method reaches at most 72.4%, and the random selection method is even lower, reaching at most 49.8%. Therefore, the method in this embodiment can effectively reduce the problems of a large number of data acquisition sensors, high data transmission pressure, and large processing latency.

[0062] In summary, to overcome the problem that existing indicator channel selection techniques are geared towards single-task scenarios of state detection or time series classification, and do not meet the general requirements of large-scale computer system indicator collection for operational status monitoring, this embodiment utilizes spatiotemporal correlation analysis between multiple monitoring channels within the system itself to select out-of-band monitoring indicators based on a comprehensive consideration of representativeness and importance. This is a general indicator selection method for monitoring large-scale computer systems. This embodiment includes: 1) Initial screening of out-of-band monitoring indicators based on in-band and out-of-band correlation analysis: In-band and out-of-band indicators are attributes of nodes in a large-scale computing system, reflecting the state changes of the node at the system and hardware levels. Considering that in-band and out-of-band indicators occur synchronously and are collected through physical nodes, possessing inherent spatiotemporal (generation time) and physical node correlation, this invention proposes to evaluate the correlation between in-band and out-of-band time series channels on a node, calculate the correlation strength of each out-of-band time series on the node to all in-band time series, and rank them by cross-channel monitoring correlation strength, selecting a series of highly correlated out-of-band time series from all out-of-band time series. 2) Sensitive Indicator Selection Based on In-Band Anomalies: Initial screening of out-of-band indicators only considers their correlation with in-band indicators. The key role of monitoring is to promptly detect system anomalies. This invention proposes, based on the out-of-band indicator sequences initially screened based on correlation, to evaluate the sensitivity of out-of-band indicator time series to system anomalies by calculating the frequency of alarm triggers before the occurrence of in-band time series anomalies. By ranking the frequency of alarm triggers in advance when system anomalies occur, a series of out-of-band time series indicators with high sensitivity to system anomalies are selected. Compared to traditional industry-based indicator filtering methods based on expert-configured blacklists and whitelists, this embodiment's method completes indicator selection in an unsupervised and automated manner. Its correlation analysis, anomaly detection, and adjacent window alarm statistics all require no annotation or manual intervention. It supports controlling the selection intensity by setting the K value in the top-K algorithm, adapting to the system's data collection capabilities and needs, and avoiding excessive indicator deletion or low indicator compression.

[0063] Furthermore, this embodiment also provides an automatic selection system for monitoring indicators of a large-scale computer system, including a microprocessor and a memory interconnected, wherein the microprocessor is programmed or configured to execute the automatic selection method for monitoring indicators of the large-scale computer system. This embodiment also provides a computer-readable storage medium storing a computer program or instructions programmed or configured to execute the automatic selection method for monitoring indicators of a large-scale computer system via a processor. This embodiment also provides a computer program product, including a computer program or instructions programmed or configured to execute the automatic selection method for monitoring indicators of a large-scale computer system via a processor.

[0064] Those skilled in the art will understand that the technical solutions provided by this invention may take the form of a method, system, or computer program product. Therefore, this invention may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this invention may take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce an implementation of the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0065] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for automatically selecting monitoring indicators for large-scale computer systems, characterized in that, The process includes the following steps: S101, for time-series data sequences of M in-band monitoring indicators and N out-of-band monitoring indicators from the same computer node, quantifying the dynamic correlation strength between the out-of-band monitoring indicators and the in-band monitoring indicators, and selecting several out-of-band monitoring indicators with the highest dynamic correlation strength with the in-band monitoring indicators to obtain a set of key out-of-band monitoring indicators; S102, for the set of key out-of-band monitoring indicators, further optimizing the key out-of-band monitoring indicators by extracting in-band anomalies and performing spatiotemporal correlation analysis between in-band indicator anomalies and out-of-band hardware alarms to obtain the final selected out-of-band indicators; Step S101 includes: S201, for time-series data sequences of M in-band monitoring indicators and N out-of-band monitoring indicators from the same computer node, calculate the arbitrary i-th... Time series data of in-band indicators With the Time series data of out-of-band indicators The dynamic time warping technique is used to calculate the first... The in-band indicators and the first Correlation scores of out-of-band indicators ; S202, All related scores Construct an M×N correlation matrix; S203, calculate the dynamic correlation strength between each out-of-band indicator and the M in-band monitoring indicators based on the correlation matrix: ; in, For the first The dynamic correlation strength between one out-of-band indicator and M in-band monitoring indicators; S204: Sort the dynamic correlation strength of each out-of-band indicator with M in-band monitoring indicators in descending order, and select the top-ranked indicators by a specified number. A set of out-of-band indicators as key out-of-band monitoring indicators; Step S102 includes: S301, use the specified abnormal data identification algorithm to identify anomalies in the time series data sequence of in-band monitoring indicators; S302, for outliers of the i-th in-band index at any time t. According to the time of its occurrence Using this as a reference point, we statistically analyze its forward adjacent time windows. Within the same node, the alarm trigger frequency of each out-of-band indicator is calculated, and the abnormal correlation frequency between each out-of-band indicator and the in-band monitoring indicator in the key out-of-band monitoring indicator set is calculated. S303, sort the out-of-band indicators in the key out-of-band monitoring indicator set in descending order according to the frequency of their abnormal correlation with in-band monitoring indicator anomalies, and select a specified number of the top ones. One out-of-band indicator was selected as the final out-of-band indicator.

2. The method for automatically selecting monitoring indicators for large-scale computer systems according to claim 1, characterized in that, In step S201, the calculation of the first The in-band indicators and the first Correlation scores of out-of-band indicators The function expression is: ; in, For the first Time series data of in-band indicators The first in One time series data, For the first Time series data of out-of-band indicators The first in One time series data, This represents the optimal alignment path between two time series. , , ~ For the first Time series data of in-band indicators The first to the second One time series data, ~ For the first Time series data of out-of-band indicators The first to the second One time series data, and These represent the time series data lengths for in-band and out-of-band indicators, respectively.

3. The method for automatically selecting monitoring indicators for large-scale computer systems according to claim 1, characterized in that, The functional expression for calculating the dynamic correlation strength of each out-of-band indicator with the M in-band monitoring indicators in step S203 is as follows: ; in, For the first The dynamic correlation strength between one out-of-band indicator and M in-band monitoring indicators. This represents the total number of in-band monitoring metrics. For the first The in-band indicators and the first The correlation score of each out-of-band indicator.

4. The method for automatically selecting monitoring indicators for large-scale computer systems according to claim 1, characterized in that, The calculation function expression for the frequency of abnormal correlation between each out-of-band indicator and the in-band monitoring indicator in step S302 is as follows: ; in, For the first The frequency of abnormal correlations between out-of-band indicators and in-band monitoring indicators. A set of outliers The size of the forward adjacent time window. Let t be the time of the next forward adjacent time window. for Time of the first An out-of-band indicator triggered an alert.

5. The method for automatically selecting monitoring indicators for large-scale computer systems according to claim 1, characterized in that, Before step S101, the process also includes preprocessing the time-series data sequences of M in-band monitoring metrics and N out-of-band monitoring metrics from the same computer node, including if any... Time series data of in-band indicators Insufficient length Then, its length is padded to the required value using interpolation. If any number Time series data of out-of-band indicators Insufficient length Then, its length is padded to the required value using interpolation. ,in and These represent the time series data lengths for in-band and out-of-band indicators, respectively.

6. An automatic selection system for monitoring indicators of a large-scale computer system, comprising a microprocessor and a memory interconnected, characterized in that, The microprocessor is programmed or configured to execute the automatic selection method for monitoring metrics of a large-scale computer system as described in any one of claims 1 to 5.

7. A computer-readable storage medium storing a computer program or instructions, characterized in that, The computer program or instructions are programmed or configured to execute, via a processor, the automatic selection method for monitoring metrics of a large-scale computer system as described in any one of claims 1 to 5.

8. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instructions are programmed or configured to execute, via a processor, the automatic selection method for monitoring metrics of a large-scale computer system as described in any one of claims 1 to 5.