Server index abnormity monitoring method and system

By using server BMC event information and log information in the data center operation and maintenance scenario, combining the operation and maintenance fault database and indicator history database, the index abnormality detection algorithm is dynamically adjusted, and the server indicator abnormality detection problem is solved, which improves detection efficiency and accuracy, and reduces operation and maintenance costs.

CN120066834APending Publication Date: 2025-05-30BEIJING TIANDI HUIYUN TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510107982.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In large-scale data center operation and maintenance scenarios, it is difficult for the existing technology to effectively monitor and detect abnormal server indicators, resulting in serious false alarms and underreporting, which brings troubles and commercial risks to operation and maintenance personnel and operations.

Method used

By obtaining event information and log information in the server BMC, the weight values ​​of each indicator are calculated based on the operation and maintenance fault library analysis, and classifying the indicators according to the indicator history library, determining their type, selecting suitable indicator abnormality detection algorithms, dynamically adjusting algorithm parameters, outputting detection results, and alarm classification based on the weight value.

Benefits of technology

It improves the efficiency and accuracy of indicator abnormality detection, reduces the false alarm rate and missed rate of alarms, reduces the workload and configuration adjustment of operation and maintenance personnel, and improves operation and maintenance efficiency and business stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066834A_ABST
    Figure CN120066834A_ABST
Patent Text Reader

Abstract

The invention provides a server index abnormity monitoring method and system. The server index abnormity monitoring method comprises the following steps: acquiring event information and log information in a server BMC (Baseboard Management Controller); analyzing the event information and the log information based on an operation and maintenance fault library, and calculating a weight value of each index; classifying the indexes according to an index history library, and determining the types of the indexes; determining a corresponding index anomaly detection algorithm according to the type of the index; selecting at least one index anomaly detection algorithm from the corresponding index anomaly detection algorithms according to the weight values of the indexes; inputting the current index into the selected index anomaly detection algorithm, dynamically adjusting parameters of the selected index anomaly detection algorithm, and outputting a detection result; when the detection result is abnormal, performing alarm grading according to the weight value of the current index; the efficiency and the accuracy of index anomaly detection can be improved, and the false alarm rate and the missing alarm rate of alarms can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer operation and maintenance, and in particular to an abnormal monitoring method and system for server metrics. Background Art

[0002] In the actual operation and maintenance scenario of a computing power data center, due to the large number of hyperscale computing power servers, it is usually necessary to monitor tens of thousands to millions of metrics. Generally, the method of using fixed thresholds is usually adopted to carry out abnormal monitoring on them. Due to the large number of these data, operation and maintenance personnel cannot set reasonable upper and lower thresholds for each metric based on business experience; on the other hand, monitoring metrics have rich forms from a data perspective (data types, such as stationarity, periodicity, trendiness, etc.), and simple fixed thresholds cannot work for all forms. Therefore, there is a common phenomenon of false alarms in abnormal monitoring in this environment, which brings great troubles to operation and maintenance personnel and also brings greater business risks to the operation of the computing power center.

[0003] In the field of IT operation and maintenance, to ensure the normal operation of services, usually the first step is to monitor the objects of operation and maintenance, and mainly among them is to monitor the metrics of the objects of operation and maintenance in real time: detect the metrics in real time through set rules, and when a certain metric value does not conform to the set rules, it is determined as abnormal, and then send corresponding alarms to the alarm platform.

[0004] After receiving the alarm, the alarm platform will assign it to the corresponding operation and maintenance personnel for processing. The operation and maintenance personnel will troubleshoot the problems of the server according to the alarm information, finally locate the root cause of the failure, and repair the failure. From this process, it can be seen that the whole process is centered around alarms, so the quality of alarms is crucial. Then, how to ensure the quality of metric-based alarms? This requires using accurate and effective rules to detect abnormalities in metrics.

[0005] Problems often faced in abnormal detection of server metrics:

[0006] Large scale. For example, with the development of business in the current AIDC, large-scale systems and equipment have led to an exponential increase in the metrics to be monitored. Using the original method of fixed-threshold monitoring and multi-metric aggregation detection not only requires the business experience of operation and maintenance personnel but also increases the operation and maintenance workload.

[0007] In the actual scenario of operation and maintenance, the alarm metric data is frequent, and it is difficult for operation and maintenance personnel to grasp which are the important alarm metrics among numerous alarm metrics.

[0008] Diversified metric types, and the forms of their metrics are ever-changing. For example, point anomalies, context anomalies, and continuity anomalies. There is no single and invariant monitoring method that can cover all metric situations. Summary of the Invention

[0009] In view of this, the purpose of the present invention is to provide a method and system for abnormal monitoring of server metrics, to solve the problem of abnormal detection of server metrics, improve the efficiency and accuracy of metric abnormal detection, and reduce the false alarm rate and missed alarm rate of alarms.

[0010] In a first aspect, an embodiment of the present invention provides a method for abnormal monitoring of server metrics, and the method includes:

[0011] Obtain event information and log information in the server BMC;

[0012] Analyze the event information and log information based on the operation and maintenance fault library, and calculate the weight values of each metric;

[0013] Classify the metrics according to the metric history library, and determine the types of the metrics;

[0014] Determine the corresponding metric abnormal detection algorithm according to the type of the metric;

[0015] Select at least one metric abnormal detection algorithm from the corresponding metric abnormal detection algorithms according to the weight value of the metric;

[0016] Input the current metric into the selected metric abnormal detection algorithm, and dynamically adjust the parameters of the selected metric abnormal detection algorithm, and output the detection result;

[0017] When the detection result is abnormal, perform alarm grading according to the weight value of the current metric.

[0018] Further, analyzing the event information and log information based on the operation and maintenance fault library, and calculating the weight values of each metric, includes:

[0019] Obtain the metrics of the event information;

[0020] Count the abnormal frequency of the metrics of the event information;

[0021] Count the abnormal frequency of the metrics of the log information;

[0022] Based on the historical fault data in the operation and maintenance fault library, compare and analyze the abnormal frequency of the metrics of the event information and the abnormal frequency of the metrics of the log information, and obtain historical fault data similar to the current abnormal situation;

[0023] Calculate the failure rate caused by each metric abnormality according to the historical fault data;

[0024] Assign corresponding weight values to each metric according to the failure rate caused by each metric abnormality;

[0025] Among them, the weight value is used to represent the importance of the indicator in the overall operation and maintenance work.

[0026] Further, classify the indicator according to the indicator historical database to determine the type of the indicator, including:

[0027] Calculate the standard deviation of the time series data;

[0028] Compare the standard deviation with the first set threshold;

[0029] When the standard deviation is less than the first set threshold, the type of the indicator is a fluctuating indicator.

[0030] Further, classify the indicator according to the indicator historical database to determine the type of the indicator, including:

[0031] Smooth the time series data through EWMA;

[0032] Fit the processed time series data with a univariate linear regression curve to obtain the k value;

[0033] Compare the k value with the second set threshold;

[0034] When the k value is greater than the second set threshold, the type of the indicator is a trend indicator.

[0035] Further, classify the indicator according to the indicator historical database to determine the type of the indicator, including:

[0036] Obtain the time series data of the current day, the time series data of the previous day, and the time series data of the previous week corresponding to the time series data of the current day;

[0037] Normalize the time series data of the current day, the time series data of the previous day, and the time series data of the previous week respectively;

[0038] Calculate the first mean square error between the processed time series data of the current day and the processed time series data of the previous day;

[0039] Calculate the second mean square error between the processed time series data of the current day and the processed time series data of the previous week;

[0040] Select the minimum value from the first mean square error and the second mean square error;

[0041] Compare the minimum value with the third set threshold;

[0042] When the minimum value is less than the third set threshold, the type of the indicator is a seasonal indicator.

[0043] Further, the method further includes:

[0044] Perform a Fourier transform on the time series according to the historical index library to obtain the main frequency component;

[0045] Calculate the index period value according to the main frequency component;

[0046] Wherein, the time series is a sequence of data points arranged in the order of time occurrence, and the time interval of the time series is the frequency of the time series data.

[0047] Furthermore, the indicators of the event information include the inlet temperature of the server, the outlet temperature of the server, and the GPU sensor temperature.

[0048] In a second aspect, an embodiment of the present invention provides an abnormal monitoring system for server indicators, the system includes:

[0049] An acquisition module, configured to acquire event information and log information in the server BMC;

[0050] An analysis module, configured to analyze the event information and log information based on an operation and maintenance fault library, and calculate the weight value of each indicator;

[0051] A classification module, configured to classify the indicators according to the historical index library, and determine the type of the indicators;

[0052] A determination module, configured to determine a corresponding indicator anomaly detection algorithm according to the type of the indicator;

[0053] A selection module, configured to select at least one indicator anomaly detection algorithm from the corresponding indicator anomaly detection algorithms according to the weight value of the indicator;

[0054] An adjustment module, configured to input the current indicator into the selected indicator anomaly detection algorithm, and dynamically adjust the parameters of the selected indicator anomaly detection algorithm, and output a detection result;

[0055] An alarm module, configured to perform alarm classification according to the weight value of the current indicator when the detection result is abnormal.

[0056] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, and a computer program that can run on the processor is stored on the memory. When the processor executes the computer program, the above-mentioned method is implemented.

[0057] In a fourth aspect, an embodiment of the present invention provides a computer-readable medium having non-volatile program code executable by a processor, and the program code causes the processor to execute the above-mentioned method.

[0058] The embodiments of the present invention provide a method and a system for abnormal monitoring of server metrics, including: obtaining event information and log information in the server BMC; analyzing the event information and log information based on an operation and maintenance failure database to calculate the weight values of each metric; classifying the metrics according to a metric history library to determine the types of the metrics; determining corresponding metric abnormal detection algorithms according to the types of the metrics; selecting at least one metric abnormal detection algorithm from the corresponding metric abnormal detection algorithms according to the weight values of the metrics; inputting the current metric into the selected metric abnormal detection algorithm and dynamically adjusting the parameters of the selected metric abnormal detection algorithm to output a detection result; when an abnormal detection result exists, performing alarm grading according to the weight value of the current metric; which can improve the efficiency and accuracy of metric abnormal detection and reduce the false alarm rate and missed alarm rate of alarms.

[0059] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, the claims, and the drawings.

[0060] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, is described in detail as follows. Description of the Drawings

[0061] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0062] Figure 1 Flowchart of the method for abnormal monitoring of server metrics provided in Embodiment 1 of the present invention;

[0063] Figure 2 Indicator schematic diagram of event information provided in Embodiment 1 of the present invention;

[0064] Figure 3 Relationship schematic diagram of failure rate and weight value provided in Embodiment 1 of the present invention;

[0065] Figure 4 Morphology schematic diagram of time series data provided in Embodiment 1 of the present invention;

[0066] Figure 5 Schematic diagram of the system for abnormal monitoring of server metrics provided in Embodiment 2 of the present invention. Detailed Embodiments

[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0068] Explanation of technical terms involved in this application:

[0069] Time series: It is a sequence of data points arranged in the order of occurrence in time. The time interval of the time series is the frequency of this time series data.

[0070] Anomaly: In a time series, an anomaly refers to an unexpected change in the pattern of one or more signals.

[0071] Common single-index anomaly detection is the process of detecting anomalies in the time series data of a single variable. In the field of operation and maintenance, the anomaly detection of indicators such as the usage of resources such as gpu, cpu, mem, disk, and network traffic all belongs to single-index anomaly detection.

[0072] This application calculates the index weight value according to the server BMC event log and the failure rate caused by index anomalies, and monitors the multi-algorithm model according to the weight value to make the alarm more accurate; adopts an unsupervised method to dynamically adjust the threshold and acquisition frequency, reducing the workload of operation and maintenance personnel for configuring and adjusting the thresholds of many indicators; classifies the indicators, self-adapts to the conventional detection algorithm model, covers the detection of diverse index types, reduces the complexity of operation and maintenance detection, has stronger universality, faster detection, and more timely alarm.

[0073] To facilitate the understanding of this embodiment, the embodiments of the present invention will be introduced in detail below.

[0074] Embodiment 1:

[0075] Figure 1 This is the flow chart of the anomaly monitoring method for server indicators provided in Embodiment 1 of the present invention.

[0076] Refer to Figure 1 , this method includes the following steps:

[0077] Step S101, obtain the event information and log information in the server BMC;

[0078] Step S102, analyze the event information and log information based on the operation and maintenance failure library, and calculate the weight value of each indicator;

[0079] Specifically, based on the event information and log information in the server BMC, a method for calculating the metric weights based on the operation and maintenance failure database. By analyzing the abnormal situations in the BMC event logs and combining with the historical failure data in the operation and maintenance failure database, the failure rate caused by each metric abnormality is calculated, and then the weight value of each metric is obtained. This weight value can be used to represent the importance of the metric in the overall operation and maintenance work, so as to give more optimizations in the subsequent monitoring and dynamic adjustment processes to ensure the accuracy of the detection of important metrics. In the actual scenario of operation and maintenance, the metric data of alarms is frequent, and it is necessary to determine which abnormal metrics need to be focused on and perform operation and maintenance processing in a timely manner after an alarm is issued.

[0080] Step S103: Classify the metrics according to the metric history library and determine the types of the metrics.

[0081] Here, aiming at the problem of diverse metric types. The forms of the metric time series data are various (as Figure 4 shown), and the server monitoring metric time series data is classified through an effective classification method.

[0082] For data metric classification, in an actual business system, usually tens of thousands to millions of metrics need to be monitored. Generally, enterprises usually use the method of fixed thresholds to carry out abnormal monitoring on them. Due to the large number of these data, it is impossible for operation and maintenance personnel to set reasonable upper and lower thresholds for each metric based on business experience; on the other hand, from the data perspective, the monitoring metrics have rich forms, and simple fixed thresholds cannot work for all form types.

[0083] Reasonable metric data classification helps in the selection of abnormal detection algorithms. Each type of algorithm has its applicable metric data type. If the metric data type is not discriminated and simply input into the algorithm for abnormal detection, it is very likely that the detection result will be incorrect. Therefore, metric data classification is crucial for the selection of abnormal detection algorithms. On the one hand, on the premise of fully understanding the characteristics of various abnormal detection algorithms, if the metric data can be reasonably classified, then various algorithms can be accurately matched to improve the accuracy of abnormal detection; on the other hand, the metric classification based on algorithm adaptability can also avoid the algorithm selection process of blindly trying different algorithms on the same metric data, greatly reducing the calculation amount.

[0084] By classifying the metric data and matching different abnormal detection algorithms, the accuracy of abnormal detection can be improved. In the actual implementation process, the algorithm can be adjusted and optimized, but the metric data cannot be selected and changed. Therefore, when performing metric classification, in this application, by considering the metric characteristics, after classification, existing algorithms are matched for each type of data or new algorithms are developed, rather than classifying the metric data based on the data characteristics that existing algorithms are good at processing.

[0085] Step S104: Determine the corresponding metric abnormal detection algorithm according to the type of the metric.

[0086] Step S105: Select at least one index anomaly detection algorithm from the corresponding index anomaly detection algorithms according to the weight value of the index.

[0087] Step S106: Input the current index into the selected index anomaly detection algorithm, and dynamically adjust the parameters of the selected index anomaly detection algorithm, and output the detection result.

[0088] Step S107: When the detection result is abnormal, perform alarm classification according to the weight value of the current index.

[0089] Here, the detection algorithm model is adapted according to the type of the index, covering the detection of diverse index types, reducing the complexity of operation and maintenance detection. For the large-scale index data, the collection period or threshold is dynamically adjusted in an unsupervised manner, reducing the workload of operation and maintenance personnel for configuring and adjusting the thresholds of numerous indexes. This makes the unsupervised algorithm have greater advantages than the supervised algorithm, and the simple model has greater advantages than the complex model: saving manpower, having strong universality, and high detection efficiency.

[0090] Specifically, adjust the data collection of the index according to the index classification and weight value. For example, for the index with a large weight, increase the collection frequency of the index. For the periodic index, adjust the comparative analysis of the model data according to the index period value.

[0091] Determine the corresponding index anomaly detection algorithm according to the type of the index. For example, for the fluctuating type index, algorithms such as EWMA smoothing, CUSUM, and box plot are used for detection; for the trend type index, algorithms such as long-term and short-term month-on-month ratio, polynomial regression, and isolation forest are used for detection; for the seasonal type index, algorithms such as year-on-year ratio, year-on-year amplitude, and linear regression are used for detection.

[0092] Select at least one index anomaly detection algorithm from the corresponding index anomaly detection algorithms according to the weight value of the index. For example, for the index with a large weight, 2 or more algorithm models need to be selected for detection, and a voting decision is made on the results of multiple algorithm model detections to improve the detection accuracy.

[0093] Input the current index into the selected index anomaly detection algorithm, and dynamically adjust the parameters of the selected index anomaly detection algorithm. For example, for the trend type index, the detection calculation is performed according to the index period trend prediction value.

[0094] When the detection result is abnormal, perform alarm classification according to the weight value of the index. The higher the weight, the higher the alarm level, prompting the operation and maintenance personnel to focus on it and discover and solve the fault in time.

[0095] Furthermore, step S102 includes the following steps:

[0096] Step S201: Obtain the index of the event information.

[0097] Step S202, count the abnormal frequency of the metrics of the event information;

[0098] Here, referring to Figure 2 , the metrics of the event information include the inlet temperature of the server, the outlet temperature of the server, and the GPU sensor temperature, and then count the abnormal frequency of these monitoring metrics.

[0099] Step S203, count the abnormal frequency of the metrics of the log information;

[0100] Step S204, based on the historical fault data in the operation and maintenance fault library, conduct a comparative analysis on the abnormal frequency of the metrics of the event information and the abnormal frequency of the metrics of the log information to obtain historical fault data similar to the current abnormal situation;

[0101] Step S205, calculate the failure rate caused by the abnormality of each metric according to the historical fault data;

[0102] Step S206, assign corresponding weight values to each metric according to the failure rate caused by the abnormality of each metric; among them, the weight value is used to represent the importance of the metric in the overall operation and maintenance work.

[0103] Specifically, associate the abnormal situation in the analyzed BMC event log with the operation and maintenance fault library. The operation and maintenance fault library is a database that records historical fault information and solutions. Through comparative analysis, historical fault case data similar to the current abnormal situation can be found. According to the historical fault data in the operation and maintenance fault library, calculate the failure rate caused by the abnormality of each metric.

[0104] According to the historical fault data in the operation and maintenance fault library, calculate the failure rate caused by the abnormality of each metric, and assign a weight value to each metric according to the calculated failure rate. This weight value can be used to represent the importance of the metric in the overall operation and maintenance work. For example, after the statistics of the alarm generated by the outlet temperature (Outet_Temp) of the server in the server event log are relatively high, and when the number of faults in the operation and maintenance fault library is relatively large (such as Figure 3 shown), the calculated weight value of this metric is large.

[0105] Furthermore, step S103 includes the following steps:

[0106] Step S301, calculate the standard deviation of the time series data;

[0107] Step S302, compare the standard deviation with the first set threshold;

[0108] Step S303, when the standard deviation is less than the first set threshold, the type of the metric is a fluctuating metric.

[0109] Further, step S103 includes the following steps:

[0110] Step S401, smoothing the time series data through EWMA;

[0111] Step S402, fitting the processed time series data with a univariate linear regression curve to obtain the k value;

[0112] Step S403, comparing the k value with the second set threshold;

[0113] Step S404, when the k value is greater than the second set threshold, the type of the indicator is a trend type indicator.

[0114] Further, step S103 includes the following steps:

[0115] Step S501, obtaining the time series data of the current day, the time series data of the previous day, and the time series data of the previous week corresponding to the time series data of the current day;

[0116] Step S502, normalizing the time series data of the current day, the time series data of the previous day, and the time series data of the previous week respectively;

[0117] Step S503, calculating the first mean square error between the processed time series data of the current day and the processed time series data of the previous day;

[0118] Step S504, calculating the second mean square error between the processed time series data of the current day and the processed time series data of the previous week;

[0119] Step S505, selecting the minimum value from the first mean square error and the second mean square error;

[0120] Step S506, comparing the minimum value with the third set threshold;

[0121] Step S507, when the minimum value is less than the third set threshold, the type of the indicator is a seasonal type indicator.

[0122] Further, the method further includes the following steps:

[0123] Step S601, performing a Fourier transform on the time series according to the indicator history library to obtain the main frequency component;

[0124] Step S602, calculating the indicator period value according to the main frequency component;

[0125] Wherein, the time series is a sequence of data points arranged in the order of time occurrence, and the time interval of the time series is the frequency of the time series data.

[0126] Specifically, the Fourier transform is a mathematical operation that decomposes a complex signal into a set of simpler sine and cosine waves. The Fourier transform is widely used in signal processing, communication, image processing, and many other scientific and engineering fields. It allows signals to be analyzed and manipulated in the frequency domain, which is often a more natural and intuitive way to understand and process signals than in the time domain. Specifically refer to the following procedure:

[0127] from scipy.fft import fft

[0128] #Calculate the Fourier transform

[0129] yf = np.fft.fft(record_by_date)

[0130] xf = np.linspace(0.0, 1.0 / (2.0), len(record_by_date) / / 2)

[0131] #Find the dominant frequency

[0132] #We have to drop the first element of the fft as it corresponds to

[0133] the

[0134] #DC component or the average value of the signal

[0135] idx = np.argmax(np.abs(yf[1:len(record_by_date) / / 2]))

[0136] freq = xf[idx]

[0137] period = (1 / freq)

[0138] print(f"The period of the time series is {period}")

[0139] Among them, the output is The period of the time series is 7.030927835051545. Therefore, the period of the time series is 7 days.

[0140] This application has the following beneficial effects:

[0141] 1) Improve the operation and maintenance efficiency of hardware failures: Through an efficient abnormal detection mechanism for metrics, fault events can be detected and alerted in a timely manner. This rapid response ability enables the operation and maintenance team to take actions within a short period after a failure occurs, thus significantly improving the operation and maintenance efficiency of hardware failures.

[0142] 2) Reduce resource waste and costs: Abnormal detection of metrics helps avoid computing power waste and cluster idling caused by hardware failures. For example, during AI training, a GPU failure may lead to training interruption, computing waste, and resource idling. Through rapid detection and handling, these wastes can be minimized, resource utilization can be improved, and thus the operating costs of enterprises can be reduced.

[0143] 3) Ensure business stability: By classifying metrics and their weights and self - adapting corresponding abnormal detection algorithms, the cumbersome workload of threshold setting in operation and maintenance work can be reduced, the quality of abnormal detection can be effectively improved, and the false alarm and misreport rates caused by the past static threshold method are greatly reduced.

[0144] 4) Reduce the operation and maintenance workload: There are various abnormal detection algorithms, including statistical algorithms, machine learning algorithms, and deep learning algorithms, etc., each with its own advantages and disadvantages. The metric data in the operation and maintenance field also has rich forms. Facing a large number of metric data with different types, there is currently no general abnormal detection method that can meet the abnormal detection requirements of various metrics. When front - line operation and maintenance personnel configure metric monitoring, although they can select algorithms based on their own experience to obtain higher detection accuracy compared to the fixed - threshold method. However, algorithm selection and parameter tuning are still the biggest difficulties faced by operation and maintenance personnel. Through the research on metric classification and the adaptability between metrics and algorithms, not only can the accuracy of abnormal detection be improved, but also the work burden of operation and maintenance personnel can be truly reduced.

[0145] Example Two:

[0146] Figure 5 An abnormal monitoring system for server metrics provided by the second embodiment of the present invention.

[0147] Refer to Figure 5 and this system includes:

[0148] An acquisition module, used to acquire event information and log information in the server BMC;

[0149] An analysis module, used to analyze the event information and log information based on the operation and maintenance fault library and calculate the weight values of each metric;

[0150] A classification module, used to classify the metrics according to the metric history library and determine the types of the metrics;

[0151] A determination module, configured to determine a corresponding index anomaly detection algorithm according to the type of the index;

[0152] A selection module, configured to select at least one index anomaly detection algorithm from the corresponding index anomaly detection algorithms according to the weight value of the index;

[0153] An adjustment module, configured to input the current index into the selected index anomaly detection algorithm, and dynamically adjust the parameters of the selected index anomaly detection algorithm, and output a detection result;

[0154] An alarm module, configured to perform alarm classification according to the weight value of the current index when an anomaly exists in the detection result.

[0155] An embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the anomaly monitoring method for server indexes provided in the above embodiment are implemented.

[0156] An embodiment of the present invention further provides a computer-readable medium having non-volatile program code executable by a processor. A computer program is stored on the computer-readable medium. When the computer program is run by the processor, the steps of the anomaly monitoring method for server indexes in the above embodiment are executed.

[0157] The computer program product provided by the embodiment of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiment. For specific implementation, reference can be made to the method embodiment, which will not be elaborated herein.

[0158] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein.

[0159] In addition, in the description of the embodiments of the present invention, unless otherwise clearly defined and limited, the terms "installation", "connection", and "connection" shall be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0160] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0161] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0162] Finally, it should be noted that the above-mentioned embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for monitoring abnormality of server indicators, characterized in that: The method comprises: Obtaining event information and log information in the server BMC; Analyze the event information and the log information based on the operation and maintenance fault library, and calculate the weight value of each indicator; Classify the indicator according to the indicator history library to determine the type of the indicator; Determine a corresponding indicator anomaly detection algorithm according to the type of the indicator; Selecting at least one indicator anomaly detection algorithm from the corresponding indicator anomaly detection algorithms according to the weight value of the indicator; Input the current indicator into the selected indicator anomaly detection algorithm, dynamically adjust the parameters of the selected indicator anomaly detection algorithm, and output the detection result; When there is an abnormality in the detection result, an alarm classification is performed according to the weight value of the current indicator.

2. The method for monitoring abnormality of server indicators according to claim 1, characterized in that: The event information and the log information are analyzed based on the operation and maintenance fault library to calculate the weight value of each indicator, including: An indicator for obtaining information about the event; Counting abnormal frequencies of indicators of the event information; Counting the abnormal frequency of the indicators of the log information; Based on the historical fault data in the operation and maintenance fault library, the abnormal frequency of the indicator of the event information and the abnormal frequency of the indicator of the log information are compared and analyzed to obtain historical fault data similar to the current abnormal situation; Calculate the failure rate caused by the abnormality of each indicator according to the historical failure data; According to the failure rate caused by the abnormality of each indicator, a corresponding weight value is assigned to each indicator; The weight value is used to represent the importance of the indicator in the overall operation and maintenance work.

3. The method for monitoring abnormality of server indicators according to claim 1, characterized in that: The indicator is classified according to the indicator history library to determine the type of the indicator, including: Calculate the standard deviation of time series data; comparing the standard deviation with a first set threshold; When the standard deviation is less than the first set threshold, the type of the indicator is a fluctuation indicator.

4. The method for monitoring abnormality of server indicators according to claim 1, characterized in that: The indicator is classified according to the indicator history library to determine the type of the indicator, including: Smoothing time series data through EWMA; The processed time series data is fitted with a univariate linear regression curve to obtain the k value; Comparing the k value with a second set threshold; When the k value is greater than the second set threshold, the type of the indicator is a trend indicator.

5. The method for monitoring abnormality of server indicators according to claim 1, characterized in that: The indicator is classified according to the indicator history library to determine the type of the indicator, including: Obtaining the time series data of the current day and the time series data of the previous day, as well as the time series data of the previous week corresponding to the time series data of the current day; Normalizing the time series data of the current day, the time series data of the previous day, and the time series data of the previous week respectively; Calculate the first mean square error between the processed time series data of the day and the processed time series data of the previous day; Calculate the second mean square error between the processed time series data of the day and the processed time series data of the previous week; Select a minimum value from the first mean square error and the second mean square error; comparing the minimum value with a third set threshold; When the minimum value is less than the third set threshold, the type of the indicator is a seasonal indicator.

6. The method for monitoring abnormality of server indicators according to claim 1, characterized in that: The method further comprises: Performing Fourier transformation on the time series according to the indicator history library to obtain the main frequency component; Calculate the index period value according to the main frequency component; The time series is a sequence of data points arranged in chronological order, and the time interval of the time series is the frequency of the time series data.

7. The method for monitoring abnormality of server indicators according to claim 2, characterized in that: The indicators of the event information include the air inlet temperature of the server, the air outlet temperature of the server, and the GPU sensor temperature.

8. A server indicator abnormality monitoring system, characterized in that: The system comprises: An acquisition module, used to acquire event information and log information in the server BMC; An analysis module is used to analyze the event information and log information based on the operation and maintenance fault library and calculate the weight value of each indicator; A classification module, used to classify the indicator according to the indicator history library and determine the type of the indicator; A determination module, used to determine a corresponding indicator anomaly detection algorithm according to the type of the indicator; A selection module, used to select at least one indicator anomaly detection algorithm from the corresponding indicator anomaly detection algorithms according to the weight value of the indicator; An adjustment module, used to input the current indicator into the selected indicator anomaly detection algorithm, dynamically adjust the parameters of the selected indicator anomaly detection algorithm, and output the detection result; The alarm module is used to classify the alarm according to the weight value of the current indicator when there is an abnormality in the detection result.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer readable medium having a non-volatile program code executable by a processor, characterized in that: The program code enables the processor to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Hardware fault real-time detection method and system based on cooperation of CPU and BMC

    CN121008967A