A data center facility monitoring method, device, equipment, medium and product

By using predictive models and user interfaces to automatically generate adaptive thresholds in the data center facility monitoring system, the problem of low efficiency in setting data center monitoring thresholds is solved, and efficient and accurate monitoring and alarms are achieved.

CN121217829BActive Publication Date: 2026-08-25INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511784376.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-08-25
Estimated Expiration
2045-11-28

AI Technical Summary

Technical Problem

Data centers require manual intervention when setting monitoring thresholds, which leads to inefficiency and an inability to adapt to changes in massive infrastructure, diverse equipment types, and operating environments.

Method used

The predictive model generates performance metric predictions, which, combined with user-input deviation parameters, automatically set adaptive thresholds, reducing manual intervention and improving threshold setting efficiency.

Benefits of technology

It enables efficient and accurate monitoring threshold settings in large-scale heterogeneous device scenarios in data centers, improving the accuracy of monitoring alarms and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121217829B_ABST
    Figure CN121217829B_ABST
Patent Text Reader

Abstract

The application discloses a data center facility monitoring method, device, equipment, medium and product, relates to the technical field of data center monitoring, and provides a man-machine interactive threshold generation mechanism for the scene that a data center has massive infrastructure and various equipment types and operating environments, generates a performance index prediction value of a target performance index according to performance data of the target equipment at multiple moments by using a prediction model, controls a client device to display a first user interface for a user to set a first deviation parameter of the target performance index, and determines a threshold range of the performance index of the target equipment according to a system deviation parameter determined by the first deviation parameter, so that the user only needs to set a deviation value, and does not need to set the system deviation parameter one by one, thereby improving the threshold setting efficiency of the monitoring scene of the large-scale heterogeneous equipment of the data center, and avoiding the situation that only the adaptive threshold generated by the model leads to the monitoring alarm that cannot match the business scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data center monitoring technology, and in particular to a method, apparatus, equipment, medium and product for monitoring data center facilities. Background Technology

[0002] Data center infrastructure management (DCIM) platforms are used to monitor, analyze, manage, and automate the control of data center infrastructure (such as power, cooling, space, and network) and intelligent devices (such as servers, storage devices, and network equipment). Specifically, DCIM platforms collect data on various performance indicators of the infrastructure through multiple monitors and compare them with preset monitoring thresholds to detect abnormal data and trigger alarms. However, given the massive amount of infrastructure and the ever-changing types of equipment and operating environments, technical personnel are required to analyze equipment operation and set monitoring thresholds for each device individually, resulting in a huge workload and low efficiency.

[0003] Improving the efficiency of setting monitoring thresholds in data centers is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] This invention provides a method, apparatus, equipment, medium, and product for monitoring data center facilities, to at least solve the problem of low efficiency caused by the need for manual setting of monitoring thresholds when performing monitoring tasks in data centers in related technologies.

[0005] This invention provides a data center facility monitoring method, comprising: In response to an adaptive threshold setting command received by the client device, a first user interface is output to the display of the client device, and a first display object is displayed on the first user interface; In response to the user's first setting operation on the first display object, a first deviation parameter of the target performance index of the corresponding target device is determined; Determine the system deviation parameters based on the first deviation parameters; Obtain the predicted performance index value of the target performance index generated by the prediction model based on the performance data of the target device at multiple times; When monitoring the target device, if the deviation between the actual value of the target performance indicator and the predicted value of the performance indicator at the corresponding time exceeds the system deviation parameter, it is determined that the target performance indicator of the target device is in an abnormal state.

[0006] The present invention also provides a data center facility monitoring device, comprising: The display control module is configured to respond to an adaptive threshold setting command received by the client device, output a first user interface to the display of the client device, and display a first display object on the first user interface; The receiving module is configured to, in response to a user’s first setting operation on the first display object, determine a first deviation parameter of the target performance index of the corresponding target device; The first determining module is used to determine the system deviation parameter based on the first deviation parameter; The second determining module is used to obtain the performance index prediction value of the target performance index generated by the prediction model based on the performance data of the target device at multiple times. The monitoring module is used to determine that the target performance indicator of the target device is in an abnormal state if the deviation between the actual value of the target performance indicator and the predicted value of the performance indicator at the corresponding time exceeds the system deviation parameter when monitoring the target device.

[0007] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described data center facility monitoring methods.

[0008] The present invention also provides a non-volatile storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described data center facility monitoring methods.

[0009] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described data center facility monitoring methods.

[0010] This invention provides an interactive threshold generation mechanism for data centers with massive infrastructure and diverse equipment types and operating environments. Specifically, it utilizes a predictive model to generate predicted performance indicators of target performance metrics based on performance data from multiple time points of the target device, achieving data-driven adaptive threshold generation. Simultaneously, a client device displays a first user interface for the user to set a first deviation parameter for the target device's performance indicators. Based on the system deviation parameter determined by the first deviation parameter, the threshold range for the monitored target device's performance indicators is defined. This allows users to set only the deviation value, eliminating the need to set system deviation parameters individually. By integrating user business needs and risk assessments into the automation efficiency, this invention improves the efficiency of threshold setting in monitoring scenarios with large-scale heterogeneous equipment like data centers. It also avoids the situation where adaptive thresholds generated solely by models fail to match monitoring alarms with business scenarios, thus contributing to improved accuracy of data center monitoring alarms and enhancing the user experience. Attached Figure Description

[0011] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a data center facility monitoring method provided in an embodiment of the present invention; Figure 2 An architecture diagram of a data center facility monitoring system provided in an embodiment of the present invention; Figure 3 A flowchart illustrating the steps for generating an adaptive threshold, as provided in an embodiment of the present invention; Figure 4 A schematic diagram of a first user interface provided in an embodiment of the present invention; Figure 5 A schematic diagram of a second type of first user interface provided in an embodiment of the present invention; Figure 6 A method provided by an embodiment of the present invention Figure 5 A schematic diagram of the first user sub-interface of the first user interface shown. Figure 7 A flowchart illustrating a specific implementation of S104 provided in this embodiment of the invention; Figure 8 A flowchart of a device health monitoring procedure provided in an embodiment of the present invention; Figure 9 This is a flowchart of a model optimization step provided in an embodiment of the present invention. Detailed Implementation

[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0014] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0015] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0016] The embodiments of the present invention provide a data center facility monitoring method. The method is described in detail below in conjunction with the execution flow of the data center facility monitoring method.

[0017] Figure 1 A flowchart illustrating a data center facility monitoring method provided in an embodiment of the present invention; Figure 2 This is an architecture diagram of a data center facility monitoring system provided in an embodiment of the present invention.

[0018] The data center facility monitoring method provided in this embodiment of the invention may include, for example: Figure 1 S101~S105 are shown.

[0019] The data center facility monitoring method provided in this embodiment of the invention can be applied to the monitoring server of the data center, specifically a single monitoring server or a cluster of monitoring servers.

[0020] like Figure 2 As shown, this embodiment of the invention provides a data center facility monitoring system, which may include client devices 201 and monitoring servers 202. The number of monitoring servers 202 can be one or more, and multiple monitoring servers 202 can form a monitoring server cluster. The client devices 201 may include, but are not limited to, user terminals, data center large screens, etc. The target device is the monitored object device, and the type of the target device may include, but is not limited to, computing servers, storage servers, network devices, etc.

[0021] The following details each step of the data center facility monitoring method provided in this embodiment of the invention.

[0022] S101: In response to the adaptive threshold setting command received by the client device, output the first user interface on the display of the client device and display the first display object on the first user interface.

[0023] In the data center facility monitoring method provided in this embodiment of the invention, the performance index prediction value output by the prediction model is used as the performance baseline, and the first deviation parameter input by the user is received to determine the allowable deviation range between the actual value of the target device's performance index and the predicted value of the performance index.

[0024] To facilitate user input of the first deviation parameter, embodiments of the present invention provide a first user interface. That is, in embodiments of the present invention, the first user interface refers to a user interaction interface used to receive user input of the first deviation parameter.

[0025] In this embodiment of the invention, the adaptive threshold setting command can be the first request for user input received from another user interface displayed on the client device.

[0026] In this embodiment of the invention, the information presented to the user by default upon entering the first user interface is defined as the first display object.

[0027] In this embodiment of the invention, the type of the first user interface can be at least one of a command-line interface and a graphical user interface.

[0028] If the first user interface is a command-line interface, the first display object can be a command prompt. The command prompt can include a prompt string (a string of characters before the cursor, used to prompt the user with the current system status information, which in this embodiment of the invention can include the username, monitoring device name, current working directory (adaptive threshold setting)), an input cursor (a blinking underscore or square, indicating the position where the user is about to enter text), and a command line / input area (the area where the user actually enters commands, starting from the cursor).

[0029] If the first user interface is a graphical user interface, the first display object may include, but is not limited to, graphical elements such as windows, icons, menus, and buttons, which can be used to prompt the user with information such as the location and setting method of the deviation parameters for setting the performance indicators of the target device.

[0030] S102: In response to the user's first setting operation on the first display object, determine the first deviation parameter of the target performance index of the corresponding target device.

[0031] In this embodiment of the invention, depending on the type of the first user interface and the first display object, the first setting operation can be inputting command codes, or it can be selecting, dragging, or clicking on a text box to input information using a mouse or touch.

[0032] In this embodiment of the invention, the monitoring device involved in the user's first setting operation on the first display object is defined as the target device, and the performance index involved in the first setting operation is defined as the target performance index.

[0033] In this embodiment of the invention, the first deviation parameter is defined as the deviation value actually input by the user. That is to say, depending on the system settings, in some scenarios the deviation value input by the user can be directly taken, while in other scenarios it is necessary to perform numerical conversion (such as converting from percentage form to performance index value difference form), data type conversion (such as converting from decimal number to binary number), data amplification or reduction processing, etc. on the deviation value input by the user.

[0034] In practical applications, some performance metrics only require monitoring whether their actual values ​​exceed the upper limit of a threshold. In this case, the first deviation parameter only applies when the actual value of the performance metric is greater than the predicted value; it does not need to be considered when the actual value is less than the predicted value.

[0035] Similarly, for some performance indicators, it is only necessary to monitor whether the actual value of the performance indicator is less than the lower limit of the threshold. In this case, the first deviation parameter only applies when the actual value of the performance indicator is less than the predicted value, and does not need to be considered when the actual value of the performance indicator is greater than the predicted value.

[0036] Furthermore, for some performance indicators, it is necessary to monitor whether the actual value of the performance indicator is between its lower threshold and upper threshold. In this case, the first deviation parameter needs to be applied to the cases where the actual value of the performance indicator is greater than the predicted value and the cases where the actual value of the performance indicator is less than the predicted value.

[0037] The commonality among the above three types of performance indicators is that the absolute value of the difference between the actual value and the predicted value of the performance indicator is monitored using the absolute value of the first deviation parameter.

[0038] In this embodiment of the invention, the first deviation parameter can correspond to multiple target devices or multiple performance indicators of the target devices. Therefore, the user only needs to input the first deviation parameter once to achieve batch threshold settings for multiple target devices or multiple performance indicators. For example, the user can set the deviation value for the CPU utilization of multiple computing servers to 15%. CPU utilization typically only needs to monitor the upper limit of the threshold, meaning the user believes the actual value of CPU utilization should not exceed 115% of the predicted value of CPU utilization output by the prediction model. Alternatively, the system can accept user settings for deviation values ​​of multiple target performance parameters for the target devices. For example, the user can set the deviation values ​​for both CPU utilization and memory utilization of computing servers to 15%. Memory utilization typically only needs to monitor the upper limit of the threshold, meaning the actual values ​​of CPU utilization and memory utilization are allowed not to exceed 115% of the corresponding predicted values ​​output by the prediction model. Thus, batch setting of multiple deviation values ​​can be achieved, further improving the convenience of configuring system deviation parameters and the efficiency of configuring data center system deviation parameters.

[0039] The first deviation parameter can also be a specific performance indicator for a particular target device, thus enabling targeted settings. For example, a user can input only the first deviation parameter for the CPU utilization of a first server.

[0040] Therefore, by recognizing the user's first setting operation on the first display object, the corresponding target device, target performance index, and first deviation parameter are determined.

[0041] S103: Determine the system deviation parameter based on the first deviation parameter.

[0042] In this embodiment of the invention, the system deviation parameter is defined as the value actually used by the monitoring server when performing the monitoring task.

[0043] Depending on whether the performance metric requires monitoring of an upper or lower threshold, set the sign of the system deviation parameter. For performance metrics that require monitoring both upper and lower thresholds, a threshold range with the first deviation parameter as the threshold radius can be generated. For example, if the user-inputted first deviation parameter is 15%, the generated system deviation parameter can be ±15%. Alternatively, for performance metrics that require monitoring both upper and lower thresholds and where the system is set to asymmetric thresholds, the first deviation parameter can be shifted up or down according to the system setting adjustment value. For example, if the user-inputted first deviation parameter is 15%, the generated system deviation parameter can be -10% to +20%.

[0044] In some optional embodiments of the present invention, the numerical value of the system deviation parameter can be equal to the numerical value of the first deviation parameter. That is, the deviation value set by the user is directly taken.

[0045] In some optional embodiments of the present invention, the first deviation parameter input by the user can also be amplified or reduced to serve as the system deviation parameter. In practical applications, users may not frequently change the deviation value, which can lead to the fixed deviation value set by the user becoming unsuitable for the actual working environment as the data center working scenario changes, resulting in false alarms. Therefore, determining the system deviation parameter based on the first deviation parameter in S103 can also include: determining the historical prediction error data of the prediction model for the target performance index; and determining the updated system deviation parameter based on the first deviation parameter and the historical prediction error data. It should be noted that the "updated system deviation parameter" here refers to the result updated based on the existing system deviation parameters. The historical prediction error data can be calculated based on the historical predicted values ​​and historical actual values ​​of the target performance index from multiple prediction models. It can be used to represent the stability of the target device's target performance index at historical moments. Adjusting the first deviation parameter input by the user using the historical prediction error data can adapt to the stability of the target performance index, achieving a balance between timely alarms and avoiding frequent alarms.

[0046] In this embodiment of the invention, determining the historical prediction error data of the prediction model for the target performance index may include: calculating at least two types of historical prediction error parameters based on multiple sets of historical predicted values ​​and historical actual values ​​of the target performance index from multiple prediction models. To improve the accuracy of adjusting the user-defined deviation based on the historical prediction error data, at least two types of historical prediction error parameters can be used for adjustment. Furthermore, at least two types of historical prediction error parameters can mutually suppress each other, avoiding excessive adjustment of the user-defined deviation.

[0047] When the first deviation parameter is adjusted using the first historical error prediction parameter and the second historical error prediction parameter, the deviation adjustment value can be calculated based on the first historical error prediction parameter and the second historical error prediction parameter. The deviation adjustment value is then added to the first deviation parameter set by the user to obtain the system deviation parameter. The first historical error prediction parameter is positively correlated with the deviation adjustment value, and the second historical error prediction parameter is negatively correlated with the deviation adjustment value, thus inhibiting each other.

[0048] In some optional embodiments of the present invention, the first historical error prediction parameter can be the standard deviation, and the second historical error prediction parameter can be the mean absolute percentage error. The first historical error parameter can then be expressed as: The second historical error prediction parameter can be expressed as MAPE=(1 / n)×Σ(|(actual value - predicted value) / actual value|)×100%. The standard deviation is represented by , and MAPE represents the mean absolute percentage error. The actual value and the predicted value both represent the actual value at a historical moment and the predicted value of the prediction model. n is the number of samples in the historical prediction data. A historical time period can be defined as a monitoring window for calculation.

[0049] The deviation adjustment value can then be expressed as: .in, This indicates the deviation adjustment value.

[0050] In practical applications, according to the solution provided in the embodiments of the present invention, the deviation adjustment value Other formulas can also be used to calculate it.

[0051] At the same time, the adjustment direction of the first deviation parameter can be determined by comparing the magnitude of the first deviation parameter with the historical prediction error parameter. Here, the historical prediction error parameter can be one of the first historical error prediction parameter, the second historical error prediction parameter, or the other historical error prediction parameter.

[0052] Using the first historical error prediction parameter (such as standard deviation) Taking the first deviation parameter acting on the upper limit of the threshold as an example, the updated system deviation parameter is determined based on the first deviation parameter and historical prediction error data, which can be expressed as: when the first deviation parameter ≥ standard deviation If the system is determined to be highly stable, the upper limit of the user-set threshold can be lowered. Therefore, the system deviation parameter = first deviation parameter - When the first deviation parameter is less than the standard deviation If the system is found to have poor stability, the upper limit of the user-defined threshold can be increased. Then, the system deviation parameter = first deviation parameter + Therefore, by mutually suppressing the two types of historical prediction error parameters, the dynamic adjustment of the user-defined first deviation parameter is prevented from becoming unbalanced.

[0053] For cases where the first deviation parameter acts on the lower threshold, or on both the upper and lower thresholds, it can be expressed as: when the absolute value of the first deviation parameter is greater than or equal to the standard deviation. If the system is determined to be highly stable, then the absolute value of the system deviation parameter = the absolute value of the first deviation parameter - When the absolute value of the first deviation parameter is less than the standard deviation. When the system is determined to have poor stability, the absolute value of the system deviation parameter = the absolute value of the first deviation parameter + The sign of the system deviation parameter is set based on whether it applies to the upper or lower threshold. This allows the two types of historical prediction error parameters to mutually suppress each other, preventing imbalances in the dynamic adjustment of the user-defined first deviation parameter.

[0054] In this embodiment of the invention, in order to realize graded alarms for performance indicators, one performance indicator can correspond to multiple system deviation parameters, which correspond to different anomaly levels of the performance indicator. Different anomaly levels can correspond to different alarm methods. For example, they can be divided into three anomaly levels: attention, trouble, and critical. The degree of anomaly they correspond to increases from low to high, and the absolute value of the corresponding system deviation parameters increases from small to large.

[0055] At this point, determining the system deviation parameter based on the first deviation parameter in S103 may include: determining the historical prediction error data of the prediction model for the target performance index; and determining the corresponding updated system deviation parameter based on the first deviation parameter and the historical prediction error data for multiple first deviation parameters corresponding to the performance index.

[0056] Determining the historical prediction error data of the prediction model for the target performance index may also include: calculating at least two types of historical prediction error parameters based on the historical predicted values ​​and historical actual values ​​of the target performance index from multiple prediction models. For the system deviation parameter corresponding to each anomaly level, its calculation method can refer to the calculation method for the system deviation parameter in the above steps, thereby adjusting the first deviation parameter corresponding to each anomaly level of the target performance index for the user to the corresponding system deviation parameter for that anomaly level.

[0057] S104: Obtain the predicted performance index of the target performance index generated by the prediction model based on the performance data of the target device at multiple times.

[0058] Specifically, the pre-trained prediction model can be fine-tuned using performance data from the target device at multiple time points to obtain a prediction model for that performance metric. This prediction model is then used to generate predicted performance metric values ​​for future time points.

[0059] S105: When monitoring the target device, if the deviation between the actual value of the target performance index and the predicted value of the performance index at the corresponding time exceeds the system deviation parameter, it is determined that the target performance index of the target device is in an abnormal state.

[0060] Through the above steps, an adaptive threshold for the target performance index of the target device can be obtained, including using the predicted performance index value as the baseline and the system deviation parameter as the allowable deviation. Based on this adaptive threshold, the monitoring server performs monitoring of the target device, collecting the actual performance index values ​​through the monitor. If the deviation between the actual performance index value and the predicted performance index value at the corresponding time exceeds the system deviation parameter, it is determined that the target performance index is in an abnormal state.

[0061] Figure 3 This is a flowchart of an adaptive threshold generation step provided in an embodiment of the present invention.

[0062] like Figure 3 As shown, by deploying a prediction model, an error calculator, a deviation adjustment module, a threshold generator, and an alarm system on the monitoring server, the steps for the monitoring server to generate adaptive thresholds may include: S301: The prediction model provides predicted values ​​of performance indicators for the current time period.

[0063] S302: Error calculator calculates historical prediction error parameters.

[0064] S303: The error calculator sends the confidence score to the deviation adjustment module.

[0065] S304: The deviation adjustment module reads the first deviation parameter input by the user.

[0066] S305: The deviation adjustment module calculates the system deviation parameter based on the first deviation parameter and the historical prediction error parameter.

[0067] S306: The deviation adjustment module outputs system deviation parameters to the threshold generator.

[0068] S307: The threshold generator generates system deviation parameters for three anomaly levels.

[0069] S308: The threshold generator sends system deviation parameters to the alarm system.

[0070] The data center facility monitoring method provided in this invention offers an interactive threshold generation mechanism for scenarios involving massive infrastructure and diverse equipment types and operating environments in data centers. Specifically, it utilizes a predictive model to generate predicted performance indicators of target performance metrics based on performance data from multiple time points of the target device, achieving data-driven adaptive threshold generation. Simultaneously, a client device displays a first user interface for the user to set a first deviation parameter for the target device's performance indicators. Based on a system deviation parameter determined by the first deviation parameter, the threshold range for the monitored target device's performance indicators is defined, allowing the user to set only the deviation value without individually configuring system deviation parameters. This invention, by integrating user business needs and risk assessments into automated efficiency, improves the efficiency of threshold setting in monitoring scenarios with large-scale heterogeneous equipment like data centers. It also avoids the situation where adaptive thresholds generated solely by models fail to match monitoring alarms with business scenarios, thereby contributing to improved accuracy of data center monitoring alarms and enhancing the user experience.

[0071] Based on the above embodiments, this invention provides an interaction scheme in which the first user interface is a graphical user interface.

[0072] In this embodiment of the invention, the first display object may include a graphic representation of the performance indicators of the target device and a display value of the deviation parameters corresponding to the performance indicators.

[0073] Figure 4 This is a schematic diagram of a first user interface provided in an embodiment of the present invention.

[0074] like Figure 4 As shown, in some optional embodiments of the present invention, in the first user interface 401, the first display object can be in the form of a table, the first column of the table is the name of the performance index (such as performance index 1, performance index 2, performance index 3, performance index 4), and the second column is the deviation parameter (such as 15, 15, 300, 80).

[0075] In some optional embodiments of the present invention, the displayed value of the deviation parameter can be a first deviation parameter input by the user. When the user first inputs the first deviation parameter, the displayed value of the deviation parameter can be the system default deviation value or be empty. Thereafter, the first deviation parameter previously input by the user is displayed.

[0076] In some alternative embodiments of the present invention, the displayed value of the deviation parameter is determined based on the existing value of the system deviation parameter. As described in the above embodiments of the present invention, the first deviation parameter input by the user can be adjusted and used as the system deviation parameter. In this case, only the adjusted system deviation parameter can be displayed, or the first deviation parameter and the system deviation parameter can be displayed separately.

[0077] Accordingly, determining the system deviation parameter based on the first deviation parameter in S103 may include: determining the historical prediction error data of the prediction model for the target performance index; and determining the updated system deviation parameter based on the first deviation parameter and the historical prediction error data.

[0078] For performance indicator-based alarms, one performance indicator of the target device corresponds to multiple deviation parameter display values, each corresponding to a different level of abnormality in the performance indicator.

[0079] Figure 5 This is a schematic diagram of a second type of first user interface provided in an embodiment of the present invention.

[0080] like Figure 5As shown, in some optional embodiments of the present invention, in the first user interface 501, multiple system deviation parameters corresponding to anomaly levels can be set for each performance indicator. Then, determining the system deviation parameters based on the first deviation parameters in S103 may include: determining the historical prediction error data of the prediction model for the target performance indicator; and determining the corresponding updated system deviation parameters for each of the multiple first deviation parameters corresponding to the performance indicator based on the first deviation parameters and the historical prediction error data.

[0081] In the above implementation, determining the historical prediction error data of the prediction model for the target performance index may include: calculating at least two types of historical prediction error parameters based on the historical predicted values ​​and historical actual values ​​of the target performance index from multiple prediction models.

[0082] For a detailed implementation of the method for calculating the system deviation parameter based on the first deviation parameter and the historical prediction error parameter, please refer to the description in the above embodiments.

[0083] In this embodiment of the invention, in response to a user's first setting operation on a first display object, determining a first deviation parameter of the target performance index of the corresponding target device may include: in response to the first setting operation, identifying the target index identifier selected in the first setting operation, and determining the target performance index; and determining the first deviation parameter according to a second setting operation in the first setting operation on the display value of the target deviation parameter corresponding to the target performance index.

[0084] In some optional embodiments of the present invention, the first display object further includes a startup deviation setting icon corresponding to the performance indicator. Responding to the first setting operation, identifying the target indicator identifier selected in the first setting operation and determining the target performance indicator may include: identifying the target startup deviation setting icon selected in the first setting operation and determining the indicator identifier graphic corresponding to the target startup deviation setting icon as the target performance indicator. Determining the first deviation parameter based on the second setting operation of the target deviation parameter display value corresponding to the target performance indicator in the first setting operation may include: responding to the user's selection operation of the target startup deviation setting icon, converting the deviation parameter display value corresponding to the target performance indicator to an edit state, and determining the first deviation parameter based on the user-edited deviation parameter display value.

[0085] like Figure 5As shown, the first user interface 501 has a "Bias Settings" column, which displays multiple performance indicators and their corresponding deviation parameter values ​​in a table format (corresponding to concern anomalies, problem anomalies, and critical anomalies, respectively). The last column is "Edit," and the symbols in each row below it are the start deviation settings icons for each performance indicator. For example, the user can click the start deviation settings icon corresponding to performance indicator 1 to switch the deviation parameter display value corresponding to performance indicator 1 to edit mode.

[0086] In some optional embodiments of the present invention, in response to the user's selection operation on the target start deviation setting icon, the displayed value of the deviation parameter corresponding to the target performance index is converted to an edit state, and the first deviation parameter is determined according to the user's edited deviation parameter display value. This may include: in response to the user's selection operation on the target start deviation setting icon, jumping to a first user sub-interface and displaying a candidate deviation parameter graphic corresponding to the target performance index on the first user sub-interface; receiving the user's edit operation on the candidate deviation parameter graphic, and determining the first deviation parameter according to the edit operation.

[0087] Figure 6 A method provided by an embodiment of the present invention Figure 5 A schematic diagram of the first user sub-interface of the first user interface shown.

[0088] like Figure 6 As shown, the first user sub-interface 5011 can be in the form of a drop-down menu. That is, after the user clicks the edit icon corresponding to performance index 1 and selects the type of deviation value to be set (such as the deviation value corresponding to the anomaly to be monitored), the first user sub-interface 5011 will be displayed, and candidate deviation parameters such as 15, 20, and 30 will be displayed for the user to select. The first user sub-interface 5011 may also include an input box for the user to input the first deviation parameter.

[0089] In some optional embodiments of the present invention, in response to a user's selection operation on the target start deviation setting icon, converting the displayed value of the deviation parameter corresponding to the target performance indicator to an editable state, and determining the first deviation parameter based on the user-edited displayed value of the deviation parameter, may include: in response to a user's selection operation on the target start deviation setting icon, setting the corresponding displayed value of the deviation parameter as an input box graphic to receive the first deviation parameter input by the user in the input box. (See reference) Figure 5 If a user clicks the edit icon corresponding to performance index 1, the deviation parameter display values ​​of the three levels corresponding to performance index 1 can be set as input box graphics for the user to select and input.

[0090] In this embodiment of the invention, the first display object may further include the default deviation parameter display value corresponding to the performance index.

[0091] like Figure 5 As shown, in addition to the deviation parameter display value for users to set the deviation value, the first user interface 501 may also include a "default deviation value" column, which may include the default deviation values ​​corresponding to three abnormality levels: concern abnormality, problem abnormality, and serious abnormality, which are 15, 20, and 25 respectively.

[0092] In addition, such as Figure 5 As shown, the first user interface 501 may also include an adaptive threshold switch graphic. In S101, in response to the adaptive threshold setting command received by the client device, the adaptive threshold setting command can be the user's operation to enable the adaptive threshold switch graphic.

[0093] Traditional forecasting schemes typically rely on time series data to predict future values. However, in data centers, with the ever-changing nature of equipment types, performance metric types, and business scenarios, the accuracy of performance metric predictions based solely on historical data is low. To address this, this invention provides a hybrid forecasting model. In this embodiment, the performance metric prediction value generated by the forecasting model based on performance data of the target device at multiple time points can include: generating a first performance metric prediction value using a first forecasting model based on performance data of the target device at multiple time points; generating a second performance metric prediction value using a second forecasting model based on performance data of the target device at multiple time points; and determining the performance metric prediction value based on the first and second performance metric prediction values.

[0094] In this embodiment of the invention, the first prediction model can be a linear model, and the second prediction model can be a nonlinear model. For example, the first prediction model can be a Prophet model, and the second prediction model can be a Long Short-Term Memory (LSTM) network.

[0095] Determining the predicted performance index value based on the predicted values ​​of the first and second performance indices can include: calculating the first squared term of the first difference obtained by subtracting the predicted value of the first performance index from the predicted value of the second performance index; and using the first weight corresponding to the first prediction model, the second weight corresponding to the second prediction model, and the third weight corresponding to the first squared term, weighted summing of the predicted value of the first performance index, the predicted value of the second performance index, and the first squared term to obtain the predicted performance index value. Specifically, this can be expressed by the following formula: Predicted performance index value = ×Predicted value of first performance index + ×Predicted value of the second performance index + × .

[0096] Different data can be used to train the first and second prediction models. For example, for Prophet branch training, historical data plus business event labels can be input to fit long-term trends and event impacts. For LSTM branch training, time-series features plus spatiotemporal encoding can be input to learn short-term fluctuation patterns. The performance index predictions generated by the prediction models based on the performance data of the target device at multiple time points can include: using the first prediction model to output a first performance prediction value based on trend features, and using the second model to output a second performance prediction value based on fluctuation features. The trend features can include time series data of historical performance data and business event labels to fit long-term trends and event impacts. The fluctuation features can include time series data of historical performance data and spatiotemporal feature encoding to learn short-term fluctuation patterns.

[0097] To address the issue of large prediction biases when relying solely on historical data, in this embodiment of the invention, the performance index prediction value of the target performance indicator generated by the prediction model based on the performance data of the target device at multiple times in step S104 can include: encoding the performance index values ​​of the target performance indicator at multiple times, the corresponding time-series parameters, the corresponding business characteristics, and the corresponding schedule parameters, and then inputting them into the prediction model to output the performance index prediction value corresponding to the target performance indicator. Specifically, in addition to the time-series data of the historical performance data of the target performance indicator, business feature encoding and spatiotemporal feature encoding can also be embedded in the input data of the prediction model. The business feature encoding can be generated by converting business tags such as promotional events and maintenance windows into numerical influence factors. The spatiotemporal feature encoding can be generated by performing sine / cosine encoding on timestamps (such as hourly or weekly cycles) to capture periodic patterns.

[0098] Specifically, during data collection and processing, historical performance data of the equipment (CPU, memory, disk I / O data, etc.) can be obtained from the monitoring system and sampled at a fixed granularity (e.g., 5 minutes). Business event tags (promotional activities, system maintenance periods) are extracted and converted into binary feature vectors (e.g., promotional period = 1, non-promotional period = 0).

[0099] For spatiotemporal feature coding, the timestamp can be decomposed into a periodic signal, with an hourly period: hourly code = sin(2 × hours / 24); Weekday cycle: Weekday code = cos(2 × Weekday number / 7). Execution environment characteristic normalization: linearly scale data such as computer room temperature and network load to the [0,1] range.

[0100] In this embodiment of the invention, to avoid false alarms caused by jumps in the predicted performance index value, obtaining the predicted performance index value of the target performance index generated by the prediction model based on the performance data of the target device at multiple times in step S104 may include: obtaining the predicted performance index value output by the prediction model that is within a preset threshold range. That is, a threshold boundary (preset threshold range) can be set for the output of the prediction model to limit the fluctuation range of the predicted performance index value and avoid false alarms caused by jumps.

[0101] Figure 7 This is a flowchart illustrating a specific implementation of S104 provided in an embodiment of the present invention.

[0102] like Figure 7 As shown, by deploying a data collector, feature engineering module, linear model, nonlinear model, and fusion output module on the monitoring server, the step in S104 to obtain the predicted value of the target performance index generated by the prediction model based on the performance data of the target device at multiple times may include: S701: The data collector sends historical performance data and business tags to the feature engineering module.

[0103] S702: The feature engineering module generates business feature codes and spatiotemporal feature codes.

[0104] S703: The feature engineering module inputs trend features into the linear model.

[0105] S704: The feature engineering module inputs fluctuation features into the nonlinear model.

[0106] S705: Linear model outputs trend prediction values.

[0107] S706: Nonlinear model outputs predicted fluctuation values.

[0108] S707: The fusion output module performs weighted residual fusion on the trend prediction value and the fluctuation prediction value to obtain the performance index prediction value.

[0109] S708: The fusion output module outputs predicted performance metrics to an external system.

[0110] The external system here can be Figure 3 The alarm system in the middle.

[0111] Based on the above embodiments, the data center facility monitoring method provided by the present invention may further include: optimizing the prediction model based on the prediction results.

[0112] In specific implementation, the data center facility monitoring method provided in this embodiment of the invention may further include: calculating the device health score of the target device based on the predicted values ​​of the performance indicators corresponding to multiple performance indicators of the target device; determining the device health score as a system-detected device fault event when it falls within the device alarm threshold range; determining the fault alarm error value of the target device based on the actual state of the target device and the system-detected device fault event; and updating the fusion weight between the first prediction model and the second prediction model based on the fault alarm error value.

[0113] The fusion weight can be the first weight, second weight, or third weight described in the above embodiments.

[0114] The device health score of the target device can be calculated based on the predicted values ​​of the performance indicators corresponding to multiple performance indicators of the target device. This can include: calculating the device health score of the target device based on the time-series data corresponding to the multiple performance indicators of the target device, the predicted values ​​of the corresponding performance indicators, and the corresponding historical failure rates.

[0115] Specifically, the equipment health score can be obtained by weighted fusion of multiple performance indicators of the target equipment. The formula for calculating the equipment health score is: DHI = w_1i × (actual performance index value / predicted performance index value) + w2 × historical failure rate. Where DHI represents the device health score. w_1i represents the weight of the i-th performance index, i = 1 to N. w2 represents the weight of the historical failure rate of the target device.

[0116] In some optional embodiments of the present invention, the calculation step of the device health score of the target device may also include: calculating the device health score of the target device based on the actual values ​​of the performance indicators of multiple performance indicators of the target device. That is, the device health score of the target device may also be calculated using only the actual values ​​of the performance indicators of the target device.

[0117] The data center facility monitoring method provided in this embodiment of the invention may further include: issuing equipment anomaly alarms based on equipment health scores.

[0118] In practical applications, multiple device alarm thresholds can be set and triggered in stages. For example, three device alarm thresholds can be set: DHI < first device alarm threshold → no alarm (logging only); first device alarm threshold ≤ DHI < second device alarm threshold → alert level (e.g., email notification); second device alarm threshold ≤ DHI < third device alarm threshold → problem level alarm (SMS + diagnostic report); DHI ≥ third device alarm threshold → critical level alarm (telephone + automatic work order).

[0119] Figure 8This is a flowchart of a device health monitoring procedure provided in an embodiment of the present invention.

[0120] like Figure 8 As shown, the device health monitoring steps provided in this embodiment of the invention may include: S801: The monitoring agent module reports real-time load data to the device health score calculation engine through the monitoring agent module, device health score calculation engine, device alarm decision-maker, and notification system deployed on the monitoring server.

[0121] S802: The device health score calculation engine calculates the load rate of various performance indicators of the target device.

[0122] S803: The device health score calculation engine generates a weighted fusion of data to produce a device health score.

[0123] S804: The device health score calculation engine sends the device health score to the alarm decision controller.

[0124] S805: Alarm decision setter matches alarm levels.

[0125] S806: The alarm decision-maker triggers the corresponding alarm action to the notification system.

[0126] S807: The notification system pushes alarm information to maintenance personnel.

[0127] In some optional embodiments of the present invention, determining the fault alarm error value for the target device based on the actual state of the target device and the system-detected device fault events may include: obtaining the operation and maintenance data of the target device; calculating the device performance score based on the predicted performance index value and executing the corresponding device alarm task and corresponding operation and maintenance data to determine the validity of the alarm for the target device.

[0128] Operation and maintenance data can be the actual status of the equipment after an alarm is triggered (it can be the result of manual confirmation or automatic diagnosis).

[0129] Alarm validity can be represented by, but is not limited to, true positive events (TP, equipment actually failed), false positive events (FP, equipment did not fail), and false negative events (FN, equipment actually failed).

[0130] Updating the fusion weights between the first and second prediction models based on the fault alarm error value may include: obtaining the operation and maintenance data of the target device; and updating the fusion weights between the first and second prediction models when the number of false alarms reaches the first time after calculating the device performance score based on the predicted performance index value and executing the corresponding device alarm task.

[0131] The data center facility monitoring method provided in this embodiment of the invention may further include: determining a first ratio of false positive events to false negative events within the first monitoring period based on system-detected equipment failure events and actual equipment failure events occurring within the first monitoring period. Determining the system deviation parameter based on the first deviation parameter in step S103 may include: determining the system deviation parameter based on the first deviation parameter and the first ratio. That is, the deviation value can be dynamically contracted or expanded according to the FP / FN ratio, thereby constructing a false positive / false negative feedback mechanism.

[0132] In practice, precision (Precision = TP / (TP + FP)) and recall (Recall = TP / (TP + FN)) can be calculated. If the precision is lower than the corresponding target value, the bias reduction factor can be increased to reduce false positives. If the recall is lower than the corresponding target value, the bias reduction factor can be decreased to reduce false negatives. Here, the bias reduction factor can refer to the confidence value described in the above embodiments.

[0133] In addition, the health index threshold can be adjusted according to the FP / FN ratio: if FP is too high, the device alarm threshold will be increased (to reduce false alarms); if FN is too high, the device alarm threshold will be decreased (to reduce missed alarms).

[0134] Figure 9 This is a flowchart of a model optimization step provided in an embodiment of the present invention.

[0135] like Figure 9 As shown, the optimization steps for the prediction model can include the following steps through the operation and maintenance feedback module, optimization analyzer, prediction engine, and threshold generator deployed on the monitoring server: S901: The operation and maintenance feedback module sends alarm validity data to the optimization analyzer.

[0136] S902: The optimizer calculates precision and recall.

[0137] S903: The optimizer sends model adjustment instructions to the prediction engine.

[0138] S904: The prediction engine updates the fusion weights of the prediction model.

[0139] S905: The optimization analyzer sends a deviation adjustment command to the threshold generator.

[0140] S906: Threshold generator updates system deviation parameters.

[0141] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0142] Embodiments of the present invention also provide a data center facility monitoring device, comprising: a display control module, configured to output a first user interface on the display of the client device and display a first display object in response to an adaptive threshold setting command received by a client device; a receiving module, configured to determine a first deviation parameter of a target performance index of a corresponding target device in response to a first setting operation by a user on the first display object; a first determining module, configured to determine a system deviation parameter based on the first deviation parameter; a second determining module, configured to obtain a performance index prediction value of the target performance index generated by a prediction model based on performance data of the target device at multiple times; and a monitoring module, configured to determine that the target performance index of the target device is in an abnormal state if, during monitoring of the target device, the deviation between the actual value of the target performance index of the target device and the predicted value of the performance index at the corresponding time exceeds the system deviation parameter.

[0143] In this embodiment of the invention, the first display object may include a graphic representation of the performance indicators of the target device and a display value of the deviation parameters corresponding to the performance indicators.

[0144] In this embodiment of the invention, the displayed value of the deviation parameter can be determined based on the existing value of the system deviation parameter.

[0145] In this embodiment of the invention, the first determining module determines the system deviation parameter based on the first deviation parameter, which may include: determining the historical prediction error data of the prediction model for the target performance index; and determining the updated system deviation parameter based on the first deviation parameter and the historical prediction error data.

[0146] In this embodiment of the invention, one performance index of the target device can correspond to multiple deviation parameter display values, which respectively correspond to different abnormal levels of the performance index.

[0147] In this embodiment of the invention, the first determining module determines the system deviation parameter based on the first deviation parameter, which may include: determining the historical prediction error data of the prediction model for the target performance index; and determining the corresponding updated system deviation parameter for each of the multiple first deviation parameters corresponding to the performance index based on the first deviation parameter and the historical prediction error data.

[0148] In this embodiment of the invention, the first determining module determines the historical prediction error data of the prediction model for the target performance index, which may include: calculating at least two types of historical prediction error parameters based on the historical predicted values ​​and historical actual values ​​of the target performance index by multiple prediction models.

[0149] In this embodiment of the invention, the display control module, in response to a user's first setting operation on a first display object, determines a first deviation parameter of the target performance index of the corresponding target device. This may include: in response to the first setting operation, identifying the target index identifier selected in the first setting operation and determining the target performance index; and determining the first deviation parameter based on a second setting operation in the first setting operation on the display value of the target deviation parameter corresponding to the target performance index.

[0150] In this embodiment of the invention, the first display object may further include a startup deviation setting icon corresponding to the performance index; the display control module, in response to the first setting operation, identifies the target index identifier selected by the first setting operation and determines the target performance index, which may include: identifying the target startup deviation setting icon selected by the first setting operation and determining the index identifier graphic corresponding to the target startup deviation setting icon as the target performance index; the receiving module, based on the second setting operation of the target deviation parameter display value corresponding to the target performance index in the first setting operation, determines the first deviation parameter, which may include: in response to the user's selection operation of the target startup deviation setting icon, converting the deviation parameter display value corresponding to the target performance index to an edit state, and determining the first deviation parameter based on the deviation parameter display value edited by the user.

[0151] In this embodiment of the invention, the receiving module, in response to the user's selection operation on the target start deviation setting icon, converts the displayed value of the deviation parameter corresponding to the target performance indicator to an edit state, and determines the first deviation parameter based on the user-edited displayed value of the deviation parameter. This may include: in response to the user's selection operation on the target start deviation setting icon, jumping to a first user sub-interface and displaying a candidate deviation parameter graphic corresponding to the target performance indicator on the first user sub-interface; receiving the user's editing operation on the candidate deviation parameter graphic, and determining the first deviation parameter based on the editing operation.

[0152] In this embodiment of the invention, the first display object may further include the default deviation parameter display value corresponding to the performance index.

[0153] In this embodiment of the invention, the performance index prediction value of the target performance index generated by the prediction model based on the performance data of the target device at multiple times may include: generating a first performance index prediction value based on the performance data of the target device at multiple times using a first prediction model; generating a second performance index prediction value based on the performance data of the target device at multiple times using a second prediction model; and determining the performance index prediction value based on the first performance index prediction value and the second performance index prediction value.

[0154] In this embodiment of the invention, the first prediction model can be a linear model, and the second prediction model can be a nonlinear model. Determining the predicted performance index value based on the predicted value of the first performance index and the predicted value of the second performance index can include: calculating the first squared term of the first difference obtained by subtracting the predicted value of the first performance index from the predicted value of the second performance index; and using the first weight corresponding to the first prediction model, the second weight corresponding to the second prediction model, and the third weight corresponding to the first squared term to perform a weighted summation of the predicted value of the first performance index, the predicted value of the second performance index, and the first squared term to obtain the predicted performance index value.

[0155] The data center facility monitoring device provided in this embodiment of the invention may further include: a device health score calculation engine, used to calculate the device health score of the target device based on the predicted values ​​of the performance indicators corresponding to multiple performance indicators of the target device; a device alarm decision-maker, used to determine that the device health score falls within the device alarm threshold range as a system-detected device fault event; an optimization analyzer, used to determine the fault alarm error value for the target device based on the actual state of the target device and the system-detected device fault event; and a prediction engine, used to update the fusion weight between the first prediction model and the second prediction model based on the fault alarm error value.

[0156] In this embodiment of the invention, the device health score calculation engine calculates the device health score of the target device based on the predicted values ​​of the performance indicators corresponding to multiple performance indicators of the target device. This may include: calculating the device health score of the target device based on the time-series data corresponding to the multiple performance indicators of the target device, the predicted values ​​of the corresponding performance indicators, and the corresponding historical failure rate.

[0157] In this embodiment of the invention, the optimization analyzer can also be used to determine a first ratio of false positive events to false negative events within the first monitoring period, based on system-detected equipment failure events and actual equipment failure events occurring within the first monitoring period. The first determining module may further include a threshold generator, used to determine a system deviation parameter based on a first deviation parameter and the first ratio.

[0158] For a description of the features in the embodiment corresponding to the data center facility monitoring device, please refer to the relevant description in the embodiment corresponding to the data center facility monitoring method, which will not be repeated here.

[0159] Embodiments of the present invention also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described embodiments of the data center facility monitoring method.

[0160] Embodiments of the present invention also provide a non-volatile storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the data center facility monitoring method when running.

[0161] In one exemplary embodiment, the aforementioned non-volatile storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0162] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described data center facility monitoring method embodiments.

[0163] Embodiments of the present invention also provide another computer program product, including a non-volatile storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described data center facility monitoring method embodiments.

[0164] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0165] The present invention has provided a detailed description of a data center facility monitoring method, apparatus, equipment, medium, and product. Specific examples have been used to illustrate the principles and implementation methods of the invention. The descriptions of these embodiments are only intended to aid in understanding the method and core ideas of the invention. It should be noted that those skilled in the art can make various improvements and modifications to the invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the invention.

Claims

1. A method for monitoring data center facilities, characterized in that, include: In response to an adaptive threshold setting command received by the client device, a first user interface is output to the display of the client device, and a first display object is displayed on the first user interface; In response to the user's first setting operation on the first display object, a first deviation parameter of the target performance index of the corresponding target device is determined; Determine the system deviation parameters based on the first deviation parameters; Obtain the predicted performance index value of the target performance index generated by the prediction model based on the performance data of the target device at multiple times; When monitoring the target device, if the deviation between the actual value of the target performance indicator and the predicted value of the performance indicator at the corresponding time exceeds the system deviation parameter, it is determined that the target performance indicator of the target device is in an abnormal state. The first display object includes indicator graphics of the performance indicators of the target device and display values ​​of deviation parameters corresponding to the performance indicators; the display values ​​of deviation parameters are determined based on existing values ​​of the system deviation parameters. Determining the system deviation parameter based on the first deviation parameter includes: Determine the first historical prediction error parameter and the second historical prediction error parameter of the prediction model for the target performance index; The deviation adjustment value is calculated based on the first historical prediction error parameter and the second historical prediction error parameter. The first historical prediction error parameter is positively correlated with the deviation adjustment value, and the second historical prediction error parameter is negatively correlated with the deviation adjustment value. The system deviation parameter is obtained by superimposing the deviation adjustment value on the first deviation parameter. The adjustment direction of the first deviation parameter is determined by comparing the magnitude of the first deviation parameter with the first historical prediction error parameter. If the first deviation parameter is greater than or equal to the first historical prediction error parameter being compared, the system stability is determined to be strong, and the absolute value of the first deviation parameter is reduced. If the first deviation parameter is less than the first historical prediction error parameter being compared, the system stability is determined to be weak, and the absolute value of the first deviation parameter is increased.

2. The data center facility monitoring method according to claim 1, characterized in that, One performance indicator of the target device corresponds to multiple deviation parameter display values, which respectively correspond to different abnormal levels of the performance indicator.

3. The data center facility monitoring method according to claim 2, characterized in that, Determining the system deviation parameter based on the first deviation parameter includes: Determine the historical prediction error data of the prediction model for the target performance index; For each of the first deviation parameters corresponding to the performance index, the corresponding updated system deviation parameter is determined based on the first deviation parameter and the historical prediction error data.

4. The data center facility monitoring method according to claim 1 or 3, characterized in that, Determining the historical prediction error data of the prediction model for the target performance index includes: Based on the historical predicted values ​​and historical actual values ​​of the target performance index from multiple sets of prediction models, at least two types of historical prediction error parameters are calculated.

5. The data center facility monitoring method according to claim 1, characterized in that, In response to a user's first setting operation on the first display object, a first deviation parameter for the target performance index of the corresponding target device is determined, including: In response to the first setting operation, the target indicator identifier selected by the first setting operation is identified, and the target performance indicator is determined; The first deviation parameter is determined based on the second setting operation of the target deviation parameter display value corresponding to the target performance index in the first setting operation.

6. The data center facility monitoring method according to claim 5, characterized in that, The first display object also includes the startup deviation setting icon corresponding to the performance index; In response to the first setting operation, the target indicator identifier selected by the first setting operation is identified, and the target performance indicator is determined, including: Identify the target startup deviation setting icon selected in the first setting operation, and determine the indicator identification graphic corresponding to the target startup deviation setting icon as the target performance indicator; Based on the second setting operation of the first setting operation regarding the displayed value of the target deviation parameter corresponding to the target performance index, the first deviation parameter is determined, including: In response to the user's selection of the target start deviation setting icon, the displayed value of the deviation parameter corresponding to the target performance index is converted to an edit state, and the first deviation parameter is determined based on the user-edited displayed value of the deviation parameter.

7. The data center facility monitoring method according to claim 6, characterized in that, In response to the user's selection of the target start deviation setting icon, the displayed value of the deviation parameter corresponding to the target performance indicator is switched to an edit state, and the first deviation parameter is determined based on the user-edited deviation parameter displayed value, including: In response to the user's selection of the target start deviation setting icon, the system jumps to the first user sub-interface and displays the candidate deviation parameter graph corresponding to the target performance index on the first user sub-interface; The system receives user editing operations on the candidate deviation parameter graph and determines the first deviation parameter based on the editing operations.

8. The data center facility monitoring method according to claim 1, characterized in that, The first display object also includes the default deviation parameter display value corresponding to the performance index.

9. The data center facility monitoring method according to claim 1, characterized in that, The predicted performance index of the target performance index, generated using the prediction model based on performance data of the target device at multiple time points, includes: A first performance index prediction value is generated based on the performance data of the target device at multiple time points using a first prediction model. A second performance index prediction value is generated based on the performance data of the target device at multiple time points using a second prediction model. The predicted value of the performance index is determined based on the predicted value of the first performance index and the predicted value of the second performance index.

10. The data center facility monitoring method according to claim 9, characterized in that, The first prediction model is a linear model, and the second prediction model is a nonlinear model; Determining the predicted performance index based on the first predicted performance index and the second predicted performance index includes: Calculate the first squared term of the first difference obtained by subtracting the predicted value of the first performance index from the predicted value of the second performance index; Using the first weight corresponding to the first prediction model, the second weight corresponding to the second prediction model, and the third weight corresponding to the first squared term, the predicted values ​​of the first performance index, the second performance index, and the first squared term are weighted and summed to obtain the predicted value of the performance index.

11. The data center facility monitoring method according to claim 9, characterized in that, Also includes: The device health score of the target device is calculated based on the predicted values ​​of the performance indicators corresponding to multiple performance indicators of the target device. If the device health score falls within the device alarm threshold range, it is determined that the system has detected a device malfunction event. Based on the actual state of the target device and the device fault events detected by the system, the fault alarm error value for the target device is determined; The fusion weights between the first prediction model and the second prediction model are updated based on the fault alarm error value.

12. The data center facility monitoring method according to claim 11, characterized in that, The device health score of the target device is calculated based on the predicted values ​​of the performance indicators corresponding to multiple performance indicators of the target device, including: The device health score of the target device is calculated based on the time-series data corresponding to the various performance indicators of the target device, the predicted values ​​of the corresponding performance indicators, and the corresponding historical failure rates.

13. The data center facility monitoring method according to claim 11, characterized in that, Also includes: Based on the system-detected equipment failure events and actual equipment failure events that occurred during the first monitoring period, determine the first ratio of false positive events to false negative events during the first monitoring period; Determining the system deviation parameter based on the first deviation parameter includes: The system deviation parameter is determined based on the first deviation parameter and the first ratio.

14. A data center facility monitoring device, characterized in that, include: The display control module is configured to respond to an adaptive threshold setting command received by the client device, output a first user interface to the display of the client device, and display a first display object on the first user interface; The receiving module is configured to, in response to a user’s first setting operation on the first display object, determine a first deviation parameter of the target performance index of the corresponding target device; The first determining module is used to determine the system deviation parameter based on the first deviation parameter; The second determining module is used to obtain the performance index prediction value of the target performance index generated by the prediction model based on the performance data of the target device at multiple times. The monitoring module is used to determine that the target performance indicator of the target device is in an abnormal state if the deviation between the actual value of the target performance indicator and the predicted value of the performance indicator at the corresponding time exceeds the system deviation parameter when monitoring the target device. The first display object includes indicator graphics of the performance indicators of the target device and display values ​​of deviation parameters corresponding to the performance indicators; the display values ​​of deviation parameters are determined based on existing values ​​of the system deviation parameters. Determine the first historical prediction error parameter and the second historical prediction error parameter of the prediction model for the target performance index; The deviation adjustment value is calculated based on the first historical prediction error parameter and the second historical prediction error parameter. The first historical prediction error parameter is positively correlated with the deviation adjustment value, and the second historical prediction error parameter is negatively correlated with the deviation adjustment value. The system deviation parameter is obtained by superimposing the deviation adjustment value on the first deviation parameter. The adjustment direction of the first deviation parameter is determined by comparing the magnitude of the first deviation parameter with the first historical prediction error parameter. If the first deviation parameter is greater than or equal to the first historical prediction error parameter being compared, the system stability is determined to be strong, and the absolute value of the first deviation parameter is reduced. If the first deviation parameter is less than the first historical prediction error parameter being compared, the system stability is determined to be weak, and the absolute value of the first deviation parameter is increased.

15. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the data center facility monitoring method as described in any one of claims 1 to 13.

16. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the data center facility monitoring method as described in any one of claims 1 to 13.

17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the data center facility monitoring method as described in any one of claims 1 to 13.

Citation Information

Patent Citations

  • Method and system for dynamically setting performance index threshold of IT equipment

    CN105956734A