Method, device, server and medium for monitoring health of server storage system
By collecting and calculating dynamic thresholds in the server storage system, the problem of inaccurate health grading in the existing technology is solved, and the accurate quantification of the health of the server storage system and the rapid response to load fluctuations are achieved.
Patent Information
- Application Number
- CN202510935769.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Existing technologies cannot accurately quantify the health of server storage systems. The health grading accuracy is low and cannot reflect the overall status of the server storage system. In addition, the pre-set thresholds cannot meet the needs of structural changes.
By collecting the indicator values of multiple operating indicators of the server storage system in different collection cycles, dynamic thresholds are calculated based on multiple operating indicators in historical collection cycles. The thresholds are dynamically adjusted to adapt to load changes. The health of the server storage system is quantified by combining indicator levels and status parameters.
It achieves accurate quantification of the health of server storage systems, improves the accuracy of health grading, and can quickly respond to load fluctuations, ensuring the accuracy and adaptability of thresholds.
Smart Images

Figure CN120448223B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of server storage system fault diagnosis, and in particular to a method, device, server, and medium for monitoring the health of a server storage system. Background Art
[0002] With the development of information technology, server storage systems are often in high-intensity, high-load operation, which places higher demands on the stability and reliability of server storage systems. By monitoring the health of server storage systems, we can understand the operating status of server storage systems.
[0003] Currently, related technologies monitor individual components of server storage systems, relying on the experience of operations and maintenance personnel to quantify the health of server storage systems based on the indicators of individual components and pre-set thresholds. However, these technologies lack holistic indicators that can reflect the overall status of server storage systems. Furthermore, as the application structure of server storage systems becomes increasingly complex, pre-set thresholds cannot meet changing needs, resulting in an inability to accurately quantify the health of server storage systems and reducing the accuracy of health grading. Summary of the Invention
[0004] The present application provides a method, device, server and medium for monitoring the health of a server storage system, so as to at least solve the problem in the related art that the health of the server storage system cannot be accurately quantified and the accuracy of health grading is low.
[0005] This application provides a method for monitoring the health of a server storage system, comprising:
[0006] According to the preset collection period, the indicator values of multiple operating indicators of the server storage system to be monitored are collected.
[0007] In a current collection cycle, obtain indicator values of multiple operating indicators of multiple historical collection cycles before the current collection cycle.
[0008] According to the indicator values of multiple operating indicators in multiple historical collection cycles, the dynamic threshold value of each operating indicator in the current collection cycle is calculated.
[0009] The state parameters of each operating indicator are determined according to the indicator value of each operating indicator in the current collection cycle and the dynamic threshold value in the current collection cycle.
[0010] Get the indicator level of each operating indicator.
[0011] The health level of the server storage system is determined based on the status parameters and indicator levels of various operating indicators.
[0012] Outputs the health rating results for the server storage system.
[0013] The present application also provides a server storage system health monitoring device, comprising:
[0014] The collection module is used to collect the index values of multiple operating indicators of the server storage system to be monitored according to a preset collection cycle.
[0015] The first acquisition module is used to acquire, in a current acquisition cycle, indicator values of multiple operating indicators of multiple historical acquisition cycles before the current acquisition cycle.
[0016] The calculation module is used to calculate the dynamic threshold value of each operating indicator in the current collection cycle based on the indicator values of multiple operating indicators in multiple historical collection cycles.
[0017] The first determination module is configured to determine the state parameters of each operating indicator according to the indicator value of each operating indicator in the current collection period and the dynamic threshold value in the current collection period.
[0018] The second acquisition module is used to obtain the indicator level of each operating indicator.
[0019] The second determination module is used to determine the health level classification result of the server storage system according to the status parameters and indicator levels of various operating indicators.
[0020] The output module is used to output the health rating results of the server storage system.
[0021] The present application also provides a server, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned server storage system health monitoring methods when executing the computer program.
[0022] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned server storage system health monitoring methods are implemented.
[0023] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned server storage system health monitoring methods when executed by a processor.
[0024] Through this application, the index values of multiple operating indicators of the server storage system in different collection cycles are collected. During the collection process, since the load of the server storage system is constantly changing and fluctuating, as the index values of various operating indicators collected increase, it is necessary to calculate the dynamic thresholds of various operating indicators in the current collection cycle in real time based on the index values of multiple operating indicators in the historical collection cycle, and realize adaptive dynamic adjustment of the thresholds with nonlinear changes in the load to ensure the accuracy of the thresholds in the current collection cycle, providing a basis for quantifying the health of the server storage system. According to the dynamic thresholds of the current collection cycle, the status parameters of various operating indicators are determined; according to the indicator levels and status parameters of different operating indicators, the health level of the server storage system is quantified, which can improve the accuracy of the health grading. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0026] Figure 1 A schematic diagram of a scenario of a method for monitoring the health of a server storage system provided in an embodiment of the present application;
[0027] Figure 2 A flow chart of a method for monitoring the health of a server storage system provided in an embodiment of the present application;
[0028] Figure 3 A schematic diagram of the structure of a server storage system health monitoring device provided in an embodiment of the present application;
[0029] Figure 4 A schematic diagram of the structure of the server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0031] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0032] With the development of information technology, server storage systems are often in high-intensity, high-load operating states, which places higher demands on the stability and reliability of server storage systems. By monitoring the health of server storage systems, it is possible to understand the operating status of server storage systems. Currently, related technologies monitor individual components of server storage systems, relying on the experience of operation and maintenance personnel to quantify the health of server storage systems based on the indicators of individual components and pre-set thresholds. However, related technologies lack holistic indicators that can reflect the overall status of server storage systems. As the application structure of server storage systems becomes increasingly complex, pre-set thresholds cannot meet changing needs, resulting in an inability to accurately quantify the health of server storage systems and reducing the accuracy of health grading.
[0033] In order to solve the above technical problems, this application proposes the following technical ideas: collecting the indicator values of multiple operating indicators of the server storage system to be monitored in different collection cycles can reflect the overall status of the server storage system; based on the indicator values of multiple operating indicators in multiple historical collection cycles, calculating the dynamic thresholds of each operating indicator in the current collection cycle, and dynamically adjusting the thresholds to adapt to changes in demand, solving the problem that the pre-set thresholds cannot meet the changing needs; according to different indicator levels and status parameters, the health of the server storage system can be accurately quantified, and the accuracy of health grading can be improved.
[0034] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0035] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the server storage system health monitoring method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0036] refer to Figure 1 , Figure 1 Schematic diagram of a scenario of a method for monitoring the health of a server storage system provided in an embodiment of the present application. Figure 1 As shown, it includes: a receiving device 101, a processing device 102 and a display device 103.
[0037] It is understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the server storage system health monitoring method. In other feasible implementations of this application, the above architecture may include more or fewer components than shown, or combine or split certain components, or arrange the components differently. The specific configuration can be determined based on the actual application scenario and is not limited here. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0038] In a specific implementation process, the receiving device 101 may be an input / output interface or a communication interface, and may obtain indicator values of multiple operating indicators of the server storage system to be monitored.
[0039] The processing device 102 can perform a series of processing on the indicator values of multiple operating indicators of the server storage system to be monitored, obtain the health level grading result of the server storage system, and output the health level grading result.
[0040] The display device 103 can be used to display the health level grading results.
[0041] In addition, the network architecture and business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0042] Figure 2 A flow chart of a method for monitoring the health of a server storage system provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, an embodiment of the present application provides a method for monitoring the health of a server storage system. The method is described in detail as follows:
[0043] S201: Collect indicator values of multiple operating indicators of a server storage system to be monitored according to a preset collection period.
[0044] In this embodiment, before executing step S201 , it is necessary to establish an operation indicator library based on the configuration of the server storage system and the running business.
[0045] In this embodiment, the operating indicator library includes operating indicators at the hardware layer and operating indicators at the system layer.
[0046] Hardware-layer metrics are selected based on the performance of different components. For the Peripheral Component Interconnect Express (PCIe) link layer, collected performance metrics include bit error rate, LTSSM state machine transition count, and link bandwidth utilization. For Redundant Array of Independent Disks (RAID) cards, collected performance metrics include array redundancy status, cache hit rate, rebuild progress, and rebuild duration. For hard drive backplanes, collected performance metrics include SAS or SATA link performance metrics, including cyclic redundancy check error counts and power supply voltage fluctuation range. For hard drives, collected performance metrics include SMART parameters, I / O latency, and throughput. SMART parameters include remapped sector counts and media errors.
[0047] In this embodiment, a mapping relationship library between operating indicators and components is established, and the fields of the mapping relationship library include primary key, operating indicator ID, operating indicator name, component name, etc.
[0048] Among them, the operating indicators of the system layer include the operating system I / O error log frequency, CPU usage, and memory resource usage.
[0049] In this embodiment, various operating indicator data are collected according to a preset collection cycle, recorded and saved in a database respectively. The fields of the data storage structure include the server serial number, operating indicator ID, indicator value and collection time, etc.
[0050] S202: In a current collection cycle, obtain indicator values of multiple operating indicators of multiple historical collection cycles before the current collection cycle.
[0051] S203: Calculating a dynamic threshold value of each operating indicator in a current collection period according to the indicator values of the multiple operating indicators in multiple historical collection periods.
[0052] In this embodiment, server storage systems are prone to transient load spikes in high-concurrency scenarios, and static thresholds cannot quickly capture short-term abnormal behavior. Therefore, adaptive dynamic thresholds are generated by dynamically adjusting the values of multiple operating indicators over multiple historical collection periods and real-time load, addressing the problem of static thresholds being unable to cope with load fluctuations.
[0053] Specifically, based on the indicator values of multiple operating indicators in multiple historical collection cycles, the dynamic threshold value of each operating indicator in the current collection cycle is calculated using the formula:
[0054]
[0055] Where, Indicates the dynamic threshold of each operating indicator in the current collection cycle. It represents the window moving average of the indicator values of multiple operating indicators in multiple historical collection periods. represents the fluctuation tolerance, represents the normalization factor.
[0056] In this embodiment, the window moving average of the indicator values of the plurality of operating indicators in a plurality of historical collection periods reflects the steady-state baseline of each operating indicator under a normal state.
[0057] The formula for the window moving mean is:
[0058]
[0059] Where, represents the window moving mean, represents the window length of the sliding window, Indicates that the current collection cycle is A collection cycle, Indicates the indicator values of various operating indicators in the i-th collection period among multiple historical collection periods.
[0060] In this embodiment, the fluctuation tolerance is based on the standard deviation of the indicator value in the moving window and is used to dynamically adjust the tolerance of the threshold to fluctuations.
[0061] In this embodiment, the formula for standard deviation is:
[0062]
[0063] Where, Represents standard deviation.
[0064] In this embodiment, the normalization factor maps the load to the interval [0,1] through the Sigmoid function, so that the threshold can be dynamically scaled with the load. That is, when the load is high, the threshold can tolerate fluctuations of Large, otherwise it decreases.
[0065] The formula for the normalization factor is:
[0066]
[0067] Where, represents the normalization factor, Indicates the current load value of the server storage system. Indicates the maximum load value of the server storage system. is the slope parameter.
[0068] In this embodiment, the operating indicator reflecting the load can be selected according to business needs, such as CPU occupancy, storage I / O throughput, etc. For example, if the operating indicator reflecting the load is CPU occupancy, then The current CPU usage of the server storage system. The maximum CPU usage of the server storage system.
[0069] In this embodiment, the load sensitivity is adjusted by the slope parameter, which can be configured to control the load sensitivity. By setting the slope parameter reasonably, the dynamic threshold is ensured to be adaptively adjusted under different loads. Based on business needs, the normalization factor value of a specific load point needs to be accurately matched. The slope parameter is calculated by the load value and the normalization factor value of the specific load point, that is, when the load value reaches When the normalization factor value is The slope parameter is calculated using the following formula:
[0070]
[0071] S204: Determine the state parameters of each operating indicator according to the indicator value of each operating indicator in the current collection period and the dynamic threshold value in the current collection period.
[0072] The status parameters include meeting the standard and failing to meet the standard. There is no intermediate transition between meeting the standard and failing to meet the standard.
[0073] Specifically, step S204 includes S2041 to S2045:
[0074] S2041: Obtain the indicator type of each operating indicator.
[0075] In this embodiment, the indicator types include positive indicators and negative indicators.
[0076] In this embodiment, higher values for positive indicators, such as cache hit ratio and throughput, are preferred. A higher cache hit ratio results in higher data read efficiency. A cache hit ratio below the dynamic threshold may affect storage performance. That is, if the positive indicator value is greater than or equal to the dynamic threshold for the current collection cycle, the status parameter is considered met; otherwise, it is considered unmet.
[0077] S2042: If the indicator type is a positive indicator, and the indicator value of each operating indicator in the current collection period is not less than the dynamic threshold value of each operating indicator in the current collection period, then the status parameter of each operating indicator is up to standard.
[0078] S2043: If the indicator type is a positive indicator, and the indicator value of each operating indicator in the current collection period is less than the dynamic threshold value of each operating indicator in the current collection period, then the status parameter of each operating indicator is not up to standard.
[0079] In this embodiment, the smaller the value of the reverse indicator, such as the bit error rate and the number of bad tracks, the better. A bit error rate exceeding the dynamic threshold may cause data transmission errors and trigger an abnormality warning. That is, if the reverse indicator value is less than or equal to the dynamic threshold of the current acquisition cycle, the status parameter is considered to meet the standard; otherwise, it is considered to be unsatisfactory.
[0080] S2044: If the indicator type is a reverse indicator, and the indicator value of each operating indicator in the current collection period is not greater than the dynamic threshold value of each operating indicator in the current collection period, then the status parameter of each operating indicator is up to standard.
[0081] S2045: If the indicator type is a reverse indicator, and the indicator value of each operating indicator in the current collection period is greater than the dynamic threshold of each operating indicator in the current collection period, then the status parameter of each operating indicator is not up to standard.
[0082] S205: Obtain the indicator level of each operating indicator.
[0083] In this embodiment, the indicator level of each operating indicator is pre-set. Optionally, the indicator level can be divided according to the severity of the indicator.
[0084] For example, operating indicators that directly affect the availability and data security of the server storage system and may cause business interruption if they fail to meet the standards are classified as core indicators; operating indicators that affect the performance or some functions of the server storage system and may cause potential risks if they fail to meet the standards are classified as secondary indicators; operating indicators that reflect the operating environment or non-critical functions of the server storage system and are used to assist in analyzing the health status are classified as auxiliary indicators.
[0085] In this embodiment, the indicator classification can be adjusted according to the configuration and application scenario of the server storage system. For example, a database server can set the RAID card reconstruction progress as the core indicator; a common file server can set the throughput as the secondary indicator.
[0086] S206: Determine a health rating result of the server storage system based on the status parameters and indicator levels of each operating indicator.
[0087] In this embodiment, the health rating results of the server storage system include healthy, warning, abnormal, and faulty. In subsequent embodiments, the process of determining the health rating results of the server storage system will be described in detail.
[0088] S207: Output the health rating result of the server storage system.
[0089] In summary, collecting the index values of multiple operating indicators of the server storage system in different collection cycles can reflect the overall status of the server storage system. During the collection process, since the load of the server storage system is constantly changing, as the index values of various operating indicators collected increase, it is necessary to calculate the dynamic thresholds of various operating indicators in the current collection cycle in real time based on the index values of multiple operating indicators in the historical collection cycle, and realize adaptive dynamic adjustment of the thresholds with nonlinear changes in load to ensure the accuracy of the thresholds in the current collection cycle, providing a basis for quantifying the health of the server storage system. According to the dynamic thresholds of the current collection cycle, the status parameters of various operating indicators are determined; according to the indicator levels and status parameters of different operating indicators, the health of the server storage system is quantified, which can improve the accuracy of the health grading.
[0090] Based on the above embodiment, in this embodiment, different processing is performed on different health level grading results.
[0091] In this embodiment, a fault handling rule base is pre-built. The fields in the fault handling rule base include health level and handling solution. The health level is the health level classification result of the above embodiment, including healthy, warning, abnormal, and fault. The handling solution stores the handling solutions for various health levels.
[0092] For example, when the health level is healthy, the corresponding processing solution is to record the health level without intervention; when the health level is warning, the corresponding processing solution is to send an early warning to notify the operation and maintenance side of possible failures in the future; when the health level is abnormal, the corresponding processing solution is to trigger an alarm and start self-repair; when the health level is failure, an emergency circuit breaker is triggered, forcibly switching the backup node of the server storage system, and starting self-repair.
[0093] Optionally, after the self-repair is completed, it is necessary to re-monitor to obtain the latest health rating result of the server storage system. If the repair is not successful, manual intervention is triggered.
[0094] In this embodiment, different health grading results are achieved through a pre-built fault handling rule base, which triggers differentiated repair actions. This allows for rapid acquisition of processing solutions for different health grading results, thereby improving fault handling efficiency.
[0095] Based on the above embodiment, this embodiment introduces the process of determining the health classification result of the server storage system based on the status parameters and indicator levels of various operating indicators. The indicator levels include core indicators, secondary indicators, and auxiliary indicators; the status parameters include meeting the standard and failing the standard, as detailed below:
[0096] S301: Classify various operating indicators according to indicator levels to obtain a core indicator set, a secondary indicator set, and an auxiliary indicator set.
[0097] In this embodiment, if the indicator level of each operating indicator is a core indicator, then each operating indicator is divided into a core indicator set; if the indicator level of each operating indicator is a secondary indicator, then each operating indicator is divided into a secondary indicator set; if the indicator level of each operating indicator is an auxiliary indicator, then each operating indicator is divided into an auxiliary indicator set.
[0098] S302: If the status parameters of all operating indicators in the core indicator set are up to standard, the status parameters of all operating indicators in the secondary indicator set are up to standard, and the proportion of the status parameters of all operating indicators in the auxiliary indicator set that are up to standard is not less than the first proportion threshold, then the health grading result is healthy.
[0099] In this embodiment, the first ratio threshold is set in advance and may be 80%.
[0100] S303: If the status parameters of all operating indicators in the core indicator set are up to standard, and the number of status parameters of all operating indicators in the secondary indicator set that are not up to standard is the first quantity threshold; or / and the proportion of status parameters of all operating indicators in the auxiliary indicator set that are up to standard is less than the first proportion threshold, and the duration is not less than the first time threshold, then the health grading result is a warning.
[0101] In this embodiment, there are two conditions for determining the health grading result as a warning. The first condition is that the status parameters of each operating indicator in the core indicator set are all up to standard, and the number of status parameters of each operating indicator in the secondary indicator set that do not meet the standard is a first quantity threshold. The second condition is that the proportion of status parameters of each operating indicator in the auxiliary indicator set that meet the standard is less than a first proportion threshold, and the duration is not less than a first time threshold. If either of these two conditions is met, the health grading result is a warning.
[0102] In this embodiment, the first quantity threshold and the first time threshold are preset, the first quantity threshold is 1, and the first time threshold is 2 minutes.
[0103] For example, if the status parameters of all operating indicators in the core indicator set are up to standard, and there is one operating indicator in the secondary indicator set whose status parameter is not up to standard; or / and the proportion of the status parameters of each operating indicator in the auxiliary indicator set that meet the standard is less than 80%, and the duration is not less than 2 minutes, then the health level grading result is a warning.
[0104] S304: If the number of status parameters of each operating indicator in the core indicator set that does not meet the standard is the first quantity threshold; or / and the number of status parameters of each operating indicator in the secondary indicator set that does not meet the standard is the second quantity threshold, and the duration is not less than the second time threshold; or / and the proportion of status parameters of each operating indicator in the auxiliary indicator set that meets the standard is less than the second proportion threshold, and the duration is not less than the third time threshold, then the health grading result is abnormal.
[0105] In this embodiment, there are three conditions for judging the health grading result as abnormal. The first condition is: the number of status parameters of each operating indicator in the core indicator set that do not meet the standard is the first quantity threshold; the second condition is that the number of status parameters of each operating indicator in the secondary indicator set that do not meet the standard is the second quantity threshold, and the duration is not less than the second time threshold; the third condition is that the proportion of status parameters of each operating indicator in the auxiliary indicator set that meet the standard is less than the second proportion threshold, and the duration is not less than the third time threshold. If any of the above three conditions is met, the health grading result is abnormal.
[0106] In this embodiment, the second quantity threshold, the second time threshold, the second ratio threshold and the third time threshold are pre-set, the second quantity threshold is 2, the second time threshold is 3 minutes, the second ratio threshold is 60%, and the third time threshold is 5 minutes.
[0107] For example, if one of the core indicator parameters is below standard, or / and two of the secondary indicator parameters are below standard for at least three minutes, or if less than 60% of the auxiliary indicator parameters meet standard for at least five minutes, the health rating is considered abnormal if any of the three conditions are met.
[0108] S305: If the number of status parameters of each operating indicator in the core indicator set that does not meet the standard is not less than the second quantity threshold; or / and the number of status parameters of each operating indicator in the core indicator set that does not meet the standard is not less than the first quantity threshold, and the number of status parameters of each operating indicator in the secondary indicator set that does not meet the standard is not less than the third quantity threshold, then the health grading result is fault.
[0109] In this embodiment, there are two conditions for determining a health grading result as a fault. The first condition is that the number of status parameters of each operating indicator in the core indicator set that do not meet the standard is not less than a second threshold. The second condition is that the number of status parameters of each operating indicator in the core indicator set that do not meet the standard is not less than a first threshold, and the number of status parameters of each operating indicator in the secondary indicator set that do not meet the standard is not less than a third threshold. If either of these two conditions is met, the health grading result is a fault.
[0110] For example, if the number of status parameters of each operating indicator in the core indicator set that does not meet the standard is not less than 2; or / and the number of status parameters of at least one operating indicator in the secondary indicator set that does not meet the standard, and the number of status parameters of each operating indicator in the secondary indicator set that does not meet the standard is not less than 3, then if either of the above two conditions is met, the health level is classified as faulty.
[0111] Alternatively, if the server storage system is completely unavailable, such as the API of the server storage system is completely unable to respond to requests, all calls fail, and the success rate is 0, then the health level rating result is failure.
[0112] In this embodiment, the determination conditions of the health level grading result can be flexibly configured according to the configuration and application scenario of the server storage system.
[0113] In summary, the health grading results are divided into four levels: healthy, warning, abnormal, and fault. Different health levels are determined according to different criteria. The criteria comprehensively consider the indicator level and status parameters, achieve accurate quantification of the health level, and improve the accuracy of health grading.
[0114] Based on the above embodiment, in this embodiment, the process of fault location based on the health level classification result is introduced, and the details are as follows:
[0115] S401: According to the health grading result, various operating indicators whose status parameters are not up to standard are obtained.
[0116] In this embodiment, if the health grading result is healthy, it means that there may be operating indicators in the auxiliary indicators whose status parameters are not up to standard; if the health grading result is warning, abnormal or faulty, there must be operating indicators whose status parameters are not up to standard.
[0117] S402: According to various operating indicators whose status parameters are not up to standard, information about faulty components of the server storage system is obtained from a pre-built root cause rule location library.
[0118] In this embodiment, the fields in the pre-built root cause rule location library include primary key, operation indicator set ID and component name. The data in the operation indicator set ID is all operation indicator IDs whose status parameters are not up to standard.
[0119] In this embodiment, according to various operating indicators whose status parameters do not meet the standards, a match is performed in a pre-built root cause rule location library to obtain a matching component name, that is, fault component information of the server storage system.
[0120] Optionally, if there are no matching items in the pre-built root cause rule location library, each operating indicator with substandard status parameters is matched against the mapping relationship library in the above embodiment to obtain the component names corresponding to each operating indicator with substandard status parameters, i.e., the fault component information of the server storage system. If there are no matching items in the mapping relationship library, fault diagnosis is performed vertically on each component on the storage link, in the order of PCIe link layer, RAID card, hard disk backplane, and hard disk.
[0121] In summary, by matching various operating indicators with substandard status parameters with the pre-built root cause rule location library to locate the fault, manual troubleshooting time can be reduced and the time for fault location can be improved.
[0122] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0123] Figure 3 This is a schematic diagram of the structure of the server storage system health monitoring device provided in the embodiment of the present application. Figure 3 As shown, an embodiment of the present application also provides a health monitoring device for a server storage system, including: a collection module 301, a first acquisition module 302, a calculation module 303, a first determination module 304, a second acquisition module 305, a second determination module 306 and an output module 307.
[0124] The collection module 301 is used to collect indicator values of multiple operating indicators of the server storage system to be monitored according to a preset collection period.
[0125] The first acquisition module 302 is configured to acquire, in a current acquisition cycle, indicator values of a plurality of operating indicators of a plurality of historical acquisition cycles before the current acquisition cycle.
[0126] The calculation module 303 is configured to calculate a dynamic threshold value of each operating indicator in a current collection period based on the indicator values of multiple operating indicators in multiple historical collection periods.
[0127] The first determining module 304 is configured to determine the state parameters of each operating indicator according to the indicator value of each operating indicator in the current collection period and the dynamic threshold value in the current collection period.
[0128] The second acquisition module 305 is used to obtain the indicator level of each operating indicator.
[0129] The second determining module 306 is configured to determine a health rating result of the server storage system according to the status parameters and indicator levels of various operating indicators.
[0130] The output module 307 is used to output the health rating result of the server storage system.
[0131] In a possible implementation, the status parameter includes meeting the standard and failing to meet the standard; the first determination module 304 includes:
[0132] The first acquisition unit is used to acquire the indicator type of each operating indicator.
[0133] The first determining unit is configured to determine that if the indicator type is a positive indicator and the indicator value of each operating indicator in the current collection period is not less than the dynamic threshold value of each operating indicator in the current collection period, then the status parameter of each operating indicator is up to standard.
[0134] The second determining unit is configured to determine that if the indicator type is a positive indicator and the indicator value of each operating indicator in the current collection period is less than the dynamic threshold value of each operating indicator in the current collection period, then the status parameter of each operating indicator is not up to standard.
[0135] The third determining unit is configured to determine that if the indicator type is a reverse indicator and the indicator value of each operating indicator in the current collection period is not greater than the dynamic threshold of each operating indicator in the current collection period, then the status parameter of each operating indicator is up to standard.
[0136] The fourth determining unit is configured to determine that if the indicator type is a reverse indicator and the indicator value of each operating indicator in the current collection period is greater than the dynamic threshold of each operating indicator in the current collection period, then the status parameter of each operating indicator is not up to standard.
[0137] In one possible implementation, the indicator levels include core indicators, secondary indicators, and auxiliary indicators; the status parameters include meeting the standards and failing to meet the standards; and the second determination module 306 includes:
[0138] The classification unit is used to classify various operating indicators according to the indicator level to obtain the core indicator set, secondary indicator set and auxiliary indicator set.
[0139] The fifth determination unit is used to determine that if the status parameters of each operating indicator in the core indicator set are all up to standard, the status parameters of each operating indicator in the secondary indicator set are all up to standard, and the proportion of the status parameters of each operating indicator in the auxiliary indicator set that are up to standard is not less than the first proportion threshold, then the health grading result is healthy.
[0140] The sixth determination unit is used to determine that if the status parameters of each operating indicator in the core indicator set are all up to standard, and the number of status parameters of each operating indicator in the secondary indicator set that are not up to standard is a first quantity threshold; or / and the proportion of status parameters of each operating indicator in the auxiliary indicator set that are up to standard is less than the first proportion threshold, and the duration is not less than the first time threshold, then the health grading result is a warning.
[0141] The seventh determination unit is used to determine that if the number of status parameters of each operating indicator in the core indicator set that does not meet the standard is a first quantity threshold; or / and the number of status parameters of each operating indicator in the secondary indicator set that does not meet the standard is a second quantity threshold, and the duration is not less than the second time threshold; or / and the proportion of status parameters of each operating indicator in the auxiliary indicator set that meets the standard is less than the second proportion threshold, and the duration is not less than the third time threshold, then the health grading result is abnormal.
[0142] The eighth determination unit is used to determine that if the number of status parameters of each operating indicator in the core indicator set that does not meet the standard is not less than the second quantity threshold; or / and the number of status parameters of each operating indicator in the core indicator set that does not meet the standard is not less than the first quantity threshold, and the number of status parameters of each operating indicator in the secondary indicator set that does not meet the standard is not less than the third quantity threshold, then the health grading result is fault.
[0143] In a possible implementation, the server storage system health monitoring device further includes a third acquisition module. The third acquisition module includes:
[0144] The second acquisition unit is used to acquire various operating indicators whose status parameters are substandard according to the health grading result.
[0145] The third acquiring unit is configured to acquire, according to various operating indicators whose status parameters are substandard, information about faulty components of the server storage system from a pre-built root cause rule location library.
[0146] In a possible implementation, the formula of the calculation module 303 is:
[0147]
[0148] Where, Indicates the dynamic threshold of each operating indicator in the current collection cycle. It represents the window moving average of the indicator values of multiple operating indicators in multiple historical collection periods. represents the fluctuation tolerance, represents the normalization factor.
[0149] In one possible implementation, the formula for the window moving mean is:
[0150]
[0151] Where, represents the window moving mean, represents the window length of the sliding window, Indicates that the current collection cycle is A collection cycle, Indicates the indicator values of various operating indicators in the i-th collection period among multiple historical collection periods.
[0152] In one possible implementation, the formula for the normalization factor is:
[0153]
[0154] Where, represents the normalization factor, Indicates the current load value of the server storage system. Indicates the maximum load value of the server storage system. is the slope parameter.
[0155] For descriptions of features in the embodiments corresponding to the apparatus for monitoring the health of a server storage system, reference may be made to the descriptions of features in the embodiments corresponding to the method for monitoring the health of a server storage system, which will not be detailed here.
[0156] Figure 4 This is a schematic diagram of the structure of the server provided in the embodiment of the present application. Figure 4 As shown, the server provided in this embodiment includes: at least one processor 401 and a memory 402. Optionally, the server also includes a communication component 403. The processor 401, the memory 402 and the communication component 403 are connected via a bus.
[0157] During the specific implementation process, at least one processor 401 executes the computer-executable instructions stored in the memory 402, so that the at least one processor 401 executes the above-mentioned embodiment of the health monitoring method for the server storage system.
[0158] The specific implementation process of the processor 401 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.
[0159] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.
[0160] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.
[0161] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.
[0162] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned server storage system health monitoring method embodiments when running.
[0163] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0164] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned server storage system health monitoring method embodiments are implemented.
[0165] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned server storage system health monitoring method embodiments.
[0166] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0167] The above is a detailed introduction to the health monitoring method, device, server and medium of a server storage system provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for monitoring the health of a server storage system, characterized in that: include: Collect the index values of multiple operating indicators of the server storage system to be monitored according to the preset collection cycle; In a current collection cycle, obtaining indicator values of multiple operating indicators of multiple historical collection cycles before the current collection cycle; Calculating a dynamic threshold value of each operating indicator in the current collection period according to the indicator values of the multiple operating indicators in the multiple historical collection periods; Where, Indicates the dynamic thresholds of the various operating indicators in the current collection cycle. represents the window moving average of the indicator values of the multiple operating indicators in the multiple historical collection periods, represents the fluctuation tolerance, represents the normalization factor, represents the current load value of the server storage system, Indicates the maximum load value of the server storage system, is the slope parameter; Determining the state parameters of the various operating indicators according to the indicator values of the various operating indicators in the current collection period and the dynamic thresholds in the current collection period; Obtaining the indicator level of each of the operating indicators; Determining a health rating of the server storage system based on the status parameters and indicator levels of the various operating indicators; The health rating result of the server storage system is output.
2. The method according to claim 1, characterized in that The status parameters include meeting the standard and failing to meet the standard; Accordingly, determining the state parameters of the various operating indicators according to the indicator values of the various operating indicators in the current collection period and the dynamic thresholds in the current collection period includes: Obtain the indicator type of each operating indicator; If the indicator type is a positive indicator, and the indicator value of each operating indicator in the current collection period is not less than the dynamic threshold value of each operating indicator in the current collection period, then the status parameter of each operating indicator is up to standard; If the indicator type is a positive indicator, and the indicator value of each operating indicator in the current collection period is less than the dynamic threshold value of each operating indicator in the current collection period, then the status parameter of each operating indicator is not up to standard; If the indicator type is a reverse indicator, and the indicator value of each operating indicator in the current collection period is not greater than the dynamic threshold value of each operating indicator in the current collection period, then the status parameter of each operating indicator is up to standard; If the indicator type is a reverse indicator, and the indicator value of each operating indicator in the current collection period is greater than the dynamic threshold of each operating indicator in the current collection period, then the status parameter of each operating indicator is not up to standard.
3. The method according to claim 1, characterized in that The indicator levels include core indicators, secondary indicators and auxiliary indicators; the status parameters include meeting the standards and failing to meet the standards; Accordingly, determining the health rating result of the server storage system according to the status parameters and indicator levels of the various operating indicators includes: Classify the various operating indicators according to the indicator levels to obtain a core indicator set, a secondary indicator set, and an auxiliary indicator set; If the status parameters of the various operating indicators in the core indicator set are all up to standard, the status parameters of the various operating indicators in the secondary indicator set are all up to standard, and the proportion of the status parameters of the various operating indicators in the auxiliary indicator set that are up to standard is not less than a first proportion threshold, then the health level classification result is healthy; If the status parameters of the various operating indicators in the core indicator set are all up to standard, and the number of the status parameters of the various operating indicators in the secondary indicator set that are not up to standard is equal to the first quantity threshold; or / and the proportion of the status parameters of the various operating indicators in the auxiliary indicator set that are up to standard is less than the first proportion threshold, and the duration is not less than the first time threshold, then the health level classification result is a warning; If the number of status parameters of the operating indicators in the core indicator set that do not meet the standards is the first quantity threshold; or / and the number of status parameters of the operating indicators in the secondary indicator set that do not meet the standards is the second quantity threshold, and the duration is not less than the second time threshold; or / and the proportion of status parameters of the operating indicators in the auxiliary indicator set that meet the standards is less than the second proportion threshold, and the duration is not less than the third time threshold, then the health level grading result is abnormal; If the number of status parameters of each operating indicator in the core indicator set that does not meet the standard is not less than the second quantity threshold; or / and the number of status parameters of each operating indicator in the core indicator set that does not meet the standard is not less than the first quantity threshold, and the number of status parameters of each operating indicator in the secondary indicator set that does not meet the standard is not less than the third quantity threshold, then the health grading result is fault.
4. The method according to claim 3, characterized in that After outputting the health rating result of the server storage system, the method further includes: According to the health grading result, obtaining various operating indicators whose status parameters are not up to standard; According to the various operating indicators whose status parameters are not up to standard, the fault component information of the server storage system is obtained from a pre-built root cause rule location library.
5. The method according to claim 1, wherein The formula for the window moving mean is: Where, represents the window moving mean, represents the window length of the sliding window, Indicates that the current collection cycle is A collection cycle, Indicates the indicator values of various operating indicators in the i-th collection period among the multiple historical collection periods.
6. A server storage system health monitoring device, characterized in that: include: The collection module is used to collect the index values of multiple operating indicators of the server storage system to be monitored according to a preset collection period; A first acquisition module is configured to acquire, in a current acquisition cycle, index values of a plurality of operating indicators of a plurality of historical acquisition cycles before the current acquisition cycle; a calculation module, configured to calculate a dynamic threshold value of each operating indicator in the current collection period based on the indicator values of the multiple operating indicators in the multiple historical collection periods; Where, Indicates the dynamic thresholds of the various operating indicators in the current collection cycle. represents the window moving average of the indicator values of the multiple operating indicators in the multiple historical collection periods, represents the fluctuation tolerance, represents the normalization factor, represents the current load value of the server storage system, Indicates the maximum load value of the server storage system, is the slope parameter; A first determining module is configured to determine the state parameters of the various operating indicators according to the indicator values of the various operating indicators in the current collection period and the dynamic thresholds in the current collection period; A second acquisition module is used to obtain the indicator level of each operating indicator; A second determining module is configured to determine a health rating result of the server storage system based on the status parameters and indicator levels of the various operating indicators; An output module is used to output the health grading result for the server storage system.
7. A server, characterized in that: include: memory for storing computer programs; A processor is configured to implement the steps of the server storage system health monitoring method according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the server storage system health monitoring method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Electric quantity value display method and device, storage medium and electronic equipment
CN119988133A
Power equipment data acquisition and processing system and method
CN120017736A