System availability monitoring and measuring method

By collecting various status data and combining them with redundancy design, a weighted average is calculated to monitor system availability in real time, solving the problem of incomplete system availability monitoring in existing technologies and improving the stability and reliability of the system.

CN120910679APending Publication Date: 2025-11-07LINKER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510776604.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies lack accurate, real-time, and comprehensive system availability monitoring mechanisms, making it impossible to fully measure the system's fault recovery capabilities and actual business reach, resulting in insufficient system stability and reliability.

Method used

The system collects various status data through embedded monitoring tools, combines redundancy and fault tolerance mechanisms, calculates the weighted average of multiple performance indicators, monitors and evaluates system availability in real time, and combines historical data to predict and warn of faults.

Benefits of technology

It enables precise measurement of system availability, timely detection of potential problems, improvement of system stability and reliability, reduction of failure rate, and provision of optimization suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 1J8BCRGPJ8U3X7VUOON6WGCP92SDJM4RGKXAMB61
    Figure 1J8BCRGPJ8U3X7VUOON6WGCP92SDJM4RGKXAMB61
  • Figure 24VG8EUZKINQHDTGBFAAFE2EIUVFDGJATNCFFWJX
    Figure 24VG8EUZKINQHDTGBFAAFE2EIUVFDGJATNCFFWJX
  • Figure 6Z09X6MDRXRESKDUMRZIFAZ5CVCXHWLGHDCKTSUV
    Figure 6Z09X6MDRXRESKDUMRZIFAZ5CVCXHWLGHDCKTSUV
Patent Text Reader

Abstract

The invention discloses a system availability monitoring and measuring method. The method comprises the following steps: acquiring various state data and redundancy and fault-tolerant mechanism working states of a system through an embedded monitoring tool of the system; calculating performance indexes based on redundancy design and a fault-tolerant mechanism, comprehensively considering reliability of components and services, and calculating system availability in different fault modes; calculating the availability rate of the system and adding fault recovery and switching duration; performing fault prediction and early warning by combining the data and the analysis result with historical data; and according to the real-time monitoring result, automatically generating availability and availability rate reports and providing optimization suggestions. According to the method, the system availability can be comprehensively and truly monitored and measured, faults are predicted in advance, early warning is performed, targeted optimization suggestions can be provided, and the overall performance of the system is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer systems and network services, and more particularly to a system availability monitoring and measuring method. BACKGROUND

[0002] In the field of information technology, system availability and availability rate are important indicators to measure the stability of computer systems or services. Traditionally, the availability of the system mainly depends on periodic manual inspection and simple online state detection, lacking accurate, real-time and comprehensive monitoring mechanism. In addition, the calculation of availability rate usually only focuses on the "online" state, ignoring the recovery ability of the system when partial failure occurs and whether the actual business is fully reached. Therefore, it is of great significance to develop a method that can not only comprehensively monitor system availability and timely discover and repair faults, but also quantify the availability rate. SUMMARY

[0003] In view of the deficiencies in the prior art, the purpose of the present application is to provide a system availability monitoring and measuring method, which can monitor the running state of the system in real time and comprehensively evaluate its availability.

[0004] To achieve the above-mentioned purpose, the present application provides the following technical scheme: a system availability monitoring and measuring method, comprising the following steps: Step one, through the system embedded monitoring tool, collect various state data of the system, including but not limited to hardware device running state, operating system performance, network connectivity, service response time, at the same time, collect the working state of redundancy and fault tolerance mechanism; Step two, based on the redundancy design and fault tolerance mechanism of the system, calculate the performance indicators such as response time, error rate and resource utilization, consider the reliability of each component and service through these indicators, and calculate the availability of the system under different failure modes; Step three, calculate the availability rate of the system according to the monitoring result, and add the fault recovery time and switching time in the availability rate calculation, so that the availability rate can more truly reflect the actual availability of the system; Step four, use the collected data and analysis results, combine with historical data to predict the fault of the system, predict the possible fault of the system, and give early warning; Step five, automatically generate availability and availability rate report according to real-time monitoring result, and provide targeted optimization suggestions.

[0005] As a further improvement of the present application, the way of considering the reliability of each component and service through these indicators in step two is: assigning different weight coefficients to each performance indicator, and then calculating the comprehensive evaluation value by weighted average method, specifically: assuming that there are n performance indicators, the weight of each indicator is respectively (i=1,2,...,n), the corresponding performance value is The weighted average formula is: Wherein, C: indicates the weighted average number, the comprehensive availability index, : indicates the weight coefficient of the i-th element, : indicates the value of the i-th performance index, n: indicates the total number of elements.

[0006] As a further improvement of the present application, the availability of the system in step three is obtained by the following steps: First, calculate the conventional system availability: Then consider the specific reach rate of the traffic volume of a single time slice to calculate the traffic reach rate: The threshold of the preset traffic reach rate is set, when the traffic reach rate is lower than the threshold, the time slice availability is considered to be 0, and when the traffic reach rate is higher than the threshold, the time slice availability is considered to be 1, and then the system availability is calculated based on the traffic reach rate: .

[0007] As a further improvement of the present application, the early warning mode in step four is: setting a minimum availability threshold Then, whether to perform early warning is judged by the following formula: .

[0008] As a further improvement of the present application, the resource utilization rate in step two is calculated by the following method, specifically: assuming is the CPU usage rate, is the memory usage rate, and the resource utilization rate of the system can be represented as: Wherein, is the overall resource utilization rate.

[0009] As a further improvement of the present application, the step five also has a real-time monitoring step, specifically: for the performance evaluation of each time point The recursive algorithm (such as weighted moving average) can be used to smooth short-term fluctuations: Wherein, is the smoothing factor, is the performance index at the current time, is the comprehensive evaluation result of the last time.

[0010] The present application has the advantages that the present application comprehensively monitors and accurately measures system availability by integrating multiple key indicators, can evaluate the health status of the system in real time, discovers potential problems in time, and provides early warning mechanism. Compared with traditional monitoring methods, the present application can more accurately reflect the actual availability of the system, reduce the system failure rate, and improve the stability and reliability of the system. DETAILED DESCRIPTION

[0011] The following embodiments will be further described to further illustrate the present application.

[0012] The system availability monitoring and measuring method of the present embodiment includes the following five processes: 1. Data collection Through the monitoring tool embedded in the system, various state data of the system are collected, including but not limited to hardware device running state, operating system performance, network connectivity, service response time, etc. At the same time, the working state of the redundancy and fault tolerance mechanism (such as the enabling condition of the standby server, the load balancing condition, etc.) is collected.

[0013] 2. Availability calculation Based on the redundancy design and fault tolerance mechanism of the system, the reliability of each component and service is comprehensively considered to calculate the availability of the system under different failure modes. The availability is dynamic and adjusted in combination with the current state. Specifically, the availability not only calculates the "online" time, but also includes the case of partial failure recovery. For example, when a subsystem fails, if the backup system takes over in time, the system is considered available.

[0014] 3. Availability rate calculation According to the monitoring results, the availability rate of the system is calculated. The calculation is based on the ratio of the normal running time of the system to the total time, considering the recovery ability of the system failure and the redundancy mechanism, the failure recovery time and the switching time are added in the availability rate calculation, so that the availability rate more truly reflects the actual availability of the system.

[0015] 4. Failure prediction and early warning Using the collected data and analysis results, combined with historical data, the system is predicted for failure, the possible failure of the system is predicted, and early warning is performed. The system administrator can take measures before the failure occurs to ensure that the system can continue to operate efficiently.

[0016] 5. Report and optimization suggestion The system automatically generates availability and availability rate reports according to real-time monitoring results, and provides targeted optimization suggestions. For example, based on the monitored failure mode, the system can suggest to increase redundant components or optimize load balancing strategy.

[0017] In the process of calculating the availability in step 2, a plurality of indexes are comprehensively reflected, and in the embodiment, compared with the conventional availability monitoring method which usually depends on a single performance index such as response time or error rate, the health condition of the system cannot be comprehensively reflected. However, the application adopts a plurality of key performance indexes (such as response time, error rate and resource utilization rate), and the indexes are weighted and comprehensively evaluated, so that the comprehensive availability of the system can be more accurately measured, and by considering the hardware resource usage of the system, the resource overload problem can be found in advance, and the system unavailability caused by resource limitation can be avoided.

[0018] Further, the application assigns different weight coefficients to each performance index, so that the influence degree of each index on the comprehensive availability index is more reasonable. The use of the weighted average method enhances the accuracy of the evaluation result, and makes the system availability evaluation more in line with the actual situation.

[0019] In the weighted average method, each performance index has a corresponding weight coefficient, and the comprehensive evaluation value can be calculated by weighted summation. Assuming that there are n performance indexes, the weight of each index is (i = 1, 2,..., n), and the corresponding performance value is The weighted average formula is: C: represents the weighted average number, the comprehensive availability index. : represents the weight coefficient of the i-th element. : represents the value of the i-th performance index. n: represents the total number of elements.

[0020] Further, in the embodiment, the availability rate is calculated in the following way: A minimum availability threshold is set, and when the comprehensive availability index of the system is lower than the threshold, the system can automatically trigger an alarm.

[0021] The conventional system availability calculation method is: However, when a system provides continuous business capability (such as an operator's short message service, which needs to measure the number of short messages planned to be sent per minute and the number of short messages actually sent successfully, and only when the success rate is high enough is it defined as available, and not that the current minute has a successful sending), the above-mentioned availability index can no longer truly measure the system business reach. What index can be used to measure whether the system is running normally? Therefore, the specific reach rate of the single time slice should be considered, and taking a single time slice of 1 minute as an example, it is represented as: A threshold is defined for the reach rate, when the reach rate is below the threshold, the time slice availability is considered as 0, and above the threshold, the time slice availability is considered as 1, then the availability of the system is calculated based on the business reach rate: .

[0022] Further, the threshold triggered way is used to realize the fault prediction and early warning, as follows: A minimum availability threshold is set, when the comprehensive availability index of the system is below the threshold, the system can automatically trigger the early warning mechanism, and timely alarm to the operation and maintenance personnel. This mechanism effectively enhances the real-time monitoring and fault response capability of the system, and helps to prevent and handle potential problems in advance.

[0023] In order to judge whether the system meets the availability requirement, a minimum availability threshold is set , when the comprehensive availability index C is below the threshold, the early warning mechanism is triggered. The formula is expressed as: This formula shows that when the system availability index is below the preset threshold, alarm or repair measures need to be taken.

[0024] Further, the resource utilization rate of the embodiment can be generally measured by calculating the usage rate of resources, such as CPU utilization rate, memory occupancy rate, etc. Assuming is the CPU usage rate, is the memory usage rate, the resource utilization rate of the system can be expressed as: Here, is the overall resource utilization rate, which can be used as an important index in comprehensive evaluation.

[0025] Further, the monitoring and measuring method of the system availability of the embodiment emphasizes real-time, which can monitor the running state of the system in real time and dynamically evaluate its availability. When the availability decreases, the system can immediately respond and take corresponding measures, which improves the efficiency and flexibility of system management.

[0026] For dynamic system state evaluation, the changes of each performance index can be tracked in time series, and the evaluation model can be adjusted according to historical data. For example, for performance evaluation at each time point , a recursive algorithm (such as weighted moving average) can be used to smooth short-term fluctuations: Where, is the smoothing factor, is the performance index at the current time, The comprehensive evaluation result of the last moment. In this way, the evaluation result not only depends on the current performance value, but also considers historical data, which can better reflect the overall trend of the system.

[0027] In summary, the system availability monitoring and measurement method of the embodiment can collect data at different levels (such as hardware, software, network status, etc.), and dynamically calculate and measure the availability and availability rate of the system by combining the design redundancy and fault tolerance mechanism of the system. Through intelligent analysis of the system, a quick response can be made when a fault or performance degradation occurs to ensure efficient operation of the system.

[0028] The above is only the preferred embodiment of the present application, and the protection scope of the present application is not limited to the above-mentioned embodiments. Any technical solution that belongs to the idea of the present application is within the protection scope of the present application. It should be noted that for ordinary technical personnel in the technical field, some improvements and refinements without departing from the principles of the present application are also considered within the protection scope of the present application.

Claims

1. A method of monitoring and measuring system availability, characterized by: Comprising the following steps: Step one, through the system embedded monitoring tools, collect various state data of the system, including but not limited to hardware device running state, operating system performance, network connectivity, service response time, at the same time, collect the working state of redundancy and fault tolerance mechanism; Step two, based on the redundancy design and fault tolerance mechanism of the system, calculate the performance indicators such as response time, error rate and resource utilization, consider the reliability of each component and service through these indicators, and calculate the availability of the system under different failure modes; Step three, according to the monitoring result, calculate the availability of the system, and add the fault recovery time and switching time in the availability calculation, so that the availability can more truly reflect the actual availability of the system; Step four, use the collected data and analysis results, combined with historical data to predict the system failure, predict the possible failure of the system, and give early warning; Step five, according to the real-time monitoring result, automatically generate availability and availability report, and provide targeted optimization suggestions.

2. The method for monitoring and measuring system usability according to claim 1, c h a r a c t e r i z e d b y: The way of considering the reliability of each component and service by the indexes in step two is: different weight coefficients are assigned to each performance index, and then the comprehensive evaluation value is calculated by the weighted average method, specifically: assuming that there are n performance indexes, the weight of each index is (i = 1, 2,..., n), and the corresponding performance value is The weighted average formula is: ; Wherein, C: represents the weighted average number, comprehensive availability index, : represents the weight coefficient of the i th element, : represents the value of the i th performance indicator, n: represents the total number of elements.

3. The method of monitoring and measuring system usability according to claim 1 or 2, c h a r a c t e r i z e d b y: The availability of the system in step three is calculated by the following steps: First, calculate the conventional system availability: ; Then consider the specific reach rate of the single time slice traffic to calculate the business reach rate: ; The threshold value of the preset business reach rate is set, when the business reach rate is lower than the threshold value, the time slice availability is considered as 0, and when the business reach rate is higher than the threshold value, the time slice availability is considered as 1, then the system availability is calculated based on the business reach rate: 。 4. The method for monitoring and measuring system usability according to claim 2, wherein: The manner of early warning in the fourth step is to set a minimum availability threshold Then, it is determined whether to perform early warning by the following formula: 。 5. The method of monitoring and measuring system usability according to claim 1 or 2, c h a r a c t e r i z e d b y: The resource utilization rate in the step two is calculated in the following way, specifically: assuming is the CPU usage rate, is the memory usage rate, and the resource utilization rate of the system can be expressed as: ; wherein, is the overall resource utilization.

6. The method of monitoring and measuring system usability according to claim 1 or 2, c h a r a c t e r i z e d b y that: The step five also has a real-time monitoring step, specifically: for each time point performance evaluation Short-term fluctuations can be smoothed using a recursive algorithm (such as a weighted moving average): ; wherein, is a smoothing factor, is a performance indicator at the current time, is the comprehensive evaluation result at the previous time.