A multi-queue system health monitoring method based on multiple indicators
Through the multi-index health monitoring method, the health of multiple key indicators is calculated using linear and exponential formulas, and the comprehensive health is calculated by weighted summing, the shortcomings of the traditional single index monitoring method are solved, and the accuracy and flexibility of system health assessment are achieved.
Patent Information
- Application Number
- CN202510246887.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-03-04
AI Technical Summary
Traditional health monitoring methods based on single indicators cannot fully reflect the overall health of the system, and linear formulas are difficult to flexibly adapt to the nonlinear characteristics of indicator changes, resulting in high false alarm rates and missed rates.
The health monitoring method of multi-queue system based on multi-index is adopted, and the data loss is filled with linear interpolation method to calculate the health of multiple key indicators, including queue environmental health, message middleware health and queue running state health. The health is calculated using linear and exponential formulas, and the comprehensive health is calculated through weighted summing.
A comprehensive evaluation of the system's multi-dimensional indicators has been achieved, the accuracy and flexibility of health assessment has been improved, the false alarm rate and missed rate have been reduced, and the timeliness and accuracy of system monitoring has been ensured.
Smart Images

Figure CN119718843B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of queue system health monitoring, and in particular to a multi-queue system health monitoring method based on multiple indicators. Background Art
[0002] With the rapid development of information technology, the scale of systems continues to expand, playing an increasingly important role in modern society. The smooth and healthy operation of the system is crucial, because once a failure occurs, it will greatly affect the reliability of the service and the user experience, and may even cause property losses. Therefore, it is necessary to monitor the health of the system in real time and detect system anomalies in real time. Indicators, as a valuable resource for recording the operation status of the system, play a key role in this process. However, due to the large scale of data, it is unrealistic to conduct real-time monitoring manually, and an automated monitoring method must be adopted. Traditional indicator-based health monitoring methods usually use linear formulas to calculate the health of a single indicator and set thresholds for alarms, but this method cannot fully reflect the overall health of the system, and linear formulas are difficult to flexibly adapt to the nonlinear characteristics of indicator changes. For indicators that change dramatically under high load or abnormal conditions, linear formulas may underestimate the potential risks brought by anomalies, resulting in high false alarm and false negative rates in actual applications, affecting the accuracy and timeliness of system monitoring. Summary of the invention
[0003] In order to help solve the above technical problems, the present application provides a multi-queue system health monitoring method based on multiple indicators, which adopts the following technical solutions:
[0004] A multi-queue system health monitoring method based on multiple indicators, wherein the method comprises:
[0005] Step S1: Fill missing values of the original data collected from the multi-queue system by linear interpolation;
[0006] Step S2: Evaluate the health of key indicators in the multi-queue system and output the evaluation results. The key indicators include the health of the queue environment. The health of the queue environment is calculated in the following way:
[0007] Step S211: Generate input data under various load conditions through scripts and input them into the system, record queue environment indicators and system performance data, draw a line graph of system performance data changing over time, and use the queue environment indicator corresponding to a sudden change in system performance data as a preset queue environment indicator threshold;
[0008] Step S212: input the current queue environment index. If the current queue environment index is less than the preset queue environment index threshold, the queue environment health is calculated using the queue environment health linear formula:
[0009] ;
[0010] Among them, α is the cutoff value that determines whether the component is in a healthy state, x is the current queue environment indicator value, threshold is the preset queue environment indicator threshold, and 100 is the maximum value of the queue environment health;
[0011] Step S213: If the current queue environment index is greater than or equal to the preset queue environment index threshold, the queue environment health index formula is used to calculate the queue environment health index:
[0012] ;
[0013] Among them, α is the cutoff value that determines whether the component is in a healthy state, x is the current queue environment indicator value, threshold is the preset queue environment indicator threshold, and slope is the set steepness, which is used to reflect the sensitivity of changes in queue environment health to changes in queue environment indicator values;
[0014] Step S3: Calculate the comprehensive health according to the evaluation result outputted in step S2, and generate warning information.
[0015] Preferably, the step S2 includes: the key indicator includes the health of the message middleware, and the health of the message middleware is calculated by the following method:
[0016] Step S221: Monitor and record the sending, persistence and receiving throughput rates of the message middleware at n moments when the system is normal, and calculate the ratio t of the sending throughput rate to the receiving throughput rate at each moment i. 1,i , the ratio of the persistence throughput rate to the receiving throughput rate t 2,i , get the throughput ratio T at that moment i = [t 1,i , t 2,i ], the throughput ratio T for all time i Take the average and get the ideal throughput ratio T = [t1, t2];
[0017] Step S222: input the current sending throughput rate, persistence throughput rate and receiving throughput rate of the message middleware, calculate the ratio k1 of the sending throughput rate to the receiving throughput rate, and the ratio k2 of the persistence throughput rate to the receiving throughput rate, and obtain the current throughput ratio K = [k1, k2];
[0018] Step S223: Input K and T, and calculate the Euclidean distance between K and T in the following way:
[0019] ;
[0020] Step S224: Substitute D(K, T) into the message middleware health calculation formula, which is as follows:
[0021] ;
[0022] Where C is the maximum acceptable Euclidean distance.
[0023] Preferably, step S2 includes: the key indicator includes the health of the message queue, the health of the message queue includes the health of the queue throughput rate and the health of the queue operation status, and the queue throughput rate health is calculated by the following method:
[0024] Step S231: Monitor and record the system throughput, input queue throughput, output queue throughput and aggregation queue throughput at n moments when the system is normal, and calculate the input queue throughput ratio t at each moment i. 1,i ’ , output queue throughput ratio t 2,i ’ And the aggregation queue throughput ratio t 3,i ’ , get the throughput ratio t at that moment i ’ = [t 1,i ’ ,t 2,i ’ , t 3,i ’ ], the throughput ratio T for all time i ’ Take the average and get the ideal throughput ratio T ’ = [t1 ’ , t2 ’ , t3 ’ ], input queue throughput ratio t 1,i ’ is the ratio of input queue throughput to system throughput, and the output queue throughput is t 2,i ’ is the ratio of the output queue throughput to the system throughput, and the aggregation queue throughput is t 3,i ’ is the ratio of the aggregation queue throughput to the system throughput;
[0025] Step S232: Input the current system throughput, input queue throughput, output queue throughput and aggregation queue throughput, and calculate the input queue throughput ratio k1 respectively. ’ , output queue throughput ratio k2 ’ And the aggregation queue throughput ratio k3 ’ , and get the current throughput ratio K = [k1 ’ , k2 ’ , k3’ ], input queue throughput ratio k1 ’ is the ratio of input queue throughput to system throughput, and the output queue throughput is k2 ’ is the ratio of the output queue throughput to the system throughput, and the aggregation queue throughput ratio k3 ’ is the ratio of the aggregation queue throughput to the system throughput;
[0026] Step S233: Input K ’ , T ’ , calculate K by ’ With T ’ The Euclidean distance between:
[0027] ;
[0028] Step S234: input D(K',T') into the queue throughput health calculation formula, which is as follows:
[0029] ;
[0030] Where C is the maximum acceptable Euclidean distance.
[0031] Preferably, step S2 includes: the queue includes an input queue, an output queue, an aggregation queue and a transaction queue, and the health of the queue operation status is calculated by the following method:
[0032] Step S241: Write a script to simulate the process of the queue gradually filling up. After each data is written, record the queue environment index and system performance data, draw a line graph of the system performance data changing over time, and use the queue occupancy rate corresponding to the sudden change in the system performance data as the preset queue environment index threshold;
[0033] Step S242: If the current queue occupancy is less than the preset queue occupancy threshold, the queue operation status health is calculated using the queue operation status health linear formula:
[0034] ;
[0035] Where α is the cutoff value for determining whether a component is in a healthy state, x is the current queue occupancy rate, and threshold is the preset queue occupancy rate threshold;
[0036] Step S243: If the current queue occupancy is greater than or equal to the preset queue occupancy threshold, the queue operation status health index formula is used to calculate the queue operation status health index:
[0037] ;
[0038] Where α is the cutoff value for determining whether a component is in a healthy state, x is the current queue occupancy rate, threshold is the preset queue occupancy rate threshold, and slope is the steepness;
[0039] Step S244: weighted sum of the operational status health of the input queue, output queue, aggregation queue, and transaction queue is performed in the following manner:
[0040] ;
[0041] Among them, h 11 Indicates the health of the input queue operation status, h 12 Indicates the health of the output queue operation status, h 13 Indicates the health of the aggregation queue operation status, h 14 Indicates the health of the transaction queue operation status, w 11 , w 12 , w 13 , w 14 are the corresponding weight coefficients respectively.
[0042] Preferably, the step S3 further comprises: calculating the comprehensive health by:
[0043] ;
[0044] Among them, h 21 Indicates the health of the queue environment, h 22 Indicates the health of the message middleware, h 23 Indicates the health of the queue operation status, h 24 Indicates the health of the queue throughput, w 21 , w 22 , w 23 , w 24 are the corresponding weight coefficients respectively.
[0045] Preferably, the step S3 comprises: generating warning information by:
[0046] Alarms include overall alarms and scenario alarms. When the health of either the queue environment health or the message middleware health is lower than the preset overall alarm threshold, an overall alarm is generated; when the health of either the queue operation status health or the queue throughput health is lower than the preset scenario alarm threshold, a scenario alarm is generated.
[0047] In summary, this application has the following advantages:
[0048] (1) Taking into account the health of multiple indicators and calculating the overall health by weighted summation can fully reflect the health status of the system, thereby ensuring the accuracy and comprehensiveness of the evaluation results.
[0049] (2) The status of each component of the system is converted into a quantitative value through the health calculation formula, making the assessment of the system health status more accurate.
[0050] (3) By quantifying the health of the system, a solid foundation can be laid for subsequent prediction and timing analysis, so that the future state of the system can be predicted and monitored. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 A schematic block diagram of an embodiment of a multi-queue system health monitoring method based on multiple indicators of the present application;
[0052] Figure 2 A schematic block diagram of an embodiment of a multi-queue system of the present application. DETAILED DESCRIPTION
[0053] The present application is further described below in conjunction with the accompanying drawings, and the structure and principle of the present application are very clear to people in the field. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0054] The software system client health monitoring method in the prior art is aimed at the client application, and its execution steps include four parts: (1) establishing a health rule base; (2) collecting the health data of the client; (3) analyzing the health data and generating warning results: when the operation fails or is abnormal, it is listed as a first-level warning; when the operation time exceeds the threshold, it is listed as different levels of abnormalities; when an abnormality occurs and is recovered within a short time, the warning is eliminated; when multiple abnormalities occur continuously, it is listed as a special-level warning; when the same type of warning reaches a specified number of times, it is automatically merged; (4) sending the warning result.
[0055] The software system performance fault prediction method and system in the prior art include four steps: (1) using LSTM to predict multiple indicators respectively; (2) calculating the difference between the predicted value and the actual value, and if the difference is greater than a threshold, it is judged as an abnormality; (3) if one indicator is abnormal, it is judged as an abnormality; (4) multi-indicator classification to determine the type of abnormality.
[0056] The above two methods have the following disadvantages:
[0057] Relying on a single indicator and focusing on judging the entire system as abnormal when a certain indicator is abnormal, while ignoring the comprehensive judgment of multiple dimensions and multiple indicators, may lead to misjudgment.
[0058] Using rules and thresholds for alerts lacks quantitative evaluation methods and cannot accurately reflect the overall health of the system.
[0059] Figure 1 This is a schematic block diagram of an embodiment of a multi-queue system health monitoring method based on multiple indicators of the present application. Figure 2 A schematic block diagram of an embodiment of a multi-queue system of the present application.
[0060] Combination Figure 1 and Figure 2 It can be understood that the multi-queue system anomaly detection method based on multiple indicators of the present invention includes three steps.
[0061] Step S1: Data preprocessing. The main purpose of this step is to fill missing values in the original data collected from the multi-queue system to ensure that subsequent anomaly detection can be carried out effectively.
[0062] Step S2: Indicator health calculation. After data preprocessing, the indicator health calculation module is responsible for performing health calculations on various system indicators, including queue environment health, middleware health, and message queue health.
[0063] Step S3: Health alarm: calculate the overall health of the system based on the health of each indicator, and issue an alarm based on the alarm rules.
[0064] The data preprocessing module mainly includes missing value filling:
[0065] In step S1, the present application uses linear interpolation to fill in missing values. Linear interpolation is a commonly used data interpolation method for estimating the value between two known data points. It assumes that the change of data points between two points is linear, that is, the data value shows a uniform increase or decrease as the independent variable changes.
[0066] The linear difference calculation formula is as follows:
[0067] ;
[0068] Where x1 and x2 are the horizontal coordinates of the known data points, y1 and y2 are the vertical coordinates of the known data points, x is the horizontal coordinate of the interpolation point, and y is the vertical coordinate of the interpolation point. In this example, there are three sets of data (1, 0.8), (2,), (3, 0.9). The vertical coordinate of the second set of data is missing, and it is filled by aligning the first and third sets of data. Substituting the data into the formula, corresponding to x is 2, x1 and x2 are 1 and 3 respectively, y1 and y2 are 0.8 and 0.9 respectively, the second set of vertical coordinates is 0.85.
[0069] The health evaluation of key indicators in the system may include queue environment health calculation, middleware health calculation, and message queue health calculation.
[0070] The following is a detailed description of the three calculations:
[0071] The design scheme for calculating the health of the queue environment is as follows:
[0072] Step S211: Collect historical data, analyze the relationship between queue environment indicators (such as CPU occupancy, memory occupancy, disk occupancy, etc.) and system performance, find the point near where performance drops significantly, determine the threshold of queue environment indicators, generate input data under various load conditions through scripts and input them into the system, record queue environment indicators and system performance data, draw a line chart of system performance data changing over time, and use the queue environment indicator corresponding to the sudden change in system performance data as the preset queue environment indicator threshold. In this example, the queue environment indicator is disk occupancy, and the threshold is 0.8.
[0073] Specifically, in step S211, a script is used to generate input data for different load conditions, which are input into the system in order from low to high. Each condition lasts for a period of time, and the index data and system performance data in the same time period are recorded. A line graph of system performance over time is drawn to observe the trend of change. Under normal circumstances, the system performance is relatively stable with small fluctuations. If there is a large fluctuation, the index value in this case is set as the index threshold.
[0074] Step S212: Input the current queue environment occupancy rate, that is, the disk occupancy rate, and determine whether the current queue environment index is greater than 0.8. If it is less than 0.8, use the linear formula to calculate the queue environment health. Otherwise, enter step S213.
[0075] The linear formula for queue environment health is as follows:
[0076] ;
[0077] Where α is the demarcation value for determining whether a component is in a healthy state, x is the current queue environment indicator value, threshold is the set queue environment indicator threshold, and 100 is the maximum value of health. In this embodiment, α is 60 and threshold is 0.8. When the disk occupancy rate is 0.5, the calculated queue environment health is 76.
[0078] Step S213: the queue environment index value, that is, the disk occupancy rate, is greater than 0.8, and the queue environment health is calculated using an exponential formula.
[0079] The formula for the queue environment indicator health index is as follows:
[0080] ;
[0081] Where α is the demarcation value for determining whether a component is in a healthy state, x is the current queue environment indicator value, threshold is the set queue environment indicator threshold, and slope is the set steepness. In this embodiment, α is 60, slope is 20, and threshold is 0.8. When the disk occupancy rate is 0.8, the calculated queue environment health is 60.
[0082] The design scheme for calculating the health of the message middleware is as follows:
[0083] Step S221: Monitor and record the sending, persistence and receiving throughput rates of the message middleware at n moments when the system is normal, and calculate the ratio t of the sending throughput rate to the receiving throughput rate at each moment i. 1,i , the ratio of the persistence throughput rate to the receiving throughput rate t 2,i , get the throughput ratio T at that moment i = [t 1,i , t 2,i ], the throughput ratio T for all time i Taking the average, we get the ideal throughput ratio T = [t1, t2].
[0084] Step S222: Input the sending throughput rate, persistence throughput rate and receiving throughput rate of the message middleware at the current moment, calculate the ratio k1 of the sending throughput rate to the receiving throughput rate, and the ratio k2 of the persistence throughput rate to the receiving throughput rate, and obtain the current throughput ratio K = [k1, k2].
[0085] Step S223: Input K and T, and calculate the Euclidean distance between K and T in the following way:
[0086] .
[0087] Step S224: Substitute D(K, T) into the message middleware health calculation formula, which is as follows:
[0088] ;
[0089] Where C is the maximum acceptable Euclidean distance.
[0090] In this embodiment, when K is (1, 1, 0.5), substituting into the formula, the Euclidean distance between K and T is 0.25.
[0091] Where D(K, T) is the Euclidean distance calculated in step S223. In this embodiment, when D(K, T) is 2, the health of the message middleware is 90.7 when substituted into the formula.
[0092] The message queue health design includes two parts: queue operation status health and queue throughput health.
[0093] The queue throughput health is implemented as follows:
[0094] Step S231: Monitor and record the system throughput, input queue throughput, output queue throughput and aggregation queue throughput at n moments when the system is normal, and calculate the input queue throughput ratio t at each moment i. 1,i ’ , output queue throughput ratio t 2,i ’ And the aggregation queue throughput ratio t 3,i ’ , get the throughput ratio t at that moment i ’ = [t 1,i ’ ,t 2,i ’ , t 3,i ’ ], the throughput ratio T for all time i ’ Take the average and get the ideal throughput ratio T ’ = [t1 ’ , t2 ’ , t3 ’ ], input queue throughput ratio t 1,i ’ is the ratio of input queue throughput to system throughput, and the output queue throughput is t 2,i ’ is the ratio of the output queue throughput to the system throughput, and the aggregation queue throughput is t 3,i ’ It is the ratio of the aggregation queue throughput to the system throughput.
[0095] Step S232: Input the current system throughput, input queue throughput, output queue throughput and aggregation queue throughput, and calculate the input queue throughput ratio k1 respectively. ’ , output queue throughput ratio k2 ’ And the aggregation queue throughput ratio k3 ’ , and get the current throughput ratio K = [k1 ’ , k2 ’ , k3 ’ ], input queue throughput ratio k1 ’ is the ratio of input queue throughput to system throughput, and the output queue throughput is k2 ’ is the ratio of the output queue throughput to the system throughput, and the aggregation queue throughput ratio k3 ’ It is the ratio of the aggregation queue throughput to the system throughput.
[0096] Step S233: Input K ’ , T ’ , calculate K by ’ With T ’ The Euclidean distance between:
[0097] ;
[0098] Step S234: input D(K',T') into the queue throughput health calculation formula, which is as follows:
[0099] ;
[0100] Where C is the maximum acceptable Euclidean distance, and D(K',T') is the Euclidean distance between K' and T'. In this example, C is 0.027. When D(K',T') is 0.25, the queue throughput health is calculated to be 90.7.
[0101] The health of the queue operation status. In this solution, the queue occupancy rate is used to describe the queue operation status. The queues include input queues, output queues, aggregation queues, and transaction queues. The specific implementation steps are as follows:
[0102] Step S241: Collect historical data, analyze the relationship between queue occupancy and system performance, find the point near which performance drops significantly, determine the threshold, write a script to simulate the process of the queue gradually filling up, record the queue environment indicators and system performance data after each data is written, draw a line graph of the system performance data changing over time, and use the queue occupancy corresponding to the sudden change in system performance data as the preset queue environment indicator threshold.
[0103] Specifically, in step S241, a script is written to simulate the process of the queue gradually filling up. After each data is written, the system performance data is recorded for a period of time, and a line graph of the system performance over time is drawn to observe the trend of changes. If there is a large fluctuation, the queue occupancy rate at this time is set as the threshold.
[0104] Step S242: Input the current queue occupancy rate. If it is greater than the set threshold, use the linear formula to calculate the queue operation health. Otherwise, go to step S243. The linear calculation formula is as follows: ;
[0105] Where α is the cutoff value for determining whether a component is in a healthy state, x is the current queue occupancy, and threshold is the set threshold. In this example, α is 60 and threshold is 0.8. When the queue occupancy is 0.5, the queue health is calculated to be 76.
[0106] Step S243: If the queue occupancy rate is greater than the set threshold, an exponential formula is used to calculate the health of the queue operation status.
[0107] The index calculation formula is as follows:
[0108] ;
[0109] Among them, α is the cutoff value that determines whether the component is in a healthy state, x is the current queue occupancy, threshold is the set threshold, and slope is the steepness. In this example, threshold is 0.8 and slope is 20. When the queue occupancy is 0.8, the queue running health is calculated to be 60 by substituting it into the formula.
[0110] Step S244: Calculate the overall queue status health. Perform a weighted sum of the running status health of the input queue, output queue, aggregation queue, and transaction queue. The specific calculation formula is as follows:
[0111] ;
[0112] Among them, h 11 Indicates the health of the input queue operation status, h 12 Indicates the health of the output queue operation status, h 13 Indicates the health of the aggregation queue operation status, h 14 Indicates the health of the transaction queue operation status, w 11 , w 12 , w 13 , w 14 are the corresponding weight coefficients. In this embodiment, w 11 ,w 12 , w 13 , w 14 are 0.3, 0.3, 0.3, and 0.1 respectively. 11 ,h 12 ,h 13 ,h 14 When the values are 80, 70, 80, and 90 respectively, the overall queue health is 78 when substituted into the formula.
[0113] The health alarm design solution has the following specific implementation steps:
[0114] Calculate the system health by weighted sum of the health of the above indicators.
[0115] ;
[0116] Among them, h 21 Indicates the health of the queue environment, h 22Indicates the health of the message middleware, h 23 Indicates the health of the queue operation status, h 24 Indicates the health of the queue throughput, w 21 , w 22 , w 23 , w 24 are the corresponding weight coefficients. In this example, w 21 , w 22 , w 23 , w 24 are set to 0.1, 0.2, 0.4, and 0.3 respectively. 21 ,h 22 ,h 23 ,h 24 When the values are 90, 80, 90, and 70 respectively, the overall health of the system is 82.
[0117] Alarm. Alarms are divided into two parts: overall alarm and scenario alarm. Overall alarm: When the current health of the system is lower than the set threshold, an alarm is generated. Scenario alarm: When some sub-health is lower than the threshold, an alarm is generated. For example, when the queue throughput health is lower than the threshold, an alarm is generated.
Claims
1. A multi-queue system health monitoring method based on multiple indicators, characterized in that: The method comprises: Step S1: Fill missing values of the original data collected from the multi-queue system by linear interpolation; Step S2: Evaluate the health of key indicators in the multi-queue system and output the evaluation results. The key indicators include the health of the queue environment. The health of the queue environment is calculated in the following way: Step S211: Generate input data under various load conditions through scripts and input them into the system, record queue environment indicators and system performance data, draw a line graph of system performance data changing over time, and use the queue environment indicator corresponding to a sudden change in system performance data as a preset queue environment indicator threshold; Step S212: input the current queue environment index. If the current queue environment index is less than the preset queue environment index threshold, the queue environment health is calculated using the queue environment health linear formula: , Among them, α is the cutoff value that determines whether the component is in a healthy state, x is the current queue environment indicator value, threshold is the preset queue environment indicator threshold, and 100 is the maximum value of the queue environment health; Step S213: If the current queue environment index is greater than or equal to the preset queue environment index threshold, the queue environment health index formula is used to calculate the queue environment health index: , Among them, α is the cutoff value that determines whether the component is in a healthy state, x is the current queue environment indicator value, threshold is the preset queue environment indicator threshold, and slope is the set steepness, which is used to reflect the sensitivity of changes in queue environment health to changes in queue environment indicator values; Step S3: Calculate the comprehensive health according to the evaluation result outputted in step S2, and generate warning information; said step S3 also includes: calculating the comprehensive health by the following method: , Among them, h 21 Indicates the health of the queue environment, h 22 Indicates the health of the message middleware, h 23 Indicates the health of the queue operation status, h 24 Indicates the health of the queue throughput, w 21 , w 22 , w 23 , w 24 are the corresponding weight coefficients respectively.
2. The multi-queue system health monitoring method based on multiple indicators according to claim 1 is characterized in that: The step S2 includes: the key indicator includes the health of the message middleware, and the health of the message middleware is calculated by the following method: Step S221: Monitor and record the sending, persistence and receiving throughput rates of the message middleware at n moments when the system is normal, and calculate the ratio t of the sending throughput rate to the receiving throughput rate at each moment i. 1,i , the ratio of the persistence throughput rate to the receiving throughput rate t 2,i , get the throughput ratio T at that moment i = [t 1,i , t 2,i ], the throughput ratio T for all time i Take the average and get the ideal throughput ratio T = [t1, t2]; Step S222: input the current sending throughput rate, persistence throughput rate and receiving throughput rate of the message middleware, calculate the ratio k1 of the sending throughput rate to the receiving throughput rate, and the ratio k2 of the persistence throughput rate to the receiving throughput rate, and obtain the current throughput ratio K = [k1, k2]; Step S223: Input K and T, and calculate the Euclidean distance between K and T in the following way: ; Step S224: Substitute D(K, T) into the message middleware health calculation formula, which is as follows: , Where C is the maximum acceptable Euclidean distance.
3. The multi-queue system health monitoring method based on multiple indicators according to claim 2 is characterized in that: The step S2 includes: the key indicator includes the health of the message queue, the health of the message queue includes the health of the queue throughput rate and the health of the queue operation status, and the queue throughput rate health is calculated by the following method: Step S231: Monitor and record the system throughput, input queue throughput, output queue throughput and aggregation queue throughput at n moments when the system is normal, and calculate the input queue throughput ratio t at each moment i. 1,i ’ , output queue throughput ratio t 2,i ’ And the aggregation queue throughput ratio t 3,i ’ , get the throughput ratio t at that moment i ’ = [t 1,i ’ ,t 2,i ’ , t 3,i ’ ], the throughput ratio T for all time i ’ Take the average and get the ideal throughput ratio T ’ = [t1 ’ , t2 ’ , t3 ’ ], input queue throughput ratio t 1,i ’ is the ratio of input queue throughput to system throughput, and the output queue throughput is t 2,i ’ is the ratio of the output queue throughput to the system throughput, and the aggregation queue throughput is t 3,i ’ is the ratio of the aggregation queue throughput to the system throughput; Step S232: Input the current system throughput, input queue throughput, output queue throughput and aggregation queue throughput, and calculate the input queue throughput ratio k1 respectively. ’ , output queue throughput ratio k2 ’ And the aggregation queue throughput ratio k3 ’ , and get the current throughput ratio K = [k1 ’ , k2 ’ , k3 ’ ], input queue throughput ratio k1 ’ is the ratio of input queue throughput to system throughput, and the output queue throughput is k2 ’ is the ratio of the output queue throughput to the system throughput, and the aggregation queue throughput ratio k3 ’ is the ratio of the aggregation queue throughput to the system throughput; Step S233: Input K ’ , T ’ , calculate K by ’ With T ’ The Euclidean distance between: ; Step S234: input D(K',T') into the queue throughput health calculation formula, which is as follows: , Where C is the maximum acceptable Euclidean distance.
4. The multi-queue system health monitoring method based on multiple indicators according to claim 3 is characterized in that: The step S2 includes: the queues include an input queue, an output queue, an aggregation queue and a transaction queue, and the health of the queue operation status is calculated in the following manner: Step S241: Write a script to simulate the process of the queue gradually filling up. After each data is written, record the queue environment index and system performance data, draw a line graph of the system performance data changing over time, and use the queue occupancy rate corresponding to the sudden change in the system performance data as the preset queue environment index threshold; Step S242: If the current queue occupancy is less than the preset queue occupancy threshold, the queue operation status health is calculated using the queue operation status health linear formula: , Where α is the cutoff value for determining whether a component is in a healthy state, x is the current queue occupancy rate, and threshold is the preset queue occupancy rate threshold; Step S243: If the current queue occupancy is greater than or equal to the preset queue occupancy threshold, the queue operation status health index formula is used to calculate the queue operation status health index: , Where α is the cutoff value for determining whether a component is in a healthy state, x is the current queue occupancy rate, threshold is the preset queue occupancy rate threshold, and slope is the steepness; Step S244: weighted sum of the operational status health of the input queue, output queue, aggregation queue, and transaction queue is performed in the following manner: , Among them, h 11 Indicates the health of the input queue operation status, h 12 Indicates the health of the output queue operation status, h 13 Indicates the health of the aggregation queue operation status, h 14 Indicates the health of the transaction queue operation status, w 11 , w 12 , w 13 , w 14 are the corresponding weight coefficients respectively.
5. The multi-queue system health monitoring method based on multiple indicators according to claim 4 is characterized in that: The step S3 comprises: generating warning information by the following method: Alarms include overall alarms and scenario alarms. When the health of either the queue environment health or the message middleware health is lower than the preset overall alarm threshold, an overall alarm is generated; when the health of either the queue operation status health or the queue throughput health is lower than the preset scenario alarm threshold, a scenario alarm is generated.
Citation Information
Patent Citations
Health inspection method and device for message queue, equipment and medium
CN116136818A
Equipment health degree assessment method and device and monitoring system
CN118428926A