Container cluster monitoring alarm system and method
By designing a container cluster monitoring and alarm system, including node performance monitoring, microservice interactive monitoring, dynamic threshold adaptation, comprehensive analysis and alarm management modules, the problems of cross-node performance bottleneck positioning, microservice performance impact capture and alarm threshold static are solved, and efficient and accurate performance monitoring and alarm are achieved.
Patent Information
- Application Number
- CN202510217151.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Existing container cluster monitoring and alarm systems are difficult to accurately locate cross-node performance bottlenecks, and cannot fully capture the performance impact of interactions between microservices, and static alarm thresholds lead to false alarms and missed alarms.
A monitoring and alarm system for container clusters is designed, including node performance monitoring module, microservice interactive monitoring module, dynamic threshold adaptive module, comprehensive analysis module and alarm management module. By monitoring cross-node resource access latency in real time, analyzing microservice call links, dynamically adjusting thresholds, and calculating comprehensive performance health index to generate detailed alarm information.
It effectively solves the precise positioning problem of cross-node performance bottlenecks, comprehensively captures the performance impact of interactions between microservices, and reduces the alarm false alarm and omission rate through dynamic threshold adjustment, and improves the real-time and accuracy of the monitoring system.
Smart Images

Figure CN120075037A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of performance monitoring, and particularly to a monitoring and alarming system and method for a container cluster. Background Art
[0002] The monitoring and alarming system for a container cluster originated from the rapid development of container technology and distributed computing. Containerization technology provides a lightweight isolated running environment for applications, and cluster management tools enable large-scale deployment and efficient scheduling of containers. However, with the expansion of the container cluster scale, the complexity of the system running state has increased sharply, and traditional monitoring methods are difficult to meet the requirements of real-time and refinement. Modern monitoring and alarming systems have gradually integrated intelligent analysis, visual monitoring, and dynamic alarming mechanisms, effectively improving the resource utilization rate and fault response speed of the container cluster, and laying a foundation for the wide application of cloud computing and microservice architectures.
[0003] However, in the prior art, the monitoring and alarming system for a container cluster often has the following technical drawbacks:
[0004] 1. Difficult to accurately locate cross-node performance bottlenecks: Performance problems in a container cluster are often caused by resource contention between multiple nodes; existing monitoring systems lack effective correlation analysis in cross-node resource access latency and bottleneck location, resulting in low diagnostic efficiency. For example, in a multi-tenant environment, contention for shared network bandwidth may be misinterpreted as a single-node resource problem.
[0005] 2. The performance impact of interactions between microservices is not comprehensively captured: In a microservice architecture, frequent interactions between containers will generate complex performance bottlenecks, and most existing systems focus on resource monitoring of a single container and cannot effectively evaluate the overall performance of the microservice call chain, such as cumulative latency or link interruption caused by multiple calls.
[0006] 3. Static alarm thresholds lead to false alarms and missed alarms: Many monitoring systems use fixed alarm thresholds, while the resource usage of containers fluctuates violently with the load and it is difficult to adapt to dynamic changes; for example, a short-term resource peak during the peak period may trigger a false alarm, and missed alarms may occur due to too high a threshold when the load is gradually increasing. Summary of the Invention
[0007] In view of the deficiencies of the prior art, the present invention provides a monitoring and alarming system and method for a container cluster, which solves the technical drawbacks mentioned in the background art.
[0008] To achieve the above object, the present invention is realized through the following technical solutions: A monitoring and alarming system for a container cluster includes a node performance monitoring module, a microservice interaction monitoring module, a dynamic threshold adaptive module, a comprehensive analysis module, and an alarm management module;
[0009] The node performance monitoring module is used to pre-obtain the electronic topology map of node distribution within the container cluster, monitor the cross-node resource access latency in real time, and finally calculate and evaluate the cross-node contention coefficient Kcz to generate the corresponding bottleneck warning signal;
[0010] The microservice interaction monitoring module is used to deeply analyze the microservice call link in the container cluster after receiving the bottleneck warning signal, calculate and evaluate the microservice performance interference index Gwz, and finally generate the microservice performance warning signal;
[0011] The dynamic threshold adaptive module collects historical load change data and real-time resource usage data for analysis, constructs the dynamic threshold adjustment factor Ztd; based on the dynamic threshold adjustment factor Ztd, it dynamically adjusts the cross-node contention threshold Qc and the microservice performance interference threshold Kw;
[0012] The comprehensive analysis module is used to construct and analyze the container performance diagnosis model for the cross-node contention coefficient Kcz, the microservice performance interference index Gwz, and the dynamic threshold adjustment factor Ztd, then calculate and evaluate the comprehensive performance health index Phz, and issue the comprehensive performance health warning signal;
[0013] The alarm management module is used to calculate and evaluate the comprehensive performance health deviation value Apj based on the evaluation content of the comprehensive performance health index Phz; finally, it performs hierarchical management on the comprehensive performance health warning signal and generates detailed alarm information.
[0014] Preferably, the node performance monitoring module includes a node performance quantification unit and a node performance evaluation unit;
[0015] The node performance quantification unit first collects and constructs a latency data set, and based on the latency data set, calculates the cross-node contention coefficient Kcz. The specific calculation formula is as follows:
[0016]
[0017] In the formula, Tnw represents the network transmission latency duration in the latency data set, Tio represents the I / O operation latency duration in the latency data set, and Trp represents the RPC call latency duration in the latency data set.
[0018] Preferably, the node performance evaluation unit is used to preset the cross-node contention threshold Qc, evaluate the cross-node contention coefficient Kcz, judge the bottleneck intensity of resource contention in the current container cluster, and finally generate the corresponding bottleneck warning signal; the specific evaluation content is as follows:
[0019] If the cross-node contention coefficient Kcz ≤ the cross-node contention threshold Qc, it is considered that the current resource contention intensity is normal, and no bottleneck warning signal is generated at this time;
[0020] If the cross - node contention coefficient \(K_{cz}\gt\) the cross - node contention threshold \(Q_{c}\), it is considered that the current resource contention intensity is abnormal and there is a bottleneck problem. At this time, an abnormal bottleneck warning signal is generated.
[0021] Preferably, the microservice interaction monitoring module includes a link performance calculation unit and an interference evaluation unit;
[0022] The link performance calculation unit collects data related to interaction delay, cumulative delay of multiple calls, and link interruption probability of the microservice call link, and then constructs a microservice link performance index dataset; based on the microservice link performance index dataset, the single - request response duration \(T_{re}\), average call link delay duration \(T_{av}\), call retry times \(R_{re}\), and link interruption rate \(P_{br}\) in the microservice link performance index dataset are extracted. After dimensionless processing, the microservice performance interference index \(G_{wz}\) is calculated through the following formula:
[0023]
[0024] Preferably, the interference evaluation unit is used to preset the microservice performance interference threshold \(K_{w}\), and compare and evaluate the microservice performance interference index \(G_{wz}\), judge the health status of the microservice call link performance and generate a microservice performance warning signal; the specific evaluation content is as follows:
[0025] If the microservice performance interference index \(G_{wz}\leq\) the microservice performance interference threshold \(K_{w}\), it means that the call chain performance is in a healthy state and no warning is required;
[0026] If the microservice performance interference index \(G_{wz}\gt\) the microservice performance interference threshold \(K_{w}\), it means that the call chain performance is interfered and there are potential performance problems. At this time, a microservice performance warning signal is issued.
[0027] Preferably, the dynamic threshold adaptive module analyzes historical load change data and real - time resource usage data, obtains the short - term fluctuation value \(S_{bd}\), long - term trend value \(C_{qx}\), resource utilization peak value \(R_{up}\), and current stability coefficient \(S_{tb}\). After dimensionless processing, the dynamic threshold adjustment factor \(Z_{td}\) is calculated through the following formula:
[0028]
[0029] Based on the dynamic threshold adjustment factor \(Z_{td}\), the cross - node contention threshold \(Q_{c}\) and the microservice performance interference threshold \(K_{w}\) are linearly adjusted. The specific adjustment formulas are as follows:
[0030] \(Q_{c}' = Q_{c}\times(1 + Z_{td})\);
[0031] \(K_{w}' = K_{w}\times(1 + Z_{td})\);
[0032] Wherein, Qc' represents the adjusted cross - node contention threshold, and Kw' represents the adjusted microservice performance interference threshold;
[0033] When the dynamic threshold adjustment factor Ztd > 0, increase the adjusted cross - node contention threshold Qc and the microservice performance interference threshold Kw by 10%;
[0034] When the dynamic threshold adjustment factor Ztd ≤ 0, decrease the adjusted cross - node contention threshold Qc and the microservice performance interference threshold Kw by 10%;
[0035] Finally, apply the adjusted cross - node contention threshold Qc and the microservice performance interference threshold Kw to the node performance monitoring module and the microservice interaction monitoring module.
[0036] Preferably, the comprehensive analysis module includes a performance index calculation unit and a performance alarm evaluation unit:
[0037] The performance index calculation unit conducts a joint analysis on the cross - node contention coefficient Kcz, the microservice performance interference index Gwz, and the dynamic threshold adjustment factor Ztd based on the container performance diagnosis model, comprehensively extracts performance characteristic parameters, including the average resource contention intensity Rci, the microservice call link stability factor Msf, and the dynamic threshold adjustment sensitivity Zsa, and finally calculates and obtains the comprehensive performance health index Phz by combining the following formula:
[0038]
[0039] Preferably, the performance alarm evaluation unit is used to preset the comprehensive performance health threshold Zj, and then evaluate the comprehensive performance health index Phz. The specific evaluation content is as follows:
[0040] If the comprehensive performance health index Phz ≥ the comprehensive performance health threshold Zj, it indicates that the container cluster performance is in a normal state and no warning is required;
[0041] If the comprehensive performance health index Phz < the comprehensive performance health threshold Zj, it indicates that the container cluster performance is in an abnormal state, there are performance problems or bottlenecks, and at this time, a comprehensive performance health alarm signal is issued.
[0042] Preferably, the alarm management module is used to calculate and evaluate the comprehensive performance health deviation value Apj; when the cluster performance is in an abnormal state, the comprehensive performance health deviation value Apj is calculated and obtained through the following formula:
[0043]
[0044] By presetting the first performance health deviation threshold W1 and the second performance health deviation threshold W2, the comprehensive performance health deviation value Apj is evaluated, and at the same time, the comprehensive performance health warning signal is hierarchically managed, and finally detailed warning information is generated; and the first performance health deviation threshold W1 is greater than the second performance health deviation threshold W2, and the specific content is as follows:
[0045] Level 1 warning signal: When Apj ≤ W2, it means that the performance fluctuation is slight and no emergency intervention is required;
[0046] Level 2 warning signal: When W2 < Apj ≤ W1, it means that the performance fluctuation is medium, and it is prompted to optimize the performance management strategy;
[0047] Level 3 warning signal: When Apj > W1, it means that the performance fluctuation seriously deviates from the healthy state. At this time, immediate intervention measures are taken, including adjusting the resource allocation strategy, adjusting the microservice call chain, and adjusting the dynamic threshold parameters.
[0048] A monitoring and warning method for a container cluster includes the following steps:
[0049] Step 1: Obtain the node distribution electronic topology map in the container cluster in advance, and monitor the cross-node resource access delay in real time. Finally, calculate and evaluate the cross-node contention coefficient Kcz, and generate the corresponding bottleneck warning signal;
[0050] Step 2: After receiving the bottleneck warning signal, deeply analyze the microservice call link in the container cluster, calculate and evaluate the microservice performance interference index Gwz, and finally generate the microservice performance warning signal;
[0051] Step 3: Collect historical load change data and real-time resource usage data and analyze them to construct the dynamic threshold adjustment factor Ztd; based on the dynamic threshold adjustment factor Ztd, dynamically adjust the cross-node contention threshold Qc and the microservice performance interference threshold Kw;
[0052] Step 4: Construct and analyze the cross-node contention coefficient Kcz, the microservice performance interference index Gwz, and the dynamic threshold adjustment factor Ztd based on the container performance diagnosis model. Then calculate and evaluate the comprehensive performance health index Phz, and issue the comprehensive performance health warning signal;
[0053] Step 5: Calculate and evaluate the comprehensive performance health deviation value Apj according to the evaluation content of the comprehensive performance health index Phz; finally, hierarchically manage the comprehensive performance health warning signal, and finally generate detailed warning information.
[0054] The present invention provides a monitoring and warning system and method for a container cluster. It has the following beneficial effects:
[0055] (1) The monitoring and warning system and method for a container cluster effectively solve the problem that it is difficult to accurately locate the cross-node performance bottleneck. Through the node performance monitoring module, the network transmission delay duration Tnw, I / O operation delay duration Tio, and RPC call delay duration Trp in the container cluster are monitored and collected in real time, a delay data set is constructed, and the cross-node contention coefficient Kcz is calculated. Using the node performance evaluation unit, the cross-node contention coefficient Kcz is compared and analyzed with the preset cross-node contention threshold Qc, and a bottleneck warning signal is generated in a timely manner. Through the performance index calculation unit in the comprehensive analysis module, the average resource contention intensity Rci is used as one of the performance characteristic parameters, and the comprehensive performance health index Phz is calculated by combining the microservice performance interference index Gwz and the dynamic threshold adjustment factor Ztd, realizing the accurate location and rapid response of cross-node resource contention.
[0056] (2) The monitoring and warning system and method for a container cluster comprehensively capture the performance impact of interactions between microservices. Through the link performance calculation unit of the microservice interaction monitoring module, the single-request response duration Tre, average call link delay duration Tav, call retry number Rre, and link interruption rate Pbr in the microservice link performance index data set are extracted, a microservice performance interference index Gwz is constructed, and it is compared and evaluated with the preset microservice performance interference threshold Kw. When the microservice performance interference index Gwz is greater than the microservice performance interference threshold Kw, a microservice performance warning signal is generated, indicating that there may be a performance problem in the call chain. Combining the microservice call chain stability factor Msf in the comprehensive analysis module, the microservice interaction performance is quantitatively analyzed to ensure that the overall performance of the call chain is comprehensively evaluated and omissions caused by single-container monitoring limitations are avoided.
[0057] (3) The monitoring and warning system and method for a container cluster effectively solve the problems of false alarms and missed alarms caused by static alarm thresholds. Through the dynamic threshold adaptive module, historical load change data and real-time resource usage data are analyzed, the short-term fluctuation value Sbd, long-term trend value Cqx, resource utilization peak Rup, and current stability coefficient Stb are extracted, the dynamic threshold adjustment factor Ztd is calculated, and the cross-node contention threshold Qc and the microservice performance interference threshold Kw are dynamically adjusted. The adjusted thresholds are applied to the contention evaluation and interference evaluation processes to improve the adaptability of the monitoring system to load fluctuations. Combining the alarm management module, the alarm events are classified and managed through the comprehensive performance health index Phz and the comprehensive performance health deviation value Apj, and mild, moderate, and severe alarm signals are generated to ensure that the monitoring system has higher accuracy and sensitivity. Description of the Drawings
[0058] Figure 1 It is a schematic diagram of the framework structure of a monitoring and warning system for a container cluster according to the present invention;
[0059] Figure 2 This is a schematic diagram of the step flow of a method for monitoring and alarming a container cluster according to the present invention. Specific embodiments
[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0061] Embodiment 1
[0062] Please refer to Figure 1 , a monitoring and alarming system for a container cluster, including a node performance monitoring module, a microservice interaction monitoring module, a dynamic threshold adaptive module, a comprehensive analysis module, and an alarm management module;
[0063] The node performance monitoring module is used to pre-obtain the electronic topology map of node distribution in the container cluster, and perform real-time monitoring on the cross-node resource access delay, and finally calculate and evaluate the cross-node contention coefficient Kcz, and generate a corresponding bottleneck warning signal;
[0064] The microservice interaction monitoring module is used to deeply analyze the microservice call link in the container cluster after receiving the bottleneck warning signal, calculate and evaluate the microservice performance interference index Gwz, and finally generate a microservice performance warning signal;
[0065] The dynamic threshold adaptive module analyzes by collecting historical load change data and real-time resource usage data, and constructs a dynamic threshold adjustment factor Ztd; based on the dynamic threshold adjustment factor Ztd, dynamically adjusts the cross-node contention threshold Qc and the microservice performance interference threshold Kw;
[0066] The comprehensive analysis module is used to construct and analyze the cross-node contention coefficient Kcz, the microservice performance interference index Gwz, and the dynamic threshold adjustment factor Ztd based on the container performance diagnosis model, then calculate and evaluate the comprehensive performance health index Phz, and issue a comprehensive performance health warning signal;
[0067] The alarm management module is used to calculate and evaluate the comprehensive performance health deviation value Apj according to the evaluation content of the comprehensive performance health index Phz; finally, classify and manage the comprehensive performance health warning signal, and finally generate detailed alarm information.
[0068] In this embodiment, the node performance monitoring module can calculate the cross-node contention coefficient \(K_{cz}\) by collecting the network transmission delay duration \(T_{nw}\), the I / O operation delay duration \(T_{io}\), and the RPC call delay duration \(T_{rp}\), and realize the accurate positioning of cross-node resource contention and bottleneck warning by combining the cross-node contention threshold \(Q_{c}\); the microservice interaction monitoring module can construct a microservice link performance index dataset by extracting the single-request response duration \(T_{re}\), the average call link delay duration \(T_{av}\), the call retry times \(R_{re}\), and the link interruption rate \(P_{br}\), calculate the microservice performance interference index \(G_{wz}\), and compare and evaluate it with the microservice performance interference threshold \(K_{w}\) to realize the accurate monitoring of the microservice call chain performance; the dynamic threshold adaptive module analyzes the historical load change data and the real-time resource usage data, extracts the short-term fluctuation value \(S_{bd}\), the long-term trend value \(C_{qx}\), the resource utilization peak value \(R_{up}\), and the current stability coefficient \(S_{tb}\), calculates the dynamic threshold adjustment factor \(Z_{td}\), and dynamically adjusts the cross-node contention threshold \(Q_{c}\) and the microservice performance interference threshold \(K_{w}\) to improve the system's adaptability to load fluctuations and the accuracy of alarms; the comprehensive analysis module calculates the comprehensive performance health index \(P_{hz}\) by combining the average resource contention intensity \(R_{ci}\), the microservice call chain stability factor \(M_{sf}\), and the dynamic threshold adjustment sensitivity \(Z_{sa}\) to realize the accurate evaluation of the overall performance status of the container cluster; the alarm management module calculates the comprehensive performance health deviation value \(A_{pj}\) based on the evaluation result of the comprehensive performance health index \(P_{hz}\), and classifies and manages the alarm events by combining the first performance health deviation threshold \(W_{1}\) and the second performance health deviation threshold \(W_{2}\), generates detailed alarm information, and effectively improves the traceability and processing efficiency of performance problems.
[0069] Embodiment 2
[0070] The node performance monitoring module includes a node performance quantification unit and a node performance evaluation unit;
[0071] The node performance quantification unit first collects and constructs a delay dataset, and based on the delay dataset, calculates the cross-node contention coefficient \(K_{cz}\). The specific calculation formula is as follows:
[0072]
[0073] In the formula, \(T_{nw}\) represents the network transmission delay duration in the delay dataset, \(T_{io}\) represents the I / O operation delay duration in the delay dataset, and \(T_{rp}\) represents the RPC call delay duration in the delay dataset.
[0074] The node performance evaluation unit is used to preset the cross-node contention threshold \(Q_{c}\), evaluate the cross-node contention coefficient \(K_{cz}\), judge the bottleneck intensity of resource contention in the current container cluster, and finally generate a corresponding bottleneck warning signal. The specific evaluation content is as follows:
[0075] If the cross - node contention coefficient \(K_{cz}\leq\) the cross - node contention threshold \(Q_{c}\), it is considered that the current resource contention intensity is normal, and no bottleneck warning signal is generated at this time;
[0076] If the cross - node contention coefficient \(K_{cz}> \) the cross - node contention threshold \(Q_{c}\), it is considered that the current resource contention intensity is abnormal, there is a bottleneck problem, and an abnormal bottleneck warning signal is generated at this time.
[0077] The microservice interaction monitoring module includes a link performance calculation unit and an interference evaluation unit;
[0078] The link performance calculation unit collects data related to interaction delay, cumulative delay of multiple calls, and link interruption probability of the microservice call link, and then constructs a microservice link performance index dataset; based on the microservice link performance index dataset, the single - request response duration \(T_{re}\), average call link delay duration \(T_{av}\), call retry times \(R_{re}\), and link interruption rate \(P_{br}\) in the microservice link performance index dataset are extracted, and after dimensionless processing, the microservice performance interference index \(G_{wz}\) is calculated through the following formula:
[0079]
[0080] The interference evaluation unit is used to preset the microservice performance interference threshold \(K_{w}\), and compare and evaluate the microservice performance interference index \(G_{wz}\), judge the health status of the microservice call link performance and generate a microservice performance warning signal; the specific evaluation content is as follows:
[0081] If the microservice performance interference index \(G_{wz}\leq\) the microservice performance interference threshold \(K_{w}\), it means that the call chain performance is in a healthy state and no warning is required;
[0082] If the microservice performance interference index \(G_{wz}> \) the microservice performance interference threshold \(K_{w}\), it means that the call chain performance is interfered and there are potential performance problems, and a microservice performance warning signal is issued at this time.
[0083] The dynamic threshold adaptive module analyzes historical load change data and real - time resource usage data to obtain the short - term fluctuation value \(S_{bd}\), long - term trend value \(C_{qx}\), resource utilization peak value \(R_{up}\), and current stability coefficient \(St_{b}\), and after dimensionless processing, uses the following formula to calculate the dynamic threshold adjustment factor \(Z_{td}\):
[0084]
[0085] Based on the dynamic threshold adjustment factor \(Z_{td}\), linearly adjust the cross - node contention threshold \(Q_{c}\) and the microservice performance interference threshold \(K_{w}\), and the specific adjustment formula is as follows:
[0086] \(Q_{c}'=Q_{c}\times(1 + Z_{td})\);
[0087] Kw' = Kw × (1 + Ztd);
[0088] In the formula, Qc' represents the adjusted cross - node contention threshold, and Kw' represents the adjusted microservice performance interference threshold;
[0089] When the dynamic threshold adjustment factor Ztd > 0, the cross - node contention threshold Qc and the microservice performance interference threshold Kw are increased by 10%;
[0090] When the dynamic threshold adjustment factor Ztd ≤ 0, the cross - node contention threshold Qc and the microservice performance interference threshold Kw are decreased by 10%;
[0091] Finally, the adjusted cross - node contention threshold Qc and the microservice performance interference threshold Kw are applied to the node performance monitoring module and the microservice interaction monitoring module.
[0092] In this embodiment, through the collaborative work of multiple modules, the comprehensive monitoring and accurate warning of the container cluster performance are realized; the node performance monitoring module ensures that resource contention problems can be accurately located by generating bottleneck warning signals; among them, the network transmission delay duration Tnw represents the delay time of data transmission between container nodes in the network, reflecting the stability and transmission efficiency of the network link; the I / O operation delay duration Tio refers to the input / output operation delay when accessing shared storage or distributed file systems across nodes, reflecting the disk read - write speed and the response ability of the storage system; the RPC call delay duration Trp represents the response delay when microservices communicate through remote procedure call RPC, reflecting the call efficiency and interaction performance between services; the cross - node contention coefficient Kcz is used to quantify the intensity of resource contention between nodes; the higher the value, the more serious the cross - node resource competition;
[0093] The microservice interaction monitoring module realizes the dynamic performance monitoring of the call link by generating microservice performance warning signals; among them, the single - request response duration Tre is the total response time from the start to the completion of a single microservice request, reflecting the immediate processing ability of the microservice; the average call link delay duration Tav represents the average response delay of each node in the microservice call chain, reflecting the operation efficiency of the entire call chain; the call retry number Rre refers to the number of service retries due to request failure or timeout, reflecting the reliability and stability of the service; the link interruption rate Pbr represents the probability that requests in the microservice call chain are interrupted or failed, directly reflecting the reliability of service interaction; the microservice performance interference index Gwz is used to evaluate the overall health status of the microservice call chain, and the higher the value of the microservice performance interference index Gwz, the more serious the performance interference in the call chain;
[0094] The dynamic threshold adaptation module calculates the dynamic threshold adjustment factor Ztd by analyzing historical load change data and real-time resource usage data, and dynamically adjusts the cross-node contention threshold Qc and the microservice performance interference threshold Kw to enhance the system's adaptability to dynamic load changes. Among them, the short-term fluctuation value Sbd is extracted to represent the fluctuation amplitude of the system resource usage in a short period of time, reflecting the sensitivity of the system to sudden load changes; the long-term trend value Cqx describes the trend of the system resource usage over time, reflecting the long-term load change trend of the system; the resource utilization peak Rup represents the highest utilization rate of the system resources (such as CPU, memory, network, etc.) during the monitoring period, reflecting the load limit state; the current stability coefficient Stb is used to measure the running stability of the system under the current load, considering indicators such as the system failure rate and error rate.
[0095] The acquisition of each lower-level parameter can quantify different performance dimensions of the cluster operation. The calculation and evaluation of each parameter further classify and quantify the performance problems, and finally realize the real-time monitoring, alarm grading and precise optimization of the system performance.
[0096] Embodiment 3
[0097] The comprehensive analysis module includes a performance index calculation unit and a performance alarm evaluation unit:
[0098] The performance index calculation unit conducts a joint analysis on the cross-node contention coefficient Kcz, the microservice performance interference index Gwz, and the dynamic threshold adjustment factor Ztd based on the container performance diagnosis model, and comprehensively extracts performance characteristic parameters, including the average resource contention intensity Rci, the microservice call link stability factor Msf, and the dynamic threshold adjustment sensitivity Zsa. Finally, the comprehensive performance health index Phz is calculated and obtained by combining the following formula:
[0099]
[0100] The performance alarm evaluation unit is used to preset the comprehensive performance health threshold Zj, and then evaluate the comprehensive performance health index Phz. The specific evaluation content is as follows:
[0101] If the comprehensive performance health index Phz ≥ the comprehensive performance health threshold Zj, it means that the container cluster performance is in a normal state and no warning is required.
[0102] If the comprehensive performance health index Phz < the comprehensive performance health threshold Zj, it means that the container cluster performance is in an abnormal state, there are performance problems or bottlenecks, and at this time, a comprehensive performance health alarm signal is issued.
[0103] The alarm management module is used to calculate and evaluate the comprehensive performance health deviation value Apj; when the cluster performance is in an abnormal state, the comprehensive performance health deviation value Apj is calculated and obtained by the following formula:
[0104]
[0105] By presetting the first performance health deviation threshold W1 and the second performance health deviation threshold W2, the comprehensive performance health deviation value Apj is evaluated, and at the same time, hierarchical management is carried out on the comprehensive performance health warning signal, and finally detailed warning information is generated; and the first performance health deviation threshold W1 is greater than the second performance health deviation threshold W2, and the specific content is as follows:
[0106] Level 1 warning signal: When Apj ≤ W2, it indicates that the performance fluctuation is slight and no emergency intervention is required;
[0107] Level 2 warning signal: When W2 < Apj ≤ W1, it indicates that the performance fluctuation is medium, and it is prompted to optimize the performance management strategy;
[0108] Level 3 warning signal: When Apj > W1, it indicates that the performance fluctuation seriously deviates from the healthy state. At this time, immediate intervention measures are taken, including adjusting the resource allocation strategy, adjusting the microservice call chain, and adjusting the dynamic threshold parameters.
[0109] In this embodiment, through the collaborative work of the comprehensive analysis module and the warning management module, the system performance monitoring and warning capabilities are comprehensively improved; in the comprehensive analysis module, the comprehensive performance health index Phz is calculated through a formula to realize the multi-dimensional quantification of the overall performance of the cluster; the performance warning evaluation unit judges whether there are performance problems or bottlenecks in the system by evaluating the health state of the comprehensive performance health index Phz, and generates warning signals in a timely manner; among them, the average resource contention intensity Rci represents the average level of cross-node resource contention within a period of time and is used to long-term monitor the resource competition situation; the microservice call chain stability factor Msf synthesizes the performance indicators of the microservice call chain and quantifies its stability; the dynamic threshold adjustment sensitivity Zsa describes the response speed and amplitude of the system to dynamic threshold adjustment and measures the ability of the system to adapt to performance fluctuations;
[0110] The warning management module calculates the comprehensive performance health deviation value Apj to quantify the deviation degree between the current performance state and the health threshold, and reflects the severity of the system performance deviating from the normal state; combined with the first performance health deviation threshold W1 and the second performance health deviation threshold W2, hierarchical management is carried out on the warning events. The level 1 warning indicates that the performance fluctuation is slight and no emergency intervention is required, the level 2 warning indicates that the performance fluctuation is medium and the performance management strategy needs to be optimized, and the level 3 warning indicates that the performance seriously deviates from the healthy state and measures such as adjusting the resource allocation strategy, optimizing the microservice call chain or adjusting the dynamic threshold parameters need to be taken immediately; the collection and calculation of each lower-level parameter respectively quantify the key dimensions of resource contention, call chain performance and system adaptability, providing a scientific basis for the accurate positioning and efficient optimization of performance problems.
[0111] Example 4
[0112] Please refer to Figure 2 , a monitoring and warning method for a container cluster, comprising the following steps:
[0113] Step 1: Pre-obtain the electronic topology map of node distribution in the container cluster, and perform real-time monitoring on the cross-node resource access latency. Finally, calculate and evaluate the cross-node contention coefficient Kcz, and generate the corresponding bottleneck warning signal;
[0114] Step 2: After receiving the bottleneck warning signal, deeply analyze the microservice call link in the container cluster, calculate and evaluate the microservice performance interference index Gwz, and finally generate the microservice performance warning signal;
[0115] Step 3: Collect historical load change data and real-time resource usage data and analyze them to construct a dynamic threshold adjustment factor Ztd; based on the dynamic threshold adjustment factor Ztd, dynamically adjust the cross-node contention threshold Qc and the microservice performance interference threshold Kw;
[0116] Step 4: Construct and analyze based on the container performance diagnosis model for the cross-node contention coefficient Kcz, the microservice performance interference index Gwz, and the dynamic threshold adjustment factor Ztd. Then calculate and evaluate the comprehensive performance health index Phz, and issue the comprehensive performance health warning signal;
[0117] Step 5: Calculate and evaluate the comprehensive performance health deviation value Apj based on the evaluation content of the comprehensive performance health index Phz; finally, perform hierarchical management on the comprehensive performance health warning signal, and finally generate detailed warning information.
[0118] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A container cluster monitoring and alarm system, characterized in that: It includes node performance monitoring module, microservice interaction monitoring module, dynamic threshold adaptive module, comprehensive analysis module and alarm management module; The node performance monitoring module is used to pre-acquire the node distribution electronic topology map in the container cluster, and monitor the cross-node resource access delay in real time, and finally calculate and evaluate the cross-node contention coefficient Kcz to generate a corresponding bottleneck warning signal; The microservice interaction monitoring module is used to perform in-depth analysis on the microservice call link in the container cluster after receiving the bottleneck warning signal, calculate and evaluate the microservice performance interference index Gwz, and finally generate a microservice performance warning signal; The dynamic threshold adaptive module collects and analyzes historical load change data and real-time resource usage data to construct a dynamic threshold adjustment factor Ztd; Based on the dynamic threshold adjustment factor Ztd, dynamically adjust the cross-node contention threshold Qc and the microservice performance interference threshold Kw; The comprehensive analysis module is used to construct and analyze the cross-node contention coefficient Kcz, the microservice performance interference index Gwz and the dynamic threshold adjustment factor Ztd based on the container performance diagnosis model, and then calculate and evaluate the comprehensive performance health index Phz, and issue a comprehensive performance health alarm signal; The alarm management module is used to calculate and evaluate the comprehensive performance health deviation value Apj based on the evaluation content of the comprehensive performance health index Phz; Finally, the comprehensive performance health warning signals are managed in a hierarchical manner and detailed warning information is generated.
2. According to claim 1, a container cluster monitoring and alarm system is characterized in that: The node performance monitoring module includes a node performance quantification unit and a node performance evaluation unit; The node performance quantification unit first collects and constructs a delay data set, and calculates the cross-node contention coefficient Kcz based on the delay data set. The specific calculation formula is as follows: Where Tnw represents the network transmission delay in the delay data set, Tio represents the I / O operation delay in the delay data set, and Trp represents the RPC call delay in the delay data set.
3. A container cluster monitoring and alarm system according to claim 2, characterized in that: The node performance evaluation unit is used to preset the cross-node contention threshold Qc, and evaluate the cross-node contention coefficient Kcz, determine the bottleneck intensity of resource contention in the current container cluster, and finally generate a corresponding bottleneck warning signal; the specific evaluation content is as follows: If the cross-node contention coefficient Kcz ≤ the cross-node contention threshold Qc, the current resource contention intensity is considered normal, and no bottleneck warning signal is generated; If the cross-node contention coefficient Kcz> the cross-node contention threshold Qc, it is considered that the current resource contention intensity is abnormal and there is a bottleneck problem, and an abnormal bottleneck warning signal is generated.
4. A container cluster monitoring and alarm system according to claim 3, characterized in that: The microservice interaction monitoring module includes a link performance calculation unit and an interference assessment unit; The link performance calculation unit collects the interaction delay related data of the microservice call link, the cumulative delay related data of multiple calls, and the link interruption probability related data, and then constructs a microservice link performance indicator data set; based on the microservice link performance indicator data set, extracts the single request response time Tre, the average call link delay time Tav, the number of call retries Rre, and the link interruption rate Pbr in the microservice link performance indicator data set, and after dimensionless processing, calculates the microservice performance interference index Gwz through the following formula:
5. A container cluster monitoring and alarm system according to claim 4, characterized in that: The interference evaluation unit is used to preset a microservice performance interference threshold Kw, and compare and evaluate the microservice performance interference index Gwz, determine the health status of the microservice call link performance and generate a microservice performance warning signal; The specific evaluation contents are as follows: If the microservice performance interference index Gwz ≤ the microservice performance interference threshold Kw, it means that the call chain performance is in a healthy state and no warning is needed; If the microservice performance interference index Gwz>microservice performance interference threshold Kw, it means that the call chain performance is disturbed and there is a potential performance problem. At this time, a microservice performance warning signal is issued.
6. A container cluster monitoring and alarm system according to claim 5, characterized in that: The dynamic threshold adaptive module analyzes historical load change data and real-time resource usage data to obtain the short-term fluctuation value Sbd, the long-term trend value Cqx, the resource utilization peak Rup and the current stability coefficient Stb, and performs dimensionless processing, and then calculates the dynamic threshold adjustment factor Ztd using the following formula: Based on the dynamic threshold adjustment factor Ztd, the cross-node contention threshold Qc and the microservice performance interference threshold Kw are linearly adjusted. The specific adjustment formula is as follows: Qc' = Qc × (1 + Ztd); Kw' = Kw × (1 + Ztd); Where Qc' represents the adjusted cross-node contention threshold, and Kw' represents the adjusted microservice performance interference threshold; When the dynamic threshold adjustment factor Ztd>0, the cross-node contention threshold Qc and the microservice performance interference threshold Kw are adjusted to increase by 10%; When the dynamic threshold adjustment factor Ztd≤0, the cross-node contention threshold Qc and the microservice performance interference threshold Kw are adjusted to be reduced by 10%; Finally, the adjusted cross-node contention threshold Qc and microservice performance interference threshold Kw are applied to the node performance monitoring module and the microservice interaction monitoring module.
7. A container cluster monitoring and alarm system according to claim 6, characterized in that: The comprehensive analysis module includes a performance indicator calculation unit and a performance alarm evaluation unit: The performance indicator calculation unit performs a joint analysis on the cross-node contention coefficient Kcz, the microservice performance interference index Gwz, and the dynamic threshold adjustment factor Ztd based on the container performance diagnosis model, and comprehensively extracts performance characteristic parameters, including the average resource contention intensity Rci, the microservice call link stability factor Msf, and the dynamic threshold adjustment sensitivity Zsa, and finally calculates the comprehensive performance health index Phz in combination with the following formula:
8. A container cluster monitoring and alarm system according to claim 7, characterized in that: The performance alarm evaluation unit is used to preset the comprehensive performance health threshold Zj, and then evaluate the comprehensive performance health index Phz. The specific evaluation content is as follows: If the comprehensive performance health index Phz ≥ the comprehensive performance health threshold Zj, it indicates that the performance of the container cluster is in a normal state and no warning is required; If the comprehensive performance health index Phz < the comprehensive performance health threshold Zj, it indicates that the performance of the container cluster is in an abnormal state, there are performance problems or bottlenecks, and at this time, a comprehensive performance health warning signal is issued.
9. A container cluster monitoring and alarm system according to claim 8, characterized in that: The alarm management module is used to calculate and evaluate the comprehensive performance health deviation value Apj; when the cluster performance is in an abnormal state, the comprehensive performance health deviation value Apj is calculated and obtained through the following formula: By presetting the first performance health deviation threshold W1 and the second performance health deviation threshold W2, the comprehensive performance health deviation value Apj is evaluated, and at the same time, the comprehensive performance health warning signal is hierarchically managed, and finally detailed warning information is generated; And the first performance health deviation threshold W1 is greater than the second performance health deviation threshold W2, and the specific content is as follows: Level 1 alarm signal: When Apj ≤ W2, it indicates that the performance fluctuation is slight and no emergency intervention is required; Level 2 alarm signal: When W2 < Apj ≤ W1, it indicates that the performance fluctuation is medium, and it is prompted to optimize the performance management strategy; Level 3 alarm signal: When Apj > W1, it indicates that the performance fluctuation seriously deviates from the healthy state, and at this time, intervention measures are immediately taken, including adjusting the resource allocation strategy, adjusting the microservice call chain, and adjusting the dynamic threshold parameters.
10. A container cluster monitoring and alarming method, applied to a container cluster monitoring and alarming system according to any one of claims 1 to 9, characterized in that: It includes the following steps: Step 1: Obtain the electronic topology map of the node distribution in the container cluster in advance, and monitor the cross-node resource access latency in real time. Finally, calculate and evaluate the cross-node contention coefficient Kcz, and generate the corresponding bottleneck warning signal; Step 2: After receiving the bottleneck warning signal, deeply analyze the microservice call link in the container cluster, calculate and evaluate the microservice performance interference index Gwz, and finally generate the microservice performance warning signal; Step 3: Collect historical load change data and real-time resource usage data and analyze them to construct the dynamic threshold adjustment factor Ztd; Based on the dynamic threshold adjustment factor Ztd, dynamically adjust the cross-node contention threshold Qc and the microservice performance interference threshold Kw; Step 4: Construct and analyze the cross-node contention coefficient Kcz, the microservice performance interference index Gwz, and the dynamic threshold adjustment factor Ztd based on the container performance diagnosis model. Then calculate and evaluate the comprehensive performance health index Phz, and issue the comprehensive performance health warning signal; Step 5: Calculate and evaluate the comprehensive performance health deviation value Apj according to the evaluation content of the comprehensive performance health index Phz; Finally, the comprehensive performance health warning signal is hierarchically managed, and finally detailed warning information is generated.
Citation Information
Patent Citations
Alarm implementation method and system based on dynamic configuration
CN111782486A
Enterprise-level informatization system based on micro-service architecture
CN112214474A
Performance monitoring alarm method and alarm system for container micro-service
CN114443435A
Micro-service cloud platform transverse network performance real-time monitoring operation and maintenance method and system
CN119383087A
Container microservice-oriented performance monitoring and alarm method and alarm system
WO2023142054A1