Dual-computer hot standby management system
Through the dual-machine hot standby management system, efficient data synchronization and automatic switching process between the main server and the standby server are realized, solving the problem that the detection and switching mechanisms are not sensitive enough during the main-standby switching process in the existing technology, and significantly improving the high availability and stability of the system.
Patent Information
- Application Number
- CN202510137517.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-27
AI Technical Summary
The detection and switching mechanisms of the prior art are not sensitive enough during the main-standby handover process, resulting in long service interruption time and high operation and maintenance costs. The health monitoring system relies on fixed threshold judgments and lacks effective automatic switching strategies, resulting in low fault detection and response efficiency.
It adopts a dual-machine hot standby management system, including the main server, the standby server and the central controller. The main server bears all business load, the standby server synchronizes the status data of the main server in real time, and the central controller integrates a health monitoring module to monitor the running status of the main server in real time and trigger an automatic switching process when an abnormal situation is detected, so that the standby server can quickly take over unfinished tasks and service requests.
It realizes efficient data synchronization and low-latency transmission between the primary server and the standby server, ensuring that when an abnormality occurs on the primary server, the standby server can take over all tasks and service requests in a very short time, significantly improving the high availability and stability of the system, while reducing operation and maintenance complexity and cost.
Smart Images

Figure CN120045386A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computer information systems, and particularly to a straw forming module. Background Art
[0002] With the accelerating advancement of enterprise digital transformation, data centers have become the key infrastructure to support business continuity. In the face of natural disasters, software and hardware failures and other emergencies, how to ensure the uninterrupted operation of services has become an important issue that needs to be solved urgently. The traditional single-server deployment mode can no longer meet the high-availability requirements of modern enterprises. Therefore, the industry generally seeks more advanced solutions to improve the stability and reliability of the system.
[0003] To address the above challenges, a commonly adopted technical means is the primary-backup architecture. In this architecture, the primary server is responsible for daily data processing and service response, while the backup server is in a standby state. Once the primary server fails, the backup server will take over the task and continue to provide services. In addition, there are some variant solutions, such as the multi-active architecture, that is, multiple nodes share the business load simultaneously, and the traffic is distributed through a load balancer to enhance the redundancy and fault tolerance of the system. Another common method is to deploy multiple data centers geographically dispersed, each with a complete backup system, and data synchronization is achieved through network connections to further improve the availability of the system.
[0004] Although these existing technologies have improved the stability of the system to a certain extent, there are still some obvious deficiencies. Especially during the primary-backup switchover process, due to the lack of sensitivity of the detection and switchover mechanisms, it often leads to a long service interruption time and high operation and maintenance costs. In addition, existing health monitoring systems usually rely on fixed threshold judgments and lack effective automatic switchover strategies, resulting in low fault detection and response efficiency and poor user experience. These problems seriously affect the overall performance and reliability of the system, and there is an urgent need for more efficient and intelligent solutions to improve them. Summary of the Invention
[0005] The purpose of this application is to overcome the above technical problems and provide a dual-active hot standby management system A dual-active hot standby management system includes a primary server, a backup server, and a central controller, where: The primary server undertakes all business loads; The backup server synchronizes the status data of the primary server in real time to maintain a high degree of consistency; The central controller integrates a health monitoring module, which is used to monitor the running status of the primary server and trigger an automatic switchover process when an abnormal situation is detected, so that the backup server can quickly take over the unfinished tasks and service requests.
[0006] By adopting the above technical solutions, efficient data synchronization and low-latency transmission between the primary server and the standby server are achieved. At the same time, the health monitoring module of the central controller can achieve real-time monitoring and rapid fault detection, ensuring that when an abnormality occurs in the primary server, the standby server can take over all tasks and service requests within an extremely short time, greatly improving the high availability and stability of the system. In addition, the system also has strong expansion capabilities and flexibility, can adapt to business requirements in different scenarios, and reduces the operation and maintenance complexity and costs.
[0007] Preferably, the primary server and the standby server are connected through a high-speed fiber channel to achieve real-time mirroring of data, ensuring high consistency and low-latency transmission of data. The specific steps include the primary server sending the latest status data packets to the standby server, and the standby server receiving and storing these data packets to ensure the complete consistency of the data between the two.
[0008] By adopting the above technical solutions, the primary server and the standby server are connected through a high-speed fiber channel to achieve real-time mirroring of data, ensuring high consistency and low-latency transmission of data. This not only improves the fault tolerance and stability of the system but also reduces the risk of service interruption caused by data asynchronization. The specific steps include the primary server sending the latest status data packets to the standby server, and the standby server receiving and storing these data packets to ensure the complete consistency of the data between the two. This design significantly improves the high availability of the system, especially in the face of sudden failures, it can quickly switch to the standby server to ensure business continuity.
[0009] Preferably, the central controller further includes an intelligent analysis engine for monitoring key indicators such as the CPU utilization rate, memory occupancy rate, and hard disk I / O rate of the primary server, and predicting the possible overload risks within the next few minutes. By analyzing historical data and real-time data, warning information is generated. The specific steps include collecting various operation data of the primary server, inputting them into the intelligent analysis engine for processing, and outputting prediction results and warning information.
[0010] By adopting the above technical solutions, all-round monitoring and intelligent analysis of the operating status of the primary server are achieved. It can not only obtain key indicators such as the CPU utilization rate, memory occupancy rate, and hard disk I / O rate in real time but also accurately predict the overload risks within the next few minutes through the comprehensive processing of historical data and real-time data. This helps to take countermeasures in advance to prevent the system from crashing due to sudden loads, thereby improving the stability and reliability of the system, reducing the workload of operation and maintenance personnel, and ensuring business continuity and the improvement of user experience.
[0011] Preferably, after receiving the instruction from the central controller, the standby server can enter the active state within seconds and seamlessly take over the role of the primary server. The specific steps include: the central controller detects a failure of the primary server and sends an activation instruction to the standby server. After receiving the instruction, the standby server initializes the system parameters, loads the latest status data, and starts processing the unfinished tasks.
[0012] By adopting the above technical solution, the standby server can quickly enter the active state within seconds after the central controller detects the failure of the primary server and sends the activation instruction, seamlessly taking over the role of the primary server. This not only significantly shortens the service interruption time, improves the high availability of the system, but also ensures business continuity and the stability of the user experience.
[0013] Preferably, the health monitoring module is based on a machine learning model, which is used to analyze historical performance data, identify abnormal behavior patterns, and early warn of potential failures. This model can self-optimize through continuously updated historical data to improve the prediction accuracy. The specific steps include collecting various operation data of the primary server in the past year, training a classification model using supervised learning algorithms, deploying the trained model to the central controller, regularly analyzing the current real-time monitoring data stream, and immediately sending an alarm signal once a data point matching the abnormal pattern is found.
[0014] By adopting the above technical solution, the health monitoring module based on the machine learning model can analyze the historical performance data and current operating status of the primary server in real time, accurately identify abnormal behavior patterns, and early warn of potential failures, thereby achieving more accurate risk assessment and preventive maintenance. This module self-optimizes through continuously updated historical data, improves the prediction accuracy, and reduces the service interruption time and operation and maintenance costs caused by failures.
[0015] Preferably, the machine learning model is trained by supervised learning algorithms, and positive and negative samples in a large-scale labeled historical database are used for training. The training process includes steps such as data cleaning, feature extraction, model selection, and verification. The specific steps include: extracting raw data from the historical database, preprocessing and cleaning it, extracting effective features, selecting a suitable model type, performing model training and cross-validation, and finally obtaining a high-precision classification model.
[0016] By adopting the above technical solution, the machine learning model can accurately identify abnormal behavior patterns during the operation of the primary server and early warn of potential failures through in-depth analysis of historical data and continuous self-optimization, thereby significantly improving the stability and reliability of the system. At the same time, the high-precision classification ability of this model helps to reduce false alarms and missed alarms, reduces the workload of operation and maintenance personnel, and realizes more intelligent health monitoring.
[0017] Preferably, the health monitoring module regularly analyzes the current real-time monitoring data stream. Once a data point matching an abnormal pattern is found, an alarm signal is immediately sent, and relevant information is recorded for subsequent analysis and processing. The specific steps include: real-time collecting the operation data of the main server, inputting it into a trained machine learning model for analysis. If an abnormal pattern is found, it immediately reports to the central controller and records detailed logs.
[0018] By adopting the above technical solution, the health monitoring module can collect the operation data of the main server in real time and use the trained machine learning model to analyze abnormal patterns. Once an abnormal data point is found, an alarm signal is immediately sent to the central controller, and detailed log information is recorded for subsequent analysis and processing. This not only improves the speed and accuracy of fault detection but also realizes timely early warning of potential risks, thus enhancing the stability and reliability of the system.
[0019] Preferably, the central controller further includes an adaptive module for dynamically adjusting the primary-backup switching threshold to ensure that the system makes the best response under different load conditions. This module dynamically calculates the appropriate switching threshold by real-time monitoring the change trend of KPIs. The specific steps include: defining a series of key performance indicators KPIs, such as average response time, maximum concurrent request number, etc., real-time collecting the actual KPIs data and comparing it with the preset target value, and dynamically adjusting the switching threshold according to the gap.
[0020] By adopting the above technical solution, the adaptive module in the central controller can real-time monitor the change trend of the key performance indicators KPIs and dynamically adjust the primary-backup switching threshold according to the gap between the actual performance and the preset target value. This ensures that the system can make the best response under different load conditions, improves the flexibility and stability of the system, reduces the need for manual intervention, simplifies the operation and maintenance work, and thus significantly improves the overall service quality.
[0021] Preferably, the adaptive module calculates an appropriate switching threshold according to the gap between the preset key performance indicators KPIs and the actual performance. The specific calculation formula considers multiple factors, including but not limited to average response time, maximum concurrent request number, etc. The specific steps include: setting an initial switching threshold, and adjusting the size of the threshold according to the deviation between the real-time KPIs data and the target value to ensure that the system always maintains an optimal state.
[0022] By adopting the above technical solution, the function of dynamically adjusting the primary-backup switching threshold is realized. Specifically, the adaptive module calculates an appropriate switching threshold according to the gap between the preset key performance indicators (KPIs) and the actual performance. This process takes into account multiple factors, including but not limited to the average response time, the maximum number of concurrent requests, etc. By setting an initial switching threshold and dynamically adjusting the threshold size according to the deviation between the real-time KPI data and the target value, it is ensured that the system always maintains an optimal state. This not only improves the flexibility and stability of the system, but also reduces the need for manual intervention and simplifies the operation and maintenance work.
[0023] Preferably, the adaptive module continuously monitors the change trend of the actual KPIs. Once it predicts that the set dynamic threshold is about to be exceeded, the system immediately triggers an emergency plan, such as dispatching additional backup resources, restricting the access frequency of customers, etc., to ensure that the service quality is always maintained at a high level. The specific steps include: real-time monitoring of the change trend of KPIs, predicting the future load situation, and if it is predicted that the set dynamic threshold is about to be exceeded, immediately execute the corresponding emergency plan, such as increasing backup resources, restricting the access frequency of the client, etc., to ensure the stability of the system and the service quality.
[0024] By adopting the above technical solution, the adaptive module can real-time monitor the change trend of the key performance indicators (KPIs) and predict the future load situation. If it is predicted that the set dynamic threshold is about to be exceeded, the system will immediately execute the corresponding emergency plan, such as increasing backup resources, restricting the access frequency of the client, etc., so as to ensure that the stability of the system and the service quality are always maintained at a high level. This not only improves the anti-pressure ability and response speed of the system, but also reduces the risk of service interruption caused by sudden high load, and further improves the user experience and the reliability of the system.
[0025] In summary, the present application includes at least one of the following beneficial technical effects: 1. The application of real-time data synchronization technology and millisecond-level fault detection mechanism makes the switching process between the primary and backup servers almost imperceptible, significantly shortening the service interruption time and improving the user experience; 2. The intelligent health monitoring system based on machine learning can identify abnormal behavior patterns by analyzing historical performance data, early warning potential faults, and improving the stability and reliability of the system; 3. The adaptive algorithm framework for dynamically adjusting the primary-backup switching threshold dynamically calculates an appropriate switching threshold according to the actual load conditions, ensuring that the system can make the best response in different situations, reducing the need for manual intervention and simplifying the operation and maintenance work. Brief Description of the Drawings
[0026] Figure 1 is a block diagram of a dual-machine hot standby management system according to an embodiment of the present application.
[0027] Figure 2 It is a block diagram of the central controller according to an embodiment of the present application. Specific embodiments
[0028] The following will Figure 1 - Figure 2 further elaborate on the present application in detail with reference to the accompanying drawings.
[0029] Embodiment 1: Referring to Figure 1 and Figure 2 A dual - machine hot - standby management system provided by an embodiment of the present application includes a primary server, a standby server, and a central controller. The primary server undertakes all business loads, and the standby server synchronizes the status data of the primary server in real time to maintain a high degree of consistency. The central controller integrates a health monitoring module, which is used to monitor the running status of the primary server and trigger an automatic switching process when an abnormal situation is detected, enabling the standby server to quickly take over unfinished tasks and service requests. Through real - time data synchronization technology and a millisecond - level fault detection mechanism, seamless connection during the service switching process is achieved, significantly improving the user experience level and reaching the ultimate pursuit of high availability.
[0030] Specifically, the primary server includes a high - performance CPU, a large - capacity memory, and a high - speed disk array, such as a high - performance processor (e.g., Intel Xeon series), DDR4 memory modules, and SSD solid - state drives. These hardware configurations ensure that the primary server can efficiently process a large amount of online transaction data. The standby server is also equipped with similar hardware configurations and is connected to the primary server through a high - speed fiber channel to achieve real - time mirroring of data. The selection of the fiber channel can be single - mode or multi - mode fiber according to actual requirements to ensure low - latency and high - reliability data transmission.
[0031] The central controller is the core part of the entire system, which integrates an intelligent analysis engine and a health monitoring module. The intelligent analysis engine can not only monitor key indicators such as the CPU utilization rate, memory occupancy rate, and hard disk I / O rate of the primary server, but also predict the potential overload risk within the next few minutes. This function generates warning information by analyzing historical data and real - time data, so as to discover potential risks in advance and take preventive measures. Specifically, the central controller can collect various running data of the primary server, input them into the intelligent analysis engine for processing, and output prediction results and warning information. For example, when it is detected that the CPU utilization rate of the primary server exceeds 90% or the memory occupancy rate reaches 80%, the central controller will immediately initiate a switching plan, notify the standby server to enter the active state, and seamlessly take over the role of the primary server.
[0032] In addition, the central controller also has an adaptive module, which is used to dynamically adjust the primary-backup switching threshold to ensure that the system can respond optimally under different load conditions. The adaptive module dynamically calculates the appropriate switching threshold by monitoring the changing trends of KPIs in real time. For example, a series of key performance indicators (KPIs) can be defined, such as the average response time, the maximum number of concurrent requests, etc. The actual KPI data is collected in real time and compared with the preset target values, and the switching threshold is dynamically adjusted according to the gap. In this way, the system can maintain the optimal working state under different load conditions, improving the overall stability and reliability.
[0033] The implementation principle of this embodiment is as follows: Through the real-time data synchronization technology, the data between the primary server and the backup server is ensured to be highly consistent, reducing the risk of switching failure caused by data inconsistency. At the same time, the intelligent analysis engine and health monitoring module of the central controller can timely detect potential risks and take preventive measures, improving the stability and reliability of the system. The introduction of the adaptive module further reduces the need for manual intervention, simplifies the operation and maintenance work, enables the system to always maintain the best performance under different load conditions, and ultimately improves the user experience level.
[0034] Embodiment 2: The difference between this embodiment and the above embodiment is that a machine learning model is further introduced to analyze historical performance data, identify abnormal behavior patterns, and give early warnings of potential faults, so as to achieve more accurate risk assessment and preventive maintenance.
[0035] Specifically, first, it is necessary to collect various operation data of the primary server in the past year, including but not limited to CPU usage rate, memory consumption, disk read and write speed, network transmission rate, etc., to form a large-scale historical database. This database can be stored in a local server or cloud storage for subsequent data processing and analysis. Next, a classification model is trained using a supervised learning algorithm to distinguish the difference between the normal operating state and the early signs of faults. A large number of positive and negative samples need to be labeled in the model training stage, and this step can be completed by manual marking or semi-automatic marking. Commonly used supervised learning algorithms include random forest, support vector machine, neural network, etc. Which algorithm to choose depends on the specific requirements of the actual scenario and the data characteristics.
[0036] Suppose the random forest algorithm is selected as the classification model, and its basic formula is as follows: \[ \text{RF}(X) =\frac{1}{T} \sum_{t=1}^{T} h_t(X) \] where \( T \) is the number of trees, and \( h_t(X) \) represents the prediction result of the \( t \)-th tree for the input feature \( X \).
[0037] The trained model is deployed to the central controller to regularly analyze the current real-time monitoring data stream. Once data points matching abnormal patterns are detected, an alarm signal is immediately sent to alert the administrator to intervene for investigation, or the resource allocation strategy is automatically adjusted to prevent the system from entering a dangerous area. For example, if the model detects frequent high CPU utilization during a specific period on the main server, it may trigger a deep inspection to determine whether there are potential software bugs or external attacks. This machine learning-based health monitoring system can greatly improve the efficiency of fault detection and response, reduce the possibility of human misjudgment, and further enhance the stability and reliability of the system.
[0038] The implementation principle of this embodiment is as follows: By introducing a machine learning model, in-depth analysis is carried out on historical performance data to identify potential abnormal behavior patterns and early warning of potential faults. This method not only improves the accuracy and timeliness of fault detection, but also can gradually improve the prediction accuracy through continuous self-optimization. Combined with the intelligent analysis engine of the central controller, a multi-level and multi-dimensional health monitoring system is formed, effectively preventing the occurrence of potential risks and greatly enhancing the stability and reliability of the system.
[0039] Embodiment 3: The difference between this embodiment and the above embodiments lies in that the design and implementation of the adaptive module are mainly introduced. By dynamically adjusting the primary-backup switching threshold, it ensures that the system makes the best response under different load conditions, reduces the need for manual intervention, and simplifies the operation and maintenance work.
[0040] Specifically, a series of key performance indicators (KPIs) are first defined, such as average response time, maximum concurrent request number, etc., which are used as standards for evaluating the system status. The selection of these KPIs should be determined according to the actual business requirements and system characteristics. For example, for a financial trading platform, the average response time and the maximum concurrent request number may be one of the most important indicators. Then, a rule-based adaptive module is designed to calculate an appropriate switching threshold based on the gap between the current actual KPI performance and the preset target value. The specific calculation formula takes into account multiple factors, including but not limited to the average response time, the maximum concurrent request number, etc. For example, if the current average response time exceeds the preset target value, the adaptive module will appropriately lower the switching threshold, making it easier for the system to trigger the switching process, thus avoiding the degradation of the user experience.
[0041] Assume that a linear function is defined to calculate the switching threshold: \[ \theta = \alpha R + \beta C \] where \( R \) is the average response time, \( C \) is the maximum number of concurrent requests, and \( \alpha \) and \( \beta \) are weight coefficients that can be adjusted according to the actual situation.
[0042] The adaptive module also needs to continuously monitor the changing trend of the actual KPIs. Once it predicts that the set dynamic threshold is about to be breached, the system immediately triggers a pre-prepared contingency plan. For example, additional backup resources can be dispatched, the access frequency of clients can be limited, etc., to ensure that the service quality is always maintained above a high standard. The specific operation steps include: monitoring the changing trend of KPIs in real time, predicting the future load situation, and if it is predicted that the set dynamic threshold is about to be breached, immediately execute the corresponding contingency plan, such as increasing backup resources, restricting the access frequency of clients, etc., to ensure the stability of the system and the service quality.
[0043] The implementation principle of this embodiment is as follows: By dynamically adjusting the primary-backup switching threshold, the adaptive module can make flexible decisions according to the actual load situation, ensuring the stable operation of the system under various conditions. This adaptive mechanism not only reduces the need for manual intervention, simplifies the operation and maintenance work, but also can quickly respond in case of emergencies, effectively preventing the decline of service quality. Combined with the intelligent analysis engine and the health monitoring module of the central controller, a multi-layer linkage management system is formed, comprehensively improving the high availability and reliability of the system.
[0044] The above are all the preferred embodiments of this application. It does not limit the protection scope of this application accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of this application should be covered within the protection scope of this application.
Claims
1. A dual-machine hot standby management system, characterized in that: It includes the main server, backup server and central controller, where: The main server bears all business loads; The standby server synchronizes the status data of the primary server in real time to maintain high consistency; The central controller integrates a health monitoring module to monitor the operating status of the primary server and trigger an automatic switching process when an abnormal situation is detected, so that the backup server can quickly take over unfinished tasks and service requests.
2. The dual-machine hot standby management system according to claim 1, characterized in that: The main server and the backup server are connected via a high-speed fiber channel to achieve real-time mirroring of data, ensuring high consistency and low-latency transmission of data. The specific steps include the main server sending the latest status data packets to the backup server, and the backup server receiving and storing these data packets to ensure that the data between the two is completely consistent.
3. The dual-machine hot standby management system according to claim 1, characterized in that: The central controller also includes an intelligent analysis engine for monitoring key indicators such as the CPU utilization, memory occupancy, and hard disk I / O rate of the main server, and predicting the overload risks that may occur in the next few minutes. It generates early warning information by analyzing historical data and real-time data. The specific steps include collecting various operating data of the main server, inputting it into the intelligent analysis engine for processing, and outputting prediction results and early warning information.
4. The dual-machine hot standby management system according to claim 1, characterized in that: After receiving the instruction from the central controller, the standby server can enter the active state within seconds and seamlessly take over the role of the main server. The specific steps include: the central controller detects the failure of the main server and sends an activation instruction to the standby server. After receiving the instruction, the standby server initializes the system parameters, loads the latest status data and starts processing unfinished tasks.
5. The dual-machine hot standby management system according to claim 1, characterized in that: The health monitoring module is based on a machine learning model and is used to analyze historical performance data, identify abnormal behavior patterns, and provide early warning of potential failures. The model can self-optimize through continuously updated historical data to improve prediction accuracy. The specific steps include collecting various operating data of the main server in the past year, using a supervised learning algorithm to train a classification model, deploying the trained model to the central controller, and regularly analyzing the current real-time monitoring data stream. Once a data point matching an abnormal pattern is found, an alarm signal is immediately issued.
6. The dual-machine hot standby management system according to claim 5, characterized in that: The machine learning model is trained by a supervised learning algorithm using positive and negative samples from a large-scale labeled historical database. The training process includes steps such as data cleaning, feature extraction, model selection and verification. The specific steps include: extracting raw data from the historical database, preprocessing and cleaning it, extracting effective features, selecting a suitable model type, performing model training and cross-validation, and finally obtaining a high-precision classification model.
7. The dual-machine hot standby management system according to claim 5, characterized in that: The health monitoring module regularly analyzes the current real-time monitoring data stream. Once a data point matching an abnormal pattern is found, an alarm signal is immediately issued and relevant information is recorded for subsequent analysis and processing. The specific steps include: real-time collection of various operating data of the main server, inputting it into the trained machine learning model for analysis, and if an abnormal pattern is found, immediately reporting to the central controller and recording detailed logs.
8. The dual-machine hot standby management system according to claim 1, characterized in that: The central controller also includes an adaptive module for dynamically adjusting the active-standby switching threshold to ensure that the system responds optimally under different load conditions. The module dynamically calculates the appropriate switching threshold by monitoring the changing trend of KPIs in real time. The specific steps include: defining a series of key performance indicators KPIs, such as average response time, maximum number of concurrent requests, etc., collecting actual KPIs data in real time and comparing it with the preset target value, and dynamically adjusting the switching threshold according to the gap.
9. The dual-machine hot standby management system according to claim 8, characterized in that: The adaptive module calculates the appropriate switching threshold based on the gap between the preset key performance indicators KPIs and the actual performance. The specific calculation formula takes into account multiple factors, including but not limited to the average response time, the maximum number of concurrent requests, etc. The specific steps include: setting the initial switching threshold, adjusting the threshold value according to the deviation between the real-time KPIs data and the target value, to ensure that the system always maintains the optimal state.
10. The dual-machine hot standby management system according to claim 8, characterized in that: The adaptive module continuously monitors the changing trend of actual KPIs. Once it is predicted that the set dynamic threshold is about to be exceeded, the system immediately triggers an emergency plan, such as adding backup resources, limiting the frequency of customer access, etc., to ensure that the service quality is always maintained at a high level. The specific steps include: real-time monitoring of the changing trend of KPIs, predicting future load conditions, and if it is predicted that the set dynamic threshold is about to be exceeded, immediately executing the corresponding emergency plan, such as adding backup resources, limiting the frequency of client access, etc., to ensure the stability of the system and service quality.
Citation Information
Patent Citations
Method and device for dynamically managing and adjusting multiple hosts
CN112035262A
Self-adaptive anomaly judgment method and device in real-time anomaly detection system
CN112508316A
Service quality monitoring method, device, equipment, medium and product
CN118798705A
Heating and ventilation equipment abnormity online monitoring system based on Internet of Things
CN118915566A
Main-standby switching method and system based on equipment synchronization
CN119011374A