Method and device for processing threshold notification information based on cluster traffic

By using a real-time collection and dynamic adjustment method for cluster traffic threshold notification information processing, the problem of false alarms and missed alarms caused by traffic switching in the financial system has been solved, enabling efficient monitoring and accurate status determination of application system clusters and improving operation and maintenance efficiency.

CN121547347APending Publication Date: 2026-02-17AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511723904.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-22
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

In financial application systems, existing technologies are prone to false alarms or missed warnings during traffic switching, failing to accurately focus on the actual production and operation status of the application system, and lacking a comprehensive assessment of the overall operation status of multiple instances, resulting in low switching efficiency and operational risks.

Method used

A threshold notification information processing method based on cluster traffic is adopted. By collecting the operating indicators of system instances in real time, the overall operating indicators and instance weights are calculated using a long short-term memory network model, and the threshold is dynamically adjusted to achieve intelligent alarm suppression and traffic switching decisions, avoiding misjudgment and false alarms.

Benefits of technology

It effectively reduced the false alarm rate of notification information, improved the accuracy of application system cluster operation status monitoring and switching efficiency, reduced interference for operation and maintenance personnel, and improved work quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121547347A_ABST
    Figure CN121547347A_ABST
Patent Text Reader

Abstract

The invention discloses a processing method and device for threshold notification information based on cluster traffic, and the method comprises the steps: collecting the instance operation indexes of a plurality of system instances contained in a system cluster in real time; under the condition that the instance operation index of any system instance does not meet the instance threshold matching condition of the system instance, obtaining an overall operation index obtained by calculating each instance operation index based on the preset instance weight of each system instance, the instance threshold matching condition is determined according to an instance operation index of the system instance collected in a historical time period; and under the condition that the overall operation index conforms to an overall threshold matching condition of the system cluster, notification information triggered when the instance operation index does not conform to the instance threshold matching condition is deleted, and the overall threshold matching condition is determined according to the overall operation index obtained in the historical time period.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular to a method and apparatus for processing threshold notification information based on cluster traffic. Background Technology

[0002] Currently, financial application systems commonly adopt a multi-site active-active high-availability architecture, such as Active-Active (AA) and Active-Active-Active-Active (AAA) architectures where all nodes remain active, and Active-Active-Standby (AAS) architectures where multiple nodes remain active and some serve as backups. However, when encountering environmental anomalies or switchover drills that require switching traffic to a different location, significant fluctuations in traffic across different locations can cause multiple indicators, including transaction volume, response time, and success rate, to exceed thresholds, triggering corresponding warnings or alarms. However, too many such warnings or alarms can cause some interference for operations and maintenance personnel, making it difficult to accurately focus on the actual production and operation status of the application system. Summary of the Invention

[0003] Therefore, this application discloses the following technical solution:

[0004] The first aspect of this application provides a method for processing threshold notification information based on cluster traffic, including:

[0005] Real-time collection of instance operation metrics from multiple system instances within the system cluster;

[0006] If the instance operation index of any of the system instances does not meet the instance threshold matching condition of the system instance, the overall operation index is obtained by calculating the instance operation index of each instance based on the preset instance weight of each system instance. The instance threshold matching condition is determined based on the instance operation index of the system instances collected in the historical period.

[0007] If the overall operating metrics meet the overall threshold matching conditions of the system cluster, delete the notification information triggered by the instance operating metrics not meeting the instance threshold matching conditions. The overall threshold matching conditions are determined based on the overall operating metrics obtained within the historical time period.

[0008] Optionally, the process of determining the instance threshold matching condition for the system instance includes:

[0009] Based on the instance operation metrics and instance weights of the multiple system instances collected during the historical period, the overall operation metrics of the system cluster during the historical period are determined.

[0010] The overall threshold matching conditions of the system cluster are determined based on the overall operating indicators of the system cluster during the historical period.

[0011] The instance threshold matching condition for each system instance is determined based on the overall threshold matching condition of the system cluster and the instance weight of each system instance.

[0012] Optionally, determining the overall threshold matching condition of the system cluster based on the overall operating indicators of the system cluster within the historical period includes:

[0013] The overall operating metrics of the system cluster within the historical period are processed according to a pre-constructed long short-term memory network model to obtain the overall threshold and overall fluctuation range of the system cluster. Based on the overall threshold and the overall fluctuation range, an overall threshold matching condition is determined. The overall threshold matching condition includes that the difference between the overall operating metrics of the system cluster and the overall threshold is within the overall fluctuation range.

[0014] Optionally, after deleting the notification triggered by the instance's running metrics failing to meet the instance threshold matching condition, the method further includes at least one of the following:

[0015] The first system instance to be expanded is processed for expansion, and the first system instance is determined among the multiple system instances based on the instance operation indicators.

[0016] Output traffic switching prompt, which is used to indicate that access traffic is switched from the second system instance to the first system instance. The second system instance is determined among the multiple system instances based on instance operation indicators, and the access traffic for the second system instance is less than the access traffic for the first system instance.

[0017] Optionally, it may also include at least one of the following:

[0018] The instance weights of the multiple system instances are determined based on their proportion of the overall access traffic of the system cluster.

[0019] The instance weights of the multiple system instances are determined based on the business type data of the multiple system instances;

[0020] The instance weights of the multiple system instances are determined based on the changing trends of their instance operation metrics during the historical period.

[0021] Based on the historical switching data of the multiple system instances, the instance weights of the multiple system instances are determined, whereby the historical switching data represents the number of traffic switching times between the multiple system instances.

[0022] A second aspect of this application provides a processing apparatus for threshold notification information based on cluster traffic, comprising:

[0023] The data acquisition unit is used to collect instance operation metrics of multiple system instances contained in the system cluster in real time.

[0024] The obtaining unit is used to obtain an overall operating index based on the instance weights of each system instance when the instance operating index of any system instance does not meet the instance threshold matching condition of the system instance. The instance threshold matching condition is determined based on the instance operating index of the system instance collected in a historical period.

[0025] The processing unit is configured to delete notification information triggered by instance operating metrics not meeting instance threshold matching conditions, provided that the overall operating metrics meet the overall threshold matching conditions of the system cluster. The overall threshold matching conditions are determined based on the overall operating metrics obtained within the historical time period.

[0026] Optionally, the obtaining unit is used to determine the instance threshold matching condition of the system instance, and the process of determining the instance threshold matching condition of the system instance includes:

[0027] Based on the instance operation metrics and instance weights of the multiple system instances collected during the historical period, the overall operation metrics of the system cluster during the historical period are determined.

[0028] The overall threshold matching conditions of the system cluster are determined based on the overall operating indicators of the system cluster during the historical period.

[0029] The instance threshold matching condition for each system instance is determined based on the overall threshold matching condition of the system cluster and the instance weight of each system instance.

[0030] Optionally, the obtaining unit determines the overall threshold matching conditions of the system cluster based on the overall operating indicators of the system cluster within the historical period, including:

[0031] The overall operating metrics of the system cluster within the historical period are processed according to a pre-constructed long short-term memory network model to obtain the overall threshold and overall fluctuation range of the system cluster. Based on the overall threshold and the overall fluctuation range, an overall threshold matching condition is determined. The overall threshold matching condition includes that the difference between the overall operating metrics of the system cluster and the overall threshold is within the overall fluctuation range.

[0032] Optionally, after deleting the notification information triggered by the instance's running metrics failing to meet the instance threshold matching condition, the processing unit is further configured to perform at least one of the following:

[0033] The first system instance to be expanded is processed for expansion, and the first system instance is determined among the multiple system instances based on the instance operation indicators.

[0034] Output traffic switching prompt, which is used to indicate that access traffic is switched from the second system instance to the first system instance. The second system instance is determined among the multiple system instances based on instance operation indicators, and the access traffic for the second system instance is less than the access traffic for the first system instance.

[0035] Optionally, it also includes a determining unit for determining the instance weights of the plurality of system instances, specifically for performing at least one of the following:

[0036] The instance weights of the multiple system instances are determined based on their proportion of the overall access traffic of the system cluster.

[0037] The instance weights of the multiple system instances are determined based on the business type data of the multiple system instances;

[0038] The instance weights of the multiple system instances are determined based on the changing trends of their instance operation metrics during the historical period.

[0039] Based on the historical switching data of the multiple system instances, the instance weights of the multiple system instances are determined, whereby the historical switching data represents the number of traffic switching times between the multiple system instances.

[0040] The beneficial effects of this solution are as follows: by weighting the instance operation metrics of multiple application system instances to obtain the overall operation metrics, when the instance operation metrics are abnormal, i.e. they do not meet the corresponding instance threshold matching conditions, the overall operation metrics are further combined to realize the collaborative status determination of multiple application system instances in the application system cluster. This avoids misjudgment caused by isolated monitoring of a single application system instance, thereby effectively reducing the notification information generated by misjudgment of a single application system instance and lowering the false alarm rate of notification information. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0042] Figure 1This is a flowchart illustrating a method for processing threshold notification information based on cluster traffic, as provided in an embodiment of this application.

[0043] Figure 2 This is a schematic diagram of a processing device for threshold notification information based on cluster traffic, provided in an embodiment of this application. Detailed Implementation

[0044] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0045] To facilitate understanding of the technical solutions of this application, some terms that may be involved in the embodiments of this application are briefly described below.

[0046] Long Short-Term Memory (LSTM) networks are a special type of recurrent neural network (RNN) that addresses the vanishing gradient and long-term dependency problems of traditional RNNs by introducing a gating mechanism. They excel at processing time-series data (such as text, speech, and sensor data). Their core functions include long-term memory retention: storing key information across time steps through memory cells; and dynamic information filtering: controlling the flow of information through forget gates, input gates, and output gates to determine which information to retain, update, or discard.

[0047] Dynamic threshold weight calculation: Dynamic threshold weight calculation is the core technology for realizing intelligent alarm optimization. Its design logic includes key links such as "weight allocation and baseline threshold setting", "dynamic floating range adaptation" and "global health status verification". Ultimately, through weight-driven dynamic threshold calculation, the system can both guarantee the performance boundary of a single instance and take into account global stability, significantly reduce the false alarm rate, and support the fine-grained operation and maintenance needs of a high-availability architecture.

[0048] Transactions Per Second (TPS) is a key metric for measuring the throughput of a business system in processing business requests.

[0049] Existing active-active architectures in the financial sector generally adopt local dual-active or geographically distributed active-active deployment modes, but they suffer from the following problems in monitoring and failover:

[0050] Firstly, static thresholds have limitations. The fixed warning thresholds set for a single instance by traditional monitoring systems cannot adapt to scenarios with dynamic traffic switching, which can easily lead to false alarms or missed alarms. For example, when traffic switches from instance A to instance B, a sudden drop in traffic on instance A may trigger an erroneous alarm, while a sudden increase in load on instance B may exceed the original threshold without timely adjustment.

[0051] Secondly, the overall indicators are fragmented. Existing solutions lack a comprehensive assessment of the overall operational status of multiple instances and only focus on single-point indicators, making it difficult to determine whether the system is truly abnormal.

[0052] Thirdly, the switching scenario response is lagging. In case of environmental anomalies or drills, manual intervention is required to adjust the threshold, resulting in low switching efficiency and operational risks.

[0053] This application aims to provide a dynamic detection method for high availability architecture of application systems in the financial field. Based on this dynamic detection method, the application system can be monitored for uninterrupted production operation status during deployment site switching, and erroneous information generated during the switching period can be avoided, thereby improving the work quality and efficiency of operation and maintenance personnel.

[0054] To address the shortcomings of the prior art, this application provides a method for processing threshold notification information based on cluster traffic, the basic principle of which is as follows.

[0055] Firstly, the application system cluster is deployed as multiple instances (referred to as system instances). The specific division method is not limited. For example, it can be divided into three instances deployed in cities X, Y, and Z, respectively, and referred to as instance X, instance Y, and instance Z. A corresponding traffic weight (or instance weight) is determined for each instance based on the instance's location, business type, and other dimensions. The weighted sum of the metrics of each instance according to the instance weight equals the overall system metric.

[0056] Example: A commercial bank deploys three instances in cities X, Y, and Z, with an instance weight ratio of 4:3:3, meaning the instance weights for X, Y, and Z are 0.4, 0.3, and 0.3 respectively. The overall TPS threshold is 100,000 transactions per second. The threshold for each individual instance can be dynamically calculated based on its instance weight. Real-time data collection is performed on each instance's transaction volume (TPS), response time (RT), and error rate (ERR), and a weighted overall metric is calculated. For example, the TPS metrics of each instance can be summed according to its instance weight. When the weight of instance X is 0.4, its TPS threshold is the overall threshold * 0.4, and the TPS metric for instance X can fluctuate within a certain range above or below its TPS threshold, for example, a fluctuation of 10%.

[0057] Secondly, an intelligent alarm suppression engine is implemented based on the thresholds and instance weights configured above.

[0058] The intelligent alarm suppression engine can detect whether the system cluster is abnormal based on a two-layer state determination model, so as to output corresponding notification information (such as the aforementioned early warning information or alarm information) when an anomaly occurs. Specifically, the intelligent alarm suppression engine can perform local anomaly detection, that is, identify whether the indicators of a single instance have changed abruptly, that is, whether they exceed the threshold corresponding to a single instance and whether the magnitude of the exceedance exceeds the allowable fluctuation range. For example, the TPS of a certain instance drops sharply and is less than 80% of its corresponding TPS threshold, or the response time (RT) of a certain instance increases and exceeds 200% of its corresponding RT threshold. Then, when a sudden change in the indicator of a single instance is detected, the intelligent alarm suppression engine can perform a global consistency check to verify whether the overall indicators are within a preset reasonable range. This preset reasonable range can be obtained by fitting historical data. For example, multiple fluctuation ranges with different confidence levels can be obtained by fitting historical data, and the fluctuation range with 95% confidence level is determined as the reasonable range.

[0059] The intelligent alarm suppression engine can also switch the scenario decision tree. For example, if the metric of instance A changes abruptly and the metric of instance B changes in the opposite direction, such as the TPS of instance A decreasing by more than 90% of the corresponding threshold, and the TPS of instance B increasing by more than 85% of the corresponding threshold, and the overall metric fluctuation of the system cluster is within 5%, then it is determined that no abnormality has occurred in the current system cluster. The TPS of the above instances is due to the triggering of active switching, so no notification information indicating abnormal conditions needs to be output.

[0060] The intelligent alarm suppression engine can also trigger dynamic threshold adjustments. Specifically, it can disable alarm rules on the side with no traffic (such as instance A); adjust the threshold on the high-load side to a dynamic range, for example, determining that the adjusted threshold = (original threshold × (1 + traffic migration ratio)), where the traffic migration ratio can be entered by relevant operations and maintenance personnel. It can also publish system status prompts indicating traffic switching, for example, outputting a system status prompt such as: "Traffic has been switched from A to B, and the application system's global traffic is normal (baseline reference value 102%)".

[0061] Thirdly, through the adaptive learning module, historical switching data is analyzed based on the LSTM model to dynamically update the threshold, optimize the threshold fluctuation range and weight allocation strategy (such as automatically expanding the threshold range by 20% during holiday traffic patterns), and optimize the weight allocation strategy in combination with historical switching data.

[0062] For details, please see Figure 1 The method for processing threshold notification information based on cluster traffic provided in this application embodiment may include the following steps.

[0063] S101 collects instance operation metrics of multiple system instances contained in the system cluster in real time.

[0064] S102, if the instance operation index of any system instance does not meet the instance threshold matching condition of the system instance, obtain the overall operation index obtained by calculating the instance operation index of each instance based on the preset instance weight of each system instance. The instance threshold matching condition is determined based on the instance operation index of the system instances collected in the historical period.

[0065] S103, if the overall operating indicators meet the overall threshold matching conditions of the system cluster, delete the notification information triggered by the instance operating indicators not meeting the instance threshold matching conditions. The overall threshold matching conditions are determined based on the overall operating indicators obtained in the historical period.

[0066] The beneficial effects of this solution are as follows: by weighting the instance operation metrics of multiple application system instances to obtain the overall operation metrics, when the instance operation metrics are abnormal, i.e. they do not meet the corresponding instance threshold matching conditions, the overall operation metrics are further combined to realize the collaborative status determination of multiple application system instances in the application system cluster. This avoids misjudgment caused by isolated monitoring of a single application system instance, thereby effectively reducing the notification information generated by misjudgment of a single application system instance and lowering the false alarm rate of notification information.

[0067] After step S102, if the overall operating indicators do not meet the overall threshold matching conditions of the system cluster, it indicates that the current system cluster is definitely in an abnormal state. At this time, the notification information triggered by the instance operating indicators not meeting the instance threshold matching conditions can be released. That is, the notification information triggered by the instance operating indicators not meeting the instance threshold matching conditions can be output to the operation and maintenance personnel through any terminal device and multiple output methods so that the operation and maintenance personnel can investigate.

[0068] In step S101, a system instance can be understood as an application used by the bank to process specific business, such as application system instances that process transactions, payments, transfers, wealth management, etc. A collection of multiple application system instances is called an application system cluster (referred to as a system cluster). Each system instance can be deployed on one or more servers.

[0069] Generally, system instances can be divided according to their location (region). For example, they can be divided into three instances deployed in cities X, Y, and Z, respectively, and referred to as instance X, instance Y, and instance Z. Each system instance has complete business processing capabilities and adopts a unitized architecture design, dividing traffic by customer location and account sharding rules (such as hash modulo).

[0070] Instance performance metrics can be collected periodically in real time, such as once every minute or once every 30 seconds, without limitation. Instance performance metrics can include any one of the following: instance transaction volume (TPS), instance response time (RT), and instance error rate (ERR), as well as other performance metrics that can characterize the processing capabilities of the corresponding system instance, without limitation.

[0071] In step S102, each system instance has a corresponding instance weight, and the instance weights of different system instances can be the same or different.

[0072] The instance weights of different system instances can be preset by the operations and maintenance personnel, or the weight input information can be processed by a pre-built and trained long short-term memory network model to obtain the instance weight of each system instance.

[0073] The weight input information can include any one or more of the following: access traffic of multiple system instances, business type data, the trend of instance operation indicators in historical periods, and historical switchover data.

[0074] Correspondingly, the method in this embodiment may further include at least one of the following methods for determining instance weights:

[0075] The first method for determining instance weights is to determine the instance weights of multiple system instances based on the proportion of access traffic of multiple system instances in the overall access traffic of the system cluster.

[0076] The second method for determining instance weights is to determine the instance weights of multiple system instances based on the business type data of multiple system instances.

[0077] The third method for determining instance weights is to determine the instance weights of multiple system instances based on the changing trends of their instance operation metrics over historical periods.

[0078] The fourth method for determining instance weights is to determine the instance weights of multiple system instances based on historical switching data of multiple system instances. The historical switching data represents the number of traffic switching between multiple system instances.

[0079] The historical period refers to a preset time period before the current moment. The specific length of the time period can be set as needed and is not limited. For example, the historical period can be the most recent month or the most recent week up to the current moment.

[0080] In the first method of determining instance weight, the access traffic of a system instance within a historical time period can be obtained. The access traffic can be expressed as the number of access requests from clients to this system instance per unit time. For example, if the access traffic is counted once per minute, the access traffic can be expressed as the number of access requests per minute. The access traffic of the system instance obtained from the historical time period for each unit time is averaged to obtain the average access traffic of the system instance. This method is used to iterate through each system instance to obtain the average access traffic of each system instance.

[0081] Then, the instance weight can be determined based on the average access traffic. The system instance with the higher the average access traffic has the higher the instance weight, and the system instance with the lower the average access traffic has the lower the instance weight.

[0082] In the second method for determining instance weights, the business type data of the system instance may include data such as which business types of access requests the system instance has processed in historical periods, and the number of access requests processed for each business type.

[0083] One method to determine the instance weight of multiple system instances based on their business type data is that the more concentrated the business types processed, the greater the instance weight of the system instance, and the more dispersed the business types processed, the smaller the instance weight of the system instance.

[0084] For example, if 90% of the access requests processed by instance A during a historical period were for business type A, and the remaining business types accounted for only 10%, while instance B processed access requests for 5 different business types, with each type accounting for nearly 20%, then instance A has a higher instance weight, and instance B has a lower instance weight.

[0085] In the third method for determining instance weights, the trend of instance performance metrics over a historical period can be measured by various statistical indicators, without any specific limitation. For example, the standard deviation of a system instance's instance TPS metric over a historical period can be used as the trend, or the fluctuation range of a system instance's instance TPS metric over a historical period can be used as the trend. The fluctuation range is equal to the difference between the maximum and minimum values ​​of the instance TPS metric over a historical period.

[0086] In the third method for determining instance weights, the more drastic the change in the system instance's performance indicators, the greater the corresponding instance weight; conversely, the more gradual the change in the system instance's performance indicators, the smaller the corresponding instance weight. Referring to the examples above, system instances with larger standard deviations and larger fluctuations have greater instance weights, while system instances with smaller standard deviations and smaller fluctuations have smaller instance weights.

[0087] In the fourth method for determining instance weights, the historical switching data of a system instance can include the number of times the system instance participated in traffic switching, the number of times it acted as an inbound endpoint when participating in traffic switching, and the number of times it acted as an outbound endpoint when participating in traffic switching within a historical period.

[0088] Traffic switching refers to routing (or forwarding) an access request to system instance A to another system instance B for processing. In this process, system instance A is the sending end and system instance B is the receiving end. Each such routing (or forwarding) is recorded as a traffic switching operation involving both system instances (system instance A and system instance B).

[0089] In the fourth method for determining instance weight, the system instance that participates in the most traffic switching times during the historical period has a higher instance weight. If two system instances participate in the same number of traffic switching times, the system instance that has been the migration endpoint more times has a higher instance weight.

[0090] In step S102, the overall threshold matching conditions can be determined first, and then the instance threshold matching conditions for each system instance can be determined based on the overall threshold matching conditions.

[0091] One way to determine the overall threshold matching conditions is to determine the overall threshold matching conditions of the system cluster based on the overall operating indicators of the system cluster over a historical period. Specifically, the overall threshold matching conditions can be determined as follows:

[0092] The system cluster's overall operating metrics within a historical period are processed using a pre-built long short-term memory network model to obtain the overall threshold and overall fluctuation range of the system cluster. Based on the overall threshold and overall fluctuation range, the overall threshold matching conditions are determined. The overall threshold matching conditions include that the difference between the overall operating metrics of the system cluster and the overall threshold is within the overall fluctuation range.

[0093] Overall performance metrics may include any one of the following: Total Transactions Per Second (TPS), Total Response Time (RT), and Total Error Rate (ERR).

[0094] Within a historical period, the overall operating metrics of the system cluster at multiple points in time can be obtained. The overall operating metrics at each point in time are equal to the weighted sum of the corresponding instance operating metrics of each system instance at that point in time, according to the instance weights.

[0095] For example, given instances X, Y, and Z with corresponding instance weights of 0.4, 0.3, and 0.3 respectively, the overall trading volume index at a given time (Time0) is calculated as follows: Overall trading volume index at Time0 = Instance trading volume index of instance X at Time0 * Instance weight of instance X (0.4) + Instance trading volume index of instance Y at Time0 * Instance weight of instance Y (0.3) + Instance trading volume index of instance Z at Time0 * Instance weight of instance Z (0.3).

[0096] Different overall threshold matching conditions can be determined based on different overall operational metrics. As some examples, based on the overall transaction volume (TPS) metric obtained over a historical period, the overall threshold matching condition can be determined as follows: the overall transaction volume metric of the system cluster should be between 100,000 transactions per second ± 10,000 transactions per second, where 100,000 transactions per second is the overall threshold and 10,000 transactions per second is the overall fluctuation range. If the overall transaction volume metric falls within this range, it meets the overall threshold matching condition; if the overall transaction volume metric falls outside this range, it does not meet the overall threshold matching condition.

[0097] Alternatively, the overall threshold matching condition could be: the overall transaction volume of the system cluster should be between 100,000 transactions per second and ±10%, where 100,000 transactions per second is the overall threshold and 10% is the overall fluctuation range. If the overall transaction volume falls within this range, it meets the overall threshold matching condition; if the overall transaction volume falls outside this range, it does not meet the overall threshold matching condition.

[0098] The construction methods and working principles of Long Short-Term Memory (LSTM) network models can be found in relevant existing technologies, and will not be elaborated here.

[0099] After determining the overall threshold matching conditions, the process of determining the instance threshold matching conditions for system instances includes:

[0100] 1. Determine the overall operating metrics of the system cluster within the historical period based on the instance operation metrics and instance weights of multiple system instances collected during the historical period.

[0101] 2. Determine the overall threshold matching conditions for the system cluster based on the overall operating indicators of the system cluster within the historical period;

[0102] 3. Determine the instance threshold matching conditions for each system instance based on the overall threshold matching conditions of the system cluster and the instance weight of each system instance.

[0103] The implementation methods for steps 1 and 2 are described in the aforementioned embodiments.

[0104] In step 3, for each system instance, the overall threshold contained in the overall threshold matching condition can be multiplied by the instance weight of the system instance, and the result can be used as the instance threshold of the system instance.

[0105] Then, if the overall fluctuation range is expressed as a percentage (e.g., 10% as mentioned above), the overall fluctuation range can be directly determined as the instance fluctuation range. If the overall fluctuation range is expressed as an absolute value (e.g., 10,000 times per second as mentioned above), the overall fluctuation range can also be multiplied by the instance weight of the system instance to obtain the instance fluctuation range of the system instance.

[0106] Finally, the instance threshold matching condition for the system instance is obtained based on the instance threshold and the instance fluctuation range of the system instance.

[0107] For example, if the overall threshold for the overall transaction volume metric is 100,000 transactions per second, and the instance weight of instance X is 0.4, then the instance threshold for instance X in the instance transaction volume metric is 10 * 0.4, which is 40,000 transactions per second. The instance fluctuation range for instance X, expressed as a percentage of the overall fluctuation range, remains at 10%. Therefore, the instance threshold matching condition for instance X in the instance transaction volume metric is:

[0108] The instance transaction volume metric for instance X should be between 40,000 transactions per second ± 10%, i.e., between 36,000 and 44,000 transactions per second. Here, 40,000 transactions per second is the instance threshold, and 10% is the instance fluctuation range. If the instance transaction volume metric falls within this range, it meets the instance threshold matching condition. If the instance transaction volume metric falls outside this range, it does not meet the instance threshold matching condition.

[0109] In some optional embodiments, the aforementioned Long Short-Term Memory Network model or the moving average method can be used to directly process the instance operation indicators of any system instance in a historical period to obtain the instance threshold and instance fluctuation range of the system instance, and then determine the instance threshold matching condition of the system instance. Alternatively, the Long Short-Term Memory Network model and / or the moving average method can be used to directly process the instance operation indicators of a system instance in a historical period, and the instance threshold matching condition of the system instance can be adjusted based on the processing results, for example, from 40,000 times per second ± 10% to 45,000 times per second ± 10%.

[0110] Optionally, multiple system instances can ensure cross-instance data consistency through asynchronous data synchronization (real-time within the same city, near real-time across different locations) and distributed transaction mechanisms (such as using Seata), thereby enhancing the data consistency guarantee of the system cluster.

[0111] The long short-term memory network model used to obtain the overall threshold and the overall fluctuation range, and the long short-term memory network model used to determine the instance weights, can be the same model or two different models.

[0112] In step S103, if the overall operating metrics meet the overall threshold matching conditions of the system cluster, and the instance operating metrics do not meet the instance threshold matching conditions and are triggered, it indicates that the current system cluster is basically in a normal operating state, with only a few system instances experiencing fluctuations. These fluctuations may be triggered by the operation and maintenance personnel's active operation, such as when actively performing traffic switching, which may cause such fluctuations. Therefore, the notification information triggered by the instance operating metrics not meeting the instance threshold matching conditions (such as the aforementioned warning information or alarm information) can be deleted. In other words, the notification information triggered in this situation will not be displayed to the user through the terminal device to avoid excessive notification information interfering with the normal operation of the system cluster and to avoid excessive notification information affecting the normal use of other functions by the operation and maintenance personnel.

[0113] Optionally, if the instance performance metrics of any system instance do not meet the instance threshold matching conditions of the system instance, but the overall performance metrics meet the overall threshold matching conditions of the system cluster, then after deleting the notification information triggered by the instance performance metrics not meeting the instance threshold matching conditions, it shall also include at least one of the following:

[0114] The first system instance to be expanded is processed for expansion. The first system instance is determined from multiple system instances based on the instance's operational metrics.

[0115] Output traffic switching prompts. Traffic switching prompts are used to indicate that access traffic is switched from the second system instance to the first system instance. The second system instance is determined among multiple system instances based on instance operation metrics, and the access traffic to the second system instance is less than the access traffic to the first system instance.

[0116] The first system instance can be determined by querying the system instance with the largest increase in the instance transaction volume indicator, or the system instance with the largest increase in the instance response time indicator, or the system instance with the largest increase in the instance error rate indicator, and then determining the system instance found in the query as the first system instance.

[0117] The second system instance can be determined by querying the system instance with the largest decrease in transaction volume, or the system instance with the largest decrease in response time, or the system instance with the largest decrease in error rate, and then determining the system instance found in the query as the second system instance.

[0118] In the above embodiments, "reduction" means that the instance performance metrics after it is discovered that the instance performance metrics of any system instance do not meet the instance threshold matching conditions of the system instance are reduced compared to the instance performance metrics before the discovery of the situation, and "increase" means that the instance performance metrics after it is discovered that the instance performance metrics of any system instance do not meet the instance threshold matching conditions of the system instance are increased compared to the instance performance metrics before the discovery of the situation.

[0119] Optionally, the alarm rules for the second system instance can be turned off for a period of time. That is, no matter how the instance operation indicators of the second system instance change in the future, the alarm rules will not respond to the change and trigger the corresponding notification information, so as to avoid misjudgment again.

[0120] The specific methods for scaling up can be found in existing technologies and will not be elaborated here. Specifically, the scaling ratio can be determined based on the historical maximum load of the first system instance, such as the historical maximum instance transaction volume or historical maximum instance response time. The higher the historical maximum load, the greater the scaling ratio.

[0121] The specific content of the traffic switching prompt is not limited, as long as it indicates that access requests to the second system instance will be forwarded to the first system instance, that is, the traffic of the second system instance will be switched to the first system instance.

[0122] The following example illustrates the execution process of the method in this embodiment:

[0123] The system cluster consists of three system instances deployed in cities X, Y, and Z, respectively, referred to as instance X, instance Y, and instance Z. Instance X has an instance weight of 0.4 (40%), instance Y has an instance weight of 0.3, and instance Z has an instance weight of 0.3 (30%), with a ratio of 4:3:3.

[0124] After processing the overall operating indicators of the system cluster based on the Long Short-Term Memory network model over historical periods, the overall threshold matching condition is obtained as follows: the overall transaction volume indicator is within 100,000 transactions per second ± 10%.

[0125] Based on the instance weight of instance X (0.4), the instance threshold matching condition for instance X is determined as follows: the instance transaction volume index is 40,000 transactions per second ± 10%.

[0126] Based on the instance weight of instance Y (0.3), the instance threshold matching condition for instance Y is determined as follows: the instance transaction volume index is 30,000 times per second ± 10%.

[0127] During the real-time collection of instance operation metrics of each system instance according to S101, at a certain moment it was found that the instance transaction volume metric of instance X fluctuated drastically, dropping by 90%, from the original 40,000 times per second to 4,000 times per second. In addition, it was found that the instance transaction volume metric of another system instance fluctuated in the opposite direction. For example, it was found that the instance transaction volume metric of instance Y surged by 85%, from the original 30,000 times per second to 55,500 times per second.

[0128] Combining the instance threshold matching conditions of instance X and instance Y mentioned above, it can be found that the instance operation indicators (i.e. instance transaction volume indicators) of the current instance X and instance Y do not meet the instance threshold matching conditions of these two system instances.

[0129] Therefore, the current overall operating indicators are obtained, namely the overall transaction volume indicators. The overall transaction volume indicators = 4,000 times per second × 40% + 55,500 times per second × 30% + 30,000 times per second × 30% = 27,250 times per second, where 30,000 times per second is the instance transaction volume indicator of the current Z instance.

[0130] The current overall operating metric of 27,250 transactions per second still meets the overall transaction volume threshold of 100,000 transactions per second and is not exceeded. Therefore, it is determined that the notification information caused by the change in transaction volume of the current X instance and Y instance is a false alarm.

[0131] Therefore, S103 is executed to delete the notification information for instance X, that is, to delete the alarm information triggered by the sudden drop in the instance transaction volume indicator of instance X.

[0132] At the same time, the Y instance with the largest increase in instance transaction volume index can be identified as the first system instance, and the X instance with the largest decrease in instance transaction volume index can be identified as the second system instance. The first system instance can be scaled up, for example, by 180%, so that its processing capacity is increased to 180% of the original capacity.

[0133] Finally, output the traffic switching prompt: "Traffic has been switched from instance X to instance Y"; afterwards, the overall threshold of 100,000 times per second (e.g., 100,000 times / second) corresponding to the overall transaction volume indicator can be used as a reference for the running status.

[0134] In some alternative embodiments, the method for determining the instance threshold matching condition and the overall threshold matching condition may also be:

[0135] First, for each system instance, refer to the method described above for determining the overall threshold and overall fluctuation range using a long short-term memory network model. Use the long short-term memory network model to process the instance operation indicators of the system instance in the historical period to obtain the instance threshold and instance fluctuation range of the system instance, where the instance fluctuation range is expressed as a percentage.

[0136] After obtaining the instance thresholds and instance fluctuation ranges of all system instances using the above method, the overall threshold is obtained by weighting all instance thresholds according to the instance weights corresponding to their respective system instances; the overall fluctuation range is obtained by weighting all instance fluctuation ranges according to the instance weights corresponding to their respective system instances.

[0137] Continuing with the previous example, assuming that for the instance transaction volume metric, the instance threshold for instance X is 4, and the instance thresholds for instances Y and Z are both 3, after weighting and summing according to the aforementioned instance weights, the overall threshold is 3.4. The instance fluctuation range for instances X, Y, and Z is 10%, so the overall fluctuation range is also 10%. Therefore, the overall threshold matching condition is that the overall transaction volume metric is between 34,000 times per second ± 10%.

[0138] At a certain moment, it was discovered that the instance transaction volume metric of instance X decreased from 40,000 times per second to 4,000 times per second, while the instance transaction volume metric of instance Y surged by 130%, from 30,000 times per second to 69,000 times per second. The instance transaction volume metric of instance Z remained at 30,000 times per second. The overall transaction volume metric was obtained by weighting and summing the instance transaction volume metrics of the three system instances according to their instance weights, i.e., the overall transaction volume metric = 0.4 * 0.4 + 6.9 * 0.3 + 3 * 0.3 = 31,300 times per second.

[0139] After comparison, it was found that the overall transaction volume indicator was still within 34,000 times per second ± 10%, which met the overall threshold matching condition. Therefore, the notification information for instance X could be deleted. Other processing methods can be found in the previous text and will not be repeated here.

[0140] This method avoids the need for post-event manual adjustments to existing monitoring rule thresholds based on historical data through a dynamic threshold adaptive adjustment mechanism, effectively improving adjustment efficiency and significantly reducing labor costs. Furthermore, it breaks down fragmented monitoring through global health status verification, achieving multi-instance collaborative status determination through weighted overall indicators, avoiding misjudgments caused by isolated monitoring. It also plays a crucial role in suppressing false alarms and improving operational efficiency, effectively reducing false warnings / alarms and enhancing the operational efficiency of frontline personnel. In addition, it optimizes resource utilization and business continuity. Through weight-driven threshold calculation, it achieves precise matching of resource allocation and traffic capacity. Dynamic adjustment strategies reduce resource waste in edge computing disaster recovery while ensuring the continuity of critical services. It also enables intelligent decision-making and scenario-based linkage; the system automatically issues a "traffic has switched from A to B" notification, using the overall threshold as an operational reference, reducing the complexity of manual intervention.

[0141] In summary, this solution addresses the core pain points of traditional multi-active architectures, such as the disconnect between local and global monitoring, high false alarm rate of static thresholds, and delayed switching response, through three major innovations: dynamic thresholds, global verification, and intelligent decision-making. It provides a quantifiable, adaptive, and integrated monitoring paradigm for financial-grade high-availability systems.

[0142] Furthermore, the method in this embodiment can effectively reduce false warnings / alarms generated during the application system switching process, and effectively improve the monitoring capability of the overall switching process.

[0143] Specifically, one approach is to analyze historical switching data based on an LSTM model and dynamically update the threshold. However, this relies heavily on historical data reserves, and insufficient data reserves will lead to deviations in the accuracy of the threshold.

[0144] Secondly, it cannot fully cover unconventional transaction fluctuations such as shutdowns, switching, and marketing that have not occurred in the past, and the dynamic adjustment rules for thresholds lack universality.

[0145] Thirdly, the mechanisms and methods that can be established based on existing application system architecture patterns are difficult to fully adapt to the rapidly changing high-availability models in the financial field, and lack flexibility.

[0146] The long short-term memory network model used in any embodiment of this application can be implemented based on mechanisms such as adaptive forget gate, gradient-aware learning rate, and time-varying parameter generation. The adaptive forget gate mechanism can dynamically adjust the forget gate bias according to the statistical characteristics of the input sequence. The gradient-aware learning rate mechanism can dynamically adjust the parameter update step size through the gradient magnitude. The time-varying parameter generation mechanism can use an auxiliary network to generate weight parameters.

[0147] Specifically, the adaptive forget gate mechanism can calculate the forget gate bias according to the following formula (1).

[0148]

[0149] Among them BF t This indicates the adjusted forget gate bias, BF t-1 The value of the initial forget gate bias (BF0) can be set by the maintenance personnel when t-1 equals 0. A is the preset learnable scaling factor. N is the sliding window size, which is a preset hyperparameter. The specific value can be set as needed without limitation. t can be incremented from 1, so that the adaptive forget gate mechanism iteratively adjusts the forget gate bias according to formula (1) until the maximum value is reached. The maximum value of t can be set as needed without limitation.

[0150] |x t-i | represents the absolute value of the ti-th data in the input sequence, reflecting the recent input amplitude.

[0151] If the Long Short-Term Memory (LSTM) network model is used to determine instance weights, the input sequence includes the aforementioned weight input information obtained within the historical time period; if the LTM network model is used to obtain the overall threshold and overall fluctuation range of the system cluster, the input sequence includes the overall operating indicators of the system cluster within the historical time period.

[0152] The derivation process of the above formula (1) is shown in formula (2).

[0153]

[0154] Where ∝ indicates that the data on the left is proportional to the data on the right, and L represents the model loss of the Long Short-Term Memory network model. For specific calculation methods, please refer to existing technologies. This indicates calculating the gradient, c t and c t-1 These represent the information retained at the t-th and t-1-th iterations when iterating according to formula (1), respectively. Their values ​​are determined by the model parameters and input sequence of the Long Short-Term Memory network model. F t This represents the feature value calculated based on the (t-1)th data point and the data preceding it in the input sequence. A larger value indicates less forgetting, while a smaller value indicates more forgetting.

[0155] It can be seen that if the recent input amplitude is large when the forget gate bias is adjusted according to formula (1), it indicates that the sequence fluctuates violently and the forget gate needs to be strengthened to discard more historical information.

[0156] In the gradient-aware learning rate mechanism, the parameters of the long short-term memory network model are updated using an adaptive learning rate, and the update process is shown in formula (3).

[0157]

[0158] Theta t+1 and Theta t Let Theta0 represent the model parameters of the Long Short-Term Memory (LSTM) network model after the (t+1)th update and the model parameters of the LSM network model after the tth update (equivalent to before the (t+1)th update), respectively. The value of Theta0 can be determined through random initialization. ▽Theta represents the gradient of the LSM network model parameters, and L represents the LSM network model loss.

[0159] Ln t Let t represent the adaptive learning rate determined in the t-th iteration, which can be determined according to formula (4).

[0160]

[0161] ||▽Theta*L p || represents the mean of the historical gradient norm at the p-th iteration, L p This represents the model loss of the Length Short-Term Memory (SSM) network model at the p-th iteration. B is a preset smoothing factor, a hyperparameter whose specific value can be set by the operations and maintenance personnel as needed. Ln0 is the preset initial learning rate. exp() indicates that the value in parentheses is used as the exponent, and the base of the natural logarithm, e, is used as the base for exponentiation.

[0162] The derivation process of the above formula (4) is shown in formula (5).

[0163]

[0164] E[||▽L||] represents the expected value of the model loss of the Long Short-Term Memory (LSTM) network model. It can be seen that when the gradient magnitude is large, the learning rate can be reduced to avoid oscillations. The exponential decay form can ensure that Ln... t It is always greater than 0 and changes smoothly.

[0165] The time-varying parameter generation mechanism can use a lightweight auxiliary network, such as a multilayer perceptron (MLP), to generate weights. The specific generation process can be represented by the following formula (6).

[0166]

[0167] Where Pai represents the preset auxiliary network parameters, and W... t * indicates the determined instance weights. MLP is a pre-built multilayer perceptron; its specific form and construction method can be found in existing technologies. Z t The mean and standard deviation of recent inputs can be represented by the following formula (7).

[0168]

[0169] avg(x) t-k:t std(x) represents the average of the data from the tk-th data point to the t-th data point in the input sequence. t-k:t ) represents calculating the standard deviation from the tkth data point to the tth data point in the input sequence.

[0170] The derivation of the above formula (6) is as follows: the instance weights need to adapt to the local statistical characteristics of the sequence. Assuming that the input sequence segment follows a Gaussian distribution, then the instance weights W determined by formula (6) are... t * Satisfies formula (8).

[0171]

[0172] Where u is the mean of the Gaussian distribution, Sigma is the standard deviation of the Gaussian distribution, and F is the fitting function, the specific form of which and the parameters it contains are determined by the MLP.

[0173] The advantages of applying the above-mentioned long short-term memory network model are as follows: by utilizing the dynamic forgetting mechanism, the forgetting gate can be adjusted according to the input amplitude, enhancing the adaptability to mutation sequences; the gradient-aware learning rate mechanism can reduce oscillations, improve the stability of the training process, and accelerate convergence; the time-varying parameter generation mechanism helps to adapt to the non-stationary characteristics of sequences and improve the accuracy of the long short-term memory network model output.

[0174] If the Long Short-Term Memory (LSTM) network model is used to determine instance weights, the model output includes the instance weights of the system instances; if the LTM network model is used to obtain the overall threshold and overall fluctuation range of the system cluster, the model output includes the overall threshold and overall fluctuation range of the system cluster.

[0175] This application also provides a processing device for threshold notification information based on cluster traffic. Please refer to [link to relevant documentation]. Figure 2 The device may include the following units.

[0176] The acquisition unit 201 is used to collect instance operation metrics of multiple system instances contained in the system cluster in real time;

[0177] The obtaining unit 202 is used to obtain the overall operation index obtained by calculating the operation index of each instance based on the preset instance weight of each system instance when the instance operation index of any system instance does not meet the instance threshold matching condition of the system instance. The instance threshold matching condition is determined based on the instance operation index of the system instance collected in the historical period.

[0178] The processing unit 203 is used to delete notification information triggered by instance operating indicators not meeting instance threshold matching conditions, provided that the overall operating indicators meet the overall threshold matching conditions of the system cluster. The overall threshold matching conditions are determined based on the overall operating indicators obtained in the historical period.

[0179] Optionally, the process of obtaining unit 202 to determine the instance operation metrics of the system instance and to determine the instance threshold matching conditions of the system instance includes:

[0180] Based on the instance operation metrics and instance weights of multiple system instances collected during historical periods, the overall operation metrics of the system cluster during historical periods are determined.

[0181] Determine the overall threshold matching conditions for the system cluster based on the overall operating indicators of the system cluster within a historical period.

[0182] The instance threshold matching conditions for each system instance are determined based on the overall threshold matching conditions of the system cluster and the instance weight of each system instance.

[0183] Optionally, the obtaining unit 202 determines the overall threshold matching conditions of the system cluster based on the overall operating indicators of the system cluster within the historical period, including:

[0184] The system cluster's overall operating metrics within a historical period are processed using a pre-built long short-term memory network model to obtain the overall threshold and overall fluctuation range of the system cluster. Based on the overall threshold and overall fluctuation range, the overall threshold matching conditions are determined. The overall threshold matching conditions include that the difference between the overall operating metrics of the system cluster and the overall threshold is within the overall fluctuation range.

[0185] Optionally, after deleting the notification information triggered by the instance's running metrics failing to meet the instance threshold matching condition, the processing unit 203 is also configured to perform at least one of the following:

[0186] The first system instance to be expanded is processed for expansion. The first system instance is determined from multiple system instances based on the instance's operational metrics.

[0187] Output traffic switching prompts. Traffic switching prompts are used to indicate that access traffic is switched from the second system instance to the first system instance. The second system instance is determined among multiple system instances based on instance operation metrics, and the access traffic to the second system instance is less than the access traffic to the first system instance.

[0188] Optionally, the apparatus of this embodiment may further include a determining unit 204, which determines the instance weights of a plurality of system instances, specifically for performing at least one of the following:

[0189] The instance weights of multiple system instances are determined based on the proportion of access traffic of multiple system instances in the overall access traffic of the system cluster.

[0190] The instance weights of multiple system instances are determined based on the business type data of multiple system instances;

[0191] The instance weights of multiple system instances are determined based on the changing trends of their instance operation metrics over historical periods.

[0192] Based on historical switching data of multiple system instances, the instance weights of multiple system instances are determined, and the historical switching data represents the number of traffic switching between multiple system instances.

[0193] The specific working principle of the threshold notification information processing device based on cluster traffic can be found in the relevant steps of the threshold notification information processing method based on cluster traffic in the foregoing embodiments, and will not be repeated here.

[0194] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0195] For ease of description, the above systems or devices are described separately as various modules or units based on their functions. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware components.

[0196] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0197] Finally, it should be noted that in this document, relational terms such as first, second, third, and fourth are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0198] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method of processing information for threshold notification based on cluster traffic, characterized by, The method comprises: collecting instance running indexes of a plurality of system instances in a system cluster; in a case where an instance running index of any of the system instances does not meet an instance threshold matching condition of the system instance, obtaining an overall running index calculated based on preset instance weights of the system instances and the instance running indexes of the system instances, the instance threshold matching condition being determined according to the instance running indexes of the system instances collected in a historical period; in a case where the overall running index meets an overall threshold matching condition of the system cluster, deleting notification information triggered by the instance running index not meeting the instance threshold matching condition, the overall threshold matching condition being determined according to the overall running index obtained in the historical period.

2. The method of claim 1, wherein, The method for determining the instance threshold matching condition of the system instance comprises: determining an overall running index of the system cluster in the historical period according to the instance running indexes of the plurality of system instances collected in the historical period and the instance weights of the plurality of system instances; determining an overall threshold matching condition of the system cluster according to the overall running index of the system cluster in the historical period; determining an instance threshold matching condition of each of the system instances according to the overall threshold matching condition of the system cluster and the instance weight of each of the system instances.

3. The method of claim 2, wherein, The determining of the overall threshold matching condition of the system cluster according to the overall running index of the system cluster in the historical period comprises: processing the overall running index of the system cluster in the historical period according to a long short-term memory network model constructed in advance to obtain an overall threshold and an overall floating range of the system cluster, so as to determine the overall threshold matching condition based on the overall threshold and the overall floating range, the overall threshold matching condition comprising that a difference between the overall running index of the system cluster and the overall threshold is within the overall floating range.

4. The method of claim 1, wherein, After the deleting of the notification information triggered by the instance running index not meeting the instance threshold matching condition, the method further comprises at least one of the following: performing expansion processing on a first system instance to be expanded, the first system instance being determined among the plurality of system instances based on the instance running index; outputting a traffic switching prompt, the traffic switching prompt being used to indicate that access traffic is switched from a second system instance to the first system instance, the second system instance being determined among the plurality of system instances based on the instance running index, and an access traffic for the second system instance being less than an access traffic for the first system instance.

5. The method of claim 1, wherein, The method further comprises at least one of the following: determining the instance weights of the plurality of system instances according to a proportion of access traffics of the plurality of system instances in overall access traffic of the system cluster; determining the instance weights of the plurality of system instances according to business type data of the plurality of system instances; determining the instance weights of the plurality of system instances according to a change trend of the instance running indexes of the plurality of system instances in the historical period; determining the instance weights of the plurality of system instances according to historical switching data of the plurality of system instances, the historical switching data representing a number of traffic switching times between the plurality of system instances.

6. A processing device for threshold notification information based on cluster traffic, characterized in that, The method comprises: The collection unit is configured to collect instance running indexes of a plurality of system instances in a system cluster in real time; The obtaining unit is configured to, in a case where the instance running index of any system instance does not meet an instance threshold matching condition of the system instance, obtain an overall running index calculated based on preset instance weights of each system instance and each instance running index, the instance threshold matching condition being determined according to the instance running index of the system instance collected in a historical period; The processing unit is configured to, in a case where the overall running index meets an overall threshold matching condition of the system cluster, delete notification information triggered by the fact that the instance running index does not meet the instance threshold matching condition, the overall threshold matching condition being determined according to the overall running index obtained in the historical period.

7. The apparatus of claim 6, wherein, The obtaining unit is configured to determine the instance threshold matching condition of the system instance, and the process of determining the instance threshold matching condition of the system instance includes: determining an overall running index of the system cluster in the historical period according to the instance running indexes of the plurality of system instances and the instance weights of the plurality of system instances collected in the historical period; determining an overall threshold matching condition of the system cluster according to the overall running index of the system cluster in the historical period; determining an instance threshold matching condition of each system instance according to the overall threshold matching condition of the system cluster and the instance weight of each system instance.

8. The apparatus of claim 7, wherein, The obtaining unit determines the overall threshold matching condition of the system cluster according to the overall running index of the system cluster in the historical period, including: processing the overall running index of the system cluster in the historical period according to a long short-term memory network model constructed in advance to obtain an overall threshold and an overall floating range of the system cluster, so as to determine the overall threshold matching condition based on the overall threshold and the overall floating range, the overall threshold matching condition including that a difference between the overall running index of the system cluster and the overall threshold is within the overall floating range.

9. The apparatus of claim 6, wherein, After the processing unit deletes the notification information triggered by the fact that the instance running index does not meet the instance threshold matching condition, the processing unit is further configured to perform at least one of the following: performing expansion processing on a first system instance to be expanded, the first system instance being determined among the plurality of system instances based on the instance running index; outputting a traffic switching prompt, the traffic switching prompt being used to instruct to switch access traffic from a second system instance to the first system instance, the second system instance being determined among the plurality of system instances based on the instance running index, and the access traffic for the second system instance being less than the access traffic for the first system instance.

10. The apparatus of claim 6, wherein, The determination unit is further configured to determine the instance weights of the plurality of system instances, and specifically configured to perform at least one of the following: determining the instance weights of the plurality of system instances according to a proportion of access traffic of the plurality of system instances in overall access traffic of the system cluster; determining the instance weights of the plurality of system instances according to business type data of the plurality of system instances. determining an instance weight of the multiple system instances according to a variation trend of the instance running indexes of the multiple system instances in the historical period; determining an instance weight of the multiple system instances according to historical switching data of the multiple system instances, the historical switching data representing a number of times of traffic switching between the multiple system instances.