Self-adaptive flow control method and device
By collecting and analyzing real-time monitoring data of the basic system, combining resource redundant data and historical water level lines, dynamically adjusting the flow control strategy, solving the problem that the flow control strategy cannot be flexibly adjusted in the existing technology, achieving efficient and automated flow control, and improving system operation and maintenance efficiency and customer experience.
Patent Information
- Application Number
- CN202410742919.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-11
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, the traffic control strategy cannot be flexibly adjusted to adapt to changes in the system's operating status, and lacks perception of underlying resources, cannot perform emergency responses based on resource conditions, and cannot adjust the traffic control strategy based on customer feelings.
By collecting real-time monitoring data of multiple basic systems, determining the system type, using preset triggering rules to make threshold and trend judgments. If flow control is triggered, flow control will be performed based on resource redundant data and historical water level lines.
Dynamic flow control is realized, resource utilization efficiency is improved, traffic control strategies are timely optimized, customer experience is improved, and the completeness and automation of flow control is improved, and system operation and maintenance efficiency is improved.
Smart Images

Figure CN120050233A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of software system control and intelligent operation and maintenance, and particularly to an adaptive traffic control method and device. Background Art
[0002] With the continuous enrichment of banking services, the continuous increase in transaction volume, and the combination of distributed transformation, the related service monitoring capabilities have become the basic capabilities of the system to ensure the overall stability of the system. However, the current traffic control strategies are basically based on fixed thresholds for traffic control and cannot be flexibly adjusted according to the changes in the operating state of the system. Secondly, there is a lack of awareness of the underlying resources and no way to perform emergency handling according to the situation of the underlying resources. In addition, there is also the problem of lacking the ability to adjust the traffic control strategy according to the actual feelings of customers. Summary of the Invention
[0003] Aiming at the problems existing in the prior art, the main purpose of the embodiments of the present invention is to provide an adaptive traffic control method and device, which can realize dynamic traffic control following the actual operation situation and improve the resource utilization efficiency.
[0004] To achieve the above purpose, the embodiments of the present invention provide an adaptive traffic control method, and the method includes:
[0005] Collect real-time monitoring data of multiple basic systems and determine the system types corresponding to the basic systems;
[0006] Use a preset trigger rule to perform threshold determination and trend determination on the real-time monitoring data to obtain a trigger determination result;
[0007] If the trigger determination result is to trigger traffic control, then according to the system types corresponding to the basic systems, use the real-time monitoring data to determine resource redundancy data, and use the historical monitoring data of the basic systems to determine the corresponding historical water level lines;
[0008] Perform traffic control on each basic system according to the resource redundancy data and the historical water level lines.
[0009] Optionally, in an embodiment of the present invention, collecting real-time monitoring data of multiple basic systems includes:
[0010] Collect access logs of the log center, performance indicators of the distributed service system, performance indicators of the PAAS console, and performance monitoring data of the performance monitoring platform, and use the access logs, performance indicators of the distributed service system, performance indicators of the PAAS console, and performance monitoring data as real-time monitoring data.
[0011] Optionally, in an embodiment of the present invention, using a preset trigger rule, threshold determination and trend determination are performed on real-time monitoring data, and the obtained trigger determination result includes:
[0012] Using a preset trigger rule, trigger current-limiting anomaly threshold determination, concurrent number average determination, and throughput average determination are performed on real-time monitoring data to obtain a threshold determination result;
[0013] Using a preset trigger rule, trend determination results are obtained for real-time monitoring data, and based on the threshold determination result and the trend determination result, a trigger determination result is obtained.
[0014] Optionally, in an embodiment of the present invention, the system types include channel application systems and background service application systems.
[0015] Optionally, in an embodiment of the present invention, according to the system type corresponding to each basic system, resource redundancy data is determined using real-time monitoring data, and the corresponding historical water level line is determined using the historical monitoring data of each basic system, including:
[0016] If the system type of the basic system is a channel application system, then using real-time monitoring data, the current node state and the current system state are determined; wherein, the current node state includes CPU usage rate, memory usage rate, and throughput, and the current system state includes virtual machine capacity and physical machine capacity;
[0017] According to the current node state and the current system state, resource redundancy data is determined; wherein, the resource redundancy data includes CPU redundancy, virtual machine redundancy, and physical machine redundancy;
[0018] Obtain the historical monitoring data of each basic system within a preset time period, and use the historical monitoring data to determine the historical water level line.
[0019] Optionally, in an embodiment of the present invention, according to the system type corresponding to each basic system, resource redundancy data is determined using real-time monitoring data, and the corresponding historical water level line is determined using the historical monitoring data of each basic system, including:
[0020] If the system type of the basic system is a background service application system, then using real-time monitoring data, the current node state and the current system state are determined; wherein, the current node state includes CPU usage rate, memory usage rate, and throughput, and the current system state includes virtual machine capacity and physical machine capacity;
[0021] According to the current node state and the current system state, resource redundancy data is determined; wherein, the resource redundancy data includes CPU redundancy, virtual machine redundancy, and physical machine redundancy;
[0022] Obtain the historical monitoring data of each basic system within a preset time period, and use the historical monitoring data to determine the historical water level line;
[0023] Track the associated call chains of each basic system, and determine the corresponding resource redundancy data in parallel for all services and nodes on the associated call chains.
[0024] Optionally, in an embodiment of the present invention, performing traffic control on each basic system according to the resource redundancy data and the historical water level line includes:
[0025] Compare the real-time monitoring data with the historical water level line to obtain a comparison result;
[0026] Perform traffic control on each basic system according to the resource redundancy data and the comparison result; wherein, the traffic control includes expanding buffer processing.
[0027] An embodiment of the present invention further provides an adaptive traffic control device, and the device includes:
[0028] A monitoring data module, configured to collect real-time monitoring data of multiple basic systems and determine the corresponding system types of each basic system;
[0029] A trigger determination module, configured to use a preset trigger rule to perform threshold determination and trend determination on the real-time monitoring data to obtain a trigger determination result;
[0030] A redundancy data module, configured to, if the trigger determination result is to trigger traffic control, determine resource redundancy data according to the system types corresponding to each basic system by using the real-time monitoring data, and determine the corresponding historical water level line by using the historical monitoring data of each basic system;
[0031] A traffic control module, configured to perform traffic control on each basic system according to the resource redundancy data and the historical water level line.
[0032] Optionally, in an embodiment of the present invention, the monitoring data module is further configured to collect access logs of the log center, performance metrics of the distributed service system, performance metrics of the PAAS console, and performance monitoring data of the performance monitoring platform, and use the access logs, performance metrics of the distributed service system, performance metrics of the PAAS console, and performance monitoring data as real-time monitoring data.
[0033] Optionally, in an embodiment of the present invention, the trigger determination module includes:
[0034] A threshold determination unit, configured to use a preset trigger rule to perform trigger current limiting anomaly threshold determination, concurrent number average value determination, and throughput average value determination on the real-time monitoring data to obtain a threshold determination result;
[0035] A trend determination unit is used to determine a trend determination result for real-time monitoring data by using a preset trigger rule, and obtain a trigger determination result according to the threshold determination result and the trend determination result.
[0036] Optionally, in an embodiment of the present invention, the system types include a channel application system and a background service application system.
[0037] Optionally, in an embodiment of the present invention, the redundant data module includes:
[0038] A first status unit is used to determine the current node status and the current system status by using real-time monitoring data if the system type of the basic system is a channel application system; wherein, the current node status includes CPU usage rate, memory usage rate, and throughput, and the current system status includes virtual machine capacity and physical machine capacity;
[0039] A first redundancy unit is used to determine resource redundancy data according to the current node status and the current system status; wherein, the resource redundancy data includes CPU redundancy, virtual machine redundancy, and physical machine redundancy;
[0040] A first water level unit is used to obtain historical monitoring data of each basic system within a preset time period, and determine a historical water level by using the historical monitoring data.
[0041] Optionally, in an embodiment of the present invention, the redundant data module further includes:
[0042] A second status unit is used to determine the current node status and the current system status by using real-time monitoring data if the system type of the basic system is a background service application system; wherein, the current node status includes CPU usage rate, memory usage rate, and throughput, and the current system status includes virtual machine capacity and physical machine capacity;
[0043] A second redundancy unit is used to determine resource redundancy data according to the current node status and the current system status; wherein, the resource redundancy data includes CPU redundancy, virtual machine redundancy, and physical machine redundancy;
[0044] A second water level unit is used to obtain historical monitoring data of each basic system within a preset time period, and determine a historical water level by using the historical monitoring data;
[0045] An associated call chain unit is used to trace the associated call chain of each basic system, and determine corresponding resource redundancy data in parallel for all services and nodes on the associated call chain.
[0046] Optionally, in an embodiment of the present invention, the traffic control module includes:
[0047] A comparison result unit is used to compare the real-time monitoring data with the historical water level to obtain a comparison result;
[0048] A flow control unit is used to perform flow control on each basic system according to resource redundancy data and comparison results; wherein, the flow control includes expanding buffer processing.
[0049] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above method is implemented.
[0050] The present invention also provides a computer-readable storage medium, which stores a computer program for the computer to execute the above method.
[0051] The present invention also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0052] Through dynamic flow control following the actual operation situation, the present invention improves the resource utilization efficiency, timely optimizes the flow control strategy, quickly adjusts the system service ability within the scope of safety and controllability, enhances the customer experience, and also enhances the completeness and automation degree of flow control, improving the efficiency of system operation and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0054] Figure 1 It is a flowchart of an adaptive flow control method according to an embodiment of the present invention;
[0055] Figure 2 It is a flowchart of obtaining a trigger determination result in an embodiment of the present invention;
[0056] Figure 3 It is a flowchart of determining resource redundancy data in an embodiment of the present invention;
[0057] Figure 4 It is a flowchart of determining resource redundancy data in another embodiment of the present invention;
[0058] Figure 5 It is a flowchart of flow control in an embodiment of the present invention;
[0059] Figure 6 It is a schematic structural diagram of a system applying the adaptive flow control method in an embodiment of the present invention;
[0060] Figure 7 Schematic diagram of the service level in the embodiment of the present invention;
[0061] Figure 8 Schematic diagram of data analysis of the channel application system in the embodiment of the present invention;
[0062] Figure 9 Schematic diagram of data analysis of the background service application system in the embodiment of the present invention;
[0063] Figure 10 Schematic diagram of the structure of an adaptive traffic control device in the embodiment of the present invention;
[0064] Figure 11 Schematic diagram of the structure of the trigger determination module in the embodiment of the present invention;
[0065] Figure 12 Schematic diagram of the structure of the redundant data module in the embodiment of the present invention;
[0066] Figure 13 Schematic diagram of the structure of the redundant data module in another embodiment of the present invention;
[0067] Figure 14 Schematic diagram of the structure of the traffic control module in the embodiment of the present invention;
[0068] Figure 15 Schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed implementation manners
[0069] The embodiment of the present invention provides an adaptive traffic control method and device, which can be used in the fields of intelligent operation and maintenance, finance, and other fields. It should be noted that the adaptive traffic control method and device of the present invention can be used in the fields of intelligent operation and maintenance and finance, and can also be used in any field other than the fields of intelligent operation and maintenance and finance. The application fields of the adaptive traffic control method and device of the present invention are not limited.
[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0071] In the technical solution of the present invention, the information collected is information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations, and standards of the relevant countries and regions. Necessary confidentiality measures are taken, which do not violate public order and good customs, and a corresponding operation entry is provided for the user to choose to authorize or reject.
[0072] Provide a corresponding operation entry for the user to choose to agree or reject the automated decision-making result; if the user chooses to reject, then enter the expert decision-making process.
[0073] Such as Figure 1 As shown in the flowchart of an adaptive traffic control method according to an embodiment of the present invention, the execution subject of the adaptive traffic control method provided by the embodiment of the present invention includes but is not limited to a computer. The present invention improves the resource utilization efficiency by performing dynamic traffic control following the actual operation situation, timely optimizes the traffic control strategy, quickly adjusts the system service capacity within a safe and controllable range, enhances the customer experience, and also improves the completeness and automation degree of traffic control, and improves the efficiency of system operation and maintenance. The method shown in the figure includes:
[0074] Step S1, collect the real-time monitoring data of multiple basic systems, and determine the system type corresponding to each basic system.
[0075] Among them, the real-time monitoring data of each basic system is obtained in a conventional manner from basic management platforms such as the log center, distributed service monitoring, PAAS console, and performance monitoring platform, that is, basic systems. Specifically, after obtaining the authorization of the customer, collect the access logs of the customer from the log center, and count the success rate and failure rate of the customer's access; from the distributed service monitoring, monitor the performance indicators of the service; from the PAAS (Platform as a Service) console, obtain the performance indicators of the container and the corresponding relationship between the container and the underlying host; obtain the performance monitoring data at the virtualization layer and the physical machine layer, as well as the relevant indicators of resource contention from the performance monitoring platform.
[0076] Furthermore, the system type includes channel application systems and background service application systems. Specifically, there are differences in the process of analyzing real-time monitoring data between channel application systems and background service application systems. The background service application system has an associated call chain more than the channel application system to perform the process of iterative data analysis steps.
[0077] Step S2, use a preset trigger rule to perform threshold determination and trend determination on the real-time monitoring data to obtain a trigger determination result.
[0078] Among them, the preset trigger rules include a threshold determination rule and a trend determination rule. Specifically, the threshold determination rule includes three types of threshold determination strategies: the threshold for triggering traffic limiting anomalies, the average concurrency, and the average throughput.
[0079] Specifically, traffic limiting anomaly means that due to a large influx of customer requests, traffic limiting fails. When the proportion of traffic limiting failures exceeds a preset proportion of the total number of service processing transactions, it is determined that the service quality has declined due to traffic limiting, triggering the subsequent traffic control process, that is, the threshold determination result is to trigger traffic control. The average concurrency and the average throughput are based on the service monitoring capabilities of the distributed service framework to monitor and collect the concurrency and throughput metrics of the service for a preset number of hours. During the period when the service is externally provided, if (average concurrency ÷ concurrency upper limit) and (average throughput ÷ throughput upper limit) are higher than the preset proportion for three consecutive collection cycles (generally, the collection cycle is one day), it is determined that the capacity exceeds the warning value, triggering the subsequent traffic control process, that is, the threshold determination result is to trigger traffic control.
[0080] Furthermore, the trend determination rule is mainly based on the monitored metrics of concurrency and throughput. Through the service monitoring capabilities of the distributed service framework, the concurrency and throughput metrics of the service are monitored and collected for a preset number of hours. Several representative time points are selected from each trading day for linear regression from both vertical and horizontal dimensions. Taking the analysis process of concurrency as an example, the method of linear regression is to construct a linear regression model with the time dimension and the concurrency dimension. Specifically, time is used as the X horizontal axis and concurrency is used as the vertical axis, and data at multiple time nodes are collected for regression fitting. If the fitted line slopes upward linearly, the subsequent traffic control process is triggered, that is, the trend determination result is to trigger traffic control. The trend determination of throughput is the same.
[0081] Furthermore, the trigger determination result is obtained based on the threshold determination result and the trend determination result. Specifically, if any one of the threshold determination result and the trend determination result is to trigger traffic control, the trigger determination result is to trigger traffic control.
[0082] Step S3, if the trigger determination result is to trigger traffic control, then according to the system types corresponding to each basic system, the resource redundancy data is determined using the real-time monitoring data, and the corresponding historical water level line is determined using the historical monitoring data of each basic system.
[0083] Among them, after traffic control is triggered, if the system type is a channel application system, the real-time monitoring data is first used to determine the current node status and the current system status. Specifically, the current node status mainly refers to the performance capacity status of the corresponding node container, including but not limited to CPU usage, memory usage, disk and network IO throughput, etc. The representative data selected for evaluation are all peak usage indicators within a preset number of minutes. Specifically, the specific methods are divided into two categories. One is the judgment of thresholds. For example, if the TPS of disk IO does not exceed the preset threshold, it passes and enters the evaluation process of the next stage. The second is the judgment of redundant capacity. For example, the difference between the peak CPU usage of the current node within the preset number of minutes and the warning value of CPU production is used to obtain the quantitative output of redundant capacity. After continuing the subsequent traffic control, the quantitative output of redundant capacity is obtained as the current node status and used as resource redundant data.
[0084] Furthermore, after obtaining the current node status, drill down to the virtualization layer and the physical machine layer through the current node to evaluate the capacity redundancy of the virtual machine and the physical machine to obtain the current system status. Specifically, the evaluation method includes two parts. The first part is the general capacity evaluation, which is the same as the evaluation method of the current node status. The second part is that since resource overcommitment may exist in the virtual machine layer and the physical machine layer, resulting in resource contention, it is necessary to supplement and check whether the resource contention indicators exceed the threshold. From these two parts, the current system status is obtained and used as resource redundant data.
[0085] Furthermore, the determination of the historical water level line of each basic system, that is, the judgment of the historical capacity status of the node. Analyze the system to save the water level line of the node. The data of the water level line (i.e., historical monitoring data) is updated daily through the monitoring method. Based on the maximum value of the usage of each resource on the current day, including all the indicators in the above two steps of determining the current node status and the current system status, if it is judged that it is greater than the maximum value of the historical water level line, the historical water level line is updated; if it is less than or equal to the maximum value of the historical water level line, it is not updated. Thus, the historical water level line of each basic system can be accurately determined. In addition, based on the water level line and the warning value and threshold of the current production capacity, an analysis and judgment are made, and the judgment process is the same as the judgment method in the above current node status and current system status.
[0086] In addition, if the system type is a background service application system, after the three steps of the current node status, the current system status, and the historical water level line, it is also necessary to determine the associated adjustment points and perform iterative processing. Specifically, through the observable tracking method, the entire call chain of the service controlled by traffic is obtained, and then all the services and nodes on the call chain are associated. The three steps of determining the current node status, the current system status, and the historical water level line are performed in parallel. Considering the relevant capacity redundancy situation, if the current capacity redundancy evaluation of all services on the link is greater than the preset ratio, the traffic control parameter expansion is implemented. Then, the capacity redundancy of the historical water level line is judged. If there is a preset ratio for checking redundancy, the production operation and maintenance personnel are notified to evaluate the subsequent expansion process.
[0087] Step S4, perform traffic control on each basic system according to the resource redundancy data and the historical water level line.
[0088] Among them, according to the current node status and the current system status in the resource redundancy data, combined with the historical water level line and the judgment result of using it to determine the current capacity, traffic control can be performed on each basic system. Specifically, traffic control can be carried out by expanding the buffer. For example, if the redundancy percentage obtained from containers, virtual machines, and physical machines is less than the preset percentage, the traffic control parameters are not adjusted, and the relevant operation and maintenance personnel can be notified of the resource and traffic control situation, and it is recommended to evaluate the possibility of expansion. If it is greater than the preset percentage, the traffic control parameters are expanded according to the preset percentage. By expanding the buffer, the situation of directly rejecting users is reduced, and the customer experience is improved. At the same time, the data of the historical water level line is judged. If the redundancy is greater than the preset percentage, no processing is performed. If it is less than the preset percentage, the operation and maintenance personnel are notified of the historical water level situation, and it is recommended to evaluate the expansion.
[0089] As an embodiment of the present invention, collecting the real-time monitoring data of multiple basic systems includes: collecting the access logs of the log center, the performance indicators of the distributed service system, the performance indicators of the PAAS console, and the performance monitoring data of the performance monitoring platform, and using the access logs, the performance indicators of the distributed service system, the performance indicators of the PAAS console, and the performance monitoring data as the real-time monitoring data.
[0090] Among them, real-time monitoring data of each basic management platform, namely the basic system, such as the log center, distributed service monitoring, PAAS console, and performance monitoring platform, is obtained in a conventional manner. Specifically, after obtaining the customer's authorization, access logs of the customer are collected from the log center to count the success rate and failure rate of customer access; from the distributed service monitoring, performance indicators of the service (response time, throughput, concurrency) are monitored; from the PAAS (Platform as a Service) console, performance indicators of the container (container CPU, memory, IO, network), as well as the corresponding relationship between the container and the underlying host are obtained; from the performance monitoring platform, performance monitoring data (CPU, memory, IO, network) at the virtualization layer and physical machine level, as well as relevant indicators of resource contention (CPU steal, npf events) are obtained.
[0091] As an embodiment of the present invention, as Figure 2 shown, using a preset trigger rule, threshold determination and trend determination are performed on the real-time monitoring data, and the obtained trigger determination results include:
[0092] Step S21, using a preset trigger rule, perform trigger throttling exception threshold determination, concurrency mean value determination, and throughput mean value determination on the real-time monitoring data to obtain a threshold determination result;
[0093] Step S22, using a preset trigger rule, perform a trend determination result on the real-time monitoring data, and based on the threshold determination result and the trend determination result, obtain a trigger determination result.
[0094] Among them, the preset trigger rule includes a threshold determination rule and a trend determination rule. Specifically, the threshold determination rule includes three types of threshold determination strategies: the threshold for triggering throttling exception, the concurrency mean value, and the throughput mean value.
[0095] Specifically, throttling exception means that due to a large influx of customer requests, throttling fails. When the throttling failure ratio exceeds 0.5% of the total number of service processing transactions, it is determined that the service quality has decreased due to throttling, triggering the subsequent traffic control process, that is, the obtained threshold determination result is to trigger traffic control. The concurrency mean value and the throughput mean value are based on the service monitoring ability of the distributed service framework to monitor and collect the concurrency and throughput indicators of the service for 7*24 hours. During the period when the service is externally provided, if (concurrency mean value ÷ concurrency upper limit) and (throughput mean value ÷ throughput upper limit) are higher than 80% for three consecutive collection cycles (generally, the collection cycle is one day), it is determined that the capacity exceeds the warning value, triggering the subsequent traffic control process, that is, the obtained threshold determination result is to trigger traffic control.
[0096] Furthermore, the trend determination rule is mainly based on the monitoring metrics of concurrency and throughput. Through the service monitoring capability of the distributed service framework, the concurrency and throughput metrics of the service are monitored and collected 24 hours a day, 7 days a week. Select several representative time points (morning peak, lunch break, evening peak, prime time, daily cut-off point) from each trading day and perform linear regression from both vertical and horizontal dimensions. Taking the analysis process of concurrency as an example, the method of linear regression is to construct a linear regression model with the time dimension and the concurrency dimension. Specifically, time is used as the X-axis and concurrency is used as the Y-axis, and data of 10 time nodes (which can be customized) are collected for regression fitting. If the fitted line is linearly upward, the subsequent traffic control process is triggered, that is, the obtained trend determination result is to trigger traffic control. The trend determination of throughput is the same.
[0097] Furthermore, the trigger determination result is obtained according to the threshold determination result and the trend determination result. Specifically, if any one of the threshold determination result and the trend determination result is to trigger traffic control, the trigger determination result is to trigger traffic control.
[0098] As an embodiment of the present invention, the system types include channel application systems and background service application systems.
[0099] In this embodiment, as Figure 3 shown, according to the system types corresponding to each basic system, resource redundancy data is determined using real-time monitoring data, and historical water level lines are determined using the historical monitoring data of each basic system, including:
[0100] Step S31, if the system type of the basic system is a channel application system, use the real-time monitoring data to determine the current node status and the current system status; wherein, the current node status includes CPU usage rate, memory usage rate, and throughput, and the current system status includes virtual machine capacity and physical machine capacity;
[0101] Step S32, determine the resource redundancy data according to the current node status and the current system status; wherein, the resource redundancy data includes CPU redundancy, virtual machine redundancy, and physical machine redundancy;
[0102] Step S33, obtain the historical monitoring data of each basic system within a preset time period, and use the historical monitoring data to determine the historical water level line.
[0103] Among them, if the system type is a channel application system, the real-time monitoring data is first used to determine the current node status and the current system status. Specifically, the current node status mainly refers to the performance capacity status of the corresponding node container, including but not limited to CPU usage rate, memory usage rate, disk and network IO throughput, etc. The representative data selected for evaluation are all peak usage indicators within 5 minutes. Specifically, the specific methods are divided into two categories. One is the judgment of thresholds. For example, if the TPS of disk IO does not exceed 500, it passes and enters the evaluation process of the next stage. The second is the judgment of redundant capacity. For example, using the peak CPU usage within 5 minutes of the current node, and obtaining the quantitative output of redundant capacity (90% - peak usage within 5 minutes) based on the difference from the warning value (90%) of CPU production, and then continuing with the subsequent traffic control), thus obtaining the quantitative output of redundant capacity as the current node status and using it as resource redundant data.
[0104] Furthermore, after obtaining the current node status, drill down to the virtualization layer and the physical machine layer through the current node (usually a paas container) to evaluate the capacity redundancy of the virtual machine and the physical machine to obtain the current system status. Specifically, the evaluation method includes two parts. The first part is the general capacity evaluation, which is the same as the evaluation method of the current node status. The second part is that since there may be resource contention in the virtual machine layer and the physical machine layer due to oversubscription of resources, it is necessary to supplement and check whether the indicators of resource contention (cpu steal, npf events) exceed the threshold. From these two parts, the current system status is obtained and used as resource redundant data.
[0105] Furthermore, the determination of the historical water level line of each basic system, that is, the judgment of the historical capacity status of the node, analyzes the water level line saved by the system for the node. The data of the water level line (i.e., historical monitoring data) is updated daily through the monitoring method. Based on the maximum value of the usage of each resource on the current day, including all the indicators in the above two steps of determining the current node status and the current system status, if it is judged to be greater than the maximum value of the historical water level line, the historical water level line is updated; if it is less than or equal to the maximum value of the historical water level line, it is not updated. Thus, the historical water level line of each basic system can be accurately determined. In addition, based on the water level line and the warning values and thresholds of the current production capacity, analysis and judgment are carried out, and the judgment process is the same as the judgment method in the above current node status and current system status.
[0106] In this embodiment, as Figure 4 shown, according to the system type corresponding to each basic system, using the real-time monitoring data to determine the resource redundant data, and using the historical monitoring data of each basic system to determine the corresponding historical water level line, including:
[0107] Step S34, if the system type of the basic system is a background service application system, use the real-time monitoring data to determine the current node status and the current system status; wherein, the current node status includes CPU usage rate, memory usage rate, and throughput, and the current system status includes virtual machine capacity and physical machine capacity;
[0108] Step S35, determine the resource redundancy data according to the current node status and the current system status; wherein, the resource redundancy data includes CPU redundancy, virtual machine redundancy, and physical machine redundancy;
[0109] Step S36, obtain the historical monitoring data of each basic system within a preset time period, and use the historical monitoring data to determine the historical water level line;
[0110] Step S37, trace the associated call chain of each basic system, and determine the corresponding resource redundancy data for all services and nodes on the associated call chain in parallel.
[0111] Among them, if the system type is a background service application system, the process of determining the current node status, the current system status, and the historical water level line is the same as that of the channel application system. After these three steps of the current node status, the current system status, and the historical water level line, it is also necessary to perform associated adjustment point determination and iterative processing. Specifically, through an observable tracing method, such as the tracing ability (a basic ability of a distributed system), obtain the entire call chain of the service controlled by traffic management, and then associate all services and nodes on the call chain, and perform these three steps of determining the current node status, the current system status, and the historical water level line in parallel. Considering the relevant capacity redundancy situation, if the current capacity redundancy evaluation of all services on the link is greater than 10%, implement traffic control parameter expansion. Then judge the capacity redundancy of the historical water level line. If there is a check redundancy of 10%, notify the production operation and maintenance personnel to evaluate the subsequent expansion process.
[0112] As an embodiment of the present invention, as Figure 5 shown, performing traffic control on each basic system according to the resource redundancy data and the historical water level line includes:
[0113] Step S41, compare the real-time monitoring data with the historical water level line to obtain a comparison result;
[0114] Step S42, perform traffic control on each basic system according to the resource redundancy data and the comparison result; wherein, the traffic control includes expanding buffer processing.
[0115] Among them, according to the current node status and the current system status in the resource redundancy data, combined with the historical water level line, using the judgment result of the current capacity determined thereby, that is, the comparison result obtained by comparing the real-time monitoring data with the historical water level line, traffic control can be performed on each basic system.
[0116] Specifically, traffic control can be carried out by expanding the buffer. For example, if the redundancy percentage obtained from the container, virtual machine, or physical machine is less than 10%, the traffic control parameters are not adjusted, and relevant operation and maintenance personnel can be notified of the resource and traffic control situation, and it is recommended to evaluate the possibility of capacity expansion. If it is greater than 10%, the traffic control parameters are expanded by 10%. By expanding the buffer, the situation of directly rejecting users is reduced, and the customer experience is improved. At the same time, the historical water level data is judged. If the redundancy is greater than 10%, it is not processed. If it is less than 10%, the operation and maintenance personnel are notified of the historical water level situation, and it is recommended to evaluate the capacity expansion.
[0117] The present invention performs dynamic traffic control by following the actual operation situation, improves the resource utilization efficiency, timely optimizes the traffic control strategy, quickly adjusts the system service capacity within the scope of safety and controllability, improves the customer experience, and also improves the completeness and automation degree of traffic control, and improves the efficiency of system operation and maintenance.
[0118] In a specific embodiment of the present invention, as Figure 6 shown in the schematic diagram of the system structure applying the adaptive traffic control method in the embodiment of the present invention, the system of the present invention supports flexible customization of traffic control strategies through a componentized traffic control module. At the same time, it links with observable modules such as front-end customer access logs, transaction monitoring, container resource monitoring, and underlying system monitoring. After aggregating and analyzing the relevant data, traffic control strategies are generated and sent to the traffic control system for real-time distribution and take effect. Thus, the traffic control capabilities are uniformly and standardizedly managed, and different traffic control models are associated with different system monitoring performances. Then, the strategies are distributed through the traffic control system and take effect through hot loading to timely respond to changes in the system operation situation and ensure the system stability. The system shown in the figure includes a traffic control module, an operation data collection module, and a data analysis module.
[0119] In this embodiment, the traffic control module mainly maintains the traffic control strategies of each system, node, and service. The traffic control strategies mainly consist of a traffic control algorithm and a traffic control scalar. The traffic control algorithm includes two dimensions: concurrency and throughput. Among them, the throughput includes different algorithms, such as: fixed window, sliding window, token bucket, leaky bucket and other traffic control algorithms. Different algorithms have different characteristics in terms of control effect, customer experience, and computing power overhead. The traffic control scalar is an integer greater than or equal to 0. 0 as the boundary value represents no traffic control (equivalent to ∞), and other integers represent that: according to the calculation results of the corresponding traffic control algorithm, the concurrency and throughput of the corresponding service cannot exceed this integer value, and the requests exceeding it will be discarded and an exception will be thrown.
[0120] The operation data collection module interfaces with basic management platforms such as the log center, distributed service monitoring, PAAS console, and performance monitoring platform. It collects customer access logs from the log center and calculates the success rate and failure rate of customer access; from distributed service monitoring, it monitors service performance metrics (response time, throughput, concurrency); from the PAAS console, it obtains container performance metrics (container CPU, memory, IO, network), as well as the corresponding relationship between containers and the underlying host; from the performance monitoring platform, it obtains performance monitoring data (CPU, memory, IO, network) at the virtualization layer and physical machine level, as well as relevant metrics for resource contention (CPU steal, npf events).
[0121] The data analysis module mainly classifies the collected data in combination with the attributes of the corresponding services, analyzes them according to the corresponding rules, generates corresponding traffic control strategies, and transmits them to the traffic control module for implementation and distribution. Specifically, the analysis method is as follows:
[0122] I. Distinguish service levels: customer-facing services, middle-tier services, and back-end services. Customer-facing services are generally directly provided for access by the channel side and are directly oriented to customers, such as login and balance query. Middle-tier services are generally services oriented to business objects, such as protocol maintenance and customer information query. Back-end services are generally oriented to specific business elements, such as limit determination and credit line occupancy. Generally, the relationship among the three service levels is as Figure 7 shown.
[0123] Among them, the algorithm is selected based on the service level, as shown in Table 1.
[0124] Table 1
[0125]
[0126] II. Analyze based on the collected operation data. The analysis process is mainly divided into two parts:
[0127] 1. Monitoring and analysis trigger the traffic control process.
[0128] 2. Analysis and implementation of the traffic control process.
[0129] Monitoring and analysis trigger the traffic control process: The main goal is to collect real-time monitoring data from the daily operation process and identify when the traffic control process intervenes through methods of threshold and trend determination.
[0130] 1) The threshold - based determination process includes three types of threshold determination strategies: the threshold for triggering flow - limiting exceptions, the average concurrency, and the average throughput. A flow - limiting exception means that due to a large influx of customer requests, flow - limiting fails. When the proportion of flow - limiting failures exceeds 0.5% of the total number of service - processed transactions, it is determined that the service quality has declined due to flow - limiting, triggering the subsequent flow control strategy adjustment process. The average concurrency and the average throughput are based on the service monitoring capabilities of the distributed service framework to monitor and collect the concurrency and throughput metrics of the service for 7 * 24 hours. During the period when the service is externally provided, if (average concurrency÷concurrency upper limit) and (average throughput÷throughput upper limit) are higher than 80% for three consecutive collection cycles (the general collection cycle is one day), it is determined that the capacity exceeds the warning value, triggering the subsequent flow control strategy adjustment process.
[0131] 2) The determination process based on trend analysis: The decision - making method of trend analysis is mainly based on the monitored metrics of concurrency and throughput. Through the service monitoring capabilities of the distributed service framework, the concurrency and throughput metrics of the service are monitored and collected for 7 * 24 hours. From each trading day, several representative time points (morning peak, lunch break, evening peak, prime time, daily cut - off point) are selected for linear regression from both vertical and horizontal dimensions. Taking the analysis process of concurrency as an example, the method of linear regression is to construct a linear regression model with the time dimension and the concurrency dimension. Specifically, time is used as the X - axis and concurrency is used as the Y - axis, and data from 10 time nodes (which can be customized) are collected for regression fitting. If the fitted line slopes upward linearly, it triggers the subsequent flow control strategy adjustment process. The decision - making method for throughput is the same.
[0132] In this embodiment, the analysis and implementation of the flow control process: After triggering the flow control process, it is necessary to calculate the resource redundancy of each node in the complete service link according to the actual operating conditions. Through decision - tree analysis, an implementable flow control plan is obtained, and the plan is sent to the distributed service framework to configure relevant service governance parameters. The specific method is as follows:
[0133] As Figure 8 shown in the implementation of the trigger flow control strategy change process for the channel - type application system, it is necessary to determine whether the conditions for amplifying the traffic are met and decide to adjust specific parameter values. The condition judgment is divided into three steps:
[0134] 1. Current node status. The status mainly corresponds to the performance capacity of the node container, including but not limited to CPU usage rate, memory usage rate, disk and network IO throughput, etc. The representative data selected for evaluation are all peak usage indicators within 5 minutes. The methods are divided into two categories. One is the judgment of thresholds. For example, if the TPS of disk IO does not exceed 500, it passes and enters the evaluation process of the next stage. The second is the judgment of redundant capacity. For example, using the peak CPU usage within 5 minutes of the current node, obtaining the quantitative output of redundant capacity (90% - peak usage within 5 minutes) based on the difference from the warning value (90%) of CPU production, and then continuing the subsequent evaluation process).
[0135] 2. Drill down from the current node (usually a paas container) to the virtualization layer and the physical machine layer to evaluate the capacity redundancy of the virtual machines and physical machines. The evaluation method includes two parts. The first part is the general capacity evaluation, which is the same as the evaluation method of the node in point 1. Secondly, due to resource overcommitment in the virtual machine layer and the physical machine layer, there may be resource contention situations, and it is necessary to supplement and check whether the indicators of resource contention (cpu steal, npf events) exceed the thresholds.
[0136] 3. Judgment of the historical capacity status of the node. Analyze the water level line saved by the platform for the node. The data of the water level line is updated to the analysis platform by the monitoring platform on a daily basis. The analysis platform will be based on the maximum value of the daily usage of each resource, including all the indicators in the above two points 1 and 2. If it is judged that it is greater than the maximum value of the historical water level, the water level data will be updated. If it is less than or equal to the maximum value of the historical water level, it will not be updated. Based on the water level line and the warning values and thresholds of the current production capacity, analyze and judge. The judgment process is the same as the judgment methods in the above two points 1 and 2. After the above three inspections are completed, if the redundant percentage obtained for the container, virtual machine, and physical machine is less than 10%, the traffic control parameters will not be adjusted, and the relevant operation and maintenance personnel will be notified of the resource and traffic control situation, and it is recommended to evaluate the possibility of expansion. If it is greater than 10%, the traffic control parameters will be expanded by 10%. By expanding the buffer, the situation of directly rejecting users is reduced, and the customer experience is improved. At the same time, the judgment of the water level line data is carried out. If the redundancy is greater than 10%, it will not be processed. If it is less than 10%, the operation and maintenance personnel will be notified of the historical water level situation, and it is recommended to evaluate the expansion.
[0137] As Figure 9 shown is the implementation of the trigger traffic control strategy change process for the background service application system. Similarly, it is necessary to judge whether the conditions for amplifying the traffic are met and determine the specific parameter values to be adjusted. However, due to differences in the service governance model from the channel application system, the implementation logic is different.
[0138] Specifically, the condition judgment is divided into four steps: the first three steps are the same as the inspection and judgment process triggered by the log data. However, since most of these backend services adopt the flow limiting mode of token bucket and fixed window, the incoming traffic requests are directly transferred to the downstream services without buffering. Therefore, it is necessary to further evaluate the processing capabilities on the entire service chain, and a fourth step of judgment needs to be added.
[0139] Among them, the fourth step mainly relies on the observability capabilities of the distributed service framework. Through the tracing capabilities of observability (basic capabilities of distributed systems), the entire call chain of the services under traffic control is obtained, and then all the services and nodes on the call chain are associated. The above steps 1-3 are analyzed in parallel. Considering the relevant capacity redundancy situation, if the current capacity redundancy evaluations of all services on the link are greater than 10%, the traffic control parameter expansion is implemented. Then, the capacity redundancy of the historical water level line is judged. If there is a check redundancy of 10%, the production operation and maintenance personnel are notified to evaluate the subsequent expansion process.
[0140] The present invention dynamically performs traffic control following the actual operation situation, improves resource utilization efficiency, timely optimizes traffic control strategies, quickly adjusts the system service capabilities within a safe and controllable range, enhances the customer experience, improves the completeness and automation degree of traffic control analysis, and improves the efficiency of operation and maintenance management.
[0141] As Figure 10 shown in the structural schematic diagram of an adaptive traffic control device according to an embodiment of the present invention. The device shown in the figure includes:
[0142] A monitoring data module 10, configured to collect real-time monitoring data of multiple basic systems and determine the corresponding system types of each basic system;
[0143] A trigger determination module 20, configured to use a preset trigger rule to perform threshold determination and trend determination on the real-time monitoring data to obtain a trigger determination result;
[0144] A redundant data module 30, configured to, if the trigger determination result is to trigger traffic control, determine resource redundant data using the real-time monitoring data according to the corresponding system types of each basic system, and determine the corresponding historical water level line using the historical monitoring data of each basic system;
[0145] A traffic control module 40, configured to perform traffic control on each basic system according to the resource redundant data and the historical water level line.
[0146] As an embodiment of the present invention, the monitoring data module is further configured to collect access logs of the log center, performance metrics of the distributed service system, performance metrics of the PAAS console, and performance monitoring data of the performance monitoring platform, and use the access logs, performance metrics of the distributed service system, performance metrics of the PAAS console, and performance monitoring data as real-time monitoring data.
[0147] As an embodiment of the present invention, as Figure 11 shown, the trigger determination module 20 includes:
[0148] A threshold determination unit 21, configured to use a preset trigger rule to perform trigger throttling exception threshold determination, concurrent number average determination, and throughput average determination on the real-time monitoring data to obtain a threshold determination result;
[0149] A trend determination unit 22, configured to use a preset trigger rule to perform a trend determination result on the real-time monitoring data, and obtain a trigger determination result according to the threshold determination result and the trend determination result.
[0150] As an embodiment of the present invention, the system types include channel application systems and background service application systems.
[0151] In this embodiment, as Figure 12 shown, the redundant data module 30 includes:
[0152] A first status unit 31, configured to determine the current node status and the current system status by using the real-time monitoring data if the system type of the basic system is a channel application system; wherein, the current node status includes CPU usage rate, memory usage rate, and throughput, and the current system status includes virtual machine capacity and physical machine capacity;
[0153] A first redundancy unit 32, configured to determine resource redundancy data according to the current node status and the current system status; wherein, the resource redundancy data includes CPU redundancy, virtual machine redundancy, and physical machine redundancy;
[0154] A first water level unit 33, configured to obtain historical monitoring data of each basic system within a preset time period, and determine the historical water level by using the historical monitoring data.
[0155] In this embodiment, as Figure 13 shown, the redundant data module 30 further includes:
[0156] A second status unit 34, configured to determine the current node status and the current system status by using the real-time monitoring data if the system type of the basic system is a background service application system; wherein, the current node status includes CPU usage rate, memory usage rate, and throughput, and the current system status includes virtual machine capacity and physical machine capacity;
[0157] A second redundancy unit 35, configured to determine resource redundancy data according to the current node state and the current system state; wherein, the resource redundancy data includes CPU redundancy, virtual machine redundancy, and physical machine redundancy;
[0158] A second water level unit 36, configured to obtain historical monitoring data of each basic system within a preset time period, and determine a historical water level line by using the historical monitoring data;
[0159] An associated call chain unit 37, configured to trace the associated call chains of each basic system, and determine corresponding resource redundancy data for all services and nodes on the associated call chains in parallel.
[0160] As an embodiment of the present invention, as Figure 14 shown, the traffic control module 40 includes:
[0161] A comparison result unit 41, configured to compare the real-time monitoring data with the historical water level line to obtain a comparison result;
[0162] A traffic control unit 42, configured to perform traffic control on each basic system according to the resource redundancy data and the comparison result; wherein, the traffic control includes expanding buffer processing.
[0163] Based on the same application concept as the above-mentioned adaptive traffic control method, the present invention further provides the above-mentioned adaptive traffic control device. Since the principle of solving problems by this adaptive traffic control device is similar to that of an adaptive traffic control method, the implementation of this adaptive traffic control device can refer to the implementation of an adaptive traffic control method, and the repeated parts will not be described again.
[0164] The present invention performs dynamic traffic control by following the actual operation situation, improves resource utilization efficiency, timely optimizes the traffic control strategy, quickly adjusts the system service capacity within a safe and controllable range, enhances the customer experience, and also enhances the completeness and automation degree of traffic control, and improves the efficiency of system operation and maintenance.
[0165] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements the above method when executing the program.
[0166] The present invention further provides a computer program product, including computer programs / instructions, and the computer programs / instructions implement the steps of the above method when executed by a processor.
[0167] The present invention further provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program for the computer to execute the above method.
[0168] As Figure 15As shown, the electronic device 600 may further include: a communication module 110, an input unit 120, an audio processor 130, a display 160, and a power supply 170. It should be noted that the electronic device 600 does not necessarily have to include Figure 15 all the components shown in; in addition, the electronic device 600 may further include Figure 15 components not shown in, and reference may be made to the prior art.
[0169] As Figure 15 shown, the central processing unit 100 is sometimes also referred to as a controller or operation control, and may include a microprocessor or other processor devices and / or logic devices. The central processing unit 100 receives inputs and controls the operations of the various components of the electronic device 600.
[0170] Among them, the memory 140 may be, for example, one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. It can store the above information related to failures, and can also store programs for executing relevant information. And the central processing unit 100 can execute the programs stored in the memory 140 to achieve information storage or processing, etc.
[0171] The input unit 120 provides inputs to the central processing unit 100. The input unit 120 is, for example, a key or a touch input device. The power supply 170 is used to supply power to the electronic device 600. The display 160 is used to display display objects such as images and texts. The display may be, for example, an LCD display, but is not limited thereto.
[0172] The memory 140 may be a solid-state memory. For example, a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be such a memory that stores information even when powered off, can be selectively erased and has more data stored. Examples of such a memory are sometimes referred to as EPROMs, etc. The memory 140 may also be some other type of device. The memory 140 includes a buffer memory 141 (sometimes referred to as a buffer). The memory 140 may include an application / function storage unit 142, and the application / function storage unit 142 is used to store application programs and function programs or the processes for operating the electronic device 600 through the central processing unit 100.
[0173] The memory 140 may also include a data storage unit 143, and the data storage unit 143 is used to store data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 144 of the memory 140 may include various drivers of the electronic device for communication functions and / or for executing other functions of the electronic device (such as a messaging application, an address book application, etc.).
[0174] The communication module 110 is a transmitter / receiver that transmits and receives signals via the antenna 111. The communication module (transmitter / receiver) 110 is coupled to the central processor 100 to provide input signals and receive output signals, which can be the same as in the case of a conventional mobile communication terminal.
[0175] Based on different communication technologies, multiple communication modules 110, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc., can be provided in the same electronic device. The communication module (transmitter / receiver) 110 is also coupled to the speaker 131 and the microphone 132 via the audio processor 130 to provide an audio output via the speaker 131 and receive an audio input from the microphone 132, thereby implementing normal telecommunication functions. The audio processor 130 can include any suitable buffers, decoders, amplifiers, etc. Additionally, the audio processor 130 is also coupled to the central processor 100, so that recording can be performed on the local machine through the microphone 132, and the sound stored on the local machine can be played through the speaker 131.
[0176] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0177] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0178] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocksFigure 1 The functions specified in one or more boxes.
[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 or more boxes.
[0180] Specific embodiments are used in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only for helping to understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An adaptive flow control method, characterized in that: The method comprises: Collect real-time monitoring data of multiple basic systems and determine the system type corresponding to each basic system; Using preset trigger rules, threshold determination and trend determination are performed on the real-time monitoring data to obtain a trigger determination result; If the trigger determination result is to trigger flow control, then according to the system type corresponding to each basic system, the resource redundancy data is determined using the real-time monitoring data, and the corresponding historical waterline is determined using the historical monitoring data of each basic system; The flow of each basic system is controlled according to the resource redundancy data and the historical water level.
2. The method according to claim 1, characterized in that Collect real-time monitoring data of multiple basic systems including: The access logs of the log center, the performance indicators of the distributed service system, the performance indicators of the PAAS console and the performance monitoring data of the performance monitoring platform are collected, and the access logs, the performance indicators of the distributed service system, the performance indicators of the PAAS console and the performance monitoring data are used as the real-time monitoring data.
3. The method according to claim 1, characterized in that Using the preset trigger rules, threshold determination and trend determination are performed on the real-time monitoring data, and the trigger determination results obtained include: Using the preset trigger rules, the real-time monitoring data is subjected to trigger current limiting abnormal threshold determination, concurrent number mean determination and throughput mean determination, and a threshold determination result is obtained; The real-time monitoring data is subjected to trend determination results using preset trigger rules, and the trigger determination result is obtained based on the threshold determination result and the trend determination result.
4. The method according to claim 1, characterized in that The system types include channel application systems and background service application systems.
5. The method according to claim 4, characterized in that According to the system type corresponding to each basic system, the resource redundancy data is determined by using the real-time monitoring data, and the corresponding historical waterline is determined by using the historical monitoring data of each basic system, including: If the system type of the basic system is a channel application system, the current node status and the current system status are determined by using the real-time monitoring data; wherein the current node status includes CPU usage, memory usage and throughput, and the current system status includes virtual machine capacity and physical machine capacity; Determine the resource redundancy data according to the current node state and the current system state; wherein the resource redundancy data includes CPU redundancy, virtual machine redundancy and physical machine redundancy; The historical monitoring data of each basic system within a preset time period is obtained, and the historical water level is determined using the historical monitoring data.
6. The method according to claim 4, characterized in that According to the system type corresponding to each basic system, the resource redundancy data is determined by using the real-time monitoring data, and the corresponding historical waterline is determined by using the historical monitoring data of each basic system, including: If the system type of the basic system is a background service application system, the real-time monitoring data is used to determine the current node status and the current system status; wherein the current node status includes CPU usage, memory usage and throughput, and the current system status includes virtual machine capacity and physical machine capacity; Determine the resource redundancy data according to the current node state and the current system state; wherein the resource redundancy data includes CPU redundancy, virtual machine redundancy and physical machine redundancy; Obtaining historical monitoring data of each basic system within a preset time period, and using the historical monitoring data to determine a historical water level; Track the associated call chains of each basic system, and determine the corresponding resource redundancy data for all services and nodes in the associated call chains in parallel.
7. The method according to claim 1, characterized in that According to the resource redundancy data and the historical waterline, flow control of each basic system includes: Compare the real-time monitoring data with the historical water level line to obtain a comparison result; According to the resource redundancy data and the comparison result, flow control is performed on each basic system; wherein the flow control includes expanding buffer processing.
8. An adaptive flow control device, characterized in that: The device comprises: The monitoring data module is used to collect real-time monitoring data of multiple basic systems and determine the system type corresponding to each basic system; A trigger determination module, used to perform threshold determination and trend determination on the real-time monitoring data using preset trigger rules to obtain a trigger determination result; A redundant data module, for determining resource redundant data using the real-time monitoring data according to the system type corresponding to each basic system if the trigger determination result is to trigger flow control, and determining the corresponding historical waterline using the historical monitoring data of each basic system; The flow control module is used to perform flow control on each basic system according to the resource redundancy data and the historical water level line.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for causing a computer to execute the method according to any one of claims 1 to 7.