A Hierarchical Rate Limiting Method and Device for a Distributed System Based on Real-Time Load Awareness

The real-time load-aware tiered speed limiting method in distributed systems addresses inefficiencies in static limit speed rules by dynamically adjusting TPS and differentiating read/write operations, enhancing system stability and efficiency.

CN119697123BActive Publication Date: 2025-07-15BEIJING FENYANG TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510203095.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-15
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing distributed system speed limiting technology cannot be dynamically adjusted according to real-time load conditions, resulting in waste or insufficient resources, unable to distinguish the resource consumption differences of read and write operations, lack of differentiated processing of business requests, and it is difficult to deal with burst traffic, affecting system stability and efficiency.

Method used

Through real-time load perception, the TPS gear of the transaction number per second is dynamically adjusted, combined with the token bucket algorithm and differentiated speed limit control, the read and write ratio and business priority are monitored to achieve a dynamic and accurate speed limit strategy.

Benefits of technology

The stability of the distributed system and the improvement of business processing capabilities are achieved, excessive consumption or idle resources are avoided, the continuity and stability of business requests are ensured, resource allocation is optimized, and the system's adaptability is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119697123B_ABST
    Figure CN119697123B_ABST
Patent Text Reader

Abstract

The present invention provides a hierarchical speed limit method and device for a distributed system based on real-time load perception, relating to the technical field of traffic control. It includes: an initialization step of configuring system basic parameters and a load data collection period; receiving service requests, identifying and differentiating read and write requests and service priorities in the service requests; periodically collecting load data of distributed nodes, calculating a standardized index threshold; judging the load status in combination with the threshold; dynamically adjusting the transactions per second (TPS) gear according to the load status; using the token bucket algorithm to dynamically adjust the current maximum number of tokens in the current token bucket of the system according to the current gear and allocate tokens as needed, and checking whether there are enough tokens in the current token bucket based on the received service requests. If so, allow the service requests to pass; otherwise, reject the service requests or queue the service requests; monitoring and analyzing the read-write ratio in the service requests, and performing differential speed limit control in combination with the read-write ratio and service priorities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic control, and particularly to a hierarchical speed limit method and device for a distributed system based on real-time load perception. Background Art

[0002] In a distributed system, speed limit technology is an important means to ensure the stable operation of the system. Its core goal is to prevent performance degradation or service interruption caused by system overload by reasonably allocating system resources. However, with the increasing complexity and dynamism of business scenarios, the existing distributed system speed limit technologies face a series of significant challenges. These challenges not only affect the overall efficiency of the system but also limit the system's ability to respond to sudden traffic.

[0003] Traditionally, distributed system speed limit technologies mostly adopt static configuration of speed limit rules. Although this method is simple and easy to implement, it is unable to cope with the dynamically changing load state. Since static rules cannot automatically adjust according to the real-time load situation, it often leads to overprotection of resources when the load is light, resulting in waste; while when the load is heavy, it may cause system performance bottlenecks or even crashes due to insufficient protection. This "one-size-fits-all" speed limit strategy clearly cannot meet the high requirements of modern distributed systems for flexibility and self-adaptability.

[0004] Furthermore, existing speed limit technologies often ignore the differential impact of different types of requests on system resources. Read and write operations, as two basic operation types in a distributed system, have significant differences in their consumption and impact on system resources. However, traditional speed limit technologies fail to distinguish between these two operations and instead adopt a unified speed limit standard. This processing method not only fails to fully exploit the potential of system resources but may also cause key read and write requests to be blocked due to uneven resource allocation, thereby affecting the efficient operation of the overall business. In addition, in practical applications, different business requests often have different priorities and read-write ratios. For example, some key business requests may require higher response speeds and more stable resource guarantees, while some non-critical business requests can be processed when resources are sufficient. However, traditional speed limit technologies lack in-depth understanding and flexible response to these business characteristics and cannot dynamically adjust the speed limit strategy according to the priority and read-write ratio of requests, resulting in unreasonable system resource allocation and limited business processing efficiency.

[0005] In addition, the lack of real-time perception of node load status and precise flow control is also a major defect of the existing technology. In the case of sudden traffic scenarios, the system often has difficulty quickly responding and adjusting the speed limit strategy, resulting in local node overload, which in turn triggers a chain reaction in the entire system, causing service interruption or a sharp decline in performance. The root cause of this problem lies in the simplicity and rigidity of the flow control algorithm, which cannot adapt to the diverse requirements in complex business scenarios.

[0006] In summary, due to the defects of existing distributed system speed limit technologies, such as static configuration, lack of differential processing, and insufficient real-time perception ability, it has become difficult to meet the requirements of modern distributed systems for efficient, stable, and adaptive operation. Therefore, there is an urgent need for a new technical solution to overcome the above defects and achieve more intelligent, flexible, and accurate speed limit control to cope with the increasingly complex and changing business scenarios and load challenges. Summary of the Invention

[0007] The present invention aims to provide a hierarchical speed limit method and device for a distributed system based on real-time load perception, so as to solve the problems existing in the prior art, protect the stability of the distributed system, achieve more intelligent, flexible, and accurate speed limit control, and minimize the impact on business requests.

[0008] In order to solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0009] A hierarchical speed limit method for a distributed system based on real-time load perception, comprising:

[0010] An initialization step for configuring the system basic parameters and the load data collection period;

[0011] Receiving a business request, identifying and distinguishing read requests, write requests, and business priorities in the business request;

[0012] Periodically collecting load data of distributed nodes and calculating a standardized index threshold;

[0013] Judging the load status in combination with the standardized index threshold;

[0014] Dynamically adjusting the transactions per second (TPS) gear according to the load status;

[0015] Using the token bucket algorithm to dynamically adjust and allocate tokens as needed the current maximum number of tokens in the current token bucket of the system according to the current TPS gear, and checking whether there are enough tokens in the current token bucket based on the received business request. If there are enough tokens, allow the business request to pass; otherwise, reject the business request or queue the business request;

[0016] Monitoring and analyzing the read / write ratio in the business request, and performing differential speed limit control in combination with the read / write ratio and the business priority.

[0017] Further, the system basic parameters include: the number of CPU cores, the TPS gear, and the initial speed limit parameters, etc.

[0018] Further, the TPS gears include the following levels: 120, 60, 40, 30, 24, 20, 17, 15, 12, 10, 8, 6, 4, 2, 1.

[0019] Further, the key parameters of the token bucket algorithm include: the maximum number of tokens in the token bucket and the number of tokens obtained each time. The maximum number of tokens in the token bucket is 120, and the number of tokens obtained each time is selected from the set {2, 3, 4, 5, 6, 7, 8, 10, 12, 15, 20, 30, 60, 120}.

[0020] Further, periodically call the system API or dedicated monitoring interface to obtain load data, and calculate the standardized metric thresholds for 1 minute, 5 minutes, and 10 minutes.

[0021] Further, the standardized metric thresholds are calculated as follows: 1-minute metric threshold = 1-minute load / number of CPU cores; 5-minute metric threshold = 5-minute load / number of CPU cores; 10-minute metric threshold = 10-minute load / number of CPU cores.

[0022] Further, the load status includes: healthy status when the 5-minute metric threshold < 1 and the 10-minute metric threshold < 1; sub-healthy status when the 5-minute metric threshold ≥ 1 or the 10-minute metric threshold ≥ 1; unhealthy status when the 5-minute metric threshold > 1 and the 10-minute metric threshold > 1.

[0023] Further, the dynamic adjustment of the TPS gears includes: increasing the TPS gear by 1 level when the load status is the healthy status; decreasing the TPS gear by 1 level when the load status is the sub-healthy status; decreasing the TPS gear by 2 levels when the load status is the unhealthy status.

[0024] Further, it includes performing more stringent speed limit control on the write requests.

[0025] A distributed system hierarchical speed limit device based on real-time load perception, comprising:

[0026] An initialization module, which configures the system basic parameters and the load data collection period;

[0027] A receiving module, which receives service requests, identifies and differentiates read requests, write requests, and service priorities in the service requests;

[0028] A load data collection module, which periodically collects load data of distributed nodes and calculates standardized metric thresholds;

[0029] A load status judgment module that judges the load status in combination with the standardized index threshold;

[0030] A dynamic adjustment module that dynamically adjusts the TPS (Transactions Per Second) gear according to the load status;

[0031] A speed limit execution module that uses the token bucket algorithm to dynamically adjust the current maximum number of tokens in the current token bucket of the system according to the current TPS gear and allocate tokens as needed, and checks whether there are enough tokens in the current token bucket based on the received service request. If there are enough tokens, the service request is allowed to pass; otherwise, the service request is rejected or queued;

[0032] The speed limit execution module further monitors and analyzes the read-write ratio in the service request, and performs differential speed limit control in combination with the read-write ratio and the service priority.

[0033] A computer-readable storage medium that stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned distributed system hierarchical speed limit method based on real-time load perception.

[0034] The distributed hierarchical speed limit method and device based on real-time load perception proposed by the present invention significantly improve the stability and service processing ability of the distributed system by innovatively integrating real-time data collection, intelligent state judgment and dynamic flow control strategies, and specifically exhibit the following beneficial technical effects:

[0035] 1. Real-time dynamic perception and response:

[0036] Timely and efficiently collect key load data of each node in the distributed system, including but not limited to the number of CPU cores, TPS gear, initial speed limit parameters, etc., and realize real-time and accurate monitoring of the node operation status.

[0037] Based on the collected data, the component can quickly judge the health status (healthy, sub-healthy, unhealthy) of the node, and dynamically adjust the request processing rate accordingly to ensure that the system can still operate stably under high load, while avoiding over-consumption or idleness of resources.

[0038] 2. Smooth speed limit adjustment, reducing business impact:

[0039] Adopt a multi-gear TPS (Transactions Per Second) control mechanism to realize delicate adjustment of the speed limit strategy, and avoid service interruption or performance jitter caused by sudden drop or rise of the speed limit strategy in traditional speed limit technologies.

[0040] This smooth adjustment method not only protects the system from the impact of sudden traffic, but also ensures the continuity and stability of business requests, enhancing the user experience.

[0041] 3. Read-write separation strategy to optimize resource allocation:

[0042] In view of the different characteristics of resource consumption between read operations and write operations in a distributed system, a speed limit strategy of read-write separation is innovatively introduced.

[0043] More stringent speed limit rules are implemented for write operations, effectively preventing system resource exhaustion or performance degradation caused by overly frequent write operations, while ensuring the smooth progress of read operations, improving the overall processing efficiency and response speed of the system. Brief Description of the Drawings

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0045] Figure 1 Flowchart of the distributed hierarchical speed limit method based on real-time load perception provided by the embodiments of the present invention;

[0046] Figure 2 Structure diagram of the distributed hierarchical speed limit device based on real-time load perception provided by the embodiments of the present invention. Detailed Embodiments

[0047] To make the above objects, features, and beneficial effects of the present invention more clearly understood, the specific embodiments of the present invention include various specific details to assist in understanding, but these specific details are only considered exemplary. Therefore, those of ordinary skill in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present invention. Additionally, descriptions of well-known functions and structures may be omitted for clarity and conciseness.

[0048] The terms and words used in the following description and claims are not limited to their written meanings. It is clear to those skilled in the art that the following description of the various embodiments of the present invention is only for the purpose of illustration and not for the purpose of limiting the present invention as defined by the appended claims and their equivalents.

[0049] Those skilled in the art know that the blocks of a flowchart (or sequence diagram) and combinations of flowcharts can be represented and executed by computer program instructions. These computer program instructions can be loaded onto the processors of a general-purpose computer, a special-purpose computer, or a programmable data processing device. When the loaded program instructions are executed by the processor, they create a means for performing the functions described in the flowchart. Since computer program instructions can be stored in a computer-readable memory available in a special-purpose computer or a programmable data processing device, an article for performing the functions described in the flowchart can also be created. Since computer program instructions can be loaded onto a computer or a programmable data processing device, when executed as a process, they can perform the operations of the functions described in the flowchart.

[0050] The blocks of a flowchart can correspond to modules, segments, or code containing one or more executable instructions for implementing one or more logical functions or can correspond to a part thereof. In some cases, the functions described by the blocks may be executed in an order different from the listed order. For example, two blocks listed in sequence can be executed simultaneously or in the reverse order.

[0051] In this specification, terms such as "unit" and "module" can refer to software components or hardware components, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC) capable of performing functions or operations. However, "unit" etc. are not limited to hardware or software. Units etc. can be configured to reside in an addressable storage medium or drive one or more processors. Units etc. can refer to software components, object-oriented software components, class components, task components, processes, functions, attributes, processes, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, or variables. The functions provided by components and units can be combinations of smaller components and units and can be combined with other components to form larger components and units. Components and units can be configured to drive devices in a secure multi-media card or one or more processors.

[0052] The following specifically describes the embodiments of the present invention in conjunction with the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0053] As Figure 1 shown, it shows a hierarchical speed limit method for a distributed system based on real-time load awareness, including the following steps: Step 1, an initialization step, for configuring system basic parameters and a load data collection period;

[0054] Specifically, the basic parameters include basic information such as the number of CPU cores, the TPS gear per second, and the initial speed limit parameter.

[0055] In a distributed system, to obtain the number of CPU cores, different methods can be adopted according to different operating systems. Since the obtaining method belongs to the category of existing technologies, the obtaining method of the number of CPU cores will not be introduced in detail here. After obtaining the number of CPU cores, it is used as the reference value for calculating subsequent standardized index thresholds.

[0056] In the distributed system speed limit technology, the TPS (Transactions Per Second) gear is an important indicator to measure the system processing capacity. It represents the number of transactions that the system can process within a unit time, and is of great significance for evaluating the system performance and formulating speed limit policies.

[0057] Generally speaking, the higher the TPS gear, the stronger the processing capacity of the system, and the more transactions can be processed in a shorter time. In a distributed system, through a reasonable speed limit policy, the request rate of the client can be restricted according to the TPS gear of the system, so as to ensure the stability and availability of the system.

[0058] For example, if the TPS gear of a certain distributed system is 1000, then when designing the speed limit policy, it can be ensured that the number of requests processed per second does not exceed 1000 to avoid system overload and performance degradation. At the same time, by monitoring the TPS gear of the system, performance bottlenecks can be discovered and processed in a timely manner, and the system performance and speed limit policy can be further optimized.

[0059] The embodiments of the present invention set a series of TPS gears, such as: 120, 60, 40,..., 2, 1, to meet the processing requirements under different loads.

[0060] Set the load data collection period, and use the timer function of the system or configure a scheduled task to ensure that the load data is automatically collected every 30 seconds.

[0061] Step 2, receive a service request, and identify and distinguish read requests, write requests, and service priorities in the service request.

[0062] Step 3, periodically collect the load data of the distributed node and calculate the standardized index threshold.

[0063] Periodically call the system API, such as the proc / loadavg file or a dedicated monitoring interface, to obtain the system load data. Taking the PostgreSQL database as an example, after installing the system_stat plugin, executing the SQL query SELECT * FROM pg_sys_os_info() can obtain the detailed load information of the database server.

[0064] After obtaining the system load data, calculate the following standardized index thresholds according to the obtained load data and the number of CPU cores:

[0065] 1-minute metric threshold: threshold_1m = load_1m / number of CPU cores;

[0066] 5-minute metric threshold: threshold_5m = load_5m / number of CPU cores;

[0067] 10-minute metric threshold: threshold_10m = load_10m / number of CPU cores.

[0068] These metric thresholds will be used for subsequent load status judgment. load_1m, load_5m, and load_10m are the loads of the server in the most recent 1 minute, 5 minutes, and 10 minutes respectively.

[0069] Step 4, judge the load status in combination with the standardized metric thresholds.

[0070] First, define the load status, which is defined as three situations, namely healthy status, sub-healthy status, and unhealthy status. The judgment criteria for each status are as follows:

[0071] Healthy status: threshold_5m < 1 and threshold_10m < 1, indicating that the system load is low and resources are abundant;

[0072] Sub-healthy status: threshold_5m ≥ 1 or threshold_10m ≥ 1, indicating that the system load is rising and needs to be closely monitored;

[0073] Unhealthy status: threshold_5m > 1 and threshold_10m > 1, indicating that the system load is too high and may face resource bottlenecks.

[0074] Automatically judge the current status of the system according to the calculated load metric thresholds.

[0075] Step 5, dynamically adjust the TPS (Transactions Per Second) gear according to the load status.

[0076] Dynamically adjusting the TPS gear according to the load status specifically includes:

[0077] When the load status is healthy, automatically increase the TPS gear by 1 level to make full use of system resources and improve throughput.

[0078] When the load status is sub-healthy, lower the TPS gear by 1 level to relieve system pressure and prevent the load from rising further.

[0079] When the load status is in an unhealthy state, immediately lower the TPS gear by 2 levels to quickly reduce the system load and ensure the stable operation of core functions.

[0080] Step 6: Use the token bucket algorithm to dynamically adjust the current maximum number of tokens in the current token bucket of the system according to the current TPS gear and allocate tokens as needed. Check whether there are enough tokens in the current token bucket based on the received service request. If there are enough tokens, allow the service request to pass; otherwise, reject the service request or queue the service request.

[0081] The principle of the token bucket algorithm is to control the data traffic rate through a "bucket". A certain number of "tokens" are stored in the bucket, and each token represents the permission to send a data packet or the permission to process a request. Tokens are generated at a fixed rate and added to the bucket, while the sending of data packets or the processing of requests requires consuming a corresponding number of tokens. If there are not enough tokens in the bucket, the data packets or requests will be rejected or queued until there are enough tokens available.

[0082] In traffic control, the advantages of the token bucket algorithm mainly include the following points:

[0083] Allow burst traffic: Since tokens can accumulate in the bucket, the system can allow the transmission of burst traffic when needed, effectively coping with traffic fluctuations in actual scenarios. This feature makes the token bucket algorithm more flexible in handling burst requests or data packets.

[0084] Efficient resource utilization: The token bucket algorithm ensures the effective utilization of resources by generating tokens at a constant rate and accumulating them in the bucket. When there are fewer requests or data packets, the tokens will accumulate in the bucket, providing sufficient processing capacity for subsequent burst traffic or peak periods, thus avoiding waste of resources.

[0085] Limit the maximum traffic rate: By setting the token generation rate and the capacity of the bucket, the token bucket algorithm can effectively limit the maximum traffic rate and prevent the system from crashing due to overload. This limitation ensures the stability and reliability of the system.

[0086] Do not need to release resources: Compared with some rate-limiting mechanisms that need to acquire and release resources, the token bucket algorithm does not need to release resources after the request is processed, thus avoiding rate-limiting problems caused by failed resource release. This feature makes the token bucket algorithm more stable and reliable in dealing with complex business scenarios.

[0087] The token bucket algorithm can achieve precise control of network traffic by setting the capacity of the token bucket and the token generation rate, avoiding network congestion and crashes.

[0088] Specifically, controlling traffic using the token bucket algorithm includes: setting the maximum capacity of the token bucket to 120 tokens. According to the TPS gear, the number of tokens that can be obtained for each business request is predefined, and the range of the number of tokens is selected from the set {2, 3, 4, 5, 6, 7, 8, 10, 12, 15, 20, 30, 60, 120}. According to the adjusted TPS gear in step 5, the maximum number of tokens in the token bucket is dynamically updated and tokens are allocated as needed. When a business request is received, it is checked whether there are enough tokens in the current token bucket. If there are enough tokens, the corresponding number of tokens is taken out and the business request is processed. If there are insufficient tokens, the business request is processed according to a preset policy (such as queuing or rejection).

[0089] Step 6 further includes monitoring and analyzing the read-write ratio in the business request, and performing differential speed limit control in combination with the read-write ratio and the business priority.

[0090] Analyzing the read-write ratio in the business area request based on the read requests and write requests in the business request identified in step 2, and performing differential speed limit control in combination with the read-write ratio and the business priority specifically means: implementing more stringent speed limit control for the write requests than for the read requests. The speed limit control for the write requests can be dynamically adjusted according to factors such as system load, business priority, and read-write ratio to achieve more refined traffic control and resource allocation.

[0091] Such as Figure 2 shown, which shows a distributed hierarchical speed limit device based on real-time load perception, including:

[0092] An initialization module, which is configured to set the system basic parameters and the load data collection period.

[0093] Specifically, the basic parameters include basic information such as the number of CPU cores, the TPS gear of transactions per second, and the initial speed limit parameters.

[0094] In a distributed system, the number of CPU cores can be obtained using different methods according to different operating systems. Since the obtaining method belongs to the category of existing technologies, the obtaining method of the number of CPU cores will not be introduced in detail here. After obtaining the number of CPU cores, it is used as the reference value for subsequent calculation of the standardized index threshold.

[0095] Embodiments of the present invention set a series of TPS gears, such as: 120, 60, 40,..., 2, 1, to meet the processing requirements under different loads.

[0096] Set the load data collection period, and use the timer function of the system or configure a timed task to ensure that the load data is automatically collected every 30 seconds.

[0097] A receiving module that receives service requests, identifies and differentiates read requests, write requests, and service priorities in the service requests.

[0098] A load data acquisition module that periodically acquires load data of distributed nodes and calculates standardized metric thresholds.

[0099] Periodically call the system API, such as the proc / loadavg file or a dedicated monitoring interface, to obtain system load data. Taking the PostgreSQL database as an example, after installing the system_stat plugin, executing the SQL query SELECT * FROM pg_sys_os_info() can obtain detailed load information of the database server.

[0100] After obtaining the system load data, calculate the following standardized metric thresholds based on the obtained load data and the number of CPU cores:

[0101] 1-minute metric threshold: threshold_1m = load_1m / number of CPU cores;

[0102] 5-minute metric threshold: threshold_5m = load_5m / number of CPU cores;

[0103] 10-minute metric threshold: threshold_10m = load_10m / number of CPU cores.

[0104] These metric thresholds will be used for subsequent load status judgments. load_1m, load_5m, and load_10m are the loads of the server in the most recent 1 minute, 5 minutes, and 10 minutes respectively.

[0105] A load status judgment module that judges the load status in combination with the standardized metric thresholds.

[0106] First, define the load status, which is defined as three situations, namely healthy status, sub-healthy status, and unhealthy status. The judgment criteria for each status are as follows:

[0107] Healthy status: threshold_5m < 1 and threshold_10m < 1, indicating that the system load is low and resources are abundant;

[0108] Sub-healthy status: threshold_5m ≥ 1 or threshold_10m ≥ 1, indicating that the system load is rising and needs to be closely monitored;

[0109] Unhealthy state: threshold_5m > 1 and threshold_10m > 1, indicating that the system load is too high and may face resource bottlenecks.

[0110] Automatically judge the current state of the system according to the calculated load index threshold.

[0111] Dynamic adjustment module, which dynamically adjusts the TPS (Transactions Per Second) gear according to the load status.

[0112] Dynamically adjusting the TPS gear according to the load status specifically includes:

[0113] When the load status is healthy, automatically raise the TPS gear by 1 level to make full use of system resources and improve throughput.

[0114] When the load status is sub-healthy, lower the TPS gear by 1 level to relieve system pressure and prevent the load from rising further.

[0115] When the load status is unhealthy, urgently lower the TPS gear by 2 levels to quickly reduce the system load and ensure the stable operation of core functions.

[0116] Rate limiting execution module, which uses the token bucket algorithm to dynamically adjust the current maximum number of tokens in the current token bucket of the system according to the current TPS gear and allocate tokens as needed. Based on the received service request, check whether there are enough tokens in the current token bucket. If there are enough tokens, allow the service request to pass; otherwise, reject the service request or queue the service request.

[0117] Specifically, controlling traffic using the token bucket algorithm includes: setting the maximum capacity of the token bucket to 120 tokens. According to the TPS gear, pre-define the number of tokens that can be obtained for each service request, and the number of tokens ranges from the set {2, 3, 4, 5, 6, 7, 8, 10, 12, 15, 20, 30, 60, 120}. According to the adjusted TPS gear in step 5, dynamically update the maximum number of tokens in the token bucket and allocate tokens as needed. When a service request is received, check whether there are enough tokens in the current token bucket. If there are enough tokens, take out the corresponding number of tokens and process the service request. If there are not enough tokens, process the service request according to a preset policy (such as queuing or rejection).

[0118] The rate limiting execution module further monitors and analyzes the read-write ratio in the service request, and performs differentiated rate limiting control in combination with the read-write ratio and service priority.

[0119] Analyze the read-write ratio in the service area request based on the read requests and write requests in the service request identified by the receiving module, and perform differential rate limiting control in combination with the read-write ratio and service priority. Specifically: Implement stricter rate limiting control for the write requests than for the read requests. The rate limiting control for the write requests can be dynamically adjusted according to factors such as system load, service priority, and read-write ratio to achieve more refined traffic control and resource allocation.

[0120] The above-mentioned distributed hierarchical rate limiting method and system based on real-time load perception can respond in real time, and the rate limiting rules can be adjusted according to the real-time state of the backend resources to ensure the efficient utilization and stable operation of the system. At the same time, it has strong adaptability. In the case of low load, the rate limiting policy will automatically relax the request limit to increase the throughput; in the case of high load, the system will tighten the flow control in a timely manner to avoid resource exhaustion. In addition, it can also achieve finer-grained rate limiting control, support multi-gear dynamic rate limiting configuration, and combine service characteristics (such as read-write ratio) to achieve differential processing of different types of requests.

[0121] Through this dynamic traffic control mechanism, the system not only effectively alleviates the pressure on the backend resources caused by high concurrency, but also can provide a more stable and reliable service experience for users. It not only effectively alleviates the pressure on the backend resources caused by high concurrency, but also can provide a more stable and reliable service experience for users.

[0122] In particular, according to the embodiments of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart.

[0123] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any one of the methods in this embodiment.

[0124] For the computer-readable storage medium in the embodiments of the present invention, those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to the computer program. The foregoing computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.

[0125] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A hierarchical speed limit method for a distributed system based on real-time load perception, characterized in that Including: An initialization step to configure the basic parameters of the distributed system and the load data collection period; Receiving a service request, identifying and differentiating read requests, write requests, and service priorities in the service request; Periodically collecting the load data of distributed nodes and calculating the standardized metric thresholds; Judging the load status in combination with the standardized metric thresholds; Dynamically adjusting the Transactions Per Second (TPS) gear according to the load status; Using the token bucket algorithm to dynamically adjust and allocate tokens as needed the current maximum number of tokens in the current token bucket of the distributed system according to the current TPS gear, and checking whether there are enough tokens in the current token bucket based on the received service request. If there are enough tokens, allow the service request to pass; otherwise, reject the service request or queue the service request; Monitoring and analyzing the read / write ratio in the service request, and performing differential rate limiting control in combination with the read / write ratio and the service priority; Periodically calling the API or dedicated monitoring interface of the distributed system to obtain load data and calculating the standardized metric thresholds for 1 minute, 5 minutes, and 10 minutes; The standardized metric thresholds are calculated in the following way: 1-minute metric threshold = 1-minute load / number of CPU cores; 5-minute metric threshold = 5-minute load / number of CPU cores; 10-minute metric threshold = 10-minute load / number of CPU cores; The load status includes: healthy status where the 5-minute metric threshold < 1 and the 10-minute metric threshold < 1; sub-healthy status where the 5-minute metric threshold ≥ 1 or the 10-minute metric threshold ≥ 1; unhealthy status where the 5-minute metric threshold > 1 and the 10-minute metric threshold > 1; The dynamic adjustment of the Transactions Per Second (TPS) gear includes: increasing the TPS gear by 1 level when the load status is the healthy status; decreasing the TPS gear by 1 level when the load status is the sub-healthy status; decreasing the TPS gear by 2 levels when the load status is the unhealthy status.

2. The method according to claim 1, characterized in that, The basic parameters of the distributed system include: number of CPU cores, Transactions Per Second (TPS) gear, and initial rate limiting parameters, etc.

3. The method according to claim 1, wherein The TPS gear includes the following levels: 120, 60, 40, 30, 24, 20, 17, 15, 12, 10, 8, 6, 4, 2, 1.

4. The method according to claim 1, wherein The key parameters of the token bucket algorithm include: the maximum number of tokens in the token bucket and the number of tokens obtained each time. The maximum number of tokens in the token bucket is 120, and the number of tokens obtained each time is selected from the set {2, 3, 4, 5, 6, 7, 8, 10, 12, 15, 20, 30, 60, 120}.

5. The method according to claim 1, further comprising performing more stringent rate limiting control on the write request.

6. A distributed system hierarchical speed limit device based on real-time load perception, characterized in that, Including: An initialization module that configures the basic parameters of the distributed system and the load data collection period; A receiving module that receives a service request, identifies and differentiates read requests, write requests, and service priorities in the service request; Load data acquisition module, which periodically acquires the load data of distributed nodes and calculates the standardized metric thresholds; Load status judgment module, which judges the load status in combination with the standardized metric thresholds; Dynamic adjustment module, which dynamically adjusts the TPS (Transactions Per Second) gear according to the load status; Rate limiting execution module, which uses the token bucket algorithm to dynamically adjust and allocate tokens as needed for the current maximum token quantity in the current token bucket of the distributed system according to the current TPS gear, and checks whether there are sufficient tokens in the current token bucket based on the received service request. If there are sufficient tokens, the service request is allowed to pass; otherwise, the service request is rejected or queued; The rate limiting execution module further monitors and analyzes the read / write ratio in the service request, and performs differential rate limiting control in combination with the read / write ratio and the service priority; Periodically call the API or dedicated monitoring interface of the distributed system to obtain load data and calculate the standardized metric thresholds for 1 minute, 5 minutes, and 10 minutes; The standardized metric thresholds are calculated as follows: 1-minute metric threshold = 1-minute load / number of CPU cores; 5-minute metric threshold = 5-minute load / number of CPU cores; 10-minute metric threshold = 10-minute load / number of CPU cores; The load status includes: healthy status where the 5-minute metric threshold < 1 and the 10-minute metric threshold < 1; sub-healthy status where the 5-minute metric threshold ≥ 1 or the 10-minute metric threshold ≥ 1; unhealthy status where the 5-minute metric threshold > 1 and the 10-minute metric threshold > 1; The dynamic adjustment of the TPS gear includes: increasing the TPS gear by 1 level when the load status is the healthy status; decreasing the TPS gear by 1 level when the load status is the sub-healthy status; decreasing the TPS gear by 2 levels when the load status is the unhealthy status.

Citation Information

Patent Citations

  • Object storage distributed service quality optimization method, server and storage device

    CN113296717A

  • Client service speed limiting method, device and equipment and storage medium

    CN114020209A