Token bucket-based distributed task concurrency control method
By using the token bucket structure to manage the concurrent number of tasks in a distributed system, the problem of inaccurate concurrent number control in the existing technology is solved, and precise control and flexible adjustment of the concurrent number of tasks in a distributed system is realized, and the stability and controllability of the system are improved.
Patent Information
- Application Number
- PCT/CN2023/138458
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
The existing task concurrency control method is difficult to accurately reduce the concurrency after reaching the maximum concurrency number, which may lead to loss of task data or inability to achieve complete concurrency control.
The distributed task concurrency control method based on token bucket is adopted, and the token bucket list structure is used to initialize and manage the concurrent number of tasks. Before the task is executed, the token is obtained from the right side of the token bucket, and after execution, the token is returned from the left side, thereby achieving precise control of the concurrent number.
It realizes precise control of the concurrency number of tasks in distributed system, avoids the problem of overloading of website servers caused by excessive concurrency, and supports flexible adjustment of concurrency number, enhancing the controllability and stability of the system.
Smart Images

Figure CN2023138458_19062025_PF_FP_ABST
Abstract
Description
A method for concurrent control of distributed tasks based on token bucket Technical Field
[0001] The present invention relates to the field of control methods, in particular to a token bucket-based distributed task concurrency control method. Background Art
[0002] Website organizers typically rent virtual servers from cloud service providers or use servers in their own computer rooms to provide access services. Neither method allocates significant bandwidth, nor does it typically equip powerful servers to host website services. Current mainstream information collection systems are unable to effectively control the pressure of concurrent access to the collected websites, and may even cause website server overloads, impacting the website's ability to provide services. For example, in a government website census, to monitor content updates on government websites, traditional collection systems use distributed multi-node collection systems to continuously access the homepage of each column on the government website to determine whether new content has been updated. Uncontrolled concurrent access can easily lead to government websites being unable to provide public services due to excessive concurrent traffic, especially if the website has a large number of columns.
[0003] With existing task concurrency control, if the maximum number of concurrent tasks has been reached, reducing the maximum number of concurrent tasks requires either abruptly terminating the currently executing task or failing to reduce the number of concurrent tasks at all. The former results in task data loss, while the latter prevents complete concurrency control of the system.
[0004] Summary of the Invention
[0005] The purpose of the present invention is to provide a method for concurrent control of distributed tasks based on token buckets.
[0006] To achieve the purpose of the present invention, the technical solution is: a method for distributed task concurrency control based on a token bucket, comprising the following steps: (1) creating a task, initializing the number of tokens in a token bucket according to the concurrency number, wherein the token bucket is a list structure, tokens are put into the list from the left side, and tokens are taken out from the right side;
[0007] (2) When executing a task, first try to get a token from the right side of the list. If the token is successfully obtained, the task will start to execute. If the token cannot be obtained, wait for the set time and then try to get the token again.
[0008] (3) First take a token from the right side of the list. If no token can be obtained, wait for a while before trying to obtain a token again;
[0009] (4) When the task is completed, return the token obtained at the start of the task from the left side of the list to the token bucket.
[0010] Further, the identifier of the list described in step (1) is the task name.
[0011] Further, step (3) also includes: when the task execution fails, return the token obtained when the task starts to execute to the left side of the current task name.
[0012] Further, when the task execution times out or the node executing the task fails due to some reason, mark the task as a "zombie task", and return the token obtained at the start of the task from the left side of the list to the token bucket.
[0013] Further, the method for increasing the concurrency is as follows: obtain the current maximum concurrency of the task; subtract the current maximum concurrency from the new value to obtain the number of tokens n to be increased, and return n tokens from the left side of the list corresponding to the current task; modify the current maximum concurrency of the current task to the new value.
[0014] Further, the method for decreasing the concurrency is as follows: obtain the current maximum concurrency of the task; subtract the new value from the current maximum concurrency to obtain the number of tokens n to be decreased, and obtain n tokens from the right side of the list corresponding to the current task; modify the current maximum concurrency of the current task to the new value.
[0015] Further, if the x - th token acquisition fails, calculate y = n - x, and insert y tokens on the left side of the token ledger corresponding to the current task; where x < n.
[0016] Further, when the token ledger is synchronized, obtain the list of tokens to be synchronized from the token ledger; repeatedly attempt to call the token acquisition interface to attempt to obtain a token from the token bucket with the key being the current task name; if the token acquisition fails, rest for 10 milliseconds and then continue to attempt to obtain the token; if the token acquisition is successful, record the number of successfully obtained tokens, assumed to be z; if z < y, continue to attempt to obtain the token, if z = y, it means that all the tokens to be synchronized have been synchronized, and exit the loop.
[0017] Compared with the prior art, the significant advantage of this invention is that: this invention provides an implementation method for distributed task concurrency control based on a token bucket, enabling the distributed system to more precisely control the concurrency of tasks. It can be used in various distributed information collection systems and has a wide range of application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is the overall flowchart of the distributed task concurrency control system technology based on a token bucket in an embodiment of this invention;
[0019] Figure 2 is the flowchart for obtaining a token in an embodiment of this invention;
[0020] FIG3 is a flowchart of returning a token in an embodiment of the present invention;
[0021] FIG4 is a flowchart of a timeout token process according to an embodiment of the present invention;
[0022] FIG5 is a flow chart of increasing the number of concurrent connections according to an embodiment of the present invention;
[0023] FIG6 is a flowchart of reducing the number of concurrent connections according to an embodiment of the present invention;
[0024] FIG7 is a flowchart of token ledger synchronization in an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0026] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0027] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0028] Example
[0029] Technical solution: The present invention discloses a method for implementing distributed task concurrency control based on token bucket.
[0030] Refer to Figure 1, token bucket and concurrency:
[0031] 1. Use Redis (an in-memory database) list structure to represent the token bucket. The key (identifier) of the list is the task name. When creating a task, the number of tokens in the token bucket is initialized according to the number of concurrent tasks.
[0032] 2. Put tokens from the left side of the list and take tokens from the right side of the list. The length of the list represents the number of concurrent tasks.
[0033] 3. When executing a task in a distributed system, you need to first obtain a token from the right side of the list before executing the task. If you cannot obtain a token, it means that the number of concurrent tasks in the current system has reached the maximum limit. You can wait for a while before trying to obtain a token again.
[0034] 4. When a task in a distributed system is completed or ends due to an error during task execution, the token obtained at the beginning of the task should be returned to the token bucket from the left side of the list;
[0035] 5. Since the size of the token bucket represents the maximum number of tasks that can be executed simultaneously, the number of concurrent tasks in the entire distributed system can be accurately controlled.
[0036] Refer to Figure 2 to obtain the token:
[0037] 1. Try to get a token from the right side of the List whose key is the current task name;
[0038] 2. If the acquisition fails, it will sleep for a while and then try to obtain the token again;
[0039] 3. Once the token is successfully obtained, the task will begin to execute;
[0040] Refer to Figure 3 to return the token:
[0041] 1. When the task is completed, the token obtained when the task starts is returned to the left side of the List whose key is the current task name;
[0042] 2. If the task fails to execute, the token obtained when the task starts is returned to the left side of the List whose key is the current task name;
[0043] Referring to Figure 4, timeout token processing:
[0044] 1. If the task execution times out or the node executing the task crashes for some reason, the "zombie task" management program will mark the task as a "zombie task" and return the token;
[0045] 2. "Zombie tasks" will not continue to execute tasks even after the node is restored, and there is no need to return tokens;
[0046] Refer to Figure 5, the process of increasing the number of concurrent connections:
[0047] 1. Get the current maximum number of concurrent connections for the task from the task information hash (a collection-type data structure in Redis) in Redis;
[0048] 2. Subtract the current value from the new value to get the number of tokens to be added, assuming it is n;
[0049] 3. Find the list in Redis according to the task name, call the return token interface, and return n tokens from the left side of the list;
[0050] 4. Modify the current maximum concurrency of the task in the task information hash in Redis to the new value;
[0051] Referring to FIG. 6, the process of reducing the concurrency number:
[0052] 1. Obtain the current maximum concurrency number of the task from the task information hash (Hash, a set - type data structure in Redis) in Redis;
[0053] 2. Subtract the new value from the current value, and the result is the number of tokens to be reduced, assumed to be n;
[0054] 3. Find the list in Redis according to the task name, and loop to call the token - obtaining interface to obtain n tokens from the right side of the list;
[0055] 4. If the x - th (x < n) token - obtaining fails, calculate the value of n - x (the computer counts from 0), assumed to be y, and insert y tokens on the left side of the token ledger (Redis list, List data structure) with the key being the current task name for the subsequent "token ledger synchronization service" to use. This step is simply referred to as bookkeeping;
[0056] 5. Modify the current maximum concurrency number of the task in the task information hash in Redis to the new value;
[0057] Referring to FIG. 7, the token ledger synchronization service (reconciliation):
[0058] 1. Obtain the list of tokens to be synchronized from the token ledger (similar to obtaining the amount to be reconciled);
[0059] 2. Loop to try to call the token - obtaining interface to try to obtain a token from the token bucket with the key being the current task name;
[0060] 3. If the token - obtaining fails, rest for 10 milliseconds and then continue to try to obtain the token;
[0061] 4. If the token - obtaining is successful, record the number of tokens that have been successfully obtained, assumed to be z;
[0062] 5. If z
Claims
1. A method for distributed task concurrent control based on a token bucket, characterized in that: It includes the following steps: (1) Create a task, initialize the number of tokens in the token bucket according to the concurrency number. The token bucket is a list structure. Tokens are put in from the left side of the list and taken out from the right side of the list; (2) When executing the task, first try to take out a token from the right side of the list. If the token is successfully obtained, start executing the task; if the token cannot be obtained, wait for the set time and then try to take out the token again; (3) First take out a token from the right side of the list. If the token cannot be obtained, wait for a period of time and then try to take the token again; (4) When the task is executed, return the token obtained at the start of the task from the left side of the list to the token bucket.
2. The method for distributed task concurrent control based on a token bucket according to claim 1, characterized in that: In step (1), the identifier of the list is the task name.
3. The method for distributed task concurrent control based on a token bucket according to claim 1, characterized in that: Step (3) also includes: when the task execution fails, return the token obtained when the task starts to the left side of the current task name.
4. The method for distributed task concurrent control based on a token bucket according to claim 1, characterized in that: When the task execution times out or the node executing the task fails due to reasons, mark the task as a "zombie task", and return the token obtained at the start of the task from the left side of the list to the token bucket.
5. The method for distributed task concurrent control based on a token bucket according to claim 1, characterized in that: The method for increasing the concurrency number is as follows: obtain the current maximum concurrency number of the task; subtract the current maximum concurrency number from the new value to obtain the number of tokens n to be increased, and return n tokens from the left side of the list corresponding to the current task; Modify the current maximum concurrency number of the current task to the new value.
6. The method for distributed task concurrent control based on a token bucket according to claim 1, characterized in that: The method for decreasing the concurrency number is as follows: obtain the current maximum concurrency number of the task; subtract the new value from the current maximum concurrency number to obtain the number of tokens n to be decreased, and obtain n tokens from the right side of the list corresponding to the current task; Modify the current maximum concurrency number of the current task to the new value.
7. The method for distributed task concurrent control based on a token bucket according to claim 6, characterized in that: If the xth token acquisition fails, calculate y = n - x, and insert y tokens on the left side of the token ledger corresponding to the current task; where x < n.
8. The method for distributed task concurrent control based on a token bucket according to claim 7, characterized in that: When the token ledger is synchronized, obtain the list of tokens to be synchronized from the token ledger; loop and try to call the token acquisition interface to try to obtain a token from the token bucket with the key as the current task name; If the token acquisition fails, rest for 10 milliseconds and then continue to try to acquire the token; If the token acquisition is successful, record the number of tokens that have been successfully obtained, assumed to be z; if z < y, continue to try to acquire the token, if z = y, it means that all the tokens to be synchronized have been synchronized, and exit the loop.
Citation Information
Patent Citations
Token management method and device, storage medium and electronic device
CN111314238A
Data traffic control method and device, electronic equipment and storage medium
CN114064219A
Distributed system current limiting method and system
CN114745334A
Distributed task concurrency control method based on token bucket
CN117082001A
Token bucket with active queue management
US20230198910A1