A method and device for limiting the traffic of file downloads

Through the token bucket algorithm combined with the priority queue, the shortcomings of the existing current limiting algorithm in handling burst traffic and high-priority requests are solved, and the reasonable allocation and stability of server resources are achieved.

CN115914119BActive Publication Date: 2025-07-11INSPUR QILU SOFTWARE IND
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211553606.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-06
Publication Date
2025-07-11
Estimated Expiration
2042-12-06

AI Technical Summary

Technical Problem

The existing current limiting algorithms are difficult to effectively handle high-priority requests when facing burst traffic, and cannot reasonably allocate server resources, resulting in a decrease in server stability and response efficiency.

Method used

The token bucket algorithm is used to combine the priority queue, and token buckets of different sizes and rates are configured to process download requests of different priority levels, and to adjust the priority to prioritize high priority requests when large traffic bursts.

Benefits of technology

It realizes reasonable load allocation for different download requests, improves the stability and response efficiency of the server, prevents traffic waste caused by blocking of high-priority tasks, and maintains the smoothness of download tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115914119B_ABST
    Figure CN115914119B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of server traffic limiting, and specifically provides a method for traffic limiting of file downloads. First, the token bucket is configured, setting the size of the token bucket and the rate at which tokens are placed in the token bucket. Then, the priority queue is configured and the file download is segmented. The requested file to be downloaded is stored in the MinIO cluster. After the server receives a request for segmented download, it will obtain a piece of data from MinIO and cache it locally. Compared with the prior art, the present invention can correctly allocate the load of the server for different download requests, greatly improving the stability of the server program in the Xinchuang environment. While optimizing the token bucket, the traffic limiting function of the token bucket and the load capacity for instantaneous large traffic are maintained, and the ability to hierarchically process requests is coupled, having good popularization value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of server traffic limiting, and specifically provides a file download traffic limiting method and device. Background Art

[0002] In recent years, the country has strongly supported the development of domestic software and hardware with independent intellectual property rights, and many basic software and hardware products with independent intellectual property rights have emerged, represented by domestic operating systems and CPUs. The ecological environment of domestic operating systems such as the Galaxy Kylin system and the Tongxin UOS operating system has become increasingly perfect, and high-end general-purpose chips with independent intellectual property rights such as Loongson and Feiteng have developed vigorously, and their technical levels have reached or approached the world advanced level of similar products.

[0003] With the vigorous development of domestic basic software and hardware, the popularization and use of domestic basic software and hardware have brought unprecedented opportunities. As a device that undertakes corresponding service requests, provides services, and guarantees services, the server has extremely high requirements in terms of performance and stability.

[0004] Server availability means that the selected server can meet the requirements of long-term stable operation. Download requests consume a large amount of traffic on the server, especially requests for downloading and distributing some large files. Limiting the traffic of the download module helps to maintain the stability of the server and prevent the impact on the server caused by the pressure of a large number of download tasks.

[0005] The current main traffic limiting algorithms include the fixed window algorithm, the sliding window algorithm, the leaky bucket algorithm, the token bucket algorithm, etc. The fixed window algorithm lacks a solution to sudden traffic; although the sliding window algorithm solves the problem of sudden traffic, it is difficult to control the request frequency in real scenarios; the leaky bucket algorithm can only limit the traffic at a constant rate when facing sudden traffic; although the token bucket algorithm can handle sudden traffic well, it cannot process requests related to priorities. When encountering large traffic requests with a long duration, it can only buffer the server pressure and cannot effectively and quickly respond to high-priority requests. Summary of the Invention

[0006] The present invention aims at the deficiencies of the above-mentioned existing technologies and provides a highly practical file download traffic limiting method.

[0007] A further technical task of the present invention is to provide a file download traffic limiting device with reasonable design, safety and applicability.

[0008] The technical solution adopted by the present invention to solve its technical problems is:

[0009] A file download rate limiting method, characterized in that, first, the token bucket is configured, the size of the token bucket and the rate at which tokens are put into the token bucket are set, then the priority queue is configured and the file download is segmented. The requested file is stored in the MinIO cluster. After the server receives the segmented download request, it will obtain a piece of data from MinIO and cache it locally.

[0010] Further, the size of the token bucket, that is, the capacity of the token bucket, is determined by the maximum instantaneous traffic that the server can handle. When the token bucket is filled, the capacity of the bucket determines the size of the rate limiting;

[0011] The rate at which tokens are put into the token bucket is used to limit the average request rate.

[0012] Further, when configuring the priority queue, after the server obtains the request, it first enters the priority queue. After dequeuing, it enters the token bucket corresponding to different priorities;

[0013] The higher the priority, the smaller the size of the corresponding token bucket, and the faster the rate at which tokens are put in;

[0014] When the token bucket corresponding to the low-priority request is blocked, a delay time is set for the request. If the request fails multiple times, the priority can be adjusted appropriately.

[0015] Further, in the file download segmentation, when the token bucket is running, one request corresponds to one token, and the maximum amount of data for a single download is restricted to control the overall traffic.

[0016] Further, when downloading in segments, by setting a threshold, files larger than the threshold size need to be segmented, and files smaller than the threshold are downloaded normally;

[0017] The lowest priority queue uses a large bucket, and the low token input rate is designed to handle ordinary download requests received by the server;

[0018] The token bucket used by the medium priority queue should have a capacity lower than that of the lowest priority token bucket, and the input rate should be similar to but slightly lower than that of the lowest priority token bucket. The token bucket used by the medium priority queue is used to share the pressure on the token bucket of the low priority queue;

[0019] The high priority token bucket should use a small bucket and a high token input rate to handle urgent download requests.

[0020] Further, under normal circumstances, except for the marked high priority tasks, all tasks enter the token bucket from the low priority queue. After obtaining the token, the server processes the download request.

[0021] Further, when a large amount of traffic bursts, the tokens in the token bucket of the low-priority queue are emptied, and the server is in a nearly fully loaded state. The overflow requests are blocked on the server side and wait for a period of time. If new tokens are added to the token bucket within this time interval, the overflow requests will be responded to in order.

[0022] If the overflow requests are not responded to during this period, a timeout will be triggered, an HTTP code 429 "Too Many Requests" will be returned, and the estimated waiting time interval will be calculated based on the low-priority token bucket rate.

[0023] Further, in the case of continuous large traffic, the number of download requests is lower than that in the case of bursty large traffic, but it can still cause the token bucket of the low-priority queue to overflow and will continue for a period of time.

[0024] Due to continuous large numbers of requests, the timeout waiting time returned by the low-priority token bucket cannot be accurately estimated. After the second request for a timeout request still fails, the priority will be raised to medium priority.

[0025] The medium priority has a faster token bucket generation rate and can quickly share the requests that the token bucket of the low-priority queue cannot handle.

[0026] After the request rises from the low-priority queue to the medium-priority queue, it generally will not continue to increase in priority; the high-priority queue is generally empty and only processes important download requests or the prerequisite requests for important download requests.

[0027] A file download traffic limiting device includes: at least one memory and at least one processor;

[0028] The at least one memory is used to store machine-readable programs;

[0029] The at least one processor is used to call the machine-readable program and execute a file download traffic limiting method.

[0030] Compared with the prior art, a file download traffic limiting method and device of the present invention have the following outstanding beneficial effects:

[0031] The present invention can correctly allocate the load of the server for different download requests, greatly improving the stability of the server program in the information and communication technology (ICT) innovation environment. While optimizing the token bucket, it maintains the traffic limiting function of the token bucket and the load capacity for instantaneous large traffic, and couples the ability to hierarchically process requests.

[0032] By dividing the priorities of download requests, the smoothness of the overall download task can be maintained, preventing waste of traffic caused by blocking of high-priority tasks.

[0033] By adjusting the token bucket and priority queue, it adapts to different server environments, and the implementation method is simple and effective. Brief Description of the Drawings

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0035] Appendix Figure 1 is a basic schematic diagram of token bucket optimization in a file download rate limiting method;

[0036] Appendix Figure 2 is a basic schematic diagram of token bucket optimization in a file download rate limiting method. Detailed Embodiments

[0037] In order to enable those skilled in the art of this technology to better understand the solution of the present invention, the following further details the present invention in combination with specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0038] The following gives an optimal embodiment:

[0039] As Figure 1 、 2 shown, a file download rate limiting method in this embodiment, in actual development, the process of file download generally consists of a file storage system, a server download module, a rate limiting module, and a client receiving file module. The file storage system stores the files to be downloaded, and it is necessary to ensure the reliability and stability of the entire file system; the download module mainly cooperates with the client receiving module to complete the main operations of file download, including special processing such as segmented downloading of large files and resume from breakpoint; the rate limiting module restricts the traffic of file download to prevent excessive traffic from impacting the server and the file storage system.

[0040] Common algorithms include fixed window algorithm, sliding window algorithm, leaky bucket algorithm, and token bucket algorithm.

[0041] The fixed window algorithm uses the number of accesses within a time window as the basis for flow control, and the counter is automatically reset every time a time window passes. The problem with the fixed window is that if a large number of requests flood in within a very short period before and after the window resets, then the server has to bear double the load during this time period. For example, if the fixed window reset time interval is 1 second and the maximum load within 1 second is 10,000 requests, and 10,000 requests are received respectively within 100 ms before the end of the current window and 100 ms after the start of the next window, this means that the server has to process 20,000 requests within 200 ms. This exponential growth may far exceed the server's capacity.

[0042] The sliding window solves the problem of the fixed window to a certain extent. The sliding window further divides the fixed window into smaller intervals and keeps the sum unchanged within the time interval of the entire fixed window. Taking a time window of 1 second as an example, the 1-second window is equally divided into 10 small windows, and each window slides forward by 0.1 second. Each time it slides, the request volume received within the 0.1-second small window that has slid out is recycled, so as to maintain the stability of the total number of requests within a large window. The pain point of this algorithm is that it cannot control the request frequency in the real scenario.

[0043] The basic idea of the leaky bucket algorithm is that traffic continuously enters the leaky bucket, and requests are processed at a constant speed at the bottom. If the rate of traffic entering is higher than the rate of requests at the bottom and the traffic in the bucket exceeds the bucket size, the traffic will overflow. The leaky bucket algorithm has a wide inlet and a strict outlet. No matter how large the request rate is, the processing at the bottom is carried out uniformly, which is somewhat similar to the processing mechanism of a message queue. However, in the face of bursty traffic, we often expect the system to be stable while processing requests faster, rather than continuing to work step by step.

[0044] The token bucket algorithm is similar to the leaky bucket, except that the leaky bucket processes requests at a constant speed at the bottom, while the token bucket inserts tokens into the bucket at a constant speed, and requests will only be processed by the server if they obtain a token. For sudden large traffic situations, as long as the traffic does not exceed the total capacity of the token bucket, smooth processing can be carried out. However, the token bucket lacks a priority mechanism.

[0045] The specific steps in the present invention are as follows:

[0046] S1. Token bucket configuration;

[0047] There are mainly two settings for the token bucket algorithm. One is the size of the token bucket, that is, the bucket capacity, which is determined by the maximum instantaneous traffic that the server can handle. When the token bucket is filled, the bucket capacity determines the size of the flow control.

[0048] The other is the rate at which tokens are put into the token bucket. This rate is mainly used to limit the average request rate.

[0049] S2. Priority queue configuration;

[0050] Optimize the token bucket using a priority queue. The priority queue is established based on a normal queue. The priority queue maintains the first-in, first-out feature of the normal queue. After the server obtains a request, it first enters the priority queue. After leaving the queue, it enters the token buckets corresponding to different priorities. The higher the priority, the smaller the size of the corresponding token bucket, and the faster the rate of putting tokens.

[0051] When the token bucket corresponding to a low-priority request is blocked, set a delay time for the request. If the request fails multiple times, the priority can be adjusted appropriately.

[0052] S3. Segment the file download;

[0053] When the token bucket is running, one request corresponds to one token, and the maximum data volume of a single download is restricted to achieve the purpose of overall traffic control.

[0054] S4. Store the requested file for download in the MinIO cluster;

[0055] MinIO is a distributed file system that uses erasure coding to protect data. It splits the data into segments, expands and encodes the redundant data blocks, and stores them in different locations. It can be understood as splitting the object file and then encoding it to prevent any two pieces of data from being lost.

[0056] After the server receives a request for segmented download, it will obtain a segment of data from MinIO and cache it locally to reduce the load pressure on MinIO.

[0057] Segmented download is achieved by setting a threshold. Files larger than this threshold size need to be segmented, and other files can be downloaded normally.

[0058] The lowest-priority queue uses a large bucket and a low token input rate design, mainly for handling ordinary download requests received by the load server. The low token input rate is relative to the other two high-priority queues. The token input rate of the token bucket should still meet the download requests under normal circumstances.

[0059] The token bucket used by the medium-priority queue should have a capacity lower than that of the lowest-priority token bucket, and the input rate should be similar to but slightly lower than that of the lowest-priority token bucket. The token bucket used by the medium-priority queue is mainly used to share the pressure on the token bucket of the lowest-priority queue.

[0060] The high-priority token bucket should use a small bucket and a high token input rate. It is mainly used to handle urgent download requests.

[0061] Under normal circumstances, except for the marked high-priority tasks, all tasks enter the token bucket from the low-priority queue. After obtaining a token, the server processes the download request.

[0062] When a large traffic burst occurs, the tokens in the token bucket of the low-priority queue are emptied, and the server is in a near-full load state. The overflow requests are blocked at the server side and wait for a period of time. If new tokens are added to the token bucket within this time interval, the overflow requests are correspondingly processed in order; if the overflow requests are not correspondingly processed during this period, a timeout is triggered, and httpcode 429 Too Many Requests is returned, and the estimated waiting time interval is calculated based on the low-priority token bucket rate.

[0063] In this case, the number of download requests is lower than the large traffic burst, but it can still cause the token bucket of the low-priority queue to overflow and will continue for a period of time.

[0064] In this case, due to the continuous large number of requests, the timeout waiting time returned by the low-priority token bucket cannot be accurately estimated. After the second request of the timeout request still fails, the priority will be raised to medium priority. The medium priority has a faster token bucket generation rate and can quickly share the requests that the token bucket of the low-priority queue cannot handle. However, due to the smaller size of the token bucket of the medium-priority queue, the total number of requests in the load will not be too large, which can ensure the health of the server and play a role in traffic limiting.

[0065] After the request rises from the low-priority queue to the medium-priority queue, it generally will not continue to increase the priority. The high-priority queue is generally empty and only processes important download requests or the prerequisite requests of important download requests.

[0066] A file download traffic limiting device includes: at least one memory and at least one processor;

[0067] The at least one memory is used to store machine-readable programs;

[0068] The at least one processor is used to call the machine-readable program and execute a file download traffic limiting method.

[0069] The above specific implementation manners are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific implementation manners. Any appropriate changes or substitutions made by any person of ordinary skill in the art in the technical field that conform to the claims of a file download traffic limiting method and device of the present invention shall fall within the patent protection scope of the present invention.

[0070] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirits of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for limiting the traffic of file downloads, characterized in that, First, configure the token bucket, set the size of the token bucket and the rate at which tokens are placed in the token bucket. Then, configure the priority queue and segment the file download. The requested file is stored in the MinIO cluster. After the server receives a request for segmented download, it fetches a segment of data from MinIO and caches it locally; The size of the token bucket, i.e., the capacity of the token bucket, is determined by the maximum instantaneous traffic that the server can handle. When the token bucket is full, the capacity of the bucket determines the size of the flow limit; The rate at which tokens are placed in the token bucket is used to limit the average request rate; When configuring the priority queue, after the server obtains a request, it first enters the priority queue. After leaving the queue, it enters the token bucket corresponding to the different priorities; The higher the priority, the smaller the size of the corresponding token bucket, and the faster the rate at which tokens are placed; When the token bucket corresponding to a low-priority request is blocked, set a delay time for the request. If the request fails multiple times, adjust the priority appropriately; In the file download segmentation, when the token bucket is running, one request corresponds to one token, and the maximum amount of data for a single download is limited to control the overall traffic; During segmented download, by setting a threshold, files larger than the threshold size need to be segmented, and files smaller than the threshold are downloaded normally; The lowest-priority queue uses a large bucket, and a low token input rate is designed to handle ordinary download requests received by the load server; The token bucket used by the medium-priority queue should have a capacity lower than that of the lowest-priority token bucket, and the input rate should be similar to but slightly lower than that of the lowest-priority token bucket. The token bucket used by the medium-priority queue is used to share the pressure on the token bucket of the low-priority queue; The high-priority token bucket should use a small bucket and a high token input rate to handle urgent download requests; Normally, except for the marked high-priority tasks, all tasks enter the token bucket from the low-priority queue. After obtaining a token, the server processes the download request; During a large traffic burst, the tokens in the token bucket of the low-priority queue are emptied, and the server is in a nearly fully loaded state. The overflow requests are blocked at the server and wait for a period of time. If new tokens are added to the token bucket during this time interval, the overflow requests are responded to in order; If the overflow requests are not responded to during this period, a timeout is triggered, and an HTTP code 429 Too Many Requests is returned, and the estimated waiting time interval is calculated based on the rate of the low-priority token bucket; In the case of continuous large traffic, the number of download requests is lower than that in the case of a large traffic burst, but it can still cause the token bucket of the low-priority queue to overflow and will continue for a period of time; Due to continuous large numbers of requests, the timeout waiting time returned by the low-priority token bucket cannot be accurately estimated. After the second request for a timeout request still fails, the priority is raised to the medium priority; The medium priority has a faster token bucket generation rate and can quickly share the requests that the token bucket of the low-priority queue cannot handle; After a request rises from the low-priority queue to the medium-priority queue, the priority will not continue to rise; The high-priority queue only processes important download requests or the prerequisite requests for important download requests.

2. A file download rate limiting device, characterized in that, Including: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is configured to call the machine-readable program to execute the method according to claim 1.

Citation Information

Patent Citations

  • Method and device for identifying congestion status of data transmission channel

    CN101267382A

  • Queue scheduling method based on multiple priorities

    CN104079501A

  • Flow control method of FC switching network based on priority

    CN114338544A

  • Method and system for packaging and downloading multiple files supporting user-defined file names

    CN115334070A