A token bucket-based traffic shaping control method, device, medium and product

CN122698534APending Publication Date: 2026-09-04BINZHOU WEIQIAO NATIONAL SCIENCE & TECHNOLOGY ADVANCED TECHNOLOGY RESEARCH INSTITUTE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610696589.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-09-04

AI Technical Summary

Technical Problem

但服务器的背景负载是动态变化的,致使资源利用不合理

Benefits of technology

[0011]The technical solution of this invention determines the load pressure index of the communication service scenario based on the load parameters and communication service scenario in network communication; based on the load pressure index, it determines the actual global token generation rate, the actual session-level micro-token generation rate, and the actual maximum token storage capacity of the token bucket to generate tokens; in network communication, it acquires user protocol data packets and retrieves tokens from the token bucket according to the user protocol data packets; when global tokens and session-level micro-tokens are acquired, the user protocol data packets are written to the application layer receive queue, and the business thread retrieves the data packets from the queue for processing; otherwise, the user protocol data packets are processed for packet loss. This solves the traffic shaping and control problem in network communication. By evaluating the communication load pressure in the communication service scenario, the token generation speed can be dynamically adjusted to achieve elastic traffic shaping and congestion control, ensuring reliable data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122698534A_ABST
    Figure CN122698534A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a token bucket-based traffic shaping control method and device, a medium and a product. The method comprises: determining a load pressure index of a communication service scenario according to a load parameter and the communication service scenario in network communication; determining an actual global token generation rate, an actual session-level micro token generation rate and an actual maximum token storage amount of a token bucket according to the load pressure index to generate tokens; in network communication, obtaining a user protocol data packet and obtaining tokens in the token bucket according to the user protocol data packet; when the global token and the session-level micro token are obtained, writing the user protocol data packet into an application layer receiving queue and taking out the data packet from the queue by a service thread for processing; otherwise, performing packet loss processing on the user protocol data packet. By evaluating the communication load pressure in the communication service scenario, the token generation speed is dynamically adjusted, flexible traffic shaping and congestion control are realized, and reliable data transmission is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication technology, and in particular to a token bucket-based flow shaping control method, device, medium, and product. Background Technology

[0002] In network communication and high-throughput data processing, User Datagram Protocol (UDP) is widely used in high-frequency, real-time data packet transmission scenarios due to its characteristics of low header overhead, no connection establishment / release process, no retransmission acknowledgment mechanism, low transmission latency, high transmission efficiency, and low resource overhead.

[0003] However, in gigabit (1000Mbps) / 10-gigabit (10Gbps) network environments, traditional UDP receiving and flow control solutions reveal numerous intractable flaws when faced with massive concurrent data transmissions from numerous terminals, severely restricting their reliability in critical business scenarios. For example, existing application-layer rate limiting typically uses a fixed-rate token bucket. However, the server's background load is dynamically changing, leading to inefficient resource utilization. Furthermore, the "thundering herd effect" frequently occurs in industrial settings, where all devices come online simultaneously after a power outage, generating instantaneous traffic spikes that directly overwhelm application-layer queues, causing abnormal packet processing.

[0004] Therefore, there is an urgent need to provide a traffic shaping control method that combines communication load pressure to dynamically adjust the token generation speed, thereby achieving elastic traffic shaping and congestion control. Summary of the Invention

[0005] This invention provides a token bucket-based flow shaping control method, device, medium, and product to dynamically adjust the token generation speed, achieve elastic flow shaping and congestion control, and ensure reliable data transmission.

[0006] According to one aspect of the present invention, a token bucket-based flow shaping control method is provided, the method comprising: Based on the load parameters in network communication and the communication service scenario, determine the load pressure index of the communication service scenario; Based on the load pressure index, determine the actual global token generation rate, actual session-level micro-token generation rate, and actual maximum token storage capacity of the token bucket to generate tokens; In network communication, user protocol data packets are acquired, and tokens are obtained from the token bucket according to the user protocol data packets; When a global token and a session-level micro token are obtained, the user protocol data packet is written into the application layer receive queue, and the business thread retrieves the data packet from the queue for processing; otherwise, the user protocol data packet is dropped.

[0007] According to another aspect of the present invention, a token bucket-based flow shaping control device is provided, the device comprising: The load pressure index determination module is used to determine the load pressure index of the communication service scenario based on the load parameters in the network communication and the communication service scenario. The token generation parameter determination module is used to determine the actual global token generation rate, actual session-level micro-token generation rate, and actual maximum token storage amount of the token bucket based on the load pressure index, so as to generate tokens. The token acquisition module is used to acquire user protocol data packets in network communication and acquire tokens from the token bucket according to the user protocol data packets. The data packet processing module is used to write the user protocol data packet into the application layer receive queue when a global token and a session-level micro token are obtained, and the business thread retrieves the data packet from the queue for processing; otherwise, the user protocol data packet is processed for packet loss.

[0008] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and a memory communicatively connected to said at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the token bucket-based traffic shaping control method according to any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the token bucket-based flow shaping control method according to any embodiment of the present invention.

[0010] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the token bucket-based flow shaping control method described in any embodiment of the present invention.

[0011] The technical solution of this invention determines the load pressure index of the communication service scenario based on the load parameters and communication service scenario in network communication; based on the load pressure index, it determines the actual global token generation rate, the actual session-level micro-token generation rate, and the actual maximum token storage capacity of the token bucket to generate tokens; in network communication, it acquires user protocol data packets and retrieves tokens from the token bucket according to the user protocol data packets; when global tokens and session-level micro-tokens are acquired, the user protocol data packets are written to the application layer receive queue, and the business thread retrieves the data packets from the queue for processing; otherwise, the user protocol data packets are processed for packet loss. This solves the traffic shaping and control problem in network communication. By evaluating the communication load pressure in the communication service scenario, the token generation speed can be dynamically adjusted to achieve elastic traffic shaping and congestion control, ensuring reliable data transmission.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart of a flow shaping control method based on a token bucket according to Embodiment 1 of the present invention; Figure 2 This is a flowchart of a flow shaping control method based on a token bucket according to Embodiment 2 of the present invention; Figure 3 This is an application flowchart of a token bucket-based flow shaping control method provided in Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of a flow shaping control device based on a token bucket according to Embodiment 3 of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device that implements the token bucket-based flow shaping control method according to an embodiment of the present invention. Detailed Implementation

[0015] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0017] Example 1 Figure 1 This is a flowchart of a token bucket-based traffic shaping control method according to Embodiment 1 of the present invention. This embodiment is applicable to traffic shaping and congestion control in network communication. The method can be executed by a token bucket-based traffic shaping control device, which can be implemented in hardware and / or software. This token bucket-based traffic shaping control device can be configured in electronic devices such as mobile phones, computers, servers, etc. Figure 1 As shown, the method includes: Step 110: Determine the load pressure index of the communication service scenario based on the load parameters in network communication and the communication service scenario.

[0018] The load parameters may include at least one of the following: processor utilization, memory utilization, application layer receive queue level, and packet loss rate. In this embodiment of the invention, a load stress index can be determined based on the load parameters. For example, the load stress index can be obtained by quantifying each load parameter, setting weights, and performing a weighted summation.

[0019] To further improve the adaptability of the load stress index to network communication scenarios, the load stress index can be determined under different communication service scenarios. Communication service scenarios can include computationally sensitive services, memory-sensitive services, and load balancing services. There are several ways to determine communication service scenarios. For example, they can be determined based on business experience. For instance, encryption / decryption or encoding / decoding data processing can be considered computationally sensitive services; massive state caching or large file transfer processing can be considered memory-sensitive services. Alternatively, service scenarios can be evaluated based on business parameters to determine the communication service scenario.

[0020] In different communication service scenarios, the weights of each load parameter can be different, so a load pressure index corresponding to the communication service scenario can be obtained through weighted processing.

[0021] Optionally, based on the load parameters and communication service scenarios in network communication, the load pressure index of the communication service scenario is determined, including: calculating the first correlation coefficient between the number of packets received per second and the processor utilization rate, and the second correlation coefficient between the number of packets received per second and the memory utilization rate within a sliding window; determining the communication service scenario of the network communication based on the first correlation coefficient, the second correlation coefficient, and a preset correlation coefficient threshold; and determining the load pressure index of the communication service scenario based on the processor utilization rate, memory utilization rate, application layer receive queue level, packet loss rate, and communication service scenario in network communication.

[0022] The sliding window length can be a value within the range of 3 to 30 seconds, such as 5 seconds. Using a sliding window can avoid the influence of momentary jitter on the judgment. The correlation coefficient can be Pearson correlation coefficient, Spearman's rank correlation coefficient, or Kendall's rank correlation coefficient, etc. For example, within the time sliding window, the number of packets received per second (QPS), processor utilization (such as CPU), and memory utilization of network communication over N (e.g., 30) sampling periods can be obtained, respectively... , , A sliding window allows newly acquired data to overwrite the oldest data when the queue is full. This can be achieved using a formula. Determine the correlation coefficient between two variable sequences.

[0023] First, determine the sequence of packets received per second. Whether the standard deviation approaches 0 indicates whether the traffic is constant. If the traffic remains unchanged, correlation cannot be calculated, and the judgment result of the previous period can be maintained, or the correlation coefficient can be recorded as 0. When the traffic changes, the correlation coefficient can be calculated using the correlation coefficient formula. The sequence of packets received per second... and CPU utilization sequence Input the data into the correlation coefficient calculation formula to obtain the first correlation coefficient. The sequence of packets received per second. and memory utilization sequence Input the data into the correlation coefficient calculation formula to obtain the second correlation coefficient. The first and second correlation coefficients are in the range of [-1, 1]. In this embodiment of the invention, only positive correlation is considered, and negative correlation coefficients can be treated as 0.

[0024] The preset correlation coefficient thresholds can include a difference judgment threshold and a correlation judgment threshold. When the difference between the first and second correlation coefficients is greater than the difference judgment threshold, and the first correlation coefficient is greater than the correlation judgment threshold, the network communication service scenario can be determined to be a computationally sensitive service. When the difference between the second and first correlation coefficients is greater than the difference judgment threshold, and the second correlation coefficient is greater than the correlation judgment threshold, the network communication service scenario can be determined to be a memory-sensitive service. For example, the correlation judgment threshold can be a value within the range [0.6, 0.7], and the difference judgment threshold can be 0.1. A correlation coefficient greater than 0.6 indicates a strong correlation between the two. If the correlation coefficient is less than 0.6, it indicates that resource fluctuations are not primarily caused by QPS changes (they may be due to background garbage collection (GC), scheduled tasks, or I / O interface waiting), and the weights of load parameters should not be blindly adjusted in this case. By using the difference judgment threshold, when there is a significant difference between the first and second correlation coefficients, the communication service scenario is switched to ensure the reliability of the determined load pressure index.

[0025] For example, the formula for determining the load pressure index can be: SPI is a value within the range [0,1], where 0 indicates the load is completely idle and 1 indicates the load is completely overloaded. In the formula, CPU_Usage and Mem_Usage are the CPU and memory utilization rates within the current time window, respectively, with values ​​ranging from [0,1]; Queue_Len is the current queue length of the application layer receive queue, and Max_Queue is the maximum capacity of the application layer receive queue. This represents the current queue level; Drop_Count represents the number of packets lost in the previous period; and Total_Recv represents the total number of packets received. Packet loss rate; , , ,as well as The weights for processor utilization, memory utilization, queue level, and packet loss rate are respectively. The weights of the load parameters can be determined empirically, for example, by default configuration. , , , The values ​​are 0.3, 0.2, 0.3, and 0.3 respectively.

[0026] In this embodiment of the invention, to improve the reliability of the load pressure coefficient, the weights of load parameters can be adjusted according to the communication service scenario. For example, for computationally sensitive services, the weight of CPU utilization, such as the weight of load parameters, can be increased. , , , These can be 0.5, 0.1, 0.2, and 0.2 respectively. For memory-sensitive applications, the weight of memory utilization can be increased, such as the weight of load parameters. , , , The values ​​can be 0.1, 0.5, 0.2, and 0.2 respectively. For other business scenarios, this can be used as a load balancing function, and the weights of the load parameters can be set to the default configuration.

[0027] In this embodiment of the invention, by determining the load pressure coefficient under multi-dimensional load parameters and communication service scenarios, the load pressure in network communication can be accurately assessed in real time, thereby laying the foundation for reliable token generation and realizing dynamic shaping control of application traffic.

[0028] Step 120: Based on the load pressure index, determine the actual global token generation rate, the actual session-level micro-token generation rate, and the actual maximum token storage amount of the token bucket to generate tokens.

[0029] Based on the real-time determined load pressure index, the actual global token generation rate, actual session-level micro-token generation rate, and actual maximum token storage can be dynamically determined, thereby generating tokens according to the token generation parameters.

[0030] For example, when the load pressure index increases, the actual global token generation rate can be reduced; when the load pressure index decreases, the actual global token generation rate can be increased. When reducing the actual global token generation rate, an exponential function can be used for rapid overload reduction, avoiding the risk of load collapse caused by linear reduction. Alternatively, one or more load thresholds can be set, and the adjustment method for the actual global token generation rate can be determined by comparing the load pressure index with the load thresholds.

[0031] Optionally, the actual global token generation rate is determined based on the load pressure index, a first load threshold, a second load threshold, a baseline global token generation rate, and a rate adjustment factor. For example, this can be achieved using the formula... Determine the actual global token generation rate. Here, Base_Rate is the baseline global token generation rate (e.g., the system's rated processing capacity); Thres_Low and Thres_High are the first and second load thresholds, respectively; for example, the first load threshold is set to 0.3 and the second load threshold to 0.8. k1 and k2 are rate adjustment factors, which can be determined empirically; for example, k1 is 1.5 and k2 is 2.0. SPI is the load stress exponent. e is a natural constant. When the system load is extremely high (SPI > Thres_High), the token generation rate decreases exponentially, forcing the system to "breathe" and preventing deadlock caused by CPU overheating.

[0032] By sensing the actual processing capacity of the server load in real time, i.e. the load pressure index, and based on the load pressure index, the actual global token generation rate of the token bucket is dynamically adjusted non-linearly to avoid rate lag or overshoot caused by linear adjustment, ensuring that the rate limiting rate matches the system processing capacity in real time and achieving elastic rate limiting effect.

[0033] The actual session-level microtoken generation rate can be negatively correlated with the load stress index. Optionally, the actual session-level microtoken generation rate can be determined based on the baseline session-level microtoken generation rate, the load stress index, and the session history reputation coefficient. For example, this can be achieved using the formula... Determine the actual session-level microtoken generation rate. Among these, The baseline session-level microtoken generation rate can be the standard processing capacity pre-allocated by the system for a single connection. For example, the baseline session-level microtoken generation rate for each UDP client is preset to 10 Mbps. A coefficient that is negatively correlated with the load pressure index, such as the global load suppression coefficient. The session history reputation score reflects the historical behavioral performance during a session, for example... It can be a value selected within the range (0, 1.2] based on historical behavior.

[0034] By sensing the actual processing capacity of the server load in real time, i.e. the load pressure index, and based on the load pressure index and session history performance, the actual session-level micro-token generation rate of the token bucket is dynamically adjusted to ensure that the rate limiting rate matches the session processing capacity in real time, thus achieving an elastic rate limiting effect.

[0035] The actual maximum token storage capacity can quickly absorb instantaneous traffic spikes, avoiding the "thundering herd effect" and ensuring the stability of the business system. When the load is idle, the actual maximum token storage capacity can be increased to improve the ability to absorb sudden traffic surges; when the load is overloaded, the actual maximum token storage capacity can be decreased to avoid the subsequent processing pressure caused by a large accumulation of tokens. For example, one or more load thresholds can be set, and the adjustment method for the actual maximum token storage capacity can be determined by comparing the load pressure index with the load thresholds.

[0036] Optionally, the actual maximum token storage is determined based on the load stress index, a first load threshold, a second load threshold, a baseline maximum token storage, and storage adjustment parameters. For example, when the load stress index is less than the first load threshold, the actual maximum token storage is determined using the formula... Determine the actual maximum token storage. When the load pressure index exceeds the second load threshold, use the formula... Determine the actual maximum token storage. When the load pressure index is between the first load threshold and the second load threshold, the baseline maximum token storage is used as the actual maximum token storage. In the formula, Base_Burst is the baseline maximum token storage, for example, it defaults to 1 / 10 of the baseline global token generation rate Base_Rate. In the formula, "0.5" is a storage adjustment parameter that can be dynamically adjusted.

[0037] In scenarios of continuous overload, by dynamically adjusting the token generation rate and the actual maximum token storage, unnecessary requests are smoothly rejected, avoiding drastic fluctuations in traffic that could lead to unstable system resource usage and ensuring stable system operation. By designing independent two-tier rate limiting strategies for user protocol data packets—global tokens and session-level micro-tokens—excessive system resource consumption by "malicious" or faulty senders is prevented, ensuring fair access for all normal senders and preventing single points of failure from impacting the overall service.

[0038] Step 130: In network communication, obtain user protocol data packets and obtain tokens from the token bucket according to the user protocol data packets.

[0039] In network communication, the server's network interface card (NIC) can receive high-concurrency user protocol data packets and transmit them to the application layer's receive buffer via kernel-mode zero-copy or fast forwarding mechanisms. The token bucket dynamically generates tokens based on the actual global token generation rate, the actual session-level micro-token generation rate, and the actual maximum token storage capacity. The application layer can then retrieve tokens from the token bucket based on user protocol data packets.

[0040] Optionally, obtaining a token from the token bucket based on the user protocol data packet includes: parsing the user protocol data packet to obtain the source address and source port of the user protocol data packet; determining the virtual session object corresponding to the user protocol data packet based on the source address and source port; and obtaining a token from the token bucket based on the virtual session object.

[0041] The source address and source port of user protocol data packets can form a tuple, which is mapped to a virtual session object in memory using consistent hashing. If a corresponding virtual session object does not exist in memory, it can be created or replaced based on a Least Recently Used (LRU) policy. Simultaneously, the application layer can perform traffic statistics using a sliding window to determine the real-time transmission rate of the session. When acquiring tokens from the token bucket, a global token can be obtained, and session-level micro-tokens can be acquired based on the virtual session object to perform traffic shaping and congestion control for a single session through secondary rate limiting.

[0042] Step 140: When the global token and session-level micro token are obtained, the user protocol data packet is written to the application layer receive queue, and the business thread retrieves the data packet from the queue for processing; otherwise, the user protocol data packet is dropped.

[0043] First, attempt to acquire a global token. If successful, then attempt to acquire a session-level micro token. If both tokens are successfully acquired: the traffic is deemed compliant, congestion control is skipped, and the user protocol data packet is written to the application layer receive queue. The business thread then retrieves the data packet from the queue for processing. After processing, the number of packets discharged from the queue per unit time is counted, and the relevant metrics of the system's processing capacity are updated accordingly.

[0044] If a global token or session-level micro-token cannot be obtained, user protocol data packets are dropped. A dynamic packet loss strategy can be employed, such as using packet loss probability. In a dynamic packet loss strategy, when it is determined that user protocol data packets should be discarded, resources are released, an audit log containing SPI status and session information is recorded, and the process ends. Conversely, when it is determined that user protocol data packets should be retained, the data packets are marked as "degraded retention," written to the application layer receive queue, and then retrieved from the queue by the business thread for processing.

[0045] By designing independent two-level rate limiting strategies for each individual session, we prevent "malicious" or faulty senders from consuming excessive system resources, ensuring fair access for all normal senders and avoiding single points of failure affecting the overall service. Dynamic packet loss handling based on packet loss probability enables differentiated intelligent packet loss, avoiding direct packet loss processing and achieving smoother service quality assurance.

[0046] The technical solution of this embodiment determines the load pressure index of the communication service scenario based on the load parameters and communication service scenario in the network communication. Based on the load pressure index, it determines the actual global token generation rate, the actual session-level micro-token generation rate, and the actual maximum token storage capacity of the token bucket to generate tokens. In network communication, user protocol data packets are acquired, and tokens are retrieved from the token bucket according to the user protocol data packets. When global tokens and session-level micro-tokens are acquired, user protocol data packets are written to the application layer receive queue, and the business thread retrieves the data packets from the queue for processing. Otherwise, packet loss processing is performed on user protocol data packets. This solves the traffic shaping and control problem in network communication. By evaluating the communication load pressure in the communication service scenario, the token generation speed can be dynamically adjusted to achieve elastic traffic shaping and congestion control, ensuring reliable data transmission.

[0047] Example 2 Figure 2 This is a flowchart of a token bucket-based flow shaping control method according to Embodiment 2 of the present invention. This embodiment is a further refinement of the above technical solution, and the technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 2 As shown, the method includes: Step 210: Determine the load pressure index of the communication service scenario based on the load parameters in network communication and the communication service scenario.

[0048] Optionally, based on the load parameters and communication service scenarios in network communication, the load pressure index of the communication service scenario is determined, including: calculating the first correlation coefficient between the number of packets received per second and the processor utilization rate, and the second correlation coefficient between the number of packets received per second and the memory utilization rate within a sliding window; determining the communication service scenario of the network communication based on the first correlation coefficient, the second correlation coefficient, and a preset correlation coefficient threshold; and determining the load pressure index of the communication service scenario based on the processor utilization rate, memory utilization rate, application layer receive queue level, packet loss rate, and communication service scenario in network communication.

[0049] Step 220: Determine the actual global token generation rate based on the load pressure index, the first load threshold, the second load threshold, the baseline global token generation rate, and the rate adjustment factor.

[0050] When determining the actual global token generation rate, the rate adjustment factor is an adjustable parameter. The rate adjustment factor may include an acceleration factor and a deceleration factor. Optionally, the actual global token generation rate is determined based on the load pressure index, a first load threshold, a second load threshold, a baseline global token generation rate, and the rate adjustment factor, including: determining the acceleration factor based on the actual processing speed of the business thread, the baseline processing speed, and a preset acceleration parameter; determining the deceleration factor based on the rate of change of the load pressure index within a preset time period and a preset deceleration parameter; determining the actual global token generation rate based on the load pressure index, the first load threshold, the acceleration factor, and the baseline global token generation rate when the load pressure index is less than the first load threshold; and determining the actual global token generation rate based on the load pressure index, the second load threshold, the deceleration factor, and the baseline global token generation rate when the load pressure index is greater than the second load threshold; wherein, the first load threshold is less than the second load threshold.

[0051] The escalation factor can be used to increase the token generation rate when the system load is idle. A larger k1 results in a faster token generation rate increase when the load is idle, allowing for quicker processing of backlogged data. The value of k1 can depend on the queue's drain rate. For example, this can be achieved through the formula... Determine the acceleration factor. This represents the actual processing speed of the business thread. The base processing speed is set to 1.0, which can be adjusted as needed. When the business thread processes extremely quickly and the system is idle, it indicates that the backend has a strong processing capacity, and k1 will increase, allowing the token generation rate to rise rapidly and quickly consume the backlog of data.

[0052] The descent factor is used to control the token generation rate decrease when the system is overloaded. A larger k2 value results in a more rapid decrease in the token generation rate under overload conditions, quickly reducing system load pressure. The value of k2 can depend on the rising slope of the SPI. For example, it can be achieved through the formula... Determine the deceleration factor. In the formula, It is an adjustable parameter. Preset time period The rate of change of the internal load pressure index is preset to a reduction parameter of 2.0, which can be adjusted according to needs. If the SPI spikes from 0.2 to 0.8 in a short period of time (with a very steep slope), it indicates that the system has encountered a sudden malicious attack or a serious failure. At this time, k2 increases significantly (e.g., from 2.0 to 10.0), triggering an "emergency brake" that instantly reduces the token generation rate to a minimum to prevent a cascading failure.

[0053] By using an acceleration factor to increase the actual global token generation rate when the load is idle, and a deceleration factor to decrease the actual global token generation rate when the load is overloaded, the baseline global token generation rate can be maintained when the load is normal. This achieves dynamic control of traffic, making full use of resources while avoiding putting pressure on the load.

[0054] Step 230: Determine the actual session-level microtoken generation rate based on the baseline session-level microtoken generation rate, load pressure index, and session history reputation coefficient.

[0055] Optionally, the actual session-level micro-token generation rate is determined based on the baseline session-level micro-token generation rate, load pressure index, and session historical reputation coefficient. This includes: determining the global load suppression coefficient based on the load pressure index and a preset pressure threshold; determining the reputation score of the virtual session object based on packet loss and / or rate limiting parameters of the virtual session object in a preset statistical period; determining the reward score of the virtual session object based on the reputation score of the virtual session object in multiple preset statistical periods; determining the session historical reputation coefficient of the virtual session object based on the reputation score and reward score; and determining the actual session-level micro-token generation rate based on the baseline session-level micro-token generation rate, the global load suppression coefficient, and the session historical reputation coefficient.

[0056] The global load suppression coefficient is negatively correlated with the load stress index. For example, when the load stress index is less than 0.5, the global load suppression coefficient is set to 1 (full speed); when the load stress index is greater than or equal to 0.5 and less than 0.9, the global load suppression coefficient adopts a linear decay, such as... =1.0 - (SPI - 0.5); When the load stress index is greater than 0.9, the global load suppression coefficient is set to 0.1 (only keep-alive traffic is retained). Alternatively, for more precise token generation control based on the load stress coefficient, the formula can be used... Determine the global load suppression coefficient. A preset pressure threshold is set, for example, to 0.8. Smooth load suppression can be achieved through the global load suppression coefficient. When the load pressure index SPI < 0.8, the global load suppression coefficient is close to 1; when SPI > 0.8, the global load suppression coefficient drops rapidly.

[0057] The reputation score of a virtual session object can be set to an initial value, such as 100. If the virtual session object does not trigger packet loss or rate limiting within a preset statistical period and has a certain amount of reserved traffic, its reputation score can be rewarded, for example, the reputation score is S=min(100,S+1). If the virtual session object triggers global token bucket exhaustion within the preset statistical period (not due to its own fault, but due to the general trend), no points are deducted. If the virtual session object triggers session-level token bucket exhaustion within the preset statistical period (due to its own rate exceeding the limit), its reputation score can be penalized, such as S=S-5. If the virtual session object triggers packet loss within the preset statistical period and is identified as a high burst source (e.g., traffic burst factor Burst_Factor greater than 1), its reputation score can be penalized, such as S=S-20 (severe penalty).

[0058] Based on the reputation score of a virtual session object over multiple preset statistical periods, a reward score can be generated for the virtual session object. For example, if the reputation score of a virtual session object remains at its initial value for T consecutive preset statistical periods, a reward score of 0.2 can be awarded; otherwise, the reward score is 0. For instance, this can be achieved through a formula... Determine the session history reputation score of the virtual session object. S / 100 reflects the rate "penalty" due to historical violations; the lower the score, the less bandwidth is allocated.

[0059] The actual session-level microtoken generation rate is determined based on the baseline session-level microtoken generation rate, the global load suppression coefficient, and the session history reputation coefficient. For example, the baseline session-level microtoken generation rate... The target is 1000 QPS. The current SPI is 0.6; determine the global load suppression factor. The value is 0.9. The corresponding virtual session object A has historically exhibited stable transmission without exceeding the speed limit, and its reputation score is S=100. Therefore, the actual session-level micro-token generation rate is determined to be... =1000 × 0.9 × 1.0 = 900 QPS, slightly affected by system load, normal transmission. The corresponding virtual session object A has historically triggered rate limiting and packet loss multiple times in the past few seconds, and its reputation score has been reduced to S=40. The actual session-level micro-token generation rate is determined to be... =1000×0.9×0.4=360 QPS. The bandwidth of this object was automatically and significantly compressed by the system, freeing up resources for normal object A.

[0060] By dynamically adjusting the baseline session-level micro-token generation rate using the global load suppression coefficient and session history reputation coefficient, dynamic balancing of multiple virtual session objects can be achieved, avoiding the problem of the server being easily dragged down by a single point of failure device (a device that constantly retransmits data).

[0061] Step 240: Determine the actual maximum token storage based on the load pressure index, the first load threshold, the second load threshold, the baseline maximum token storage, and the storage adjustment parameters.

[0062] Step 250: Generate tokens based on the actual global token generation rate of the token bucket, the actual session-level micro-token generation rate, and the actual maximum token storage.

[0063] Step 260: In network communication, obtain user protocol data packets and obtain tokens from the token bucket according to the user protocol data packets.

[0064] Optionally, obtaining a token from the token bucket based on the user protocol data packet includes: parsing the user protocol data packet to obtain the source address and source port of the user protocol data packet; determining the virtual session object corresponding to the user protocol data packet based on the source address and source port; and obtaining a token from the token bucket based on the virtual session object.

[0065] Step 270: Upon obtaining the global token and session-level micro token, the user protocol data packet is written into the application layer receive queue, and the business thread retrieves the data packet from the queue for processing.

[0066] If a global token or session-level micro token cannot be obtained, proceed to step 280 for differentiated packet loss handling.

[0067] Step 280: Determine the packet loss probability of the user protocol data packet based on the priority of the user protocol data packet, the traffic burst factor, the queue level of the application layer receive queue, and the sensitivity index.

[0068] The payload header of the user protocol data packet can reserve 1 byte as a priority field. For example, five priority levels can be defined from P1 to P5, with P1 being the highest priority and P5 the lowest. Priorities can be pre-set based on data importance, such as P1 for industrial fault alarms, P3 for routine operation data, and P5 for debugging logs. When determining the probability of packet loss, the probability of packet loss is negatively correlated with the packet priority; that is, the higher the priority, the lower the probability of packet loss. Different weights can be set for different packet priorities; for example, the weights for P1-P5 are 10, 8, 5, 3, and 1 respectively. The higher the weight, the lower the probability of packet loss.

[0069] The traffic burst factor characterizes the real-time transmission rate of a virtual session object. For example, when the real-time transmission rate exceeds the session rate limit, a traffic burst can be identified; the greater the exceedance, the larger the traffic burst factor. The traffic burst factor is positively correlated with the packet loss probability; that is, the larger the traffic burst factor, the greater the packet loss probability.

[0070] Optionally, a traffic burst factor can be determined based on the real-time session transmission rate and the session rate limit. For example, this can be achieved using a formula... Determine the traffic burst factor. For example, maintain a time window for each virtual session object (e.g., 1 second divided into 10 100ms segments), count the total number of bytes within the window, and calculate the session real-time transmission rate, Real_Rate. If the session real-time transmission rate of the virtual session object is the session rate limit... It is twice the normal packet loss probability, with a burst factor of 4.

[0071] Queue water level The queue level can be determined by the ratio of queue length to the maximum queue capacity, i.e., Queue_Ratio = Queue_Len / Max_Queue. The queue level is positively correlated with the probability of packet loss; that is, the higher the queue level, the greater the probability of packet loss.

[0072] The sensitivity index can be used to adjust the queue level. For example, by using a power function with the sensitivity index, the impact of the queue level on the packet loss probability can be amplified, resulting in an extremely low packet loss probability when the queue level is low and a sharp increase in the packet loss probability when the queue level is close to the upper limit. The sensitivity index determines the steepness of the packet loss probability curve. The sensitivity index can be a default value, such as 3, or it can be dynamically adjusted. For example, the sensitivity index can be determined based on the rate of change of the queue level over a preset time period.

[0073] Optionally, the sensitivity index is determined based on the rate of change of the queue water level within a preset time period, preset adjustment parameters, and preset sensitivity parameters. For example, it can be determined using a formula... Determine the sensitivity index. In the formula, This represents the current queue level. Before the current moment The queue water level at that time The rate of change of the queue water level within a preset time period; These are preset adjustment parameters. The preset sensitivity parameter is set to the default value of 3. When the queue surges (rapid filling speed), the sensitivity index value increases, causing the packet loss probability curve to enter the high packet loss zone earlier, thus suppressing traffic in advance; when the queue is stable, the sensitivity index returns to the default value of 3, maintaining smooth packet loss.

[0074] For example, the formula for determining the packet loss probability is as follows: . Let be the priority weight of the i-th user protocol data packet.

[0075] When determining the packet loss rate of user protocol data packets, an upper limit for packet loss probability can also be considered. Optionally, the packet loss probability of user protocol data packets can be determined based on the priority of the user protocol data packets, the traffic burst factor, the queue level of the application layer receive queue, and the sensitivity index. This includes: determining the sensitivity index based on the rate of change of the queue level within a preset time period, preset adjustment parameters, and preset sensitivity parameters; determining the traffic burst factor based on the session real-time transmission rate and session rate limiting value; determining the upper limit for the packet loss probability of user protocol data packets based on the priority of the user protocol data packets; and determining the packet loss probability of user protocol data packets based on the priority of the user protocol data packets, the traffic burst factor, the queue level of the application layer receive queue, the sensitivity index, and the upper limit for packet loss probability.

[0076] For example, in passing If the packet loss probability is determined to be greater than the upper limit of the packet loss probability, the upper limit of the packet loss probability is used as the final packet loss probability; otherwise, the packet loss probability calculated by the formula is used as the final packet loss probability. For example, Queue_Ratio=0.8, n=3, Burst_Factor=1. For packets with priority P1 (Priority_i=10), the calculated... The upper limit for the packet loss probability of P1 priority data packets is 5%, and the final packet loss probability of this data packet is 5%. For P5 priority data (Priority_i=1), Prob_Drop is calculated to be 0.512 / 1=51.2%, which does not reach the upper limit for the packet loss probability of P5 priority data packets, and the packet loss probability of this data packet is 51.2%. This achieves effective protection for high-priority data packets and avoids the high probability of packet loss of high-priority data packets.

[0077] Step 290: Perform packet loss processing on user protocol data packets based on the packet loss probability.

[0078] When handling packet loss, a packet loss probability can be used to differentiate the handling of lost user protocol data packets. For example, a random number is generated in the range [0,1]. If the random number is less than the packet loss probability, the corresponding user protocol data packet is discarded; otherwise, the user protocol data packet is allowed to enter the application layer receive queue.

[0079] The technical solution of this invention involves determining the load pressure index of a communication service scenario based on load parameters and the communication service scenario; determining the actual global token generation rate based on the load pressure index, a first load threshold, a second load threshold, a baseline global token generation rate, and a rate adjustment factor; determining the actual session-level micro-token generation rate based on the baseline session-level micro-token generation rate, the load pressure index, and the session history reputation coefficient; determining the actual maximum token storage based on the load pressure index, the first load threshold, the second load threshold, the baseline maximum token storage, and a storage adjustment parameter; and generating tokens based on the actual global token generation rate, the actual session-level micro-token generation rate, and the actual maximum token storage. During communication, user protocol data packets are acquired, and tokens are obtained from the token bucket based on these packets. When a global token or a session-level micro-token is acquired, the user protocol data packet is written to the application layer receive queue, and the business thread retrieves the data packet from the queue for processing. When neither a global token nor a session-level micro-token is acquired, the packet loss probability of the user protocol data packet is determined based on its priority, traffic burst factor, the queue level of the application layer receive queue, and the sensitivity index. Packet loss processing is performed on the user protocol data packet based on the packet loss probability, solving the traffic shaping and control problem in network communication. By assessing the communication load pressure under communication service scenarios, the token generation speed can be dynamically adjusted to achieve elastic traffic shaping and congestion control, ensuring reliable data transmission.

[0080] Specifically, by constructing a multi-dimensional load pressure index to quantify system load, the actual processing capacity of the server can be perceived in real time. The load pressure index dynamically and non-linearly adjusts the token generation speed, avoiding rate lag or overshoot caused by linear adjustment, ensuring that the rate limiting rate matches the system's processing capacity in real time, and achieving elastic rate limiting. Through the "exponential deceleration" mechanism, the system can automatically reduce its processing capacity when facing traffic storms caused by DDoS attacks or equipment failures, ensuring that core services do not crash. By dynamically adjusting the token generation rate, idle resources are fully utilized and overload is avoided, achieving peak shaving and valley filling of resources. By dynamically adjusting the maximum token storage, instantaneous traffic spikes are dynamically smoothed and shaped, avoiding queue overflow, large-scale packet loss, resource contention, and system latency jitter, ensuring system stability. The system operates by employing a lightweight virtual session mapping mechanism, enabling accurate identification and status tracking of each sender's data stream. A cascaded structure of a global token bucket and a session-level micro-token bucket is designed, requiring data packets to acquire dual tokens simultaneously. The global token bucket handles system-level SPI overload, while the session-level micro-token bucket addresses single-IP rate anomalies. An independent two-level rate limiting strategy is designed for each session to prevent malicious or faulty senders from consuming excessive system resources, ensuring fair access for all normal senders and preventing single-point anomalies from impacting global services. The sensitivity index is dynamically adjusted based on the queue growth rate to determine packet loss probability. When any token bucket is exhausted, it is not directly discarded but undergoes a secondary judgment through differentiated packet loss processing, ensuring reliable transmission of core business data and guaranteeing a high survival rate for high-priority data in congested scenarios.

[0081] Figure 3 This is an application flowchart of a token bucket-based flow shaping control method according to Embodiment 2 of the present invention. Figure 3As shown, in network communication, the server's network interface card (NIC) transmits received user protocol data packets to the application layer. The application layer parses the user protocol data packets to obtain their source address and source port. Based on the source address and source port, it determines the virtual session object corresponding to the user protocol data packet. It updates the session's real-time transmission rate based on the virtual session object and retrieves a token from the token bucket. The token is generated based on the actual global token generation rate, the actual session-level micro-token generation rate, and the actual maximum token storage. When a global token or session-level micro-token cannot be obtained, the packet loss probability of the user protocol data packet is determined based on the user protocol data packet's priority, traffic burst factor, the application layer's receive queue level, and the sensitivity index. A random number is generated; if the random number is less than the packet loss probability, the user protocol data packet is discarded. If the random number is greater than or equal to the packet loss probability, the user protocol data packet is written to the application layer's receive queue, and the business thread retrieves the data packet from the queue for processing. When a global token and session-level micro-token are obtained, the user protocol data packet is written to the application layer's receive queue, and the business thread retrieves the data packet from the queue for processing. Update the actual processing speed of the business thread and the queue level based on the processing status of the business thread.

[0082] Among them, such as Figure 3 As shown, in network communication, the load pressure index of the communication service scenario is determined based on load parameters and the communication service scenario; the acceleration factor is determined based on the actual processing speed of the service thread, the baseline processing speed, and the preset acceleration parameters; the deceleration factor is determined based on the rate of change of the load pressure index within a preset time period and the preset deceleration parameters; the actual global token generation rate is determined based on the load pressure index, the first load threshold, the second load threshold, the baseline global token generation rate, the acceleration factor, and the deceleration factor; the actual session-level micro-token generation rate is determined based on the baseline session-level micro-token generation rate, the load pressure index, and the session historical reputation coefficient; the actual maximum token storage is determined based on the load pressure index, the first load threshold, the second load threshold, the baseline maximum token storage, and the storage adjustment parameters; and tokens are generated based on the actual global token generation rate, the actual session-level micro-token generation rate, and the actual maximum token storage.

[0083] Through such Figure 3 The process shown can automatically sense the sensitivity of business resources, dynamically adjust flow control parameters, and implement intelligent flow control with two-level cascaded flow limiting.

[0084] Example 3 Figure 4 This is a schematic diagram of a flow shaping control device based on a token bucket according to Embodiment 3 of the present invention. Figure 4As shown, the device includes: a load pressure index determination module 410, a token generation parameter determination module 420, a token acquisition module 430, and a data packet processing module 440. Wherein: The load pressure index determination module 410 is used to determine the load pressure index of the communication service scenario based on the load parameters in the network communication and the communication service scenario. The token generation parameter determination module 420 is used to determine the actual global token generation rate, the actual session-level micro-token generation rate, and the actual maximum token storage amount of the token bucket based on the load pressure index, so as to generate tokens. The token acquisition module 430 is used to acquire user protocol data packets in network communication and acquire tokens from the token bucket according to the user protocol data packets. The data packet processing module 440 is used to write the user protocol data packet into the application layer receive queue when a global token and a session-level micro token are obtained, and the business thread retrieves the data packet from the queue for processing; otherwise, the user protocol data packet is lost.

[0085] Optionally, the load stress index determination module 410 includes: The correlation coefficient determination unit is used to calculate the first correlation coefficient between the number of packets received per second and the processor utilization, and the second correlation coefficient between the number of packets received per second and the memory utilization within a sliding window; The communication service scenario determination unit is used to determine the communication service scenario of network communication based on the first correlation coefficient, the second correlation coefficient, and the preset correlation coefficient threshold. The load pressure index determination unit is used to determine the load pressure index of the communication service scenario based on the processor utilization, memory utilization, queue level of the application layer receive queue, packet loss rate, and communication service scenario in network communication.

[0086] Optionally, the token generation parameter determination module 420 includes: The actual global token generation rate determination unit is used to determine the actual global token generation rate based on the load pressure index, the first load threshold, the second load threshold, the baseline global token generation rate, and the rate adjustment factor. The actual session-level microtoken generation rate determination unit is used to determine the actual session-level microtoken generation rate based on the baseline session-level microtoken generation rate, load pressure index, and session history reputation coefficient. The actual maximum token storage unit is used to determine the actual maximum token storage based on the load pressure index, the first load threshold, the second load threshold, the baseline maximum token storage, and the storage adjustment parameters.

[0087] Optionally, the actual global token generation rate determination unit includes: The ramp-up factor determination subunit is used to determine the ramp-up factor based on the actual processing speed of the business thread, the baseline processing speed, and the preset ramp-up parameters. The deceleration factor determination subunit is used to determine the deceleration factor based on the rate of change of the load pressure index within a preset time period and the preset deceleration parameters. The first unit for determining the actual global token generation rate is used to determine the actual global token generation rate based on the load pressure index, the first load threshold, the acceleration factor, and the benchmark global token generation rate when the load pressure index is less than the first load threshold. The second unit for determining the actual global token generation rate is used to determine the actual global token generation rate based on the load pressure index, the second load threshold, the throttling factor, and the baseline global token generation rate when the load pressure index is greater than the second load threshold. The first load threshold is less than the second load threshold.

[0088] Optionally, the actual session-level microtoken generation rate determination unit includes: The global load suppression coefficient determination subunit is used to determine the global load suppression coefficient based on the load pressure index and the preset pressure threshold. The reputation score determination subunit is used to determine the reputation score of a virtual session object based on packet loss and / or rate limiting parameters of the virtual session object in a preset statistical period. The reward score determination subunit is used to determine the reward score of a virtual session object based on the reputation score of the virtual session object in multiple preset statistical periods. The Session History Reputation Coefficient Determination Subunit is used to determine the session history reputation coefficient of a virtual session object based on its reputation score and reward score. The actual session-level microtoken generation rate determination subunit is used to determine the actual session-level microtoken generation rate based on the baseline session-level microtoken generation rate, the global load suppression coefficient, and the session history reputation coefficient.

[0089] Optionally, the packet processing module 440 includes: The packet loss probability determination unit is used to determine the packet loss probability of user protocol data packets based on the priority of user protocol data packets, traffic burst factor, queue level of application layer receive queue, and sensitivity index. The packet loss processing unit is used to process user protocol data packets based on the packet loss probability.

[0090] Optionally, the packet loss probability determination unit includes: The sensitivity index determination subunit is used to determine the sensitivity index based on the rate of change of the queue water level within a preset time period, preset adjustment parameters, and preset sensitivity parameters. The traffic burst factor determination subunit is used to determine the traffic burst factor based on the session real-time transmission rate and the session rate limit value. The packet loss probability upper limit determination subunit is used to determine the upper limit of the packet loss probability of user protocol data packets based on the priority of user protocol data packets; The packet loss probability determination subunit is used to determine the packet loss probability of user protocol data packets based on the priority of user protocol data packets, traffic burst factor, queue level of application layer receive queue, sensitivity index, and upper limit of packet loss probability.

[0091] Optionally, the token acquisition module 430 includes: The data packet parsing unit is used to parse user protocol data packets to obtain the source address and source port of the user protocol data packets; The virtual session object determination unit is used to determine the virtual session object corresponding to the user protocol data packet based on the source address and source port. The token acquisition unit is used to acquire tokens from the token bucket based on the virtual session object.

[0092] The token bucket-based flow shaping control device provided in this embodiment of the invention can execute the token bucket-based flow shaping control method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0093] Example 4 Figure 5 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0094] like Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) or random access memory (RAM), communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from the storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. Input / output (I / O) interfaces are also connected to the bus 14.

[0095] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0096] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as token bucket-based flow shaping control methods.

[0097] In some embodiments, the token bucket-based traffic shaping control method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the token bucket-based traffic shaping control method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the token bucket-based traffic shaping control method by any other suitable means (e.g., by means of firmware).

[0098] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0099] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0100] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0101] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0102] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0103] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0104] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0105] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A flow shaping control method based on token bucket, characterized in that, include: Based on the load parameters in network communication and the communication service scenario, determine the load pressure index of the communication service scenario; Based on the load pressure index, determine the actual global token generation rate, actual session-level micro-token generation rate, and actual maximum token storage capacity of the token bucket to generate tokens; In network communication, user protocol data packets are acquired, and tokens are obtained from the token bucket according to the user protocol data packets; When a global token and a session-level micro token are obtained, the user protocol data packet is written into the application layer receive queue, and the business thread retrieves the data packet from the queue for processing; otherwise, the user protocol data packet is dropped.

2. The method according to claim 1, characterized in that, Based on the load parameters and communication service scenarios in network communication, determine the load pressure index of the communication service scenario, including: Calculate the first correlation coefficient between the number of packets received per second and processor utilization, and the second correlation coefficient between the number of packets received per second and memory utilization within a sliding window; Based on the first correlation coefficient, the second correlation coefficient, and the preset correlation coefficient threshold, the communication service scenario of the network communication is determined; The load pressure index of the communication service scenario is determined based on the processor utilization, memory utilization, application layer receive queue level, packet loss rate, and the communication service scenario in the network communication.

3. The method according to claim 1, characterized in that, Based on the load pressure index, determine the actual global token generation rate, actual session-level micro-token generation rate, and actual maximum token storage capacity of the token bucket to generate tokens, including: The actual global token generation rate is determined based on the load pressure index, the first load threshold, the second load threshold, the baseline global token generation rate, and the rate adjustment factor. The actual session-level microtoken generation rate is determined based on the baseline session-level microtoken generation rate, the load pressure index, and the session history reputation coefficient. The actual maximum token storage is determined based on the load pressure index, the first load threshold, the second load threshold, the baseline maximum token storage, and the storage adjustment parameters.

4. The method according to claim 3, characterized in that, The actual global token generation rate is determined based on the load pressure index, the first load threshold, the second load threshold, the baseline global token generation rate, and the rate adjustment factor, including: The acceleration factor is determined based on the actual processing speed of the business thread, the baseline processing speed, and the preset acceleration parameters. The deceleration factor is determined based on the rate of change of the load pressure index within a preset time period and the preset deceleration parameters. When the load pressure index is less than the first load threshold, the actual global token generation rate is determined based on the load pressure index, the first load threshold, the acceleration factor, and the baseline global token generation rate. When the load pressure index is greater than the second load threshold, the actual global token generation rate is determined based on the load pressure index, the second load threshold, the throttling factor, and the baseline global token generation rate. Wherein, the first load threshold is less than the second load threshold.

5. The method according to claim 3, characterized in that, The actual session-level microtoken generation rate is determined based on the baseline session-level microtoken generation rate, the load pressure index, and the session history reputation coefficient, including: The global load suppression coefficient is determined based on the load pressure index and the preset pressure threshold. The reputation score of the virtual session object is determined based on the packet loss and / or rate limiting parameters of the virtual session object in a preset statistical period; The reward score of the virtual session object is determined based on the reputation score of the virtual session object in multiple preset statistical periods; The session history reputation coefficient of the virtual session object is determined based on the reputation score and the reward score. The actual session-level microtoken generation rate is determined based on the baseline session-level microtoken generation rate, the global load suppression coefficient, and the session history reputation coefficient.

6. The method according to claim 1, characterized in that, The packet loss handling of the user protocol data packets includes: The probability of packet loss of the user protocol data packet is determined based on the priority of the user protocol data packet, the traffic burst factor, the queue level of the application layer receiving queue, and the sensitivity index. The user protocol data packets are processed for packet loss based on the packet loss probability.

7. The method according to claim 6, characterized in that, The probability of packet loss of the user protocol data packet is determined based on the priority of the user protocol data packet, the traffic burst factor, the queue water level of the application layer receive queue, and the sensitivity index, including: The sensitivity index is determined based on the rate of change of the queue water level within a preset time period, preset adjustment parameters, and preset sensitivity parameters. The traffic burst factor is determined based on the real-time transmission rate of the session and the session rate limit. Based on the priority of the user protocol data packets, determine the upper limit of the packet loss probability of the user protocol data packets; The packet loss probability of the user protocol data packet is determined based on the priority of the user protocol data packet, the traffic burst factor, the queue level of the application layer receive queue, the sensitivity index, and the upper limit of the packet loss probability.

8. The method according to claim 1, characterized in that, Retrieving a token from the token bucket according to the user protocol data packet includes: The source address and source port of the user protocol data packet are obtained by parsing the user protocol data packet. Based on the source address and the source port, determine the virtual session object corresponding to the user protocol data packet; The virtual session object retrieves a token from the token bucket.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to said at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the token bucket-based traffic shaping control method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the token bucket-based flow shaping control method according to any one of claims 1-8.

11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the token bucket-based flow shaping control method according to any one of claims 1-8.