Global interface dynamic flow limiting method based on reinforcement learning
Through the global interface dynamic current limiting method based on reinforcement learning, using the request simulator and deep neural network training agent, the problem that the existing current limiting method cannot be dynamically adjusted is solved, and efficient resource utilization and network security improvement is achieved.
Patent Information
- Application Number
- CN202510305238.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The existing current limiting methods cannot be dynamically adjusted based on real load conditions and interface characteristics, resulting in inefficient system resource utilization, poor user experience and insufficient network security.
Using a global interface dynamic current limiting method based on reinforcement learning, we can realize intelligent decisions on whether to limit or release the current request by building a request simulator, collecting status information, establishing action return functions and deep neural network training agents.
It improves the system throughput, reduces the requested access delay, improves system stability and network security, and can make dynamic decisions based on interface characteristics and load conditions to avoid overloading.
Smart Images

Figure CN120281718A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of interface traffic control, and particularly relates to a global interface dynamic current-limiting method based on reinforcement learning. Background Art
[0002] In Internet applications, interface current-limiting is an extremely crucial strategy. Its main purposes are as follows: First, to ensure the stability of the system. When a large number of requests flood in within a short period of time and exceed the threshold that the system can bear, it may cause the system to crash. Interface current-limiting can effectively control the number of requests within a unit time, avoiding the system from being paralyzed due to overload. Second, to improve the user experience. Without an interface current-limiting mechanism, the requests of some users may experience long waits or timeouts due to resource competition. The current-limiting measure can ensure that user requests are processed promptly and efficiently, thus providing users with a smooth and stable service. Third, to ensure network security. Current-limiting can effectively resist attacks by malicious users sending a large number of invalid requests, thereby maintaining the security and health of the entire application ecosystem.
[0003] Existing current-limiting methods generally set static rules and cannot perform dynamic current-limiting according to the actual load conditions and characteristics of the interface. In order to improve the overall throughput of resources and reduce the average delay of the interface, it is urgent to propose a global interface dynamic current-limiting method based on reinforcement learning. Summary of the Invention
[0004] Aiming at the above deficiencies in the prior art, the global interface dynamic current-limiting method based on reinforcement learning provided by the present invention solves the problem that static rules cannot sense the load conditions of the interface to perform current-limiting, so as to achieve the effect of maximizing the throughput of the current-limiting logic and reducing the request access delay.
[0005] In order to achieve the above invention purpose, the technical solution adopted by the present invention is: A global interface dynamic current-limiting method based on reinforcement learning, including the following steps:
[0006] S1. Construct a request simulator, and send simulated requests to the intelligent agent through the request simulator;
[0007] S2. Collect state information, make decisions by the intelligent agent according to the current state information, and establish an action reward function;
[0008] S3. Randomly select actions by the intelligent agent, obtain a reward value according to the selected action through the action reward function, and establish a Q function;
[0009] S4. Train the Q function through a deep neural network, and obtain a trained intelligent agent according to the optimal Q function;
[0010] S5. Make decisions on the simulation requests by the trained agent.
[0011] Further: In S1, the method for the request simulator to generate simulation requests is specifically as follows:
[0012] Capture http traffic from the online environment through tcpdump, use Tshark to convert the http traffic into a json file, parse the json file and create simulation requests.
[0013] Further: In S2, the method for collecting status information is specifically as follows:
[0014] Obtain status feature information, including cpu utilization rate, memory usage rate, disk usage rate, network bandwidth usage rate, stack space usage rate, the set of interfaces to be accessed currently, and the average response millisecond duration of all interfaces within a preset time interval;
[0015] Convert the status feature information into a vector, and perform regularization processing on the values of the vector to obtain the status information.
[0016] Further: In S2, the specific expression for establishing the action reward function R is as follows:
[0017]
[0018] In the formula, n is the number of simulation requests within the time interval Δ from the previous moment to this moment, r t is the reward value of the i-th simulation request; i For the i-th simulation request;
[0019]
[0020] In the formula, res(i) is the response duration of each interface, i is the sequence number of each request within Δ t , avg(t - 1) is, t is the time.
[0021] Further: In S3, the method for updating the Q function Q(S,A) is specifically as follows:
[0022] Q(S,A)' ← Q(S,A) + α[R + γmax Q(S',A') - Q(S,A)]
[0023] In the formula, Q(S,A)' is the updated Q function, α is the first fixed value, γ is the second fixed value, S' is the next state, A' is the next action, A is the action, S is the state, including release and traffic limiting, and max Q(S',A') is the value with the maximum expected reward among all possible actions taken in the next state S'.
[0024] Further: In S4, the method for training the Q function through a deep neural network is specifically as follows:
[0025] S41. Define a neural network;
[0026] S42. Obtain the state information of the Q function for each round of training;
[0027] S43. Update the network parameters using the gradient descent algorithm according to the current Q function;
[0028] S44. Set the termination flag for each round of training;
[0029] S45. Use cgroup to limit the amount of resources that the business program can use, and simulate scenarios with different resource amounts for training.
[0030] Furthermore: In S41, the neural network includes an evaluation network and a target network. The evaluation network and the target network have the same structure, and both include:
[0031] The input layer is: 1000 input nodes and 64 output nodes. The number of input nodes is set according to the maximum concurrency. The input of the input layer is the state vector collected in step S2;
[0032] The hidden layer is: 64 input nodes and 64 output nodes;
[0033] The hidden layer is: 64 input nodes and 64 output nodes;
[0034] The output layer is: 64 input nodes and 2 output nodes. The number of output nodes is set according to the action type. The action type includes release and flow limiting, that is, the action to be executed for the last request in the same batch of inputs is output;
[0035] Specifically, S43 is: Update the weight parameters of the evaluation network and the target network at different frequencies;
[0036] Among them, the mean squared error is used as the loss function, and the gradient descent algorithm is used to update the network parameter θ;
[0037]
[0038] In the formula, L(θ) is the loss function. The goal of training is to find a set of optimal parameters θ* such that L(θ) obtains the minimum value. E represents the mathematical expectation, y is the target Q value, Q(S; A; θ) is the Q value estimation of the evaluation network for action A in state S, and θ is the parameter of the evaluation network.
[0039] The beneficial effects of the present invention are:
[0040] (1) The present invention provides a global interface dynamic current limiting method based on reinforcement learning. The interface current limiting logic based on reinforcement learning enables the agent to learn to make intelligent decisions on whether to limit the current request or let it pass under different loads, different concurrent requests, and when there are potential statistical laws in the access sequence of the interface.
[0041] (2) Traditional rule-based current limiting methods can only be applied to a single interface, while this method is applied to global current limiting and can maximize the global throughput.
[0042] (3) Traditional rule-based current limiting methods cannot make decisions dynamically according to the system load, while this method can make dynamic decisions based on the resource load situation.
[0043] (4) Traditional rule-based current limiting methods cannot perform dynamic current limiting dynamically according to the interface characteristics. This method is an end-to-end method that can learn the interface characteristics to make decisions.
[0044] (5) This method can learn the potential statistical laws of the interface access sequence and can avoid overload in advance when approaching the critical point. Description of the Drawings
[0045] Figure 1 It is a flowchart of a global interface dynamic current limiting method based on reinforcement learning according to the present invention. Detailed Embodiments
[0046] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the inventive concept of the present invention are within the scope of protection.
[0047] As Figure 1 shown, in an embodiment of the present invention, a global interface dynamic current limiting method based on reinforcement learning includes the following steps:
[0048] S1. Construct a request simulator and send simulated requests to the agent through the request simulator;
[0049] S2. Collect status information, and the agent makes decisions based on the current status information to establish an action reward function;
[0050] S3. The agent randomly selects an action, and the action reward function obtains a reward value according to the selected action to establish a Q function;
[0051] S4. Train the Q function through a deep neural network, and obtain a trained agent according to the optimal Q function;
[0052] S5. Make decisions on simulated requests through the trained agent.
[0053] In this embodiment, the basic idea of the present invention is that the requests of users for interfaces generally have certain statistical laws. For example, if a user first accesses the registration function, then it is highly probable that the user will access the login function next. Based on this phenomenon, when multiple users request simultaneously, it is possible to dynamically decide whether to allow certain requests according to the pressure on the interface caused by the access.
[0054] Suppose at a certain moment, there is the following upcoming access sequence:
[0055] The access sequence of User 1 is Interface 1 -> Interface 2, and Interface 2 has a relatively high service pressure and takes a long time; the access sequence of User 2 is Interface 3 -> Interface 4, and the service pressures of Interface 3 and Interface 4 are relatively small; the access sequence of User 3 is Interface 4 -> Interface 5, and the service pressures of Interface 4 and Interface 5 are relatively small.
[0056] When the current server pressure is about to reach the threshold and is insufficient to support the access of 3 requests, in order to improve the overall throughput, at this time, the throttling logic should throttle the request for Interface 1 of User 1 and allow the requests of User 2 and User 3. In this way, the request for Interface 2 of User 1 will not occur either, and the subsequent requests of User 2 and User 3 can pass normally.
[0057] The present invention creates an agent that can make an optimal decision on whether to throttle the current request according to the environmental information, and this decision needs to consider the current state and possible future states.
[0058] In S1, the method for the request simulator to generate simulated requests is specifically as follows:
[0059] Capture http traffic from the online environment through tcpdump, use Tshark to convert the http traffic into a json file, parse the json file and create simulated requests.
[0060] In this embodiment, the time required for the agent to be created from the initial state to a state where it can make a good decision is relatively long, and the direct online training cost is relatively high. Therefore, the present invention creates an offline request simulator to simulate the online simulated requests to train the agent.
[0061] In S2, the method for collecting state information is specifically as follows:
[0062] Obtain status feature information, including CPU utilization rate, memory usage rate, disk usage rate, network bandwidth usage rate, stack space usage rate, the set of interfaces to be accessed currently, and the average response millisecond duration of all interfaces within a preset time interval;
[0063] Convert the status feature information into a vector, and normalize the values of the vector. For example, multiply the CPU utilization rate of 30% by 100 to get 30, and for the value of the URL type, only take the prefix part to establish a dictionary mapping and identify it with numbers to obtain the status information.
[0064] In S2, the specific expression for establishing the action return function R is:
[0065]
[0066] In the formula, n is the simulated request volume within the time interval Δ of the calculation cycle from the previous time to this moment, r t is the return value of the i-th simulated request; i In the formula, res(i) is the response duration of each interface, i is the serial number of each request within Δ
[0067]
[0068] In the formula, avg(t - 1) is, t is the time. If the interface is released and the response time is shorter than the previous response time, the return value is 1; if it is directly rate-limited, the return value is -1. t In this embodiment, the present invention establishes an action return function based on the idea of reinforcement learning to obtain the return values when taking various actions in the current state. The global goal is to obtain a decision sequence with the maximum cumulative return value.
[0069] In this embodiment, the present invention establishes an action return function based on the idea of reinforcement learning to obtain the return values when taking various actions in the current state. The global goal is to obtain a decision sequence with the maximum cumulative return value.
[0070] In S3, the method for updating the Q function Q(S, A) is specifically:
[0071] Q(S, A)' ← Q(S, A) + α[R + γmax Q(S', A') - Q(S, A)]
[0072] In the formula, Q(S, A)' is the updated Q function, α is the first fixed value, γ is the second fixed value, S' is the next state, A' is the next action, A is the action, S is the state, including release and rate-limiting, and max Q(S', A') is the value with the maximum expected return among all possible actions taken in the next state S'.
[0073] In this embodiment, during the training phase, the agent selects different actions, calculates the reward value according to the action reward function, and thus establishes the Q function. The Q function is the result of the agent's training. That is, after obtaining the Q function, it is possible to judge the maximum benefit that can be obtained after taking an action based on the current state, and make a decision on which action to execute based on this benefit.
[0074] In S4, since most of the input vectors are continuous non-discrete data and representing the Q function with a matrix is too sparse, the Q function is represented by a neural network. The specific definition and training process are as follows:
[0075] S41. Define the neural network;
[0076] The input layer has: 1000 input nodes and 64 output nodes. The number of input nodes is set according to the maximum concurrency. The input of the input layer is the state vector collected in step S2.
[0077] The hidden layer has: 64 input nodes and 64 output nodes;
[0078] The hidden layer has: 64 input nodes and 64 output nodes;
[0079] The output layer has: 64 input nodes and 2 output nodes. The number of output nodes is set according to the action type. The action types include allowing passage and restricting flow, that is, the action that the last request in the same batch of inputs needs to execute is output.
[0080] S42. Obtain the state information of the Q function for each round of training;
[0081] Simulate the random access requests of multiple users to generate training data. The specific generation method is to randomly select the access sequences of m different users for each round of training, and use the traffic generator to generate simulated requests using multi-threading according to the selected concurrency and access interval, and then generate the state information.
[0082] S43. Update the network parameters using the gradient descent algorithm according to the current Q function;
[0083] In order to update the parameter weights of the Q network, the method of using a target network is used to simulate the effect of temporal difference. The mean square error is used as the loss function, and the gradient descent algorithm is used to update the network parameter θ;
[0084] L(θ) = E[(y - Q(S; A; θ)) 2
[0085] In the formula, L(θ) is the loss function. The goal of training is to find a set of optimal parameters θ* such that L(θ) obtains the minimum value. E represents the mathematical expectation, y is the target Q value, Q(S; A; θ) is the Q value estimation of the evaluation network for action A in state S, and θ is the parameter of the evaluation network;
[0086] In this embodiment, there are two networks existing simultaneously, Q (the evaluation network) and Q π (the target network). The two networks update their weight parameters at different frequencies, that is, Q π is updated in each batch iteration, while Q is updated once every 5 (a hyperparameter that can be adjusted) batches (that is, copy Q π to Q).
[0087] To prevent the agent from prematurely falling into a local optimal strategy, the ε-greedy strategy is used here to select exploration actions, that is, with a probability of 1 - ε, the action is determined by the Q function, and with a probability of ε, other actions are randomly selected. ε gradually decreases as the iteration progresses, that is, the more towards the end, it tends to use the result of the Q function.
[0088] Each execution of an action will cause a change in the state of the business program, and the released interfaces will also have different response durations. The agent will monitor the response durations of all requests to calculate the reward value.
[0089] S44. Set the termination flag for each round of training;
[0090] The termination flag for each round of training: The server resource utilization rate is 100% (cpu / mem / stack), and the response duration of interfaces exceeding a preset ratio (such as 5%) exceeds a certain threshold (such as 10s), as well as the number of iterations (such as 1000 iterations) and the training time (such as 2 hours of training).
[0091] S45. Use cgroup to limit the amount of resources that the business program can use, and simulate scenarios with different resource amounts for training.
[0092] The beneficial effects of the present invention are as follows: The present invention provides a global interface dynamic current limiting method based on reinforcement learning. Based on the interface current limiting logic of reinforcement learning, the agent can learn to make intelligent decisions on whether to limit the current request or release it under different loads, different concurrent requests, and when there are potential statistical laws in the access sequence of interfaces.
[0093] Traditional rule-based current limiting methods can only be applied to a single interface, while this method is applied to global current limiting and can maximize the improvement of global throughput.
[0094] Traditional rule-based current limiting methods cannot make decisions dynamically according to the system load, while this method can make dynamic decisions based on the resource load situation.
[0095] Traditional rule-based current limiting methods cannot perform dynamic current limiting according to interface characteristics. This method is an end-to-end method and can learn interface characteristics to make decisions.
[0096] This method can learn the potential statistical laws of the interface access sequence and avoid overload in advance when approaching the critical point.
[0097] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation on the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of technical features. Therefore, the features defined by "first", "second", "third" may explicitly or implicitly include one or more of such features.
Claims
1. A global interface dynamic flow limiting method based on reinforcement learning, characterized in that It includes the following steps: S1. Construct a request simulator and send simulated requests to the agent through the request simulator; S2. Collect status information, make decisions by the agent according to the current status information, and establish an action reward function; S3. Randomly select actions by the agent, obtain reward values according to the selected actions through the action reward function, and establish a Q function; S4. Train the Q function through a deep neural network, and obtain a trained agent according to the optimal Q function; S5. Make decisions on simulated requests through the trained agent.
2. The global interface dynamic current limiting method based on reinforcement learning according to claim 1, wherein In S1, the method for the request simulator to generate simulated requests is specifically: Capture http traffic from the online environment through tcpdump, use Tshark to convert the http traffic into a json file, parse the json file and create simulated requests.
3. The global interface dynamic flow limiting method based on reinforcement learning according to claim 1, wherein In S2, the method for collecting status information is specifically: Obtain status feature information, including cpu utilization rate, memory usage rate, disk usage rate, network bandwidth usage rate, stack space usage rate, the set of interfaces to be accessed currently, and the average response millisecond duration of all interfaces within a preset time interval; Convert the status feature information into a vector, and regularize the values of the vector to obtain status information.
4. The global interface dynamic current limiting method based on reinforcement learning according to claim 3, wherein, In S2, the specific expression for establishing the action reward function R is: Where n is the number of simulated requests within the time interval Δ from the previous moment to the current moment t and r i is the return value of the i-th simulated request; Where res(i) is the response duration of each interface, i is the sequence number of each request within Δ t , avg(t - 1) is, and t is the time.
5. The global interface dynamic current limiting method based on reinforcement learning according to claim 4, wherein In S3, the method for updating the Q function Q(S,A) is specifically: Q(S,A)'←Q(S,A)+α[R+γmax Q(S',A')-Q(S,A)] In the formula, Q(S,A)' is the updated Q function, α is the first fixed value, γ is the second fixed value, S' is the next state, A' is the next action, A is the action, S is the state, including release and flow limit, and max Q(S',A') is the value with the maximum expected reward among all possible actions taken in the next state S'.
6. The global interface dynamic current limiting method based on reinforcement learning according to claim 4, characterized in that In S4, the method for training the Q function through a deep neural network is specifically: S41. Define a neural network; S42. Obtain the status information of the Q function for each round of training; S43. Update the network parameters according to the current Q function using the gradient descent algorithm; S44. Set the termination flag for each round of training; S45. Use cgroup to limit the amount of resources that the business program can use, and simulate scenarios with different resource amounts for training.
7. The global interface dynamic flow limiting method based on reinforcement learning according to claim 6, characterized in that In S41, the neural network includes an evaluation network and a target network. The structures of the evaluation network and the target network are the same, and both include: The input layer is: 1000 input nodes and 64 output nodes. The number of input nodes is set according to the maximum concurrency. The input of the input layer is the status vector collected in step S2; The hidden layer is: 64 input nodes and 64 output nodes; The hidden layer is: 64 input nodes and 64 output nodes; The output layer is: 64 input nodes and 2 output nodes. The number of output nodes is set according to the action type. The action type includes release and flow limit, that is, output the action that the last request in the same batch needs to execute; S43 is specifically: update the weight parameters of the evaluation network and the target network at different frequencies; Among them, the mean squared error is used as the loss function, and the gradient descent algorithm is used to update the network parameter θ; L(θ) = E[(y - Q(S; A; θ)) 2 In the formula, L(θ) is the loss function. The training objective is to find a set of optimal parameters θ* such that L(θ) reaches the minimum value. E represents the mathematical expectation, y is the target Q value, Q(S; A; θ) is the Q value estimation of the evaluation network for action A in state S, and θ is the parameter of the evaluation network.
Citation Information
Patent Citations
Distributed software service guarantee method based on swarm intelligence
CN114900420A
Deep learning multi-agent micro-grid cooperative control method based on double neural networks
CN115333143A
DQN-based adaptive hello interval adjustment method
CN118250764A
Information adjusting method and device, and storage medium
WO2022252546A1
Cited By
IMS (IP Multimedia Subsystem) service distribution method based on reinforcement learning
CN120768841A