Method and system for dynamically adjusting congestion control algorithm based on reinforcement learning
Through reinforcement learning algorithms, real-time monitoring of network status and dynamically selecting the optimal congestion control algorithm, the problem of improper congestion control in network transmission is solved, and network efficiency and data transmission rate are improved.
Patent Information
- Application Number
- CN202510711449.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-08
AI Technical Summary
The prior art is difficult to select the optimal congestion control algorithm in real time during network transmission, resulting in worsening of network congestion.
Through the reinforcement learning algorithm, the network state changes are monitored in real time, the loss value of the congestion control algorithm is updated, and the optimal congestion control algorithm is dynamically selected to optimize network transmission.
Improve network transmission efficiency, reduce network congestion, and improve data transmission rate.
Smart Images

Figure CN120455372A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network communication technology, and in particular to a method and system for dynamically adjusting a congestion control algorithm based on reinforcement learning. Background Art
[0002] In many network usage scenarios today, network transmission load has increased dramatically, and network congestion has become a common and urgent problem in network transmission. Therefore, quickly selecting an appropriate congestion control algorithm is crucial to ensuring efficient and stable network operation.
[0003] Existing related technologies have many shortcomings when dealing with network congestion. For example, Chinese patent application No. CN113595923A discloses a network congestion control method and device. This method obtains data through a network simulator and trains a congestion control algorithm model based on deep reinforcement learning. However, this method only trains and learns congestion control parameters and does not involve the selection of different congestion control algorithms. Chinese patent application No. CN117692396A discloses a TCP unilateral acceleration method and device in a complex network environment. Although this method can adjust the sending window according to network conditions, it cannot achieve real-time switching of congestion control algorithms. For another example, Chinese patent application No. CN114679413A discloses a congestion control method, device, equipment, and storage medium for heterogeneous networks. This method selects a smaller congestion window for data transmission and similarly does not solve the problem of real-time switching of congestion control algorithms.
[0004] On the other hand, although there are many types of network congestion control algorithms, including BBR algorithm, Cubic algorithm, DCCTCP algorithm, etc., each of these congestion control algorithms has its own applicable network environment, and it is difficult to always choose the best algorithm. Inappropriate selection will aggravate network congestion. Summary of the Invention
[0005] The technical problem to be solved by the present invention is as follows: In response to the above-mentioned problems in the prior art, a method and system for dynamically adjusting the congestion control algorithm based on reinforcement learning are provided. The present invention aims to automatically select a network congestion control algorithm suitable for the network conditions in response to the network congestion problem during network transmission to achieve more efficient data transmission.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: A method for dynamically adjusting a congestion control algorithm based on reinforcement learning comprises the following steps: obtaining a round-trip time of network transmission when the network state changes, mapping the round-trip time of network transmission to the current environmental state of the network, using the current environmental state of the network as the state s of the reinforcement learning algorithm, selecting a congestion control algorithm from a preset congestion control algorithm list as an action a of the reinforcement learning algorithm, and using the reinforcement learning algorithm to update the Q value of each congestion control algorithm in the congestion control algorithm list under the current environmental state; and when a network connection is established, sorting the congestion control algorithms in the congestion control algorithm list according to their Q values under the current environmental state, and selecting the congestion control algorithm with the smallest Q value under the current environmental state with a preset probability to perform congestion control.
[0007] Optionally, obtaining the round-trip time of network transmission when the state of the network changes includes: detecting the network state change and the corresponding IP address, port and time of connection establishment by using code written by using eBPF technology and mounted on the critical path of the sockops hook of the Linux kernel, the critical path includes a BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB node for detecting the active connection establishment of the network, a BPF_SOCK_OPS_PASSIVE_ESTABLISHED_CB node for detecting the passive connection establishment of the network, and a BPF_SOCK_OPS_STATE_CB node for detecting the state change of the network. If the state change of the network is detected by the BPF_SOCK_OPS_STATE_CB node, the IP address and port of the active and passive connection establishment records are searched according to the IP address and port of the closed network connection to determine the time of establishment of the corresponding network connection, and the time of network state change is subtracted from the time of establishment of the corresponding network connection to obtain the round-trip time of network transmission.
[0008] Optionally, mapping the round-trip time of network transmission to the current environmental state of the network includes: matching the round-trip time of network transmission with the value ranges corresponding to the preset value ranges of the three delay state levels of low delay, medium delay, and high delay, thereby obtaining the delay state level corresponding to the round-trip time of network transmission, which should be converted into the corresponding current environmental state value through a preset mapping relationship table.
[0009] Optionally, the function expression for updating the Q value of each congestion control algorithm in the congestion control algorithm list under the current environment state by using the reinforcement learning algorithm is: , in, is the Q value predicted by the congestion control algorithm under state s and action a, For update operations, is the learning rate, is the reward for executing action a in state s, which is the difference between the round-trip time and the minimum round-trip time of network transmission in state s and action a of the congestion control algorithm. is the attenuation parameter, The congestion control algorithm is in the next state and the maximum value of the Q value predicted under action a.
[0010] Optionally, before using the reinforcement learning algorithm to update the Q value of each congestion control algorithm in the congestion control algorithm list under the current environment state, the method further includes querying a preset mapping function based on the difference between the two most recent Q values of each congestion control algorithm to dynamically determine the learning rate. The greater the difference between the two most recent Q values, the greater the learning rate The larger the value of .
[0011] Optionally, when selecting the congestion control algorithm with the smallest Q value under the current environmental state with a preset probability to perform congestion control, it includes selecting the congestion control algorithm with the smallest Q value under the current environmental state with a preset first probability to perform congestion control, and randomly selecting a congestion control algorithm from a preset congestion control algorithm list with a second probability to perform congestion control, the sum of the first probability and the second probability is 1, and the first probability is greater than the second probability.
[0012] Optionally, the first probability is 95% and the second probability is 5%.
[0013] The present invention also provides a system for dynamically adjusting a congestion control algorithm based on reinforcement learning, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the method for dynamically adjusting a congestion control algorithm based on reinforcement learning.
[0014] The present invention also provides a computer-readable storage medium, which stores a computer program or instruction. The computer program or instruction is programmed or configured to execute the method of dynamically adjusting the congestion control algorithm based on reinforcement learning through a processor.
[0015] The present invention also provides a computer program product, comprising a computer program or instructions, which are programmed or configured to execute the method of dynamically adjusting the congestion control algorithm based on reinforcement learning through a processor.
[0016] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: The present invention monitors network status changes in real time and updates the loss values of various congestion control algorithms according to the latest network environment, thereby continuously optimizing the selection of congestion control algorithms, improving network transmission efficiency, and reducing network congestion. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a flow chart of a method for dynamically adjusting a congestion control algorithm based on reinforcement learning in an embodiment of the present invention.
[0018] Figure 2 This is a flowchart of a congestion control algorithm that dynamically adjusts traffic based on reinforcement learning in an embodiment of the present invention.
[0019] Figure 3 This is a flowchart of a congestion control algorithm that is dynamically adjusted based on reinforcement learning in a specific application embodiment of the present invention. DETAILED DESCRIPTION
[0020] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings in the embodiments of the present invention.
[0021] Example 1 like Figure 1 As shown, this embodiment provides a method for dynamically adjusting a congestion control algorithm based on reinforcement learning, comprising the following steps: obtaining a round-trip time of network transmission when the network state changes, mapping the round-trip time of network transmission to the current environmental state of the network, using the current environmental state of the network as the state s of the reinforcement learning algorithm, selecting a congestion control algorithm from a preset congestion control algorithm list as the action a of the reinforcement learning algorithm, and using the reinforcement learning algorithm to update the Q value of each congestion control algorithm in the congestion control algorithm list under the current environmental state; when a network connection is established, sorting the congestion control algorithms in the congestion control algorithm list according to their Q values under the current environmental state, and selecting the congestion control algorithm with the smallest Q value under the current environmental state with a preset probability to perform congestion control.
[0022] like Figure 2 As shown, reinforcement learning involves an agent and an environment. When the agent executes action a using a certain strategy in its current state, the environment changes to a new state. After calculating the resulting reward, the agent then continues to execute actions using a certain strategy based on the new state and reward, repeating the cycle until convergence. In this example, action a corresponds to the congestion control algorithm employed, the environment is the network environment, the state s represents multiple states of RTT (the total time it takes for data to travel from the sender to the receiver and back again, a measure of network latency), the reward corresponds to the difference between the current RTT and the minimum RTT, and Q is the calculated loss.
[0023] In this embodiment, obtaining the round-trip time of network transmission when the network state changes includes: detecting the network state change and the IP address, port, and time corresponding to the connection establishment by writing code using eBPF technology (Extended Berkeley Packet Filter, a technology used in multiple fields such as network and security to track and detect abnormal states) and mounting it on the critical path of the sockops hook of the Linux kernel. The critical path includes a BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB node for detecting active network connection establishment, a BPF_SOCK_OPS_PASSIVE_ESTABLISHED_CB node for detecting passive network connection establishment, and a BPF_SOCK_OPS_STATE_CB node for detecting network state changes. If the network state change is detected using the BPF_SOCK_OPS_STATE_CB node, the IP addresses and ports of the active and passive connection establishment records are searched based on the IP address and port of the closed network connection to determine the time when the corresponding network connection was established. The round-trip time of the network transmission is obtained by subtracting the time when the corresponding network connection was established from the time when the network state changed.
[0024] In this embodiment, mapping the round-trip time of network transmission to the current environmental state of the network includes: matching the round-trip time of network transmission with the value ranges corresponding to the preset value ranges of the three delay state levels of low delay, medium delay, and high delay, thereby obtaining the delay state level corresponding to the round-trip time of network transmission, which should be converted into the corresponding current environmental state value through a preset mapping relationship table.
[0025] like Figure 3 As shown, in a specific application embodiment, when the TCP (Transmission Control Protocol) connection state changes, the IP address and port of the connection are obtained, and the minimum RTT of the same IP address and port in the reinforcement learning process is obtained. The current environment state is set according to the size of the current RTT: when the RTT is less than 10ms, the current state is low latency, and the variable current_state is set to 0; when the RTT is greater than 10ms and less than 50ms, the current state is medium latency, and the variable current_state is set to 1; when the RTT is greater than 50ms, the current state is high latency, and the variable current_state is set to 2. The loss value of the current delay state is calculated using the following congestion control algorithm: The loss value of the current delay state = (current rtt-minimum rtt) / minimum rtt, The loss value of the current delayed state is then updated using the Q-learning algorithm. Q-learning is a classic model-free reinforcement learning algorithm. Its core goal is to enable an agent to learn the optimal action strategy for a specific state through interaction with the environment, thereby maximizing long-term cumulative rewards. Its core is to construct a Q-value table that records the expected reward for each state-action pair.
[0026] When TCP establishes a connection, it sorts each congestion control algorithm from smallest to largest based on their current loss values. Using the random function provided by BPF, it obtains a random value, ensuring a 95% probability of using the optimal congestion control algorithm and a 5% probability of using another random congestion control algorithm. If multiple optimal congestion control algorithms (i.e., those with the same minimum loss values) exist, the optimal one is randomly selected.
[0027] In this embodiment, the Q-learning algorithm, which dynamically adjusts the congestion control algorithm based on reinforcement learning, first sets two algorithm parameters: the exploration coefficient ϵ and the update step size α (where α ranges from (0, 1]). Next, for all possible states s (in the state space S) and actions a (in the action space A), the corresponding Q values are initialized, typically to 0. Then, iterations are performed according to the following steps: S1, based on the current Q table and greedy policy ϵ, selects an action a in the current state s; S2, executes the selected action,a, and observes the reward r obtained after execution and the new state s′ achieved; S3, update Q value; S4, updates the status to s '; S5, repeat the above steps until the termination state is reached.
[0028] In this embodiment, through this iterative process, the algorithm can learn the optimal strategy for taking different actions in different states, that is, find the action selection strategy that maximizes the long-term cumulative reward.
[0029] In this embodiment, the reinforcement learning algorithm is used to update the Q value of each congestion control algorithm in the congestion control algorithm list under the current environment state. The function expression is: , in, is the Q value predicted by the congestion control algorithm under state s and action a, For update operations, is the learning rate, is the reward for executing action a in state s. The reward is the difference between the round-trip time and the minimum round-trip time of network transmission in state s and action a of the congestion control algorithm. is the attenuation parameter, The congestion control algorithm is in the next state and the maximum value of the Q value predicted under action a.
[0030] In this embodiment, before using the reinforcement learning algorithm to update the Q value of each congestion control algorithm in the congestion control algorithm list under the current environment state, it also includes querying the preset mapping function based on the difference between the two most recent Q values of each congestion control algorithm to dynamically determine the learning rate The greater the difference between the two most recent Q values, the greater the learning rate The larger the value of .
[0031] Specifically, when the difference between the old and new Q values is less than 1000, the learning rate is 0.01; when the difference is greater than 1000 but less than 2000, the learning rate is 0.03; when the difference is greater than 2000 but less than 5000, the learning rate is 0.06; when the difference is greater than 5000 but less than 10000, the learning rate is 0.12; and when the difference is greater than 10000, the learning rate is 0.25. This ensures both rapid learning in the early stages and stable convergence in the later stages.
[0032] In this embodiment, when selecting the congestion control algorithm with the smallest Q value under the current environmental state to perform congestion control with a preset probability, it includes selecting the congestion control algorithm with the smallest Q value under the current environmental state to perform congestion control with a preset first probability, and randomly selecting a congestion control algorithm from a preset congestion control algorithm list to perform congestion control with a second probability, the sum of the first probability and the second probability is 1, and the first probability is greater than the second probability.
[0033] In this embodiment, the first probability is 95% and the second probability is 5%.
[0034] Example 2 In this embodiment, the Linux kernel's sockops hooks several key paths, such as BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB, which is called when a passive connection occurs. BPF_SOCK_OPS_PASSIVE_ESTABLISHED_CB, which is called when an active connection occurs, can be used to determine when a connection is established and set the congestion control algorithm. BPF_SOCK_OPS_STATE_CB, which is called when the TCP connection state changes, determines whether the connection is closed and can be used to calculate losses and updates.
[0035] In this embodiment, the method for dynamically adjusting the congestion control algorithm based on reinforcement learning includes the following steps: S1, adds loss calculation to the congestion control algorithm; 1) Based on eBPF technology, obtain the currently connected sock (network socket), and obtain the IP address and port according to the current sock; 2) Get the rtt of the TCP connection under current congestion control; 3) Based on the RTT size, the system is divided into three delay states: low delay, medium delay, and high delay, which correspond to the state in reinforcement learning. The delay state value of the environment at that time is set according to the RTT size; 4) Calculate the loss value based on the RTT size and the minimum RTT difference of all congestion control algorithms for the current connection.
[0036] S2, add Q-learning algorithm to update loss value; 1) Based on eBPF technology, obtain the currently connected sock, and obtain the IP address and port according to the current sock; 2) Get the old loss value of this state under the current congestion control; 3) Dynamically adjust the learning rate based on the difference between the old loss value and the new loss value calculated by S1; 4) Use Q-learning to update the loss value, where the learning rate varies based on the difference between the old and new losses. When the difference is large, the learning rate is high, and when the difference is small, the learning rate is low. This accelerates convergence when the difference is large, ensuring rapid learning in the early stages, and ensures stable convergence in the later stages when the difference is small.
[0037] S3, adds dynamic selection of congestion control algorithm; 1) Based on eBPF technology, obtain the currently connected sock; 2) Sort the congestion control algorithm by the loss size of the current environment delay state; 3) The congestion control algorithm with the minimum loss value is selected for the connection with a greater probability, and other algorithm settings are selected with a smaller probability.
[0038] This implementation classifies RTT into different delay states, uses the Q-learning algorithm to calculate the loss for each state, and dynamically adjusts the learning rate for updates, selecting the optimal congestion control algorithm settings in real time. Using this algorithm in IPERF traffic testing, we achieved a 10% rate improvement in both environments with a delay of approximately 1ms and 30ms.
[0039] This embodiment also provides a system for dynamically adjusting a congestion control algorithm based on reinforcement learning, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute a method for dynamically adjusting a congestion control algorithm based on reinforcement learning.
[0040] This embodiment also provides a computer-readable storage medium, which stores a computer program or instruction. The computer program or instruction is programmed or configured to execute a method for dynamically adjusting a congestion control algorithm based on reinforcement learning through a processor.
[0041] This embodiment also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute, through a processor, a method for dynamically adjusting a congestion control algorithm based on reinforcement learning.
[0042] Those skilled in the art should understand that the technical solution provided by the present invention may be in the form of a method, a system, or a computer program product. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the functions described in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0043] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A method for dynamically adjusting a congestion control algorithm based on reinforcement learning, characterized in that: The method comprises the following steps: obtaining the round-trip time of network transmission when the state of the network changes, mapping the round-trip time of network transmission to the current environmental state of the network, using the current environmental state of the network as the state s of the reinforcement learning algorithm, selecting a congestion control algorithm from a preset congestion control algorithm list as the action a of the reinforcement learning algorithm, and using the reinforcement learning algorithm to update the Q value of each congestion control algorithm in the congestion control algorithm list under the current environmental state; when a network connection is established, sorting the congestion control algorithms in the congestion control algorithm list according to their Q values under the current environmental state, and selecting the congestion control algorithm with the smallest Q value under the current environmental state with a preset probability to perform congestion control.
2. The method for dynamically adjusting the congestion control algorithm based on reinforcement learning according to claim 1, characterized in that: The method of obtaining the round-trip time of network transmission when the state of the network changes includes: detecting the network state change and the corresponding IP address, port and time of connection establishment by using code written by using eBPF technology and mounted on the critical path of the sockops hook of the Linux kernel, wherein the critical path includes a BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB node for detecting the active connection establishment of the network, a BPF_SOCK_OPS_PASSIVE_ESTABLISHED_CB node for detecting the passive connection establishment of the network, and a BPF_SOCK_OPS_STATE_CB node for detecting the state change of the network. If the state change of the network is detected by the BPF_SOCK_OPS_STATE_CB node, the IP address and port of the active and passive connection establishment records are searched according to the IP address and port of the closed network connection, so as to determine the time of establishment of the corresponding network connection, and the time of network state change minus the time of establishment of the corresponding network connection is subtracted to obtain the round-trip time of network transmission.
3. The method for dynamically adjusting the congestion control algorithm based on reinforcement learning according to claim 1, characterized in that: The mapping of the round-trip time of network transmission to the current environmental state of the network includes: matching the round-trip time of network transmission with the value ranges corresponding to the preset value ranges of the three delay state levels of low delay, medium delay, and high delay, thereby obtaining the delay state level corresponding to the round-trip time of network transmission, which should be converted into the corresponding current environmental state value through a preset mapping relationship table.
4. The method for dynamically adjusting the congestion control algorithm based on reinforcement learning according to claim 1, characterized in that: The function expression for updating the Q value of each congestion control algorithm in the congestion control algorithm list under the current environment state by using the reinforcement learning algorithm is: , in, is the Q value predicted by the congestion control algorithm under state s and action a, For update operations, is the learning rate, is the reward for executing action a in state s, which is the difference between the round-trip time and the minimum round-trip time of network transmission in state s and action a of the congestion control algorithm. is the attenuation parameter, The congestion control algorithm is in the next state and the maximum value of the Q value predicted under action a.
5. The method for dynamically adjusting the congestion control algorithm based on reinforcement learning according to claim 4, characterized in that: Before using the reinforcement learning algorithm to update the Q value of each congestion control algorithm in the congestion control algorithm list under the current environment state, the method also includes querying a preset mapping function based on the difference between the two most recent Q values of each congestion control algorithm to dynamically determine the learning rate. The greater the difference between the two most recent Q values, the greater the learning rate The larger the value of .
6. The method for dynamically adjusting the congestion control algorithm based on reinforcement learning according to claim 4, characterized in that: The selecting of the congestion control algorithm with the smallest Q value under the current environmental state to perform congestion control with a preset probability includes selecting the congestion control algorithm with the smallest Q value under the current environmental state to perform congestion control with a preset first probability, and randomly selecting a congestion control algorithm from a preset congestion control algorithm list to perform congestion control with a second probability, wherein the sum of the first probability and the second probability is 1, and the first probability is greater than the second probability.
7. The method for dynamically adjusting the congestion control algorithm based on reinforcement learning according to claim 6, characterized in that: The first probability is 95% and the second probability is 5%.
8. A system for dynamically adjusting a congestion control algorithm based on reinforcement learning, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the method for dynamically adjusting the congestion control algorithm based on reinforcement learning as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute, through a processor, the method for dynamically adjusting a congestion control algorithm based on reinforcement learning as recited in any one of claims 1 to 7.
10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute, through a processor, the method for dynamically adjusting a congestion control algorithm based on reinforcement learning as recited in any one of claims 1 to 7.
Citation Information
Patent Citations
Congestion algorithm adaptive control method, storage medium, congestion algorithm adaptive control equipment and congestion algorithm adaptive control system
CN112422443A
Self-adaptive congestion control method and system integrating deep reinforcement learning and BBR protocol
CN113645144A
Satellite network congestion control method and control system based on Q-learning for TCP protocol
CN118740739A
Congestion control based on network telemetry
US20220311711A1