A congestion control method for satellite network environment

Optimizing congestion window control through reinforcement learning methods and dynamic reward function, the problem of poor throughput and delay control in satellite network environment is solved, and higher throughput and lower latency are achieved.

CN116209005BActive Publication Date: 2025-07-08CHANGCHUN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211732698.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2025-07-08
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

The existing congestion control methods cannot effectively adapt to the high latency, high packet loss and dynamic periodic changes in satellite network environments, resulting in poor throughput and delay control.

Method used

The reinforcement learning method is adopted to obtain the network state through interval time sampling, and the inverse proportional increase, linear reduction of the automatic speed change type action and reward function based on historical information are designed. The congestion window adjustment is combined with the heuristic idea, and the congestion window control strategy is updated using the Q-learning algorithm.

Benefits of technology

High throughput and low round trip delay are achieved in satellite network environments, network performance is optimized, and better throughput and delay control balance is shown in dynamic satellite network simulation experiments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116209005B_ABST
    Figure CN116209005B_ABST
Patent Text Reader

Abstract

A congestion control method for satellite network environment, which relates to the field of network transmission congestion control, solves the problems that the existing technologies do not consider the characteristics of the dynamically changing network environment in the design of action space and reward function, and there are certain limitations, etc. The control method of the present invention, aiming at the characteristics of high latency, high packet loss and dynamic periodicity in satellite network scenarios, introduces the method of reinforcement learning to carry out network congestion control, and designs an automatic variable-speed inverse proportional increase, proportional decrease action combined with heuristic ideas and a reward function based on historical information. In the dynamic satellite network simulation experiment, the present invention has achieved high throughput while maintaining a low RTT. The present invention tests the performance of different congestion control methods through satellite dynamic delay scenarios, satellite handover scenarios and packet loss scenarios respectively, and good results are obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network transmission congestion control, and in particular to a congestion control method based on reinforcement learning for a dynamic satellite network environment. Background Art

[0002] Satellite networks exhibit properties completely different from traditional terrestrial networks. Firstly, due to the transmission distance, satellite networks often have higher transmission delays. Secondly, since their transmission form is wireless transmission and the transmission path passes through the atmosphere, there is a high channel error rate, and packet loss will further affect the performance of TCP. At the same time, research in the field of network routing shows that building an inter-satellite link between satellites to form an inter-satellite network is an important development trend. In this scenario, satellites are periodically visible to each other, the transmission path is periodically switched, and the transmission distance is also periodically changed.

[0003] Existing congestion control methods cannot well meet the requirements of the satellite network environment. Early congestion control methods such as TCP Reno, Cubic, etc. use packet loss events as the sign of network transmission congestion, and Vegas uses the change in round-trip transmission delay (RTT) to detect network congestion, which are only applicable to scenarios with low packet loss rates and stability. BBR is a method that directly models the network and does not rely on packet loss and delay events, but BBR will cause the problem of a high retransmission rate and has performance problems in a large buffer network.

[0004] In recent years, with the development of machine learning and reinforcement learning, data-driven congestion control methods have been successively proposed. RemyCC and Indigo use the method of offline machine learning to explore the network to realize the mapping of the environment and actions; the Pcc series of methods calculate the utility by increasing or decreasing actions in an attempt to approximate the optimal transmission window in the network environment; Mvfst, Owl, and Aurora belong to reinforcement learning-based methods, which select link or network state measurement values as the state space and set appropriate utility functions to learn the optimal model of network congestion control. These methods have certain limitations in that they do not consider the characteristics of the dynamically changing network environment in the design of the action space and the reward function.

[0005] In summary, for this satellite network scenario with high latency, high packet loss, and dynamic periodic changes, there is still room for improvement in existing methods. Summary of the Invention

[0006] In order to solve the problems that existing technologies have certain limitations in that they do not consider the characteristics of the dynamically changing network environment in the design of the action space and the reward function, the present invention provides a congestion control method for a satellite network environment.

[0007] A congestion control method for satellite network environment, which is realized through startup state, reinforcement learning state and minimum RTT exploration state during operation; this method is implemented by the following steps:

[0008] Step 1: TCP flow starts, enters the startup state, the congestion window grows, and the window is accumulated according to the length of the currently acknowledged packet;

[0009] Step 2: After the first packet loss retransmission occurs, it enters the reinforcement learning state; the specific process is as follows:

[0010] Step 2-1: Obtain the satellite network environment state by means of interval time sampling, including throughput and delay fluctuation;

[0011] Estimate the throughput of the last sampling time interval by calculating the number of acknowledged packets in the last sampling time interval, use the minimum RTT within 10 seconds to represent the true RTT of the link, and continuously use the value obtained by subtracting the minimum RTT from the instant RTT as the delay fluctuation;

[0012] Step 2-2: Initially select to increase the congestion window, decrease the congestion window or keep the congestion window unchanged according to the sampled state values and the Q table according to the ∈-greedy principle;

[0013] Step 2-3: Calculate the congestion window adjustment action to be executed according to the set automatic variable-speed inverse proportional increase and linear decrease action space;

[0014] Step 2-4: Execute the selected action to adjust the congestion window, sample the change of the network environment state in the next sampling period, and calculate the reward value Reward using the utility function;

[0015] Calculate a value Goodvalue using the smooth throughput and the softsign function. Goodvalue represents the size of the current instant throughput relative to the past smooth throughput, and is used to evaluate the quality of the current throughput; the specific calculation is as follows:

[0016]

[0017] In the formula, Thr is the currently measured instant throughput, S_Thr is the smooth throughput, and the formula for the reward value Reward is as follows:

[0018]

[0019] Wherein, α, β, and γ are used to balance Goodvalue, Transdelay and Retransmit are used to achieve different convergence goals; Transdelay is equal to the current stream's instantaneous RTT minus the estimated link RTT, which is used to represent the network fluctuation situation; Retransmit is the retransmission rate that occurred in the satellite network during the last sampling interval;

[0020] Step 25: Adopt the ∈-greedy principle for annealing and smooth throughput update. The smooth throughput update is calculated as follows:

[0021]

[0022] Step 26: Update the q value corresponding to the previous state and action according to the reward value Reward calculated in Step 24 based on the q-learning algorithm update formula;

[0023]

[0024] Wherein, Q(s,a) is the q value to be updated after the current state s executes the currently selected action a; is the maximum q value that can be obtained by the next state s' according to the current Q table, representing the expectation of future rewards, a' is the action corresponding to the maximum q value obtained in the next state; α ∈ [0,1) represents the learning rate, indicating how much new information is learned each time; γ ∈ [0,1) is the discount factor, indicating the importance of future rewards;

[0025] Step 3: When the minimum RTT does not update for more than 10s, enter the minimum RTT exploration state, reduce the congestion window size to 4 for RTT measurement, end this state after exploring the minimum RTT, and return to Step 21.

[0026] The beneficial effects of the present invention: The control method of the present invention, aiming at the characteristics of high latency, high packet loss, and dynamic periodicity in the satellite network scenario, introduces the method of reinforcement learning to perform network congestion control, and designs an automatic variable-speed "inversely proportional increase, directly proportional decrease" action combined with heuristic thinking and a reward function based on historical information. In the dynamic satellite network simulation experiment, the present invention has achieved both high throughput and low RTT.

[0027] The present invention tests the performance of different congestion control methods through three scenarios: satellite dynamic delay scenario, satellite handover scenario, and packet loss scenario, and compares them with the kernel versions of Cubic, Vegas, Reno, BBR, Hybla, and Pcc-vivace. In the satellite dynamic delay scenario, the average throughput of the present invention is the highest among the above six methods, and relatively excellent delay control is achieved. In the satellite handover scenario, although the control of delay by the present invention is not as good as that of BBR, it is better than the remaining methods, and the best throughput performance is achieved. In the packet loss scenario, although the present invention is not as good as the BBR and Pcc methods, it is much better than other event-driven methods, and good performance is maintained within a certain range of packet loss rates.

[0028] In summary, in the satellite transmission simulation experiment, the present invention achieves a better balance between throughput and delay control, achieving higher throughput than BBR and Pcc, and lower round-trip transmission delay than Cubic, Reno, and Hybla methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a schematic block diagram of a congestion control method for a satellite network environment according to the present invention;

[0030] Figure 2 is a delay change diagram of the satellite dynamic delay scenario;

[0031] Figure 3 is a throughput diagram of six methods in the satellite dynamic delay scenario;

[0032] Figure 4 is a delay statistical diagram of six methods in the satellite dynamic delay scenario;

[0033] Figure 5 is a throughput diagram of six methods in the satellite handover scenario;

[0034] Figure 6 is a delay statistical diagram of six methods in the satellite handover scenario;

[0035] Figure 7 is a throughput statistical diagram of six methods in the packet loss scenario. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] Combined with Figures 1 to 7 to illustrate this embodiment, a congestion control method for a satellite network environment. The present invention proposes: a congestion control method based on reinforcement learning. The present invention obtains the current network state as input through interval time sampling, makes a decision to select the optimal action as output, and the method learns the optimal control through the constraint of the utility function during operation.

[0037] When the present invention is running, it mainly has three states. One is the startup state, in which the method performs a fast window accumulation according to the length of the confirmed data packet to achieve a fast startup of the flow. The second is the reinforcement learning state, which is entered after the first retransmission of the flow occurs. At this time, congestion control is completely entrusted to the reinforcement learning algorithm. The third is the minimum RTT exploration state, which is entered when the minimum RTT of the flow has not changed for 10 consecutive seconds to re-obtain the true minimum RTT of the link. The transitions between different states and the interaction relationship with the reinforcement learning algorithm are as Figure 1 .

[0038] The control logic and state transition steps of the method of the present invention are as follows:

[0039] Step 1: TCP (Transmission Control Protocol) flow starts, and the method enters the startup state. The congestion window grows rapidly, and a fast window accumulation is performed according to the length of the currently confirmed data packet.

[0040] cwnd new = cwnd old + acked

[0041] Wherein, cwnd new is the newly calculated congestion window size, cwnd old is the old congestion window size, and acked represents the length of the confirmed data packet.

[0042] Step 2: Determine whether a packet loss occurs according to the status interface provided by the kernel network module. After the first packet loss retransmission occurs, the method enters the reinforcement learning state.

[0043] When it is recognized that the first packet loss retransmission occurs or the startup state remains unchanged for 2000 milliseconds,

[0044] Step 2.1: Sample to obtain the environmental state: throughput, delay fluctuation. The method estimates the throughput of the previous time interval by calculating the number of confirmed packets in the previous sampling time interval, uses the minimum RTT within a period of time (within 10 seconds) to represent the true RTT of the link, and continuously uses the instantaneous RTT (the time interval from the time when a packet is sent to the time when it is confirmed can be calculated when each ack packet is processed) minus the minimum RTT as the delay fluctuation. The delay fluctuation is used to characterize the RTT fluctuation situation of the network where the current transmission flow is located.

[0045] Step 2.2: Input the state (input the sampled state value into the reinforcement learning model (the q-learning algorithm of the present invention), that is, use the state to look up the value in the Q table), and initially select an action of increasing the congestion window, decreasing the congestion window or keeping the congestion window unchanged according to the sampled state value and the Q table according to the ε-greedy principle;

[0046] The Q-table is a high-dimensional table used to store the goodness or badness of different actions selected under different states. The Q-table of the present invention has three dimensions, where the first and second dimensions represent throughput and latency fluctuation respectively, and the third dimension represents actions. Each item in the Q-table stores the action value q = Q(s,a) corresponding to the state and action index, and the q value is used to characterize the goodness or badness of selecting this action under the current state.

[0047] The method selects the action with the largest q value according to the input throughput and latency fluctuation in accordance with the ε-greedy principle, that is, with a probability of 1 - ε, it selects the currently selected action, and with a probability of ε, it selects a random action. This is to balance exploration and exploitation.

[0048] Step 2.3: Further calculate the window adjustment action to be executed according to the "inversely proportional increase, linearly decrease" action space of the automatic speed change type designed by the present invention.

[0049] The design of the inverse proportion action helps the flow transmission to quickly converge to fairness. When the sending window is large, as the window increases, the action amplitude of the inverse proportion increasing action will become smaller and smaller, thus ensuring that the transmission will not exhibit the behavior of too large fluctuations in the sending window. On the other hand, the speed change mechanism enables the algorithm to quickly respond to network changes when the window needs to be quickly increased. When the method converges near the optimal solution, the method will oscillate between the increasing action and the decreasing action, and at the same time, the action adjustment amplitude each time will not be too large to ensure control stability.

[0050] The automatic speed change type "inversely proportional increase, linearly decrease" action space of the present invention is as follows:

[0051]

[0052] For increasing the window, the present invention designs a flag n, where the value range of n is 0 - 7, indicating a total of 8 different speed adjustment amplitudes. During the operation, the consecutive selection times up_times (0 - 15) of the increasing action are recorded. If the increasing action is selected continuously twice, then starting from the third time, n will be increased by 1 each time.

[0053] For the decreasing window action, the present invention designs a linearly decreasing window with gradients down_list = {1, 3, 5, 9, 15, 21, 33, 51}. During the operation of the method, the consecutive selection times down_times (0 - 8) of the decreasing action are recorded. The current value of the decreasing action is the value corresponding to the subscript of the current down_times in the array. If down_times is greater than 7, the decreasing action becomes halving the window.

[0054] During the running process, if the method currently selects the increase action, the down_times of the decrease window will be set to 0. If the method currently selects the decrease action, the up_times of the increase window will be set to 0. If the method selects the keep unchanged action, n will be decreased by 1.

[0055] Step 2.4: Execute the selected action to adjust the congestion window, sample the network changes in the next sampling period, calculate the reward value using the utility function, and evaluate the quality of the action selection.

[0056] The present invention proposes a dynamic reward function based on historical information, which achieves better convergence effect in a dynamically changing environment. For a dynamic satellite network, the transmission link will undergo periodic switching. Therefore, the absolute values of throughput on different transmission links may be of different orders of magnitude. Directly using throughput as the reward may lead to convergence problems for the method. The present invention uses the smoothed throughput and the softsign function to calculate a value Goodvalue to evaluate the quality of the current throughput. Goodvalue represents the magnitude relationship between the current instantaneous throughput and the past smoothed throughput. If the two values are comparable, the Goodvalue is about 50. If the instantaneous throughput is greater than the smoothed throughput, the value of Goodvalue increases as the difference amplitude increases, and vice versa. The specific calculation is as follows:

[0057]

[0058] Where Thr is the currently measured instantaneous throughput, and S_Thr is the smoothed throughput. The value range of Goodvalue is between 0 and 100.

[0059] The reward function is designed as follows:

[0060]

[0061] Where α, β, γ are used to balance Goodvalue, Transdelay, and Retransmit to achieve different convergence goals. Transdelay is equal to the RTT of the current flow minus the estimated link RTT, which is used to represent the network fluctuation situation. Retransmit is the number of retransmissions that occurred in the network during the last sampling interval.

[0062] Step 2.5: ∈ annealing and smoothed throughput update.

[0063] The present invention designs an ∈ annealing mechanism based on network change point detection to improve the learning ability and performance of the method. If the ∈ value is too large, it is not conducive to method exploration; if it is too small, it affects the convergence speed of training. Therefore, a relatively large ∈ = 0.1 is set at the beginning of training and gradually annealed and decreased. When a network change is detected, it is reset to the initial ∈. There are mainly two change point detection mechanisms: throughput detection and RTT change detection. During operation, the method records the past five throughput values and calculates their average value. If the difference between the average value and the smoothed throughput is greater than 1 / 8 of the smoothed throughput, the reset mechanism is triggered. RTT detection performs change point detection by detecting the difference between the current latency and the estimated RTT of the link. If the difference between the two is greater than 1 / 8 of the minimum RTT, the reset mechanism is triggered.

[0064] The calculation of the smoothed throughput update is as follows:

[0065]

[0066] Step 2.6: Update the q value corresponding to the previous state and action according to the reward calculated in Step 2.4 based on the q-learning update formula.

[0067]

[0068] In the formula, Q(s,a) represents the q value to be updated after executing the currently selected action a in the current state s; represents the maximum q value that can be obtained for the next state s' according to the current Q table, characterizing the expectation of future rewards. a' is the action corresponding to the maximum q value obtained for the next state; α ∈ [0,1) represents the learning rate, indicating how much new information is learned each time; γ ∈ [0,1) is the discount factor, indicating the importance of future rewards; Reward is the reward value calculated in 2.4.

[0069] During operation, if the minimum RTT does not update for a long time (10s), the method enters the minimum RTT exploration state, actively reduces the congestion window to 4 for RTT measurement, and ends this state after exploring the minimum RTT, and resumes to Step 2.

[0070] As Figures 1 to 7 shown, in this embodiment, three scenarios are designed to test the performance of different congestion control methods and compare them with the kernel versions of Cubic, Vegas, Reno, BBR, Hybla, and Pcc-vivace.

[0071] 1. Satellite dynamic delay scenario;

[0072] This scenario simulates a transmission stream with periodically changing time delays. The low Earth orbit satellite constellation is simulated using satellite motion simulation software, and a 300 ms satellite transmission link is recorded. Then, the network emulator is used to simulate this link. The variation of the RTT is as shown in Figure 2 .

[0073] The throughput and delay are as shown in Figure 3 and 4 . By comprehensively comparing six methods, the average throughput of the present invention is the highest among the six methods, and relatively excellent delay control is achieved.

[0074] 2. Satellite handover scenario;

[0075] This scenario simulates the throughput variation of two transmission paths. The RTT of both links is set to 40 ms. The throughput of one link is 20 Mbps, and the throughput of the other link is 10 Mbps. The duration of each stream is 40 s. To test the handover performance of the method, link handover is simulated every 5 s, and the bandwidth of the bottleneck link is changed.

[0076] The throughput and delay are as shown in Figure 5 and 6 . By comprehensively comparing the performance of six methods, although the control of the delay by the present invention is not as good as that of BBR, it is better than the remaining other methods, and the best throughput performance is achieved.

[0077] 3. Packet loss scenario;

[0078] This scenario simulates network scenarios with different packet loss rates. Six methods are tested at packet loss rates of 0, 0.25, 0.5, 1, 1.5, 2, 3, 4, and 5 respectively. The throughput of the bottleneck link is set to 50 Mbps, and the delay is 40 ms.

[0079] The throughput statistics of the experimental results are as shown in Figure 7 . By comparing six methods, although the present invention is not as good as BBR and Pcc methods, it is much better than other event-driven methods and maintains good performance within a certain range of packet loss rates.

[0080] In summary, in the satellite transmission simulation experiment, the present invention achieves a better balance between throughput and delay control, achieving higher throughput than BBR and Pcc and lower round-trip transmission delay than Cubic, Reno, and Hybla methods.

[0081] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0082] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A congestion control method for satellite network environment, characterized in that: This method is implemented through a startup state, a reinforcement learning state, and a minimum RTT exploration state during operation; this method is implemented by the following steps: Step 1: The TCP flow starts and enters the startup state. The congestion window grows, and the window is accumulated according to the length of the currently acknowledged packet; Step 2: After the first packet loss retransmission occurs, it enters the reinforcement learning state; The specific process is as follows: Step 2-1: Obtain the satellite network environment state by means of interval time sampling, including throughput and delay fluctuation; Estimate the throughput of the last sampling time interval by calculating the number of acknowledged packets in the last sampling time interval. Use the minimum RTT within 10 seconds to represent the true RTT of the link, and continuously use the value obtained by subtracting the minimum RTT from the instantaneous RTT as the delay fluctuation; Step 2-2: Initially select to increase the congestion window, decrease the congestion window, or keep the congestion window unchanged according to the sampled state value and the Q table according to the ε-greedy principle; Step 2-3: Calculate the congestion window adjustment action to be executed according to the set automatic variable-speed inverse proportional increase and linear decrease action space; Step 2-4: Execute the selected action to adjust the congestion window, sample the change of the network environment state in the next sampling period, and calculate the reward value Reward using the utility function; Calculate a value Goodvalue using the smoothed throughput and the softsign function. Goodvalue represents the situation of the current instantaneous throughput relative to the past smoothed throughput, and is used to evaluate the quality of the current throughput; the specific calculation is as follows: In the formula, Thr is the currently measured instantaneous throughput, S_Thr is the smoothed throughput, and the formula for the reward value Reward is as follows: In the formula, A, B, and C are used to balance Goodvalue, Transdelay and Retransmit are used to achieve different convergence goals; Transdelay is equal to the instantaneous RTT of the current flow minus the estimated link RTT, and is used to represent the network fluctuation situation; Retransmit is the retransmission rate that occurred in the satellite network in the last sampling time interval; Step 2-5: Anneal according to the ε-greedy principle and update the smoothed throughput. The calculation of the smoothed throughput update is as follows: Step 2-6: Update the q value corresponding to the previous state and action according to the q-learning algorithm update formula based on the reward value Reward calculated in Step 2-4; Where Q(s, a) is the q-value to be updated after performing the currently selected action a in the current state s; is the maximum q-value that can be obtained for the next state s' according to the current Q-table, representing the expectation of future rewards. a' is the action corresponding to the maximum q-value in the next state. α ∈ [0, 1) represents the learning rate, indicating how much new information is learned each time. γ ∈ [0, 1) is the discount factor, indicating the importance of future rewards; Step 3: When the minimum RTT does not update for more than 10s, enter the minimum RTT exploration state, reduce the congestion window size to 4 for RTT measurement, end this state after exploring the minimum RTT, and return to Step 2-1.

2. The congestion control method for a satellite network environment according to claim 1, characterized in that: In Step 1, the formula for the growth of the congestion window is as follows: cwnd new = cwnd old + acked where cwnd new is the newly calculated congestion window size, cwnd old is the old congestion window size, and acked is the length of the acknowledged data packet.

3. A congestion control method for a satellite network environment according to claim 1, characterized in that: In Step 2, judge whether packet loss occurs according to the state interface provided by the kernel network module. When it is recognized that the first packet loss retransmission occurs or the startup state remains unchanged for 2000 milliseconds, enter the reinforcement learning state.

4. A congestion control method for a satellite network environment according to claim 1, characterized in that: In Step 2-2, the Q table is a high-dimensional table used to store the quality of choosing different actions in different states; It is assumed that the Q-table has three dimensions, where the first and second dimensions represent throughput and latency fluctuation respectively, and the third dimension represents actions; each item in the Q-table stores the action value q = Q(s, a) corresponding to the state and action index, and the q value is used to characterize the goodness or badness of selecting this action in the current state; According to the input throughput and latency fluctuation, select the action with the largest q value according to the ε-greedy principle, and select a random action with a probability of ε.

5. The congestion control method for a satellite network environment according to claim 1, characterized in that: In Step 2 and Step 3, the automatic variable-speed inverse proportional increase and linear decrease action space is as follows: where cwnd new is the newly calculated congestion window size, cwnd old is the old congestion window size. For the window increase action, set the flag value n, where the value range of n is 0 - 7, indicating a total of 8 different speed adjustment amplitudes. During the operation, record the consecutive selection times up_times of the increase action. If the increase action is selected continuously twice, then starting from the third time, n will be increased by 1 each time; For the window reduction action, set the linearly decreasing window with gradients down_list = {1, 3, 5, 9, 15, 21, 33, 51}, record the continuous selection times down_times of the reduction action, and the currently reduced action value is the value of the subscript corresponding to the current down_times in the array. If down_times is greater than 7, the reduction action becomes halving the window; If the increase action is selected currently, set down_times of the window reduction to 0. If the reduction action is selected, set up_times of the window increase to 0. If the keep unchanged action is selected, decrease n by 1.

Citation Information

Patent Citations

  • Congestion control method and system based on deep reinforcement learning

    CN110581808A

  • Congestion control in network communications

    CN111919423A