A sub-flow coupling aware multipath congestion control method and computer readable medium
Patent Information
- Application Number
- CN202311843110.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-12-27
AI Technical Summary
当存在耦合子流与单路径连接竞争时,耦合子流凭借多条子流的优势容易获得更大的拥塞窗口加速度,最终导致多路径连接所获得的带宽大于单路径连接,难以保证网络传输的公平性
本发明提出基于深度强化学习和子流耦合感知的多路径拥塞算法,支持针对不同耦合特征的子流进行不同程度的拥塞控制。同时可以从与环境的交互中学习经验,以改进拥塞控制策略,增强了对变化网络环境的适应能力。采用速率耦合的方式集成了CoupledBBR,即吸收了其对于丢包不敏感,适用于无线网络环境的优点,又改善了其环境探测方式固定的不足之处,相较于原算法和其他基于深度强化学习的多路径拥塞控制算法可以达到更高的吞吐量,同时实现更好的传输公平性。
Smart Images

Figure CN117955910B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer networks, and particularly relates to a sub-flow coupling-aware multipath congestion control method and a computer-readable medium. Background Technology
[0002] With the rapid development of wireless access technology, heterogeneous wireless networks, composed of various wireless networks such as Wi-Fi, mobile communication networks (4G / 5G), and satellite communication networks, have been widely used in autonomous driving, virtual reality, and video conferencing scenarios. Terminal nodes in heterogeneous wireless networks can use multiple network interfaces simultaneously to establish multiple sub-streams for parallel multi-path transmission, thereby improving network communication throughput.
[0003] Multipath congestion control algorithms have a significant impact on network performance. In multipath transmission, multiple sub-flows often traverse the same path, resulting in coupled sub-flows. Existing multipath congestion control algorithms do not consider the actual coupling of sub-flows; some algorithms assume all sub-flows traverse different links and apply independent congestion control to each sub-flow. When coupled sub-flows compete with single-path connections, the coupled sub-flows, due to their multiple sub-flows, tend to gain a larger congestion window acceleration, ultimately leading to a higher bandwidth for multipath connections compared to single-path connections, making it difficult to guarantee network transmission fairness. Some algorithms assume all sub-flows traverse the same link, combining all sub-flows into a single coupled sub-flow for unified congestion control, limiting the sum of the congestion window growth rates of all sub-flows in a multipath connection to no more than the growth rate of the congestion window for a single-path connection. This results in underutilization of network link resources, consequently reducing multipath transmission throughput. Furthermore, traditional multipath congestion control algorithms employ relatively fixed congestion control strategies, which are unsuitable for complex heterogeneous wireless network environments with frequent node movement and dynamic changes in network quality.
[0004] Therefore, when performing congestion control in complex heterogeneous wireless network environments, it is necessary not only to continuously optimize the congestion control strategy to adapt to the dynamic network environment, but also to implement different congestion control strategies for sub-flows in different states. Thus, it is essential to study multi-path congestion control algorithms based on deep reinforcement learning and sub-flow coupling awareness. Summary of the Invention
[0005] This invention aims to propose a sub-stream coupling-aware multipath congestion control method and a computer-readable medium. It achieves sub-stream coupling state awareness by monitoring the round-trip time (RTT) trends of different sub-streams within each monitoring period, and further improves feature awareness accuracy by combining a Coupled BBR state rotation mechanism. The method concatenates various coupling features and link state features as sub-stream features to explore the transmission potential of sub-streams with different coupling features, guiding the method to fully utilize link resources to improve network throughput. An LSTM network is used to eliminate network noise, and temporal hidden information is extracted from historical sub-stream states. The output results serve as the state input for a deep reinforcement learning agent. The agent updates its policy network parameters using the Proximal Policy Optimization (PPO) algorithm and obtains action decisions based on the state input. Regarding the decision execution problem, since the action form is a sequence of transmission gain coefficients for each sub-stream, it can be directly used to adjust the sub-stream transmission gain coefficients of the Coupled BBR, thereby changing the traffic injected into each sub-stream link to achieve multipath transmission congestion control.
[0006] This algorithm consists of three stages: state extraction, decision making, and action execution. The state extraction stage includes two parts: coupling feature perception and link state extraction. Coupling feature perception involves acquiring RTT data from the environment and calculating trends to perceive the coupling features of sub-flows. The obtained coupling features are then concatenated with other link features fed back from the sub-flows to generate the state information for each sub-flow. In the link state extraction part, LSTM sequences are used to eliminate network noise, and temporal hidden information is extracted from historical state sequences to generate new state information. The decision making stage uses PPO as the agent's deep reinforcement learning algorithm. The agent can adjust the transmission gain factors of different sub-flows based on the state information of its policy network, generating an adjusted gain factor sequence, called the action decision. In this stage, the agent collects experience through interaction with the environment, learns congestion control strategies for different sub-flows under different states, and thus updates its policy network. In the action execution stage, the Coupled BBR algorithm is used as the execution unit. Based on the received action decisions, the transmission rates of different sub-flows are adjusted. By adjusting the data transmission rate in a timely manner to adapt to changes in link states, transmission congestion control is achieved.
[0007] This invention provides a sub-stream coupling-aware multipath congestion control method and a computer-readable medium, relating to heterogeneous wireless network environments, coupling feature awareness, and PPO agents, characterized by: Step 1: The PPO agent obtains the RTT information and link status information of each sub-stream from the heterogeneous wireless network environment and enters the state extraction stage. The PPO agent calculates the RTT change trend of each sub-stream and obtains real-time link quality characteristics from the heterogeneous wireless network environment. Step 2: Based on the calculated RTT change trends of each sub-stream, the coupling characteristics of the sub-stream are sensed. The sensed coupling characteristics are then concatenated with the link quality characteristics to obtain the sub-stream status. Step 3: Input the latest sub-stream state and the sub-stream states from the previous K monitoring periods into the LSTM network, select the mean squared error to calculate the loss function, use the gradient descent algorithm to update the network parameters, and use Adam as the optimizer. Step 4: The sequence of states formed by each sub-stream is used as the state input of the agent and enters the decision generation stage. The agent's policy network calculates action decisions based on this state information. Step 5: Action decisions serve as input to the Coupled BBR, controlling the transmission rate of each sub-stream in the next monitoring cycle, thereby achieving congestion control in a heterogeneous wireless network environment. Step 6: After performing an action, the agent receives a reward and obtains the next state from the environment. This information, combined with the current state, forms an experience tuple, which is stored in the experience buffer pool. Step 7: Once the experience buffer is full, the PPO agent begins training using historical experience tuples to update its policy network and value network parameters. As a preferred method, step 1 is implemented as follows: Step 1.1: By confirming the timestamp of the data packet, according to the formula... Calculate the round-trip transmission delay of each substream. ; in, It is the timestamp of the confirmation packet being sent. It is the timestamp returned by the confirmation packet. Indicates the sub-stream sequence number. Indicates the sequence number of the feedback data packet; Step 1.2: Record the average arrival time of feedback data packets within the monitoring period. and calculate respectively and ; Then, the trend of substream RTT was calculated according to the formula. , The number of sub-streams; in, The average arrival time of feedback data packets is used to represent the average arrival time of all data packets collected during the monitoring period. The average is obtained as follows: Used to represent the average RTT, by analyzing all data collected during the monitoring period. We get the average. Step 1.3: Based on real-time statistical information during transmission, obtain the link status information of each sub-stream, including throughput. ,rate Packet loss rate Congestion window size ; The real-time link quality characteristics mentioned in step 1 include: bandwidth, transmission rate, and packet loss rate; As a preferred method, step 2 is implemented as follows: Step 2.1: Compare the RTT change trend of each sub-stream with kGradientMin. When the value is greater than kGradientMin, start sub-stream coupling sensing. kGradientMin is the threshold used to determine whether to enable sub-stream coupling sensing.
[0008] Step 2.2 For each sub-flow, calculate the difference between it and other sub-flows. , believes that sub-stream Heziliu There is a possibility of coupling; This is the coupling sensing interval, where sub-flows within this interval may be coupled. Step 2.3: When there is a possibility of coupling between sub-streams, the state machine state of the Coupled BBR on each sub-stream will be used to assist in the judgment; If the state machines of both sub-streams are in the Probe RTT stage, then a coupling relationship exists between the two sub-streams. This is denoted as... ,in This indicates whether the sub-stream is in a coupled state. This represents the number of sub-streams in the coupled sub-stream.
[0009] Step 2.4: Concatenate the sensed coupling relationship with the real-time sub-stream link to obtain a representation of the sub-stream state. ; in, Represents the current monitoring cycle number. Represents the sub-stream sequence number; Furthermore, step 3 can be implemented as follows: The LSTM network described in step 3 will be used to eliminate network noise, extract hidden information from historical state sequences, and finally form new state information. Step 3.1: Concatenate the latest state obtained from the LSTM network with the sub-stream states from the previous K monitoring periods to obtain... .in, Represents the current monitoring cycle number. Represents the sub-stream sequence number.
[0010] Step 3.2: Input this state sequence into the LSTM network to calculate a new state that has eliminated network noise and contains temporal hidden information, called the new state. .in, Represents the current monitoring cycle number. Represents the sub-stream sequence number; Step 3.3: Combine the sub-streams Combine to form a state vector , used to characterize the The status of all sub-streams in each monitoring cycle Furthermore, step 4 can be implemented as follows: Step 4.1: State vectors as Actor policy networks With the input, the policy network begins to calculate and obtain the action policy; Step 4.2: Generate an action distribution based on the probability distribution of the action policy, and randomly sample actions from it, in the following form: ; in, Represents the current monitoring cycle number. Represents the sub-stream sequence number; Furthermore, step 5 can be implemented as follows: Step 5.1: The Coupled BBR receives the action generated in Step 4 as input; Step 5.2: The Coupled BBR maintains a transmission rate gain coefficient for each substream in each monitoring cycle. It adjusts its own transmission gain coefficient based on the transmission gain value of each substream in the received action. Step 5.3: Coupled BBR can control the traffic injected into the network through this rate coupling, thereby achieving congestion control of transmission; Furthermore, step 6 can be implemented as follows: Step 6.1: After the action is executed by the Coupled BBR, it will affect the environment, causing the environment to change. After going through steps 1-4 again, record the network environment state at this time. Feedback is given to the intelligent agent; Step 6.2: Simultaneously calculate the reward obtained by the agent after performing the action according to the following formula:
[0011]
[0012] in, Represents the current monitoring cycle number. Represents the sub-stream sequence number. Represents substream throughput. Represents the substream transmission delay. This represents the packet loss rate of the substream. The final reward is calculated as follows: This feedback is given to the intelligent agent. Among them... It is a reward calculation for uncoupled subflows. It is a reward calculation for coupled sub-flows; Step 6.3: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require , , , The experience tuples are combined and stored in the agent's experience buffer pool to train the agent's policy network and value network parameters. Furthermore, step 7 can be implemented as follows: Step 7.1: For each empirical tuple in the empirical buffer pool Using value network computing and value and .in, Represents the current monitoring cycle number; Step 7.2: For each empirical tuple in the empirical buffer pool ,according to Calculate timing difference error , This is the discount factor, with a value of 0.9. and For the policy network, for the state and Calculate state value, Represents the current monitoring cycle number; Step 7.3: For each empirical tuple in the empirical buffer pool Generalization advantage estimation (GAE) is employed and based on Calculate the advantage value of the previous action and current action advantage value ,in, Represents the current monitoring cycle number; Step 7.4: Set the current The parameters are copied to the old policy network with the same network structure. In the middle, make ; Step 7.5: Randomly sample M experience tuples from the experience buffer pool and input them into the policy network. and In the middle, using the formula:
[0013] Calculate the loss function of the policy network and update the network parameters using gradient descent; in, and Representing the intelligent agent in Actions taken and network status observed during the monitoring period. The function is a clipping function, where This is a correction factor used to determine the upper and lower bounds of the clipping. Step 7.6: After updating the policy network, use MSEloss as the loss function to update the parameters of the value network:
[0014] in, In order to be in The state value estimated by the Critic value network during the monitoring period. In order to be in The objectives of updating the Critic value network during the monitoring period are:
[0015] in, Represents the advantage value of the action. As a discount factor, The number of steps representing future rewards; Step 7.7: When the training iteration count is reached, clear the experience buffer pool and end the training phase; The present invention also provides a computer-readable medium storing a computer program executed by an electronic device, wherein when the computer program is run on the electronic device, it executes the steps of the sub-stream coupling-aware multipath congestion control method.
[0016] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: This invention proposes a multipath congestion algorithm based on deep reinforcement learning and sub-stream coupling awareness, supporting different levels of congestion control for sub-streams with varying coupling characteristics. It can also learn from interactions with the environment to improve congestion control strategies, enhancing adaptability to changing network environments. By integrating CoupledBBR in a rate-coupled manner, it absorbs its advantages of being insensitive to packet loss and suitable for wireless network environments, while improving upon its fixed environment detection method. Compared to the original algorithm and other deep reinforcement learning-based multipath congestion control algorithms, it achieves higher throughput and better transmission fairness. Attached Figure Description
[0017] Figure 1 : A schematic diagram of the method flow of an embodiment of the present invention; Figure 2 : A schematic diagram of the average throughput and latency test scenario of this invention embodiment; Figure 3 : A schematic diagram of a transmission fairness test scenario according to an embodiment of the present invention; Figure 4 A comparison chart of the average throughput of different multipath congestion control algorithms on different paths according to embodiments of the present invention; Figure 5 Comparison of average RTT on different paths for different multipath congestion control algorithms in this invention embodiment; Figure 6 A comparison chart of transmission fairness between different multipath congestion control algorithms and the BBR algorithm in embodiments of the present invention; Detailed Implementation The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] In specific implementation, the method proposed in the technical solution of this invention can be automatically executed by those skilled in the art using computer software technology. System devices for implementing the method, such as computer-readable storage media storing the corresponding computer program of the technical solution of this invention and computer equipment including the computer program running the corresponding computer program, should also be within the protection scope of this invention.
[0019] To verify the effectiveness of the proposed algorithm, the NS3 network simulation platform was used to simulate a complex heterogeneous network environment, testing its learning efficiency and convergence. Various scenarios were designed and built to compare the performance of LIA, Wvegas, Coupled BBR, and DRL-CC in terms of throughput and average transmission latency. The simulation environment was set up as follows: Figure 2 As shown, the transmitting end S0 and the receiving end C0 establish three sub-streams for data transmission through different wireless network interfaces. The sub-stream connected to the Satellite network interface passes through the non-shared bottleneck link Path1, while the sub-streams connected to the WiFi and Cellular network interfaces pass through the same link during transmission, sharing the bottleneck link Path2. The parameter settings for Path1 and Path2 are shown in Table 1.
[0020] Table 1: Network Parameter Settings Table
[0021] like Figure 1As shown, this embodiment of the invention provides a multipath transmission algorithm based on deep reinforcement learning and substream coupling recognition, including the following steps: Step 1: The PPO agent obtains the RTT information and link status information of each sub-stream from the heterogeneous wireless network environment and enters the state extraction stage. The PPO agent calculates the RTT change trend of each sub-stream and obtains real-time link quality characteristics from the heterogeneous wireless network environment. The real-time link quality characteristics include: bandwidth, transmission rate, and packet loss rate; Furthermore, the implementation method of step 1 is as follows: Step 1.1: By confirming the timestamp of the data packet, according to the formula... Calculate the round-trip transmission delay of each substream. ; in, It is the timestamp of the confirmation packet being sent. It is the timestamp returned by the confirmation packet. Indicates the sub-stream sequence number. Indicates the sequence number of the feedback data packet; Step 1.2: Record the average arrival time of feedback data packets within the monitoring period. and calculate respectively and ; Then, the trend of substream RTT was calculated according to the formula. , The number of sub-streams; in, The average arrival time of feedback data packets is used to represent the average arrival time of all data packets collected during the monitoring period. The average is obtained as follows: Used to represent the average RTT, by analyzing all data collected during the monitoring period. We get the average. Step 1.3: Based on real-time statistical information during transmission, obtain the link status information of each sub-stream, including throughput. ,rate Packet loss rate Congestion window size ; Step 2: Based on the calculated RTT change trends of each sub-stream, the coupling characteristics of the sub-stream are sensed. The sensed coupling characteristics are then concatenated with the link quality characteristics to obtain the sub-stream status.
[0022] Furthermore, step 2 can be implemented as follows: Step 2.1, convert each sub-flow Compared with kGradientMin=0.7, when its value is greater than kGradientMin, sub-stream coupling sensing is started. kGradientMin is the threshold used to determine whether to enable sub-stream coupling sensing.
[0023] Step 2.2, for each sub-flow, calculate the difference between it and other sub-flows. , believes that sub-stream Heziliu There is a possibility of coupling; This is the coupling sensing interval, where sub-flows within this interval may be coupled. Step 2.3: When coupling between sub-streams is possible, the state machine state of the Coupled BBR on each sub-stream will be used to assist in the judgment. If the state machines on both sub-streams are in the Probe RTT stage, then a coupling relationship is considered to exist between these two sub-streams. This is denoted as... ,in This indicates whether the sub-stream is in a coupled state. This represents the number of sub-streams in the coupled sub-stream.
[0024] Step 2.4: Concatenate the sensed coupling relationship with the real-time link of the sub-stream to obtain a representation of the sub-stream state. ; in, Represents the current monitoring cycle number. Represents the sub-stream sequence number; Step 3: Compare the latest sub-stream state with the past... =The sub-stream states within 32 time periods are input into the LSTM sequence, the mean squared error is used to calculate the loss function, the gradient descent algorithm is used to update the network parameters, and Adam is used as the optimizer; The LSTM network will be used to eliminate network noise and extract hidden information from historical state sequences, ultimately forming new state information. Furthermore, step 3 can be implemented as follows: Step 3.1, compare the latest state obtained by LSTM with the previous state. The sub-stream states within each monitoring cycle are spliced together to obtain... .in, Represents the current monitoring cycle number. Represents the sub-stream sequence number.
[0025] Step 3.2: Input this state sequence into the LSTM network to calculate a new state that has eliminated network noise and contains temporal hidden information, referred to as... .in, Represents the current monitoring cycle number. Represents the sub-stream sequence number; Step 3.3, convert each sub-stream's... Combine to form a state vector , used to characterize the The status of all sub-streams in each monitoring cycle Step 4: The sequence of states formed by each sub-stream is used as the state input of the agent and enters the decision generation stage. The agent's policy network calculates action decisions based on this state information.
[0026] Furthermore, step 4 can be implemented as follows: Step 4.1: State vectors as Actor policy networks With the input, the policy network begins to calculate and obtain the action policy; Step 4.2: Generate an action distribution based on the probability distribution of the action policy, and randomly sample actions from it, in the following form: ; in, Represents the current monitoring cycle number. Represents the sub-stream sequence number; Step 5: Action decision is used as input to Coupled BBR to control the transmission rate of each sub-stream in the next detection cycle, thereby realizing congestion control in a heterogeneous wireless network environment. Furthermore, step 5 can be implemented as follows: Step 5.1: The Coupled BBR receives the action generated in Step 4 as input; Step 5.2: The Coupled BBR maintains a transmission rate gain coefficient for each substream in each cycle. It adjusts its own transmission gain coefficient based on the transmission gain value of each substream in the receiving operation. Step 5.3: Coupled BBR can control the traffic injected into the network through this rate coupling, thereby achieving congestion control of transmission; Step 6: After performing an action, the agent will receive a reward and obtain the next state from the environment. This information, along with the current state, is combined to form an experience tuple, which is stored in the experience replay buffer. Furthermore, step 6 can be implemented as follows: Step 6.1: After the action is executed by the Coupled BBR, it will cause changes in the environment. After going through steps 1-4 again, record the network environment state at this time. Feedback is given to the intelligent agent; Step 6.2, simultaneously calculate the reward obtained by the agent after performing the action according to the following formula:
[0027]
[0028] in, Represents the current monitoring cycle number. Represents the sub-stream sequence number. Represents substream throughput. Represents the substream transmission delay. This represents the packet loss rate of the substream. The final reward is calculated as follows: This feedback is given to the intelligent agent. Among them... It is a reward calculation for uncoupled subflows. It is a reward calculation for coupled sub-flows; Step 6.3: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require , , , The experience replay tuples are combined and stored in the agent's experience replay buffer pool for training the agent's policy network and value network parameters. Step 7: Once the experience buffer is full, the PPO agent begins training using historical experience tuples to update its own policy network and value network parameters. Furthermore, step 7 can be implemented as follows: Step 7.1, for each empirical tuple in the empirical buffer pool Using value network computing and value and .in, Represents the current monitoring cycle number; Step 7.2, for each empirical tuple in the empirical buffer pool ,according to Calculate timing difference error , This is the discount factor, with a value of 0.9. and For the policy network, for the state and Calculate state value, Represents the current monitoring cycle number; Step 7.3, for each empirical tuple in the empirical buffer pool Generalization advantage estimation (GAE) is employed and based on Calculate the advantage value of the previous action and current action advantage value ,in, Represents the current monitoring cycle number; Step 7.4, set the current The parameters are copied to the old policy network with the same network structure. In the middle, make ; Step 7.5: Randomly sample M experience tuples from the experience buffer pool and input them into the policy network. and In the middle, using the formula:
[0029] Calculate the loss function of the policy network and update the network parameters using gradient descent; in, and Representing the intelligent agent in Actions taken and network status observed during the monitoring period. The function is a clipping function, where This is a correction factor used to determine the upper and lower bounds of the clipping. Step 7.6: After updating the policy network, use MSEloss as the loss function to update the parameters of the value network:
[0030] in, In order to be in The state value estimated by the Critic value network during the monitoring period. In order to be in The objectives of updating the Critic value network during the monitoring period are:
[0031] in, Represents the advantage value of the action. As a discount factor, The number of steps that represent future rewards.
[0032] Step 7.7: When the number of training iterations K is reached, the experience replay buffer is cleared and the training phase ends. Depend on Figure 2As can be seen, in the non-shared bottleneck path (Path 1), compared with coupled multipath congestion control algorithms such as LIA, Wvegas, and CoupledBBR, this algorithm achieves a significant increase in throughput with a small increase in RTT. Compared with the deep reinforcement learning-based DRL-CC algorithm, it achieves higher throughput while maintaining lower latency. In the shared bottleneck path (Path 2), the MSCP algorithm significantly improves throughput compared to LIA, Wvegas, CoupledBBR, and DRL-CC algorithms, while maintaining lower latency than the DRL-CC algorithm. In summary, this algorithm can achieve high throughput while maintaining low latency in both non-shared and shared bottleneck links.
[0033] Setting up the simulation environment, such as Figure 3 To test the transmission fairness of the algorithm, a multipath connection with two sub-streams is established between the transmitter S0 and the receiver C0 via a WiFi interface and a Cellular interface, and the two sub-streams pass through a shared bottleneck link. At the same time, a single-path connection is established between the transmitter S1 and the receiver C1, and congestion control is performed using the BBR algorithm. This single-path connection serves as background traffic, competing for bandwidth resources with the multipath connection within the shared bottleneck.
[0034] Figure 4 The average throughput of different algorithms on different paths is compared.
[0035] Figure 5 The average RTT of different algorithms on different paths is shown.
[0036] Figure 6 This paper compares the average throughput of different multipath congestion control algorithms and the BBR congestion control algorithm when competing on the same bottleneck link. Because LIA, Wvegas, and Coupled BBR cannot effectively utilize link resources in complex and heterogeneous wireless network scenarios, they are at a disadvantage when competing with single-path BBR flows, resulting in lower throughput. DRL-CC, on the other hand, focuses on improving the overall network throughput and does not consider competition with single-path connections; therefore, its throughput is significantly higher than that of single-path BBR flows. The MSCP algorithm first identifies the coupling characteristics between sub-flows, generates a reward function based on these characteristics, and incorporates fairness metrics. Therefore, the MSCP algorithm achieves a throughput more similar to the BBR algorithm on the shared bottleneck link, while also demonstrating higher fairness.
[0037] To better measure the fairness of competition between multi-path congestion control algorithms and single-path congestion control algorithms, this chapter introduces the Jain fairness index for evaluation. The Jain fairness index is widely used as a measure of fairness in resource allocation, and its definition is as follows:
[0038] in, Jain represents the throughput of a flow. The closer the Jain index is to 1, the fairer the bandwidth allocation. Table 2 shows the Jain fairness index of LIA, Wvegas, Coupled BBR, DRL-CC, and MSCP algorithms when competing for bottleneck bandwidth with single-path connected BBR flows. As can be seen from the table, the Jain fairness index of this algorithm is closer to 1, indicating higher fairness.
[0039] Table 2 shows the comparison of total throughput of different algorithms in the test scenario. The throughput of this algorithm is improved by approximately 477%, 220%, 125%, and 20% compared to LIA, Wvegas, Coupled BBR, and DRL-CC, respectively. This is mainly due to the algorithm's adaptive perception of sub-stream coupling state characteristics and link congestion state, enabling it to execute different congestion control strategies for sub-streams under different coupling states, thereby improving data transmission throughput.
[0040] Table 2: Jain Fairness Index Comparison Table
[0041] The computer-readable medium is a server workstation; The server workstation stores a computer program executed by the electronic device. When the computer program runs on the electronic device, it causes the electronic device to execute the steps of the sub-stream coupling-aware multipath congestion control method of the present invention.
[0042] It should be understood that any parts not described in detail in this specification belong to the prior art.
[0043] It should be understood that the above description of the embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art can make substitutions or modifications under the guidance of this invention without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.
Claims
1. A sub-stream coupled sensing multipath congestion control method, characterized in that: Step 1: The PPO agent obtains the RTT information and link status information of each sub-stream from the heterogeneous wireless network environment and enters the state extraction stage. The PPO agent calculates the RTT change trend of each sub-stream and obtains real-time link quality characteristics from the heterogeneous wireless network environment. Step 2: Based on the calculated RTT change trend of each sub-stream, perceive the coupling characteristics of the sub-stream; The perceived coupling features are then concatenated with the link quality features to obtain the sub-flow state; Step 3: Input the latest sub-stream state and the sub-stream states from the previous K monitoring periods into the LSTM network, select the mean squared error to calculate the loss function, use the gradient descent algorithm to update the network parameters, and use Adam as the optimizer. Step 4: The sequence of states formed by each sub-stream is used as the state input of the agent and enters the decision generation stage. The agent's policy network calculates action decisions based on this state information. Step 5: Action decisions serve as input to the Coupled BBR, controlling the transmission rate of each sub-stream in the next monitoring cycle, thereby achieving congestion control in a heterogeneous wireless network environment. Step 6: The agent will receive a reward after performing the action and will obtain the next state from the environment; This information, combined with the current state, forms an experience tuple, which is stored in the experience buffer pool; Step 7: Once the experience buffer is full, the PPO agent begins training using historical experience tuples to update its policy network and value network parameters.
2. The sub-stream coupled sensing multipath congestion control method according to claim 1, characterized in that: How to implement step 1: Step 1.1: By confirming the timestamp of the data packet, according to the formula... Calculate the round-trip transmission delay of each substream. ; in, It is the timestamp of the confirmation packet being sent. It is the timestamp returned by the confirmation packet. Indicates the sub-stream sequence number. Indicates the sequence number of the feedback data packet; Step 1.2: Record the average arrival time of feedback data packets within the monitoring period. and calculate respectively and ; Then, the trend of substream RTT was calculated according to the formula. , The number of sub-streams; in, The average arrival time of feedback data packets is used to represent the average arrival time of all data packets collected during the monitoring period. The average is obtained as follows: Used to represent the average RTT, by analyzing all data collected during the monitoring period. We get the average. Step 1.3: Based on real-time statistical information during transmission, obtain the link status information of each sub-stream, including throughput. ,rate Packet loss rate Congestion window size ; The real-time link quality characteristics mentioned in step 1 include: bandwidth, transmission rate, and packet loss rate.
3. The sub-stream coupled sensing multipath congestion control method according to claim 2, characterized in that: How to implement step 2: Step 2.1: Compare the RTT change trend of each sub-stream with kGradientMin. When its value is greater than kGradientMin, start sub-stream coupling sensing. kGradientMin is the threshold used to determine whether to start sub-stream coupling sensing. Step 2.2: For each sub-flow, calculate the difference between it and other sub-flows. , believes that sub-stream Heziliu There is a possibility of coupling; This is the coupling sensing interval, where sub-flows within this interval may be coupled. Step 2.3: When there is a possibility of coupling between sub-streams, the state machine state of the Coupled BBR on each sub-stream will be used to assist in the judgment; If the state machines of both sub-streams are in the Probe RTT stage, then the two sub-streams are considered to be coupled. Recorded as ,in This indicates whether the sub-stream is in a coupled state. This represents the number of sub-flows in the coupled sub-flow; Step 2.4: Concatenate the sensed coupling relationship with the real-time sub-stream link to obtain a representation of the sub-stream state. ; in, Represents the current monitoring cycle number. Represents the sub-stream sequence number.
4. The sub-stream coupled sensing multipath congestion control method according to claim 3, characterized in that: How to implement step 3: The LSTM network described in step 3 will be used to eliminate network noise, extract hidden information from historical state sequences, and finally form new state information. Step 3.1: Concatenate the latest state obtained from the LSTM network with the sub-stream states from the previous K monitoring periods to obtain... ; in, Represents the current monitoring cycle number. Represents the sub-stream sequence number; Step 3.2: Input this state sequence into the LSTM network to calculate a new state that has eliminated network noise and contains temporal hidden information, called the new state. ; in, Represents the current monitoring cycle number. Represents the sub-stream sequence number; Step 3.3: Combine the sub-streams Combine to form a state vector , used to characterize the The status of all sub-streams in each monitoring cycle.
5. The sub-stream coupled sensing multipath congestion control method according to claim 4, characterized in that: How to implement step 4: Step 4.1: State vectors as Actor policy networks With the input, the policy network begins to calculate and obtain the action policy; Step 4.2: Generate an action distribution based on the probability distribution of the action policy, and randomly sample actions from it, in the following form: ; in, Represents the current monitoring cycle number. Represents the sub-stream sequence number.
6. The sub-stream coupling sensing multipath congestion control method according to claim 5, characterized in that: How to implement step 5: Step 5.1: The Coupled BBR receives the action generated in Step 4 as input; Step 5.2: The Coupled BBR maintains a transmission rate gain factor for each substream in each monitoring cycle; Adjust its own transmission gain coefficient according to the transmission gain value of each substream in the received action; Step 5.3: Coupled BBR can control the traffic injected into the network through such rate coupling, thereby achieving congestion control of transmission.
7. The sub-stream coupling sensing multipath congestion control method according to claim 6, characterized in that: How to implement step 6: Step 6.1: After the action is executed by the Coupled BBR, it will affect the environment, causing the environment to change. After going through steps 1-4 again, record the network environment state at this time. Feedback is given to the intelligent agent; Step 6.2: Simultaneously calculate the reward obtained by the agent after performing the action according to the following formula: in, This represents the current monitoring cycle number, and 'i' represents the sub-stream sequence number. Represents substream throughput. Represents the substream transmission delay. Represents the packet loss rate of the substream; The final reward is calculated as follows: Feedback is given to the intelligent agent; in It is a reward calculation for uncoupled subflows. It is a reward calculation for coupled sub-flows; Step 6.3: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] , , , These are combined to form experience tuples, which are stored in the agent's experience buffer pool for training the agent's policy network and value network parameters.
8. The sub-stream coupled sensing multipath congestion control method according to claim 7, characterized in that: How to implement step 7: Step 7.1: For each empirical tuple in the empirical buffer pool Using value network computing and value and ; in, Represents the current monitoring cycle number; Step 7.2: For each empirical tuple in the empirical buffer pool ,according to Calculate timing difference error , This is the discount factor, with a value of 0.9; and For value networks, the state and Calculate state value, Represents the current monitoring cycle number; Step 7.3: For each empirical tuple in the empirical buffer pool Generalization advantage estimation (GAE) is employed and based on Calculate the advantage value of the previous action and current action advantage value ,in, Represents the current monitoring cycle number; Step 7.4: Set the current The parameters are copied to the old policy network with the same network structure. In the middle, make ; Step 7.5: Randomly sample M experience tuples from the experience buffer pool and input them into the policy network. and In the middle, using the formula: Calculate the loss function of the policy network and update the network parameters using gradient descent; in, and Representing the intelligent agent in Actions taken and network status observed during the monitoring period; The function is a clipping function, where This is a correction factor used to determine the upper and lower bounds of the clipping. Step 7.6: After updating the policy network, use MSEloss as the loss function to update the parameters of the value network: in, In order to be in The state value estimated by the Critic value network during the monitoring period. In order to be in The objectives of updating the Critic value network during the monitoring period are: in, Represents the advantage value of the action. As a discount factor, The number of steps representing future rewards; Step 7.7: When the number of training iterations is reached, clear the experience buffer pool and end the training phase.
9. A computer-readable medium, characterized in that, It stores a computer program executed by an electronic device, which, when run on the electronic device, causes the electronic device to perform the steps of the method as described in any one of claims 1-8.