A network queue management method, device, and router for live video streams

The DDPG-based reinforcement learning network dynamically adjusts queue delay thresholds for live video streaming, addressing the limitations of fixed threshold-based packet dropping to enhance real-time performance and adaptability in network queue management.

CN116389375BActive Publication Date: 2025-07-15HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310282294.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-17
Publication Date
2025-07-15
Estimated Expiration
2043-03-17

AI Technical Summary

Technical Problem

The prior art cannot flexibly adjust the packet loss strategy based on the current status of the live video stream, and cannot meet the real-time requirements of the live video stream, especially when the network environment changes.

Method used

The reinforcement learning network (such as DDPG) is used to dynamically calculate the queue delay threshold of the video stream, and adjust the packet loss probability according to the arrival rate and round trip delay of the video stream. Through the packet management step, data packets exceeding the threshold are actively discarded to achieve dynamic adjustment of queue delay.

Benefits of technology

Effectively alleviate bottleneck node congestion, reduce queue delay, improve the real-time nature of live video streams, meet the high real-time requirements of each video stream, and improve user viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116389375B_ABST
    Figure CN116389375B_ABST
Patent Text Reader

Abstract

The present invention discloses a network queue management method, device and router for live video streams, belonging to the field of network congestion control, including: when a data packet dequeues, if its queuing delay does not exceed the corresponding queuing delay threshold, forward it, otherwise, discard it; and execute at preset time intervals: calculate the arrival rate and average queuing delay of each video stream within the current time interval, input them into a reinforcement learning network, and obtain the actions output by the network, which are used to describe the proportional relationship between the packet loss probability and queuing delay of the video stream; calculate and update the queuing delay threshold of each video stream according to the actions; calculate the current reward value and feedback it to the reinforcement learning network, and update the network parameters with the goal of maximizing the cumulative reward value; the closer the arrival rate is to the ideal arrival rate and the smaller the round-trip delay is, the greater the reward value is. The present invention can relieve the congestion state of bottleneck nodes while reducing the queuing delay and improving the real-time performance of live video streams.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network congestion control, and more specifically, relates to a network queue management method, device, and router for live video streams. Background Art

[0002] With the rapid development of mobile communication and multimedia technologies, video traffic is growing rapidly and currently accounts for more than 70% of Internet traffic. At the same time, network link congestion has become increasingly serious. In the video live broadcast scenario, since video data is generated in real time at the source end and cannot be cached or transmitted in advance, in order for users to watch videos smoothly, live broadcasts have strict requirements for the real-time nature of data transmission. In addition, some live broadcast software considers that users hope to see the latest content and may abandon the lagging content and directly start playing from the latest frame when the delay between the user end and the source end is large. When the link is congested, the delay of data packets will increase significantly, and the queuing delay accounts for the main part. Therefore, the real-time nature of video data transmission cannot be guaranteed, and even some video data will not be played even if it reaches the user end, which not only seriously affects the viewing experience of live broadcast users but also causes waste of bandwidth resources.

[0003] During the process of live video stream data packets being sent from the source end to the user end, they need to pass through a series of nodes (usually routers) for forwarding data packets. When the data packets reach these nodes, they will first be cached in the buffer within the node and organized through a queue. The data packets at the head of the queue will be taken out and forwarded in sequence. When the amount of data received by the node exceeds its data forwarding capacity, the node will become a bottleneck node. The network queue management algorithm deployed on the bottleneck node alleviates the congestion of the bottleneck node by actively discarding data packets. However, the prior art rarely considers the characteristics of upper-layer applications and usually designs a fixed strategy to discard data packets. For example, in the patent document CN113037697A, when the round-trip delay of a data packet is greater than a preset round-trip delay threshold, it is determined that the network status information meets the network congestion condition. Discarding data packets based on a fixed strategy during network congestion can effectively alleviate the congestion of the network. However, its parameter configuration cannot be changed in real time and it makes no distinction between heterogeneous flows. In the live broadcast scenario, such a fixed packet loss strategy is difficult to simultaneously meet the high real-time requirements of each video stream and has insufficient adaptability to the diverse network environment.

[0004] As mentioned in the patent document CN108965151A, only using static threshold to judge network congestion cannot provide accurate congestion feedback information for data center networks with high variability and complexity. At the same time, it is not applicable to multi-queue scheduling schemes and cannot distinguish multiple different priority queues. Accordingly, an explicit congestion control method based on queuing delay is proposed in this document. This method dynamically calculates a new threshold at the receiver according to the end-to-end (sender to receiver) queuing delay, which can effectively handle congestion caused by changes in application and network states, and provides differentiated thresholds for applications with different priorities based on the average queuing delay of different applications. However, the dynamic threshold calculated by this method can only be used as the judgment basis for the next transmission round, which has a certain lag and still cannot meet the real-time requirements of live broadcasts. Moreover, the update calculation of the threshold is based on the end-to-end delay and can only adapt to the overall state of the network, but cannot flexibly adjust the threshold according to the state of the current video stream.

[0005] Generally speaking, how to flexibly adjust the packet loss strategy according to the state of the current video stream is of great significance. Summary of the Invention

[0006] In view of the deficiencies and improvement requirements of the prior art, the present invention provides a network queue management method, device and router for live video streams, aiming to relieve the congestion state of bottleneck nodes, reduce queuing delay and improve the real-time performance of live video streams.

[0007] To achieve the above object, according to one aspect of the present invention, a network queue management method for live video streams is provided, including: a queuing delay threshold update step and a data packet management step;

[0008] The queuing delay threshold update step includes performing the following steps at a preset time interval:

[0009] (S1) Calculate the arrival rate and average queuing delay of each video stream within the current time interval as the state of the video stream and input it into the reinforcement learning network, so that the reinforcement learning network outputs a corresponding action; the action is used to describe the proportional relationship between the packet loss probability and queuing delay of the video stream;

[0010] (S2) Calculate and update the queuing delay threshold of each video stream according to the action output by the reinforcement learning network;

[0011] (S3) Calculate the current reward value and feedback it to the reinforcement learning network, and update the parameters of the reinforcement learning network with the goal of maximizing the cumulative reward value; the reward value is related to the arrival rate and round-trip delay of the video stream, and the closer the arrival rate is to the ideal arrival rate and the smaller the round-trip delay, the greater the reward value;

[0012] The data packet management steps include: when a data packet dequeues, if its queuing delay does not exceed the current queuing delay threshold of the video stream to which it belongs, forward the data packet; otherwise, discard the data packet.

[0013] Further, between step (S2) and step (S3), it also includes: determining whether the reinforcement learning network has converged. If so, wait until the end of the current time interval; if not, transfer to step (S3).

[0014] Further, step (S2) includes:

[0015] For any i-th video stream, calculate the ideal achievable rate of this video stream according to and calculate the coefficient K according to and then calculate the queuing delay threshold of this video stream according to

[0016]

[0017] where I is the total number of video streams, C is the total link bandwidth, k represents the action output by the reinforcement learning model; τ i represents the current round-trip delay of the i-th video stream, and s i (t) represents the current actual arrival rate of the i-th video stream.

[0018] Further, the calculation formula for the reward value is:

[0019] where R represents the reward value, s i represents the actual arrival rate of the i-th video stream,

[0020] represents the ideal arrival rate of the i-th video stream, and λ represents the coefficient of the actual arrival rate offset. Further,

[0021]

[0022] where e is the natural base.

[0023] Further, the video stream to which the data packet belongs is determined based on the five-tuple information in the data packet header. Data packets with the same five-tuple information belong to the same video stream.

[0023] Further, the reinforcement learning network is a DDPG network.

[0024] According to another aspect of the present invention, there is provided a network queue management device for live video streams, including:

[0025] A computer-readable storage medium for storing a computer program;

[0026] And a processor for reading a computer program in a computer-readable storage medium and executing the above-mentioned network queue management method for live video streams provided by the present invention.

[0027] According to another aspect of the present invention, a router is provided, including the above-mentioned network queue management device for live video streams provided by the present invention.

[0028] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0029] (1) According to the preset time interval, the present invention completes the dynamic update of the queuing delay threshold of each video stream by means of a reinforcement learning network, and actively discards the data packets exceeding the queuing delay threshold. Since the input of the reinforcement learning network is the arrival rate and average queuing delay of each video stream in the current time interval, and the output action is the key parameter for calculating the queuing delay threshold, that is, the ratio of the packet loss probability to the queuing delay of the video stream, and the reward value is related to the arrival rate and round-trip delay of the video stream, and the closer the arrival rate is to the ideal arrival rate and the smaller the round-trip delay, the greater the reward value. Therefore, the queuing delay threshold calculated for each video stream by the present invention is adapted to the current state of the corresponding video stream. Based on this queuing delay threshold, the data packets with too long queuing delay are discarded, which can effectively relieve the congestion state of the bottleneck node, reduce the queuing delay, and at the same time meet the real-time requirements of each video stream as much as possible.

[0030] (2) In the preferred solution of the present invention, when the reinforcement learning network converges, only the action output by the reinforcement learning network is used for the dynamic update of the queuing delay threshold of each video stream, and the reward value is no longer calculated, nor is the parameter update of the reinforcement learning network performed. Thus, while ensuring the calculation accuracy of the queuing delay threshold, the calculation amount can be reduced, and overfitting of the reinforcement learning network can be prevented.

[0031] (3) In the preferred solution of the present invention, calculate the queuing delay threshold of each video stream according to where K is a coefficient calculated according to the action k output by the reinforcement learning network. This queuing delay threshold calculation formula is obtained by analyzing based on the fluid model. Based on this calculation formula, when calculating the queuing delay threshold of each video stream using the action output by the reinforcement learning network, the calculation accuracy can be improved.

[0032] (4) In the preferred solution of the present invention, specifically calculate the reward value according to This calculation formula comprehensively reflects the comprehensive network performance index of all video streams in the current time interval. Using this as the calculation formula for the reward value can optimize the network performance through network queue management. In its further preferred solution, λ adopts the form of a hyperbolic tangent function, and its value changes with the change of the actual arrival rate offset, that is Let \(e\) be the base of the natural logarithm. By calculating the coefficient \(\lambda\) in the reward value in this way, it can better meet the changing trend that the closer the arrival rate is to the ideal arrival rate and the smaller the round-trip delay is, the greater the reward value.

[0033] (5) According to the quintuple information in the data packet header as the identifier of the corresponding video stream, the present invention can identify different video streams more accurately and quickly, and flexibly manage the network queue according to the states of different video streams.

[0034] (6) In the preferred solution of the present invention, the reinforcement learning network is specifically a DDPG (Deep Deterministic Policy Gradient) network. This network can solve continuous control problems and can be well applied to the video stream scenario. Moreover, the DDPG network finally definitely outputs only one action. In the present invention, within each time interval, the parameter of the ratio of the packet loss probability to the queuing delay of the video stream is shared by each video stream. Therefore, when using this parameter as the action output by the reinforcement learning network, specifically using DDPG can more accurately complete the prediction of relevant information. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a flowchart of the network queue management method for live video streams provided by an embodiment of the present invention;

[0036] Figure 2 is a schematic diagram of the existing DDPG network structure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0038] In the present invention, the terms "first", "second", etc. (if any) in the present invention and the drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0039] In order to solve the technical problem that the existing network queue management method cannot be flexibly adjusted according to the current state of the video stream and cannot well meet the real-time requirements of live video streams, the present invention provides a network queue management method, device and router for live video streams. The overall idea is: according to the current states of each video stream, dynamically calculate the ideal queuing delay threshold corresponding to each video stream, and actively discard the data packets whose queuing delay exceeds the corresponding queuing delay threshold.

[0040] The following are examples.

[0041] Example 1:

[0042] A network queue management method for live video streams, as Figure 1 shown, includes: a queuing delay threshold update step and a data packet management step; wherein, the queuing delay threshold update step is used to dynamically calculate the ideal queuing delay threshold corresponding to each video stream according to the current state of each video stream; the data packet management step is used to determine whether the queuing delay of the data packet exceeds the current queuing delay threshold of the video stream to which the data packet belongs when the data packet dequeues, and if so, directly discard the data packet, otherwise, forward the data packet.

[0043] It is easy to understand that at the initial moment, it is necessary to initialize the queuing delay threshold of each video stream. Optionally, in this embodiment, the queuing delay thresholds of each video stream are initialized to the same value, specifically 10 milliseconds. The packet header of each data packet contains five-tuple information, specifically the IP address, source port, destination IP address, destination port, and transport layer protocol. This five-tuple information can be used as the unique identifier of the video stream. In this embodiment, that is, the video stream to which the data packet belongs is determined according to the five-tuple information in the data packet header, and data packets with the same five-tuple information belong to the same video stream.

[0044] As Figure 1 shown, in this embodiment, the queuing delay threshold update step includes performing the following steps at a preset time interval:

[0045] (S1) Calculate the arrival rate and average queuing delay of each video stream in the current time interval as the state of the video stream and input it into the reinforcement learning network, so that the reinforcement learning network outputs the corresponding action; the action is used to describe the proportional relationship between the packet loss probability and the queuing delay of the video stream;

[0046] The arrival rate and average queuing delay of each video stream in the current time interval accurately reflect the state of each video stream in the current time interval. In this embodiment, this is used as the input of the reinforcement learning network, which can enable the reinforcement learning network to learn the state of each video stream and output an action adapted to the state of each video stream;

[0047] (S2) Calculate and update the queuing delay threshold of each video stream according to the action output by the reinforcement learning network;

[0048] (S3) Calculate the current reward value and feedback it to the reinforcement learning network, and update the parameters of the reinforcement learning network with the goal of maximizing the cumulative reward value; the reward value is related to the arrival rate and round-trip delay of the video stream, and the closer the arrival rate is to the ideal arrival rate and the smaller the round-trip delay, the greater the reward value;

[0049] In this embodiment, the reward value of the reinforcement learning network is set in the above manner, serving as the basis for updating the parameters of the reinforcement learning network. After calculating the queuing delay threshold based on the actions output by the reinforcement learning network and actively discarding the data packets with too long queuing delays according to the calculated queuing delay, the overall network performance is optimized. While alleviating the congestion state of the bottleneck node, the queuing delay is reduced, and the real-time requirements of the live video stream are satisfied as much as possible.

[0050] To improve the calculation accuracy of the queuing delay pre-threshold, as a preferred implementation, in this embodiment, the specific calculation formula is obtained through fluid model analysis. The specific analysis process is as follows:

[0051] The dynamic value of the congestion window of each video stream can be expressed as the following non-linear differential equation:

[0052]

[0053] where τ i (t) is the round-trip delay of the i-th video stream, p i (t) is the packet loss probability of the i-th video stream, which is proportional to the queuing delay d i (t), that is, p i (t) = kd i (t), where k is a positive number. k is a key parameter for queuing delay calculation and is also very crucial for reasonably determining the queuing delay threshold of each video stream. Moreover, its actual value is difficult to directly determine. Based on this consideration, in this embodiment, the parameter k, which describes the proportional relationship between the packet loss probability and the queuing delay of the video stream, is used as the action output by the reinforcement learning network, which can ensure the accuracy of relevant calculations. At the same time, compared with directly using the reinforcement learning network to output the queuing delay threshold, in this embodiment, only the key parameter for queuing delay threshold calculation is predicted by the reinforcement learning network, and the queuing delay threshold is further calculated based on the action output by the reinforcement learning network, which can combine the reinforcement learning with the theoretical knowledge of data packet transmission and further improve the calculation accuracy of the queuing delay threshold.

[0054] The propagation delay from the source end to the bottleneck node can be ignored compared with the queuing delay, so d i (t) can be expressed as:

[0055]

[0056] where I is the total number of video streams and C is the total link bandwidth. In addition, the dynamic value of the actual arrival rate of the i-th video stream can be expressed as:

[0057]

[0058] Therefore, in the equilibrium state, that is when, the ideal arrival rate of the i-th video stream can be calculated as:

[0059]

[0060] where is the packet loss probability of the i-th video stream in the equilibrium state. Substituting p i (t) = kd i (t), we can get where represents the optimal queuing time of the i-th video stream, that is, the ideal queuing delay threshold of the i-th video stream. At the same time, since the sending window w i (t) is restricted by , the actual arrival rate of the i-th video stream can be expressed as:

[0061]

[0062] To meet the high real-time requirements of live video streams, let the ideal arrival rate of each video stream be proportional to its round-trip delay, that is Its meaning is that video streams with longer delays tend to get more bandwidth resources, thereby reducing the delay and making the delays of each video stream as equal as possible. Finally, the ideal queuing delay threshold of the i-th video stream is obtained:

[0063]

[0064] where where k is the action output by the reinforcement learning network.

[0065] Based on the above analysis, in this embodiment, step (S2) includes:

[0066] For any i-th video stream, according to calculate the ideal arrival rate of this video stream and according to calculate the coefficient K, and then according to calculate the queuing delay threshold of this video stream

[0067] In this embodiment, the selected reinforcement learning network is specifically a DDPG network, and its structure is specifically as Figure 2As shown in the figure. The DDPG network mainly consists of an action network and a value network. The action network is responsible for outputting corresponding actions according to the current state. It is composed of a convolutional layer and a fully connected layer in series, and the activation function uses the tanh function. The value network is also composed of a convolutional layer and a fully connected layer in series, and the activation function uses the relu function, which is responsible for estimating the value of the current state. To achieve stable value estimation, each network is further divided into a target network and a real network. The real network and the target network have the same structure, but the process of updating their network parameters is slightly different. The real network updates its parameters by the gradient descent method after each training, where the loss function is calculated by the target network. Then the target network performs a soft update on this basis, specifically expressed as:

[0068] θ' = βθ+(1 - β)θ'

[0069] where θ' is the parameter of the target network, θ is the parameter of the real network, and β is the soft update coefficient. The DDPG network can solve continuous control problems and can be well applied to video stream scenarios, and can complete the prediction of relevant information more accurately. Moreover, the DDPG network finally definitely outputs only one action. In the present invention, within each time interval, the parameter of the ratio of the packet loss probability to the queuing delay of the video stream is shared by each video stream. Therefore, when using this parameter as the action output by the reinforcement learning network, specifically using DDPG can more accurately complete the prediction of relevant information.

[0070] The time interval for updating the queuing delay threshold of each video stream can be set accordingly according to the state of the video stream and the network state. Optionally, in this embodiment, the time interval is specifically set to 10 milliseconds.

[0071] In this embodiment, after each update of the queuing delay threshold of each video stream, the corresponding reward value will be calculated as the basis for updating the parameters of the reinforcement learning network. As a preferred implementation manner, the calculation formula of the reward value is:

[0072]

[0073] where R represents the reward value, s i represents the actual arrival rate of the i-th video stream, represents the ideal arrival rate of the i-th video stream, and λ represents the coefficient of the actual arrival rate offset. Optionally, in this embodiment, λ adopts the form of the hyperbolic tangent function, and its value changes with the change of the actual arrival rate offset, that is e is the natural base.

[0074] Optionally, in this embodiment, when updating the parameters of the reinforcement learning network, the gradient descent method is specifically used. Through the update of the network parameters, the network performance will be continuously optimized and finally converge, that is, the feedback reward value stops growing. At this time, there is no need to update the network parameters. Based on this consideration, after completing the update of the queuing delay threshold in this embodiment and before updating the network parameters, it will be further determined whether the network has converged. If the network has converged, the gradient descent method is used to update the network parameters; otherwise, no update is performed. Correspondingly, as Figure 1 shown, in this embodiment, between step (S2) and step (S3), it further includes: determining whether the reinforcement learning network has converged. If so, wait until the end of the current time interval; if not, proceed to step (S3).

[0075] Generally speaking, in this embodiment, the queuing delay threshold of each video stream is flexibly adjusted according to its status information. Congestion is alleviated by actively discarding packets exceeding the threshold, while the queuing delay is reduced to meet the real-time requirements of the live video stream as much as possible and ensure the viewing experience quality of each user.

[0076] Embodiment 2:

[0077] A network queue management device for live video streams includes:

[0078] A computer-readable storage medium for storing computer programs;

[0079] And a processor for reading the computer programs in the computer-readable storage medium and executing the network queue management method for live video streams provided in the above Embodiment 1.

[0080] Embodiment 3:

[0081] A router includes the network queue management device for live video streams provided in the above Embodiment 2.

[0082] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A network queue management method for live video streams, characterized in that, Including: A queuing delay threshold update step and a data packet management step; The queuing delay threshold update step includes performing the following steps at a preset time interval: (S1) Calculate the arrival rate and average queuing delay of each video stream within the current time interval as the state of the video stream and input it into the reinforcement learning network, so that the reinforcement learning network outputs a corresponding action; the action is used to describe the proportional relationship between the packet loss probability and queuing delay of the video stream; (S2) Calculate and update the queuing delay threshold of each video stream according to the action output by the reinforcement learning network; (S3) Calculate the current reward value and feedback it to the reinforcement learning network, and update the parameters of the reinforcement learning network with the maximization of the cumulative reward value as the optimization goal; the reward value is related to the arrival rate and round-trip delay of the video stream, and the closer the arrival rate is to the ideal arrival rate and the smaller the round-trip delay, the greater the reward value; The data packet management step includes: when the data packet dequeues, if its queuing delay does not exceed the current queuing delay threshold of the video stream to which it belongs, forward the data packet; otherwise, discard the data packet; Between the step (S2) and the step (S3), it further includes: determining whether the reinforcement learning network converges, if so, wait until the end of the current time interval; if not, transfer to the step (S3); The step (S2) includes: For any i-th video stream, according to calculate the ideal achievable rate of this video stream and according to calculate the coefficient K, and then according to calculate the queuing delay threshold of this video stream Where I is the total number of video streams, C is the total link bandwidth, k represents the action output by the reinforcement learning model; τ i represents the current round-trip delay of the i-th video stream, s i (t) represents the current actual arrival rate of the i-th video stream.

2. The network queue management method for live video streams according to claim 1, characterized in that, The calculation formula of the reward value is: where R represents the reward value, and s i represents the actual arrival rate of the i-th video stream, represents the ideal arrival rate of the i-th video stream, and λ represents the coefficient of the actual arrival rate offset.

3. The network queue management method for live video streams according to claim 2, wherein where e is the natural logarithm base.

4. The network queue management method for live video streams according to claim 1, characterized in that, The video stream to which the data packet belongs is judged according to the five-tuple information in the data packet header, and data packets with the same five-tuple information belong to the same video stream.

5. The network queue management method for live video streams according to claim 1, characterized in that The reinforcement learning network is a DDPG network.

6. A network queue management device for live video streams, characterized in that, Including: A computer-readable storage medium for storing a computer program; And a processor for reading the computer program in the computer-readable storage medium and executing the network queue management method for live video streams according to any one of claims 1 to 5.

7. A router, characterized in that, Including the network queue management device for live video streams according to claim 6.

Citation Information

Patent Citations

  • Explicit congestion control method based on queuing delay

    CN108965151A

  • Video frame processing method and device, electronic equipment and readable storage medium

    CN113037697A