Network flow management method, electronic equipment, storage medium and program product
By building global state and global reward functions, combined with learning algorithms to optimize congestion window control strategy, the resource utilization and fairness problems of traditional congestion control algorithms in a multi-stream competition environment are solved, and efficient utilization of network resources and performance improvement under multi-stream concurrency is achieved.
Patent Information
- Application Number
- CN202510856428.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Traditional congestion control algorithms cannot effectively utilize network resources in high-latency and high-bandwidth network environments, resulting in insufficient fairness and resource utilization, and cannot achieve fair bandwidth allocation in a multi-stream competition environment.
By generating and starting multiple network flows, periodically collecting data packet statistics information, building global states and designing global reward functions, optimizing congestion window control strategy using preset learning algorithms, and dynamically adjusting the congestion window parameters of network flows to achieve efficient utilization and fair distribution of network resources.
Maintain good convergence speed, stability and fairness in high bandwidth and high latency environments, and improve the overall utilization efficiency of network resources and performance under multi-stream concurrency.
Smart Images

Figure CN120378364A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of network management, and in particular, to a method for managing network flows, an electronic device, a storage medium, and a program product. Background Art
[0002] With the rapid development of the Internet, the types and quantities of network applications have been increasing continuously, and the demand for network bandwidth has also been growing day by day. Against this background, traditional congestion control algorithms have gradually revealed some deficiencies when dealing with congestion problems in complex network environments. Specifically, in high-latency and high-bandwidth network scenarios, traditional algorithms often cannot effectively utilize network resources, resulting in a significant decline in network performance. On the other hand, since the training environment of traditional algorithms mainly focuses on the performance optimization of a single network flow and fails to fully consider the balanced allocation problem among multiple network flows, it is difficult for them to maintain consistently good performance in terms of convergence characteristics such as fairness, fast convergence, and stability. That is to say, the insensitivity of traditional algorithms to fairness leads to their inability to achieve fair bandwidth allocation in a multi-flow competition environment, thereby causing an unbalanced phenomenon where some network flows obtain excessive bandwidth resources while other network flows are difficult to obtain sufficient resources.
[0003] Therefore, providing an algorithm that can balance the bandwidth allocation of multiple network flows is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention
[0004] This application provides a method for managing network flows, an electronic device, a storage medium, and a program product to at least solve the problems of insufficient fairness and resource utilization in a multi-flow competition environment in related technologies.
[0005] This application provides a method for managing network flows, including: generating and starting multiple network flows according to flow configuration information; collecting packet statistical information of the multiple network flows every preset monitoring time period; the packet statistical information characterizing the network status of the current network flow; constructing a global state according to the multiple packet statistical information and constructing a global reward function according to the global state; using a preset learning algorithm to optimize the global reward function to obtain the optimal congestion window control strategy for each network flow; generating action values according to the optimal congestion window control strategy and adjusting the congestion window parameters of each network flow according to the action values.
[0006] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above methods for managing network flows when executing the computer program.
[0007] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any one of the above-mentioned network flow management methods.
[0008] The present application also provides a computer program product including a computer program, which, when executed by a processor, implements the steps of any one of the above-mentioned network flow management methods.
[0009] With the present application, by introducing a global state construction and global reward function optimization mechanism, the problems of insufficient fairness and resource utilization rate of traditional congestion control algorithms in a multi-flow competition environment are effectively overcome. Specifically, after generating and starting multiple network flows, the method periodically collects packet statistical information of each network flow to comprehensively characterize the current network condition, thereby constructing a unified global state view, and designing a global reward function based on this global state. This function not only considers the performance of individual network flows but also comprehensively considers the utilization rate of overall network resources and the bandwidth fairness among network flows. Further, by using a preset learning algorithm to optimize the global reward function, an optimal congestion window control strategy for multiple network flows can be dynamically obtained, realizing coordination and dynamic adjustment among flows and avoiding the uneven resource allocation phenomenon caused by only optimizing a single flow in traditional algorithms. Finally, the method adjusts the congestion window parameters of each network flow in real time according to the action values generated by the optimal strategy, enabling the network system to still maintain good convergence speed, stability, and fairness in a high-bandwidth and high-latency environment, fundamentally improving the overall utilization efficiency of network resources and the performance under multi-flow concurrency, solving the problems of insufficient fairness and resource utilization rate in a multi-flow competition environment, and achieving the effect of improving the overall utilization efficiency of network resources and the performance under multi-flow concurrency. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 It is a flowchart of a network flow management method provided by an embodiment of the present application; Figure 2 It is another flowchart of a network flow management method provided by an embodiment of the present application; Figure 3 It is a module schematic diagram of a network flow management method provided by an embodiment of the present application; Figure 4Schematic diagram of network training for a network flow management method provided by an embodiment of this application; Figure 5 Schematic diagram of an electronic device provided by an embodiment of this application; Figure 6 Schematic diagram of a computer-readable storage medium provided by an embodiment of this application. Detailed implementation manners
[0012] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0013] It should be noted that in the description of this application, the terms "including", "comprising" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0014] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be made in conjunction with the accompanying drawings and specific implementation manners.
[0015] As Figure 1 , an embodiment of this application provides a network flow management method, including: S11: Generate and start multiple network flows according to flow configuration information.
[0016] Specifically, a flow generator can be used to generate and start multiple network flows according to preset flow configuration information, with the aim of accurately simulating the communication behavior of multi-flow concurrency in an actual network in a control environment. These configuration information usually include the start time, running time, preselected congestion control algorithm type of each network flow, as well as special parameters such as whether to introduce artificial delay and bandwidth limitation, for restoring complex network operation scenarios.
[0017] With such settings, not only can a representative network load model be flexibly constructed, but it also provides a basis for diversity and controllability in subsequent state monitoring and policy optimization, contributing to the training and evaluation of the adaptability and fairness performance of different network policies in a multi-flow environment. This step ensures that the entire management method is built on the basis of realistic simulation and high configurability from the very beginning, enabling subsequent learning optimization to have stronger generalization ability and practical significance.
[0018] S12: At every preset monitoring time period, collect the packet statistical information of multiple network flows; the packet statistical information characterizes the network status of the current network flow.
[0019] Specifically, perform periodic monitoring on multiple network flows to obtain packet statistical information reflecting the current network operating state. These packet statistical information includes, but is not limited to, indicators such as the sending rate, packet loss rate, transmission delay, and number of congestion events of each network flow, which are used to comprehensively describe the communication performance and resource occupancy of each flow in the current network environment.
[0020] By collecting at every preset monitoring time period, the changing trend of the network state can be dynamically tracked, providing real-time and accurate data support for subsequent policy adjustment. This periodic monitoring mechanism not only improves the timeliness of state perception but also provides a multi-dimensional input basis for constructing the global state, thus supporting the formulation of more refined congestion control and resource scheduling policies.
[0021] S13: Construct a global state based on the multiple packet statistical information, and construct a global reward function based on the global state.
[0022] Specifically, integrate the packet statistical information collected from multiple network flows to construct a global state view that can comprehensively reflect the overall state of the current network. This global state not only includes the independent characteristics of each network flow (such as packet loss rate, delay, throughput, etc.) but also reflects the mutual influence and competition relationship between network flows, thus providing a unified description of the usage of the entire network resources. Compared with traditional methods that only rely on local information, the global state is more conducive to optimization and regulation from the system level.
[0023] On the basis of constructing the global state, further design a global reward function according to this state. This global reward function comprehensively considers multiple objectives, such as maximizing the overall throughput, minimizing the network delay, and fairness of bandwidth allocation, etc., and is used to measure the quality of the current network control policy. Through the reward function constructed in this way, the network performance objectives can be quantified, and a clear evaluation basis can be provided for subsequent optimization using learning algorithms, ensuring that the direction of policy adjustment has global optimality.
[0024] S14: Optimize the global reward function using a preset learning algorithm to obtain the optimal congestion window control strategy for each network flow.
[0025] Specifically, optimize the constructed global reward function through a preset learning algorithm. The purpose is to find a set of optimal congestion window control strategies to achieve the efficient utilization and fair allocation of network resources. The learning algorithm can adopt reinforcement learning, deep reinforcement learning, or other strategy optimization methods suitable for dealing with dynamic decision-making problems. By continuously interacting with the network environment, it learns which control strategy to adopt under different global states to obtain higher reward values.
[0026] During the optimization process, the algorithm will try various strategies and update according to the feedback of the global reward function after each adjustment, gradually converging to a control strategy with good performance in a complex network environment. This feedback-based adaptive learning method enables it to dynamically adapt to changes in network states, overcome the limitations of traditional methods such as static parameter tuning or single-flow optimization, and achieve the unified optimization goal of bandwidth fairness, congestion mitigation, and performance improvement from a global perspective.
[0027] S15: Generate action values according to the optimal congestion window control strategy and adjust the congestion window parameters of each network flow according to the action values.
[0028] Specifically, convert the optimal congestion window control strategy output by the learning algorithm into specific control instructions, that is, "action values", and dynamically adjust the congestion window parameters of each network flow accordingly. Action values usually represent the increase or decrease amplitude or adjustment direction of the window size, reflecting the optimal bandwidth usage behavior under the current global state. By mapping the strategy into an executable parameter adjustment operation, the sending rate of each network flow can be regulated in real time to respond to changes in the current network conditions.
[0029] This strategy-driven window adjustment method not only achieves fine-grained control but also has high adaptability, being able to make timely adjustments according to fluctuations in network load to avoid congestion aggravation or bandwidth waste. At the same time, the action values are generated under the guidance of the global reward function, so it can ensure more fair bandwidth allocation among multiple flows and higher overall system throughput, effectively solving the performance bottleneck problem caused by rigid strategies in traditional algorithms.
[0030] In an exemplary embodiment, a global reward function is constructed according to a global state, including: calculating a throughput indicator, a delay indicator, a packet loss rate indicator, a fairness indicator and a stability indicator according to the global state; the throughput indicator is used to characterize the network bandwidth utilization, the delay indicator is used to characterize the network transmission delay, the packet loss rate indicator is used to characterize the network transmission reliability, the fairness indicator is used to characterize the fairness of the bandwidth allocation of each network flow, and the stability indicator is used to characterize the fluctuation degree of the bandwidth allocation of each network flow; the throughput indicator, the packet loss rate indicator, the delay indicator, the fairness indicator and the stability indicator are coupled to obtain a global reward function.
[0031] Specifically, the construction of the global reward function is based on multiple key performance indicators, with the aim of comprehensively reflecting the current operating quality of the network system and providing an accurate evaluation basis for subsequent learning optimization. First, the five core indicators in the network are extracted through the global state: throughput, latency, packet loss rate, fairness, and stability (a combination of more than two of them can be selected).
[0032] Among them, the throughput index is used to measure the efficiency of the entire network in utilizing bandwidth resources. The higher the value, the stronger the network transmission capacity. The delay index reflects the time it takes for a data packet to be sent and received. The smaller the delay, the more timely the network response. The packet loss rate index measures the data loss during network transmission, which is a key manifestation of network reliability. Too high a packet loss rate will cause a serious degradation in the performance of upper-layer applications. The fairness index is used to measure the balance of bandwidth resource allocation for each network flow, preventing some flows from occupying too many resources while other flows have insufficient resources. This is a crucial performance factor in a multi-flow competition environment. The stability index focuses on the variation of bandwidth allocation under the network control strategy. Excessive fluctuations will affect user experience and system predictability, so stability is also an important evaluation dimension.
[0033] After obtaining the above indicators, this embodiment further weights and couples these indicators to construct a unified global reward function. The value of this function comprehensively reflects the overall advantages and disadvantages of the current network control strategy, and by setting the weights of different indicators, the system can flexibly adjust the optimization direction according to the focus of the actual application scenario. For example, in the video transmission scenario, the weights of throughput and stability can be enhanced, while in interactive services, the control of delay and packet loss rate can be strengthened. For example, in a specific embodiment, the expression of the global reward function is , where R is the reward value of the global reward function, c0, c1, c2, c3, c4 are the weights of the corresponding indicators, and R thr is the throughput indicator, R lat is the delay index, R loss is the packet loss rate indicator, R fair is the fairness index, R stabis the stability index. When R fair and R stab is 0, the network flow reaches the best balance of fairness and stability.
[0034] In addition, the reward range is restricted for each preset monitoring time period and scaled to between (-0.1, 0.1). Through a specially designed global reward function, the reinforcement learning training algorithm rewards high throughput, good fairness and stability, while punishing high latency and loss rate. Different trade-offs between these metrics can be achieved by adjusting the weight coefficients c0, c1, c2, c3, c4. The global reward function serves as the overall goal of congestion control schemes in various network environments. By maintaining a consistent global reward function, it is ensured that the reinforcement learning mechanism can adaptively align the current policy with the performance objectives in different scenarios.
[0035] Through this method of multi-dimensional metric fusion, the global reward function can achieve trade-offs between multiple performance objectives, provide an accurate and comprehensive optimization reference for the learning algorithm, effectively guide policy adjustment, thereby improving the overall system performance, and realizing the collaborative optimization of maximizing bandwidth utilization, enhancing fairness between network flows, and improving transmission stability.
[0036] For example Figure 2 , this is a network congestion control process based on reinforcement learning. The process starts by preparing the configuration information of the network flow as the input. Based on the configuration information of the network flow, the network flow is generated while initializing the shared control policy (the basic rule for subsequent optimization and adjustment of the policy). The packet statistics of each network flow (such as throughput, latency, etc.) are collected, and an action request (such as a decision requirement to adjust the congestion window) is initiated. The above requests are passed, and the packet statistics of all flows are aggregated into a global state (enabling the algorithm to perceive the network situation from an overall perspective). Based on the global state, a global reward function for measuring network performance (such as fairness, stability, etc.) is generated as the optimization goal. Guided by the global reward function, the control policy of the network flow is trained, and the congestion window parameters are updated (to make the network more adaptable to the current environment and optimize data transmission). After completing one round of the optimization process, subsequent continuous optimization can be performed in a loop. Briefly speaking, it is a closed-loop process that starts from configuring the network flow, collects data, constructs an overall perspective, and then uses the reward function to guide the training to finally optimize network transmission.
[0037] In an exemplary embodiment, calculating the throughput metric according to the global state includes: calculating the throughput metric according to the current throughput of each network flow and the link bandwidth; the global state includes the current throughput of each network flow and the link bandwidth.
[0038] Specifically, the calculation of the throughput metric is based on the current throughput of each network flow in the global state and the bandwidth information of the link where it is located. The current throughput represents the amount of data successfully transmitted by the network flow per unit time, while the link bandwidth is the maximum transmission capacity that the network flow can theoretically use. By comparing the current throughput with the link bandwidth, the utilization efficiency of each network flow for the available network resources can be evaluated.
[0039] Specifically, the throughput metric is usually expressed in the form of a ratio or normalization. For example, dividing the current throughput by the link bandwidth gives a utilization value between 0 and 1. The utilization rates of multiple network flows can be further averaged or weighted and integrated to obtain the throughput metric of the entire system, thereby reflecting the overall usage status of the current network resources. This metric provides a key reference for the network utilization efficiency for the optimization algorithm, helping to improve the overall bandwidth utilization and avoid resource waste.
[0040] In an exemplary embodiment, calculating the throughput metric according to the current throughput and link bandwidth of each network flow includes: according to Calculating the throughput metric; where R thr is the throughput metric, i is the serial number of the network flow, and thr i is the current throughput of the i-th network flow, and c is the link bandwidth.
[0041] The throughput metric is obtained by calculating the ratio of the total throughput of all current network flows to the total bandwidth of the network link. Specifically, first, the sum of the throughputs of all network flows at the current moment is counted, and then this sum is divided by the maximum bandwidth of the link to obtain a ratio value between 0 and 1. This ratio reflects the overall utilization efficiency of the network bandwidth. The closer the value is to 1, the more fully the network resources are utilized and the higher the throughput.
[0042] In this way, the throughput metric can intuitively quantify the transmission capacity and efficiency of the network in the current state. As a key component in the global reward function, it provides an effective performance feedback basis for optimizing the congestion control strategy, helping the system achieve more efficient bandwidth allocation and utilization.
[0043] In an exemplary embodiment, calculating the packet loss rate metric according to the global state includes: calculating the packet loss rate metric according to the packet loss rate and current throughput of each network flow; the global state includes the packet loss rate and current throughput of each network flow.
[0044] Specifically, the calculation of the packet loss rate metric is based on the packet loss rate and the current throughput information of each network flow in the global state. The packet loss rate reflects the proportion of packets lost during transmission and is an important indicator for measuring the reliability of network transmission; while the current throughput represents the actual amount of data transmitted by the network flow. Combining these two can more accurately evaluate the impact of packet loss on network performance.
[0045] Specifically, when calculating the packet loss rate metric, not only the packet loss ratio of each flow is concerned, but also the impact of the throughput on packet loss is weighted, so that the impact of packet loss in flows with larger transmission volumes on the overall metric is more significant. In this way, the packet loss rate metric can comprehensively reflect the impact of packet loss phenomena in the network on the overall transmission performance, provide a targeted optimization direction for congestion control algorithms, and help improve the reliability and transmission quality of the network.
[0046] In an exemplary embodiment, the packet loss rate metric is calculated according to the packet loss rate and the current throughput of each network flow, including: according to Calculate the packet loss rate metric; where R loss is the packet loss rate metric, n is the number of all network flows, i is the serial number of the network flow, loss i is the packet loss rate of the i-th network flow, and thr i is the current throughput of the i-th network flow.
[0047] The packet loss rate metric is obtained by calculating the ratio of the packet loss rate of all network flows to their current throughput and then taking the average of these ratios. Specifically, for each network flow, first combine its packet loss rate with the current throughput to reflect the actual impact degree of packet loss in this flow on the overall transmission, and then average the ratios of all flows to obtain a comprehensive packet loss rate metric.
[0048] This calculation method takes into account both the packet loss rate itself and the transmission scale of the flow, making the impact of packet loss in flows with larger throughputs on the metric greater, and reflecting the real impact of packet loss on the overall network performance. Through this metric, the transmission reliability of the network can be more accurately evaluated, and thus an effective basis can be provided for optimizing congestion control strategies.
[0049] In an exemplary embodiment, the delay metric is calculated according to the global state, including: calculating the delay metric according to the transmission delay, reference delay, and data sending rate of each network flow; the global state includes the transmission delay, reference delay, and data sending rate of each network flow.
[0050] Specifically, the calculation of the delay metric is based on the transmission delay, reference delay, and data sending rate of each network flow in the global state. The transmission delay represents the time it takes for a data packet to travel from the sender to the receiver, and the reference delay is the ideal or expected minimum delay. The difference between the two reflects the additional delay burden on the network. Combining with the sending rate can more accurately measure the impact of delay on the actual data transmission performance.
[0051] Specifically, the delay metric emphasizes the impact of the delay performance of high-sending-rate flows on the overall network performance by comparing the deviation between the transmission delay and the reference delay of each network flow and weighting it according to the data sending rate of that flow. The delay metric calculated in this way can comprehensively reflect the impact of network delay on transmission efficiency, provide an important reference for optimizing congestion control strategies, help reduce network response time, and improve user experience.
[0052] In an exemplary embodiment, the delay metric is calculated according to the transmission delay, reference delay, and data sending rate of each network flow, including: according to Calculate the delay metric; where, R lat is the delay metric, n is the number of all network flows, i is the serial number of the network flow, lat i is the transmission delay of the i-th network flow, d0 is the reference delay, P rate is the data sending rate, and β is the tolerance coefficient.
[0053] Specifically, the delay metric measures the deviation between the average transmission delay of the network flow and the reference delay d0, and at the same time combines the data packet sending rate P rate to adjust the penalty for high-delay situations. Specifically, when all network flows enter the stable transmission state, they are allowed to reach the maximum throughput in the dynamic link, but when the average delay exceeds the threshold (1 + β)d0, this delay inflation behavior will be punished. Otherwise, R lat is 0, and normal or slightly increased delays are not punished, thus avoiding misjudgment of smaller queue lengths.
[0054] In addition, the data packet sending rate is used as a multiplier in the metric, which means that for high-sending-rate flows, if the delay shows a significant increase, the penalty will be more severe. This design effectively suppresses the behavior of attempting to occupy more bandwidth by increasing the sending rate on high-delay links, avoids the deterioration of network congestion, promotes the balance between delay and throughput, and thus improves the overall network performance and user experience.
[0055] In an exemplary embodiment, calculating a fairness metric based on the global state includes: for each network flow, obtaining the historical throughput of w preset monitoring time periods, and calculating the average throughput based on the w historical throughputs; w is an integer greater than 1; the global state includes the historical throughputs of the w preset monitoring time periods; and calculating the fairness metric based on the average throughput.
[0056] Specifically, the calculation of the fairness metric is based on the historical throughput data of each network flow within multiple preset monitoring time periods. Specifically, the throughput of each flow is collected over w consecutive time periods, and the average of these data is calculated to reflect the stable transmission capacity and resource occupancy of the flow over a period of time.
[0057] By comparing the average throughputs of all network flows, the fairness of bandwidth allocation can be evaluated. That is, it is determined whether each flow has obtained relatively balanced resources. The smaller the value of the fairness metric, the more balanced the bandwidth allocation, which helps to avoid the problem that some flows occupy too many resources for a long time while other flows have no available bandwidth, thus promoting the harmonious and stable operation of the overall network.
[0058] In an exemplary embodiment, calculating the fairness metric based on the average throughput includes: according to , calculating the fairness metric; where R fair is the fairness metric, n is the number of all network flows, i is the serial number of the network flow, and avg_thr i is the average throughput.
[0059] Specifically, the fairness metric measures the degree of balance of bandwidth allocation based on the standard deviation of the throughputs of all active network flows within the same time period. Specifically, first, the average throughput of each flow in the past w monitoring time periods is calculated to avoid errors caused by instantaneous fluctuations, and then the standard deviation of these average throughputs is calculated. The smaller the standard deviation, the closer the throughput distribution of each flow is to being uniform, and the fairer the bandwidth allocation.
[0060] In addition, the denominator of the standard deviation in the formula contains the sum of the squares of the average throughputs of all flows. This normalization process makes the metric scale-independent and can adapt to network environments with different bandwidth scales. By averaging the fairness metric along the time axis, the fairness trend of bandwidth allocation can be stably evaluated, helping the congestion control algorithm to maintain balance among multiple flows during the optimization process and avoiding long-term restrictions on some flows due to uneven resource allocation.
[0061] In an exemplary embodiment, calculating a stability metric based on the global state includes: for each network flow, obtaining the historical throughputs of w preset monitoring time periods, and calculating the average throughput according to the w historical throughputs; w is an integer greater than 1; the global state includes the historical throughputs of the w preset monitoring time periods; calculating the stability metric according to the average throughput and the w historical throughputs.
[0062] Specifically, the stability metric is used to measure the fluctuations of the throughput of each network flow over time. First, for each network flow, collect the historical throughput data within the past w preset monitoring time periods, and calculate the average value of these data to reflect the stable transmission level of the flow over a period of time.
[0063] Next, by comparing the differences between these w historical throughputs and the average throughput (such as calculating the standard deviation or variance), quantify the fluctuation range of the throughput. The smaller the fluctuation, the more stable the transmission of the network flow; a larger fluctuation indicates unstable bandwidth allocation or network conditions.
[0064] As part of the global reward function, the stability metric helps to guide the congestion control strategy to reduce the uncertainty and fluctuations of transmission while increasing the throughput, ensuring the continuity of network performance and the smoothness of the user experience.
[0065] In an exemplary embodiment, calculating the stability metric according to the average throughput and the w historical throughputs includes: according to , calculating the stability metric; where R stab is the stability metric, n is the number of all network flows, i is the serial number of the network flow, w is the number of historical throughputs, j is the serial number of the historical throughput, t is the t-th preset monitoring time period, avg_thr i is the average throughput, and thr i,t-j is the historical throughput of the i-th network flow in the t-j-th preset monitoring time period.
[0066] Specifically, the stability metric measures the stability of bandwidth allocation based on the standard deviation of the throughput of each network flow within a historical monitoring time period of fixed length w. Specifically, for each network flow i, first calculate the sum of the squared deviations between its historical throughputs thr i,t−j , thr i,t-j within the past w time periods and the average throughput avg_thr i of the network flow, and perform normalization processing (divide by w and the square of the average throughput) to eliminate the influence of the throughput magnitude on the degree of fluctuation.
[0067] Subsequently, the normalized variances of all network flows are averaged to obtain an overall stability index. The smaller the stability index, the smoother the throughput change of the network flow and the smaller the fluctuation; a larger stability index indicates severe traffic fluctuations and unstable network performance. By averaging the stability index for each flow, the interference of transient anomalies on the evaluation results is avoided, making the index smoother and more reliable, which helps the congestion control algorithm to achieve more stable and continuous bandwidth allocation.
[0068] As Figure 3 , this process describes an agent-based network flow control mechanism. Each network flow can be understood as an agent. First, the flow configuration information is prepared, including rules such as the start time, running time, and congestion control of the network flow. The flow generator generates network flows based on the flow configuration information and hands them over to the running module for management. The running module is responsible for the actual operation logic of carrying multiple network flows. The running module collects the status information of network flows (such as throughput, latency, etc.) in real time and transmits it to the controller (the controller is a software program module in each network flow); the observer in the controller first receives and parses these statuses, and then gives the statuses to the executor; based on the status, the executor drives the agent to output actions (such as instructions to adjust the congestion window), which are passed back to the running module and act on the network flow. After the agent executes the action, new statuses and experiences are generated and synchronously fed back to the learner; based on these experiences, the learner updates the strategy of the agent to make subsequent actions more adaptable to the network environment, forming a closed-loop optimization of "status collection → action execution → experience learning → strategy update". Simply put, starting from the flow configuration information, network flows are generated and run. Through the observer and executor in the controller, the agent is linked to operate the flow, and then the learner continuously optimizes the agent to achieve dynamic and intelligent control of the network flow.
[0069] As Figure 4 , the mechanism in this application is based on the actor-critic reinforcement learning framework. The core is the collaborative optimization of local decision-making + global evaluation. Among them, the actor network focuses on local states and is responsible for quickly generating "currently available" actions (such as adjusting the congestion window of a single network flow), making decisions based on locally trained strategies to ensure the basic response speed. The critic network integrates the global state + the aggregated local state, evaluates the quality of actions from a "global perspective", and outputs "rewards" to guide the direction of policy optimization.
[0070] Figure 4 The process logic in 1. State input: Local state (such as the throughput and latency of a single network flow) → given to the actor network to make local decisions; global state (the overall data of all network flows) + aggregated local state (the summary of local information of multiple network flows) → given to the critic network to make global evaluations; 2. Action Network Decision: Based on a stable local training strategy, the action network processes the local state and directly outputs specific actions (such as increasing / decreasing the congestion window), quickly responding to network changes; 3. Critic Network Evaluation: The critic network takes the global state + the aggregated local state, comprehensively judges the long-term and global impacts of the actions of the action network, calculates the reward. A high reward indicates that the action improves the overall performance of the network; a low reward indicates the opposite.
[0071] 4. Policy Optimization: The action network implements the local action into the network flow, generating a new state; the critic network uses the global reward to reverse-guide the action network, adjusting its local policy, making the subsequent actions of the action network more in line with the global optimal goal, realizing the coordination of local decision-making and global optimization; The core of the design is to balance efficiency and global perspective. The action network makes quick decisions based on the local state to ensure real-time response of the network; the critic network uses the global perspective to ensure that the overall network is not imbalanced due to local optimality (for example, a certain network flow occupies a large amount of bandwidth crazily, squeezing other network flows). The dynamic optimization closed-loop, action execution → global evaluation → policy update iterates, enabling the action network to gradually learn actions that adapt to both local scenarios and global interests, continuously optimizing network congestion control. Among them, the relationship between the current action and the next action is: , where a t is the action at time t, a t+1 is the action at time t+1, cwnd t is the congestion window parameter at time t, cwnd t+1 is the congestion window parameter at time t+1.
[0072] Simply put, it is to enable the action network with fast local response to gradually learn smarter network control strategies under the guidance of the critic network with global supervision and evaluation, balancing real-time performance and overall performance.
[0073] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.
[0074] Such as Figure 5 , the embodiments of the present application also provide an electronic device, including a memory 101 and a processor 102. The memory 101 stores a computer program, and the processor 102 is configured to run the computer program to execute the steps in any one of the above embodiments of the network flow management method. For the introduction of the electronic device, please refer to the above embodiments, and the present application will not repeat it here.
[0075] Such as Figure 6An embodiment of the present application further provides a computer-readable storage medium 201, in which a computer program 202 is stored. The computer program 202 is configured to execute the steps in any of the above-described method embodiments for managing network traffic when running.
[0076] In an exemplary embodiment, the above computer-readable storage medium 201 may include, but is not limited to: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs that can store computer programs. For the introduction of the computer-readable storage medium 201, please refer to the above embodiments, and the present application will not repeat it here.
[0077] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps in any of the above-described method embodiments for managing network traffic.
[0078] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps in any of the above-described method embodiments for managing network traffic.
[0079] For the introduction of the computer program product, please refer to the above embodiments, and the present application will not repeat it here.
[0080] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0081] The above has provided a detailed introduction to a method, an electronic device, a storage medium, and a program product for managing network traffic provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for managing network traffic, characterized in that, Including: Generate and start multiple network flows according to the flow configuration information; Collect the packet statistical information of the multiple network flows at preset monitoring time intervals; The packet statistical information characterizes the network status of the current network flow; Construct a global state according to the multiple packet statistical information, and construct a global reward function according to the global state; Optimize the global reward function using a preset learning algorithm to obtain the optimal congestion window control strategy for each network flow; Generate action values according to the optimal congestion window control strategy, and adjust the congestion window parameters of each network flow according to the action values.
2. The network flow management method according to claim 1, wherein Constructing a global reward function according to the global state includes: Calculate throughput metrics, latency metrics, packet loss rate metrics, fairness metrics, and stability metrics according to the global state; the throughput metric is used to characterize the network bandwidth utilization rate, the latency metric is used to characterize the network transmission delay, the packet loss rate metric is used to characterize the network transmission reliability, the fairness metric is used to characterize the fairness of the bandwidth allocation of each network flow, and the stability metric is used to characterize the fluctuation degree of the bandwidth allocation of each network flow; Couple the throughput metric, the packet loss rate metric, the latency metric, the fairness metric, and the stability metric to obtain the global reward function.
3. The method for managing network flows according to claim 2, wherein Calculating the throughput metric according to the global state includes: Calculate the throughput metric according to the current throughput and link bandwidth of each network flow; the global state includes the current throughput and link bandwidth of each network flow.
4. The method for managing network flow according to claim 3, wherein, Calculating the throughput metric according to the current throughput and link bandwidth of each network flow includes: According to calculate the throughput metric; where R thr is the throughput metric, i is the sequence number of the network flow, and thr i is the current throughput of the i-th network flow, and c is the link bandwidth.
5. The management method of network flow according to claim 2, characterized in that Calculating the packet loss rate metric according to the global state includes: Calculate the packet loss rate metric according to the packet loss rate and current throughput of each network flow; The global state includes the packet loss rate and current throughput of each network flow.
6. The method for managing network flow according to claim 5, wherein Calculating the packet loss rate metric according to the packet loss rate and current throughput of each network flow includes: According to calculate the packet loss rate index; Wherein, R loss is the packet loss rate index, n is the number of all the network flows, i is the serial number of the network flow, loss i is the packet loss rate of the i-th network flow, and thr i is the current throughput of the i-th network flow.
7. The network flow management method according to claim 2, wherein Calculating the latency metric according to the global state includes: Calculate the latency metric according to the transmission delay, reference delay, and data transmission rate of each network flow; The global state includes the transmission delay, reference delay, and data transmission rate of each network flow.
8. The management method of network flow according to claim 7, characterized in that Calculating the latency metric according to the transmission delay, reference delay, and data transmission rate of each network flow includes: According to calculate the delay metric; Among them, R lat is the delay index, n is the number of all the said network flows, i is the serial number of the said network flow, lat i is the transmission delay of the i-th network flow, d0 is the reference delay, P rate is the data sending rate, is the tolerance coefficient.
9. The method for managing network traffic according to any one of claims 2-8, characterized in that, Calculating the fairness metric according to the global state includes: For each network flow, obtain the historical throughput of w preset monitoring time intervals, calculate the average throughput according to the w historical throughputs, where w is an integer greater than 1; the global state includes the historical throughput of w preset monitoring time intervals; Calculate the fairness metric according to the average throughput.
10. The method for managing network flow according to claim 9, wherein Calculating the fairness metric according to the average throughput includes: According to , calculate the fairness index; where R fair is the fairness index, n is the number of all the network flows, i is the serial number of the network flow, and avg_thr i is the average throughput.
11. The method for managing network flows according to any one of claims 2-8, characterized in that, Calculating the stability metric according to the global state includes: For each of the network flows, obtain the historical throughputs of w preset monitoring time periods, calculate the average throughput according to the w historical throughputs, where w is an integer greater than 1; the global state includes the historical throughputs of the w preset monitoring time periods; Calculate the stability index according to the average throughput and the w historical throughputs.
12. The network flow management method according to claim 11, wherein Calculating the stability index according to the average throughput and the w historical throughputs includes: According to , calculate the stability index; where R stab is the stability index, n is the number of all the network flows, i is the serial number of the network flow, w is the number of historical throughputs, j is the serial number of the historical throughput, t is the t-th preset monitoring time period, avg_thr i is the average throughput, and thr i,t-j is the historical throughput of the i-th network flow in the (t - j)-th preset monitoring time period.
13. An electronic device, characterized in that, including: A memory for storing a computer program; A processor for implementing the steps of the network flow management method according to any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the network flow management method according to any one of claims 1 to 12 when executed by a processor.
15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the network flow management method according to any one of claims 1 to 12 when executed by a processor.
Citation Information
Patent Citations
Distributed intra-network congestion control method based on QMIX
CN113315715A
Data stream mapping method and device, equipment, storage medium and product
CN118890390A
Congestion window length adjusting method and device and nonvolatile storage medium
CN119155254A
Air-space-ground network congestion control method based on deep reinforcement learning
CN119364423A
Method for automatically regulating explicit congestion notification of data center network based on multi-agent reinforcement learning
US20240080270A1