Network flow management method, electronic device, storage medium and program product

By constructing a global state and reward function and combining it with a learning algorithm to optimize the congestion window control strategy, the problems of insufficient resource utilization and fairness of traditional congestion control algorithms in a multi-flow competition environment are solved, and efficient and fair allocation of network resources and performance improvement are achieved.

CN120378364BActive Publication Date: 2025-09-12SHANDONG YINGXIN COMP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510856428.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-12
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

Traditional congestion control algorithms cannot effectively utilize network resources in high-latency, high-bandwidth network environments, resulting in degraded network performance and the inability to achieve fair bandwidth allocation in a multi-flow competition environment.

Method used

By generating and starting multiple network flows, periodically collecting packet statistics, building a global state and designing a global reward function, using a preset learning algorithm to optimize the congestion window control strategy, and dynamically adjusting the congestion window parameters of the network flow, efficient utilization and fair distribution of network resources can be achieved.

Benefits of technology

Maintain good convergence speed, stability and fairness in high-bandwidth, high-latency environments, and improve the overall utilization efficiency of network resources and performance under multi-stream concurrency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378364B_ABST
    Figure CN120378364B_ABST
Patent Text Reader

Abstract

The present application discloses a network flow management method, electronic device, storage medium and program product, and relates to the field of network management technology. After generating and starting multiple network flows, the statistical information of data packets of multiple network flows is periodically collected to fully reflect the current network status. Based on this, a global state is constructed, and a learning algorithm is used to optimize the global reward function to obtain the optimal congestion window control strategy for each network flow. The congestion window parameters of each network flow are dynamically adjusted according to the optimization results, so that it has good convergence speed, stability and fairness in a high-bandwidth, high-latency environment. The present application comprehensively considers the utilization rate of the overall network resources and the bandwidth fairness between network flows, and the coordination and dynamic adjustment between flows. It solves the problem of insufficient fairness and resource utilization in a multi-flow competition environment, and achieves the effect of improving the overall utilization efficiency of network resources and the performance under multi-flow concurrency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network management technology, and in particular to a network flow management method, electronic device, storage medium, and program product. Background Art

[0002] With the rapid development of the internet, the types and number of network applications continue to increase, and the demand for network bandwidth is also growing. Against this backdrop, traditional congestion control algorithms have gradually exposed some shortcomings when dealing with congestion problems in complex network environments. Specifically, in high-latency, high-bandwidth network scenarios, traditional algorithms often fail to effectively utilize network resources, resulting in a significant decline in network performance. Furthermore, because the training environment of traditional algorithms primarily optimizes the performance of a single network flow and fails to fully consider the balanced allocation among multiple network flows, it is difficult for them to maintain consistently good performance in terms of convergence characteristics such as fairness, rapid convergence, and stability. In other words, traditional algorithms' insensitivity to fairness makes it impossible to achieve fair bandwidth allocation in a multi-flow competitive environment, resulting in an imbalance in which some network flows receive excessive bandwidth resources while others struggle to obtain sufficient resources.

[0003] Therefore, providing an algorithm that can balance bandwidth allocation for multiple network flows is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention

[0004] The present application provides a network flow management method, electronic device, storage medium and program product to at least solve the problems of insufficient fairness and resource utilization in a multi-flow competition environment in related technologies.

[0005] The present application provides a method for managing network flows, comprising: generating and starting multiple network flows based on flow configuration information; collecting packet statistical information of the multiple network flows at preset monitoring time periods; the packet statistical information characterizing the network status of the current network flow; constructing a global state based on the multiple packet statistical information, and constructing a global reward function based on the global state; optimizing the global reward function using a preset learning algorithm to obtain an optimal congestion window control strategy for each of the network flows; generating an action value based on the optimal congestion window control strategy, and adjusting the congestion window parameters of each of the network flows based on the action value.

[0006] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned network flow management methods when executing the computer program.

[0007] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned network flow management methods are implemented.

[0008] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned network flow management methods when executed by a processor.

[0009] Through this application, due to the introduction of global state construction and global reward function optimization mechanism, the problems of insufficient fairness and resource utilization of traditional congestion control algorithms in a multi-flow competition environment are effectively overcome. Specifically, after generating and starting multiple network flows, this method periodically collects data packet statistics of each network flow to comprehensively characterize the current network status, thereby building a unified global state view, and designing a global reward function based on the global state. This function not only considers the performance of individual network flows, but also comprehensively considers the utilization of overall network resources and bandwidth fairness between network flows. Furthermore, by optimizing the global reward function using a preset learning algorithm, the optimal congestion window control strategy for multiple network flows can be dynamically obtained, and coordination and dynamic adjustment between flows can be achieved, thereby avoiding the uneven resource allocation phenomenon caused by optimizing only a single flow in traditional algorithms. Ultimately, this method adjusts the congestion window parameters of each network flow in real time based on the action value generated by the optimal strategy, so that the network system can still maintain good convergence speed, stability and fairness in a high-bandwidth, high-latency environment, fundamentally improving the overall utilization efficiency of network resources and performance under multi-flow concurrency, solving the problems of insufficient fairness and resource utilization in a multi-flow competition environment, and achieving the effect of improving the overall utilization efficiency of network resources and performance under multi-flow concurrency. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 A flowchart of a network flow management method provided in an embodiment of the present application;

[0012] Figure 2 Another flowchart of a network flow management method provided in an embodiment of the present application;

[0013] Figure 3 A schematic diagram of a module of a network flow management method provided in an embodiment of the present application;

[0014] Figure 4 A schematic diagram of network training for a network flow management method provided in an embodiment of the present application;

[0015] Figure 5 A schematic diagram of an electronic device provided in an embodiment of the present application;

[0016] Figure 6 A schematic diagram of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION

[0017] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0018] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0019] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0020] like Figure 1 , an embodiment of the present application provides a network flow management method, including:

[0021] S11: Generate and start multiple network flows according to the flow configuration information.

[0022] Specifically, a flow generator can generate and start multiple network flows based on pre-set flow configuration information, aiming to accurately simulate the concurrent communication behavior of multiple flows in a real network in a controlled environment. This configuration information typically includes the start time and run time of each network flow, the type of pre-selected congestion control algorithm, and whether to introduce special parameters such as artificial delay and bandwidth limit, allowing for the simulation of complex network operation scenarios.

[0023] This setup not only enables the flexible construction of representative network load models, but also provides a diverse and controllable foundation for subsequent status monitoring and policy optimization, facilitating the training and evaluation of the adaptability and fairness of different network policies in multi-flow environments. This step ensures that the entire management approach is based on realistic simulation and high configurability from the outset, making subsequent learning and optimization more generalizable and practical.

[0024] S12: Collecting statistical information of data packets of multiple network flows at every preset monitoring time period; the statistical information of data packets represents the network status of the current network flow.

[0025] Specifically, multiple network flows are periodically monitored to obtain packet statistics reflecting the current network operation status. This packet statistics includes, but is not limited to, metrics such as the send rate, packet loss rate, transmission delay, and number of congestion events for each network flow. This provides a comprehensive description of the communication performance and resource usage of each flow in the current network environment.

[0026] By collecting data at preset monitoring intervals, we can dynamically track changing trends in network status, providing real-time, accurate data support for subsequent policy adjustments. This periodic monitoring mechanism not only improves the timeliness of status perception but also provides a multi-dimensional input foundation for building global status, thereby supporting more refined congestion control and resource scheduling policy formulation.

[0027] S13: Build a global state based on multiple packet statistics, and build a global reward function based on the global state.

[0028] Specifically, packet statistics collected from multiple network flows are integrated to construct a global state view that comprehensively reflects the current overall network status. This global state not only includes the individual characteristics of each network flow (such as packet loss rate, latency, and throughput), but also reflects the interactions and competition between network flows, providing a unified picture of overall network resource usage. Compared to traditional methods that rely solely on local information, the global state is more conducive to system-level optimization and regulation.

[0029] Based on the global state, a global reward function is designed based on this state. This global reward function comprehensively considers multiple objectives, such as maximizing overall throughput, minimizing network latency, and ensuring fairness in bandwidth allocation, to measure the performance of the current network control strategy. This reward function quantifies network performance objectives and provides a clear evaluation basis for subsequent optimization using learning algorithms, ensuring that policy adjustments are globally optimal.

[0030] S14: Use the preset learning algorithm to optimize the global reward function and obtain the optimal congestion window control strategy for each network flow.

[0031] Specifically, a pre-defined learning algorithm optimizes the constructed global reward function to find an optimal set of congestion window control policies to achieve efficient utilization and fair allocation of network resources. The learning algorithm can employ reinforcement learning, deep reinforcement learning, or other policy optimization methods suitable for dynamic decision-making. By continuously interacting with the network environment, it learns which control policies, under different global states, yield higher rewards.

[0032] During the optimization process, the algorithm attempts multiple strategies and updates them based on feedback from the global reward function after each adjustment, gradually converging to a control strategy that performs well in complex network environments. This feedback-based adaptive learning approach enables it to dynamically adapt to changes in network conditions, overcoming the limitations of traditional static parameter adjustment or single-flow optimization methods. It achieves the unified optimization goals of bandwidth fairness, congestion relief, and performance improvement from a global perspective.

[0033] S15: Generate an action value according to the optimal congestion window control strategy, and adjust the congestion window parameters of each network flow according to the action value.

[0034] Specifically, the optimal congestion window control policy output by the learning algorithm is converted into specific control instructions, or "action values," and the congestion window parameters of each network flow are dynamically adjusted accordingly. The action value typically indicates the increase or decrease in window size or the direction of adjustment, reflecting the optimal bandwidth usage behavior under the current global state. By mapping the policy into executable parameter adjustment operations, the sending rate of each network flow can be adjusted in real time to respond to changes in current network conditions.

[0035] This policy-driven window adjustment approach not only enables fine-grained control but also exhibits high adaptability, enabling timely adjustments based on fluctuations in network load to avoid increased congestion or bandwidth waste. Furthermore, action values ​​are generated under the guidance of a global reward function, ensuring fairer bandwidth allocation across multiple flows and higher overall system throughput, effectively addressing performance bottlenecks caused by rigid policies in traditional algorithms.

[0036] In an exemplary embodiment, a global reward function is constructed based on the global state, including: calculating a throughput indicator, a delay indicator, a packet loss rate indicator, a fairness indicator, and a stability indicator based on the global state; the throughput indicator is used to characterize the network bandwidth utilization, the delay indicator is used to characterize the network transmission delay, the packet loss rate indicator is used to characterize the network transmission reliability, the fairness indicator is used to characterize the fairness of the bandwidth allocation of each network flow, and the stability indicator is used to characterize the fluctuation degree of the bandwidth allocation of each network flow; the throughput indicator, the packet loss rate indicator, the delay indicator, the fairness indicator, and the stability indicator are coupled to obtain a global reward function.

[0037] Specifically, the global reward function is constructed based on multiple key performance indicators. Its purpose is to comprehensively reflect the current operational quality of the network system and provide an accurate evaluation basis for subsequent learning optimization. First, the global state is used to extract five core network indicators: throughput, latency, packet loss rate, fairness, and stability (a combination of two or more of these can be selected).

[0038] Among them, the throughput indicator is used to measure the efficiency of the entire network's utilization of bandwidth resources. The higher the value, the stronger the network transmission capacity; the delay indicator reflects the time it takes for a data packet to be sent and received. The smaller the delay, the more timely the network response. The packet loss rate indicator measures the data loss during network transmission. It is a key indicator of network reliability. Excessive packet loss rate will cause a serious degradation in the performance of upper-layer applications. The fairness indicator is used to measure the balance of bandwidth resource allocation among various network flows, preventing some flows from occupying too many resources while others are under-resourced. This is a crucial performance factor in a multi-flow competitive environment. The stability indicator focuses on the fluctuation range of bandwidth allocation under the network control policy. Excessive fluctuations will affect the user experience and the predictability of the system. Therefore, stability is also an important evaluation dimension.

[0039] After obtaining the above indicators, this embodiment further weights and couples these indicators to construct a unified global reward function. The value of this function comprehensively reflects the overall advantages and disadvantages of the current network control strategy, and by setting the weights of different indicators, the system can flexibly adjust the optimization direction according to the focus of the actual application scenario. For example, in the video transmission scenario, the weights of throughput and stability can be enhanced, while in interactive services, the control of delay and packet loss rate can be strengthened. In a specific embodiment, the expression of the global reward function is , where R is the reward value of the global reward function, c0, c1, c2, c3, c4 are the weights of the corresponding indicators, R thr is the throughput indicator, R lat is the delay index, R loss is the packet loss rate indicator, R fair is the fairness index, R stabis a stability index. fair and R stab When it is 0, the network flow achieves the best balance between fairness and stability.

[0040] Furthermore, the reward range is constrained for each preset monitoring period and scaled to (-0.1, 0.1). Through a specially designed global reward function, the reinforcement learning training algorithm rewards high throughput, good fairness, and stability, while penalizing high latency and loss rate. Different trade-offs between these metrics can be achieved by adjusting the weight coefficients c0, c1, c2, c3, and c4. The global reward function serves as the overall objective of the congestion control scheme in various network environments. By maintaining a consistent global reward function, the reinforcement learning mechanism can adaptively align the current policy with the performance objectives in different scenarios.

[0041] Through this multi-dimensional indicator fusion method, the global reward function can achieve a trade-off between multiple performance objectives, providing an accurate and comprehensive optimization reference for the learning algorithm, effectively guiding strategy adjustments, thereby improving the overall performance of the system and achieving collaborative optimization of maximizing bandwidth utilization, enhancing fairness between network flows, and improving transmission stability.

[0042] like Figure 2 This is a network congestion control process based on reinforcement learning. The process begins by preparing network flow configuration information as input. Based on this configuration, network flows are generated, and a shared control policy is initialized (the basis for subsequent optimization and adjustment). Packet statistics (such as throughput and latency) for each network flow are collected, and action requests (such as those for adjusting the congestion window) are initiated. These requests are then passed, and packet statistics from all flows are aggregated into a global state (allowing the algorithm to perceive the network from a holistic perspective). Based on this global state, a global reward function is generated to measure network performance (such as fairness and stability) as the optimization objective. Guided by the global reward function, the network flow control policy is trained and the congestion window parameters are updated to make the network more adaptable to the current environment and optimize data transmission. This completes one round of optimization, and subsequent optimization cycles can be repeated. Simply put, this is a closed-loop process that begins with configuring network flows, collects data, builds a global perspective, and then uses the reward function to guide training, ultimately optimizing network transmission.

[0043] In an exemplary embodiment, calculating the throughput indicator according to the global state includes: calculating the throughput indicator according to the current throughput and link bandwidth of each network flow; the global state includes the current throughput and link bandwidth of each network flow.

[0044] Specifically, the throughput metric is calculated based on the current throughput of each network flow in the global state and the bandwidth of the link it is on. Current throughput represents the amount of data successfully transmitted per unit time by the network flow, while link bandwidth represents the theoretical maximum transmission capacity of the network flow. By comparing current throughput with link bandwidth, we can assess how efficiently each network flow utilizes available network resources.

[0045] Specifically, throughput metrics are typically expressed as a ratio or normalized value. For example, dividing the current throughput by the link bandwidth yields a utilization value between 0 and 1. The utilization of multiple network flows can be further averaged or weighted to produce a throughput metric for the entire system, reflecting the overall utilization of network resources. This metric provides a key reference for optimization algorithms to assess network utilization efficiency, helping to improve overall bandwidth utilization and avoid resource waste.

[0046] In an exemplary embodiment, the throughput index is calculated based on the current throughput and link bandwidth of each network flow, including: Calculate the throughput index; where R thr is the throughput index, i is the sequence number of the network flow, thr i is the current throughput of the ith network flow, and c is the link bandwidth.

[0047] Throughput is calculated by dividing the total throughput of all current network flows by the total bandwidth of the network link. Specifically, the throughput of all network flows at the current moment is summed, then divided by the maximum bandwidth of the link to produce a ratio between 0 and 1. This ratio reflects the overall efficiency of network bandwidth utilization. A value closer to 1 indicates full utilization of network resources and higher throughput.

[0048] In this way, the throughput metric can intuitively quantify the network's transmission capacity and efficiency in its current state. As a key component of the global reward function, it provides effective performance feedback for optimizing congestion control strategies and helps the system achieve more efficient bandwidth allocation and utilization.

[0049] In an exemplary embodiment, calculating the packet loss rate indicator according to the global state includes: calculating the packet loss rate indicator according to the packet loss rate and current throughput of each network flow; the global state includes the packet loss rate and current throughput of each network flow.

[0050] Specifically, the packet loss rate metric is calculated based on the packet loss rate and current throughput information for each network flow in the global state. The packet loss rate reflects the proportion of data packets lost during transmission and is an important indicator of network transmission reliability; the current throughput, on the other hand, indicates the actual amount of data transmitted by the network flow. Combining these two metrics allows for a more accurate assessment of the impact of packet loss on network performance.

[0051] Specifically, when calculating the packet loss ratio, we not only focus on the packet loss percentage of each flow, but also weight the impact of packet loss based on its throughput, making packet loss in flows with higher transmission volumes more significant on the overall metric. This allows the packet loss ratio to fully reflect the impact of packet loss on overall transmission performance, providing targeted optimization directions for congestion control algorithms and helping to improve network reliability and transmission quality.

[0052] In an exemplary embodiment, the packet loss rate indicator is calculated based on the packet loss rate and current throughput of each network flow, including: Calculate the packet loss rate indicator; where R loss is the packet loss rate indicator, n is the number of all network flows, i is the sequence number of the network flow, loss i is the packet loss rate of the i-th network flow, thr i is the current throughput of the i-th network flow.

[0053] The packet loss rate metric is calculated by calculating the ratio of the packet loss rate of all network flows to their current throughput, then averaging these ratios. Specifically, for each network flow, its packet loss rate is combined with the current throughput to reflect the actual impact of packet loss on overall transmission. These ratios are then averaged across all flows to produce a comprehensive packet loss rate metric.

[0054] This calculation method takes into account both the packet loss rate and the flow transmission scale, making packet loss in flows with higher throughput have a greater impact on the metric, reflecting the true impact of packet loss on overall network performance. This metric allows for a more accurate assessment of network transmission reliability, providing an effective basis for optimizing congestion control strategies.

[0055] In an exemplary embodiment, calculating the delay index based on the global state includes: calculating the delay index based on the transmission delay, baseline delay and data sending rate of each network flow; the global state includes the transmission delay, baseline delay and data sending rate of each network flow.

[0056] Specifically, latency metrics are calculated based on the global state of each network flow's transmission delay, baseline latency, and data sending rate. Transmission delay represents the time it takes for a data packet to travel from sender to receiver, while baseline latency represents the ideal or expected minimum latency. The difference between the two reflects the network's additional latency burden. Combined with the sending rate, this more accurately measures the impact of latency on actual data transmission performance.

[0057] Specifically, the latency metric compares the deviation of each network flow's transmission delay from a baseline delay and weights it according to the flow's data transmission rate. This emphasizes the impact of high-rate flow latency on overall network performance. The resulting latency metric comprehensively reflects the impact of network latency on transmission efficiency, providing a valuable reference for optimizing congestion control strategies, reducing network response times, and improving user experience.

[0058] In an exemplary embodiment, the delay index is calculated based on the transmission delay of each network flow, the benchmark delay and the data sending rate, including: Calculate the delay index; where R lat is the delay index, n is the number of all network flows, i is the sequence number of the network flow, lat i is the transmission delay of the i-th network flow, d0 is the reference delay, P rate is the data transmission rate, and β is the tolerance coefficient.

[0059] Specifically, the delay indicator measures the deviation between the average transmission delay of the network flow and the benchmark delay d0, and combines the packet sending rate P rate Specifically, when all network flows enter a stable transmission state, they are allowed to reach the maximum throughput in the dynamic link. However, when the average delay exceeds the threshold (1+β)d0, the delay expansion behavior will be punished. Otherwise, R lat A value of 0 does not penalize normal or slightly increased delays, thus avoiding false positives for smaller queue lengths.

[0060] Furthermore, the metric uses the packet sending rate as a multiplier, meaning that for flows with high sending rates, if latency increases significantly, the penalty will be more severe. This design effectively discourages attempts to increase bandwidth by increasing the sending rate on high-latency links, preventing worsening network congestion and promoting a balance between latency and throughput, thereby improving overall network performance and user experience.

[0061] In an exemplary embodiment, calculating a fairness indicator based on a global state includes: for each network flow, obtaining historical throughputs for w preset monitoring time periods, and calculating an average throughput based on the w historical throughputs; w is an integer greater than 1; the global state includes the historical throughputs for the w preset monitoring time periods; and calculating the fairness indicator based on the average throughput.

[0062] Specifically, the fairness index is calculated based on the historical throughput data of each network flow over multiple preset monitoring time periods. Specifically, the throughput of each flow over w consecutive time periods is collected and the average of these data is calculated to reflect the stable transmission capacity and resource utilization of the flow over a period of time.

[0063] By comparing the average throughput of all network flows, we can assess the fairness of bandwidth allocation. Specifically, we can determine whether each flow receives a relatively balanced share of resources. A smaller fairness index value indicates more balanced bandwidth allocation, helping to avoid situations where some flows chronically consume excessive resources while others starve for bandwidth, thereby promoting harmonious and stable operation of the overall network.

[0064] In an exemplary embodiment, calculating the fairness index according to the average throughput includes: , calculate the fairness index; where R fair is the fairness index, n is the number of all network flows, i is the sequence number of the network flow, avg_thr i is the average throughput.

[0065] Specifically, the fairness metric measures the balance of bandwidth allocation based on the standard deviation of the throughput of all active network flows within the same time period. Specifically, the average throughput of each flow over the past w monitoring time periods is calculated to account for errors caused by transient fluctuations. The standard deviation of these average throughputs is then calculated. A smaller standard deviation indicates a more even distribution of throughput across flows and a more equitable bandwidth allocation.

[0066] Furthermore, the denominator of the standard deviation in the formula includes the sum of the squares of the average throughput of all flows. This normalization makes the metric scale-independent and adaptable to networks with varying bandwidth sizes. By averaging the fairness metric over time, we can consistently assess the fairness trend of bandwidth allocation, helping the congestion control algorithm maintain balance among multiple flows during optimization and preventing uneven resource allocation from causing some flows to be restricted for extended periods.

[0067] In an exemplary embodiment, a stability index is calculated based on a global state, including: for each network flow, obtaining historical throughputs for w preset monitoring time periods, and calculating an average throughput based on the w historical throughputs; w is an integer greater than 1; the global state includes the historical throughputs for the w preset monitoring time periods; and calculating a stability index based on the average throughput and the w historical throughputs.

[0068] Specifically, the stability index is used to measure the fluctuation of the throughput of each network flow over time. First, for each network flow, historical throughput data from the past w preset monitoring time periods is collected and the average of these data is calculated to reflect the stable transmission level of the flow over a period of time.

[0069] Next, the throughput fluctuation is quantified by comparing the differences between these w historical throughputs and the average throughput (e.g., calculating the standard deviation or variance). Smaller fluctuations indicate more stable network flow transmission; larger fluctuations indicate unstable bandwidth allocation or network conditions.

[0070] As part of the global reward function, the stability indicator helps guide the congestion control strategy to improve throughput while reducing transmission uncertainty and fluctuations, ensuring the continuity of network performance and the smoothness of user experience.

[0071] In an exemplary embodiment, calculating the stability index based on the average throughput and w historical throughputs includes: , calculate the stability index; among them, R stab is the stability index, n is the number of all network flows, i is the sequence number of the network flow, w is the number of historical throughputs, j is the sequence number of historical throughputs, t is the tth preset monitoring time period, avg_thr i is the average throughput, thr i,t-j is the historical throughput of the i-th network flow in the tj-th preset monitoring time period.

[0072] Specifically, the stability index measures the stability of bandwidth allocation based on the standard deviation of the throughput of each network flow in a fixed-length historical monitoring period w. Specifically, for each network flow i, its historical throughput thr in the past w time periods is first calculated. i,t−j , thr i,t-j The average throughput of the network flow avg_thr i The sum of squared deviations between the two is normalized (divided by w and the square of the average throughput) to eliminate the impact of throughput size on the degree of fluctuation.

[0073] The normalized variances of all network flows are then averaged to produce an overall stability index. A smaller stability index indicates more stable network flow throughput and less fluctuation; a larger stability index indicates significant traffic fluctuations and unstable network performance. By averaging the stability index per flow, transient anomalies can be avoided, resulting in a smoother and more reliable index, helping congestion control algorithms achieve more stable and consistent bandwidth allocation.

[0074] like Figure 3 This process describes an agent-based network flow control mechanism. Each network flow can be considered an agent. Flow configuration information is first prepared, including the flow's start time, run time, congestion control rules, and so on. The flow generator generates network flows based on the flow configuration information and transfers it to the execution module, which is responsible for implementing the operational logic of multiple network flows. The execution module collects network flow status information (such as throughput and latency) in real time and transmits it to the controller (a software module within each network flow). The observer within the controller first receives and interprets this status information and then passes it to the executor. Based on this status information, the executor drives the agent to output an action (such as an instruction to adjust the congestion window), which is then transmitted back to the execution module to act on the network flow. After the agent executes an action, it generates new state and experience, which is synchronously fed back to the learner. Based on this experience, the learner updates the agent's strategy to make subsequent actions more adaptable to the network environment, forming a closed-loop optimization loop of "state collection → action execution → experience learning → strategy update." Simply put, it starts with the flow configuration information, generates and runs the network flow, links the intelligent agent operation flow through the observer and executor in the controller, and then relies on the learner to continuously optimize the intelligent agent to achieve dynamic and intelligent control of the network flow.

[0075] like Figure 4 The mechanism in this application is based on the action-critic reinforcement learning framework. Its core is the collaborative optimization of local decision-making and global evaluation. The action network focuses on local state and is responsible for quickly generating "currently available" actions (such as adjusting the congestion window of a single network flow). It relies on locally trained policies to make decisions and ensure basic responsiveness. The critic network integrates the global state and aggregated local state, evaluates the quality of actions from a "global perspective," and outputs "rewards" to guide policy optimization.

[0076] Figure 4 The process logic in is:

[0077] 1. State input: Local state (such as throughput and latency of a single network flow) → given to the action network for local decision-making; global state (overall data of all network flows) + aggregated local state (summary of local information of multiple network flows) → given to the critic network for global evaluation;

[0078] 2. Action network decision-making: The action network processes local states based on a stable local training strategy and directly outputs specific actions (such as increasing / decreasing the congestion window), quickly responding to network changes.

[0079] 3. Critic Network Evaluation: The Critic Network receives the global state and the aggregated local state, comprehensively judges the long-term and global impact of the action network's actions, and calculates the reward. A high reward indicates that the action improves the overall performance of the network; a low reward indicates the opposite.

[0080] 4. Strategy Optimization: The action network implements local actions into the network flow, generating a new state. The critic network uses global rewards to reversely guide the action network, adjusting its local strategy so that subsequent actions of the action network are more aligned with the global optimal goal, achieving synergy between local decision-making and global optimization.

[0081] The core of the design is to balance efficiency and global perspective. The action network relies on local state to make rapid decisions and ensure real-time network response. The critic network uses a global perspective to check and prevent local optimality from causing overall network imbalance (for example, a network flow excessively occupies bandwidth, squeezing out other network flows). The dynamic optimization closed loop, action execution → global evaluation → policy update, iterates in a cycle, allowing the action network to gradually learn actions that are both suitable for local scenarios and in line with global interests, continuously optimizing network congestion control. The relationship between the current action and the action at the next moment is: , where a t is the action at time t, a t+1 is the action at time t+1, cwnd t is the congestion window parameter at time t, cwnd t+1 is the congestion window parameter at time t+1.

[0082] Simply put, it allows the local fast-response action network to gradually learn smarter network control strategies under the guidance of the global supervisory evaluation critic network, balancing real-time and overall performance.

[0083] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0084] like Figure 5 An embodiment of the present application further provides an electronic device, including a memory 101 and a processor 102. The memory 101 stores a computer program, and the processor 102 is configured to execute the computer program to perform the steps of any of the above-mentioned network flow management method embodiments. For an introduction to the electronic device, please refer to the above-mentioned embodiments, and this application will not elaborate on them here.

[0085] like Figure 6 An embodiment of the present application further provides a computer-readable storage medium 201 , in which a computer program 202 is stored. The computer program 202 is configured to execute the steps of any of the above-mentioned network flow management method embodiments when running.

[0086] In an exemplary embodiment, the computer-readable storage medium 201 may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a removable hard drive, a magnetic disk, or an optical disk. For an introduction to the computer-readable storage medium 201, please refer to the above embodiment; this application will not elaborate further here.

[0087] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned network flow management method embodiments.

[0088] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, implementing the steps in any of the above-mentioned network flow management method embodiments.

[0089] For an introduction to the computer program product, please refer to the above embodiments, which will not be described in detail in this application.

[0090] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0091] The above is a detailed introduction to a network flow management method, electronic device, storage medium and program product provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A method for managing network flow, characterized in that: include: Generate and start multiple network flows according to flow configuration information; Collecting statistical information of data packets of a plurality of network flows at every preset monitoring time period; The data packet statistical information represents the network status of the current network flow; Constructing a global state based on the plurality of data packet statistics, and constructing a global reward function based on the global state, wherein the global reward function is used to measure network performance, and the global state includes a fairness indicator and a stability indicator; Optimizing the global reward function using a preset learning algorithm to obtain an optimal congestion window control strategy for each of the network flows; generating an action value according to the optimal congestion window control strategy, and adjusting the congestion window parameter of each of the network flows according to the action value; The method of optimizing the global reward function using a preset learning algorithm to obtain the optimal congestion window control strategy for each network flow includes: training the control strategy for the network flow using the global reward function as a guide, updating the congestion window parameters to make the network more adaptable to the current environment and optimize data transmission, completing one round of optimization process, and performing continuous optimization in subsequent cycles; Furthermore, the network flow management method further includes: The agent-based network flow control mechanism prepares flow configuration information. The flow generator generates network flows based on the flow configuration information and hands them over to the operation module for management. The operation module carries the operation logic of multiple network flows. Each network flow is an agent. The flow configuration information includes the start time, running time and congestion control of the network flow. The running module collects the status of the network flow in real time and passes it to the controller; the observer in the controller receives and parses the status and then gives it to the executor. The executor drives the intelligent agent to output action values ​​based on the status, and returns it to the running module to act on the network flow; after the intelligent agent executes the action value, it generates a new state and experience, which is synchronously fed back to the learner so that the learner can update the intelligent agent's strategy based on the experience and make the subsequent action values ​​adapt to the network environment, forming a closed-loop optimization of state collection, action execution, experience learning and strategy update.

2. The network flow management method according to claim 1, characterized in that: Constructing a global reward function based on the global state, including: Calculate the throughput index, delay index, packet loss rate index, fairness index and stability index based on the global state; the throughput index is used to characterize the network bandwidth utilization, the delay index is used to characterize the network transmission delay, the packet loss rate index is used to characterize the network transmission reliability, the fairness index is used to characterize the fairness of the bandwidth allocation of each network flow, and the stability index is used to characterize the fluctuation of the bandwidth allocation of each network flow; The throughput indicator, the packet loss rate indicator, the delay indicator, the fairness indicator, and the stability indicator are coupled to obtain the global reward function.

3. The network flow management method according to claim 2, characterized in that: Calculating a throughput indicator based on the global state includes: The throughput indicator is calculated according to the current throughput and link bandwidth of each network flow; the global state includes the current throughput and link bandwidth of each network flow.

4. The network flow management method according to claim 3, characterized in that: Calculating the throughput indicator according to the current throughput and link bandwidth of each of the network flows includes: according to Calculating the throughput indicator; in, is the throughput index, i is the sequence number of the network flow, is the current throughput of the ith network flow, and c is the link bandwidth.

5. The network flow management method according to claim 2, characterized in that: Calculating a packet loss rate indicator according to the global state includes: Calculating the packet loss rate indicator according to the packet loss rate and current throughput of each of the network flows; The global state includes the packet loss rate and current throughput of each of the network flows.

6. The network flow management method according to claim 5, characterized in that: Calculating the packet loss rate indicator according to the packet loss rate and current throughput of each of the network flows includes: according to Calculating the packet loss rate indicator; in, is the packet loss rate indicator, n is the number of all the network flows, i is the sequence number of the network flow, is the packet loss rate of the i-th network flow, is the current throughput of the i-th network flow.

7. The network flow management method according to claim 2, characterized in that: Calculating a delay indicator according to the global state includes: Calculating the delay index according to the transmission delay of each of the network flows, the benchmark delay and the data sending rate; The global state includes a transmission delay, a reference delay, and a data sending rate of each of the network flows.

8. The network flow management method according to claim 7, characterized in that: Calculating the delay index according to the transmission delay, the benchmark delay, and the data sending rate of each of the network flows includes: according to Calculating the delay indicator; in, is the delay index, n is the number of all the network flows, i is the sequence number of the network flow, is the transmission delay of the i-th network flow, is the baseline delay, is the data sending rate, is the tolerance factor.

9. The network flow management method according to any one of claims 2 to 8, characterized in that: Calculating a fairness indicator based on the global state, including: For each of the network flows, obtaining historical throughputs of w preset monitoring time periods, and calculating an average throughput based on the w historical throughputs, where w is an integer greater than 1; the global state includes the historical throughputs of the w preset monitoring time periods; The fairness indicator is calculated according to the average throughput.

10. The network flow management method according to claim 9, characterized in that: Calculating the fairness indicator according to the average throughput includes: according to , calculating the fairness index; in, is the fairness index, n is the number of all the network flows, i is the sequence number of the network flow, is the average throughput.

11. The network flow management method according to any one of claims 2 to 8, characterized in that: Calculating the stability index according to the global state includes: For each of the network flows, obtaining historical throughputs of w preset monitoring time periods, and calculating an average throughput based on the w historical throughputs, where w is an integer greater than 1; the global state includes the historical throughputs of the w preset monitoring time periods; The stability index is calculated according to the average throughput and w historical throughputs.

12. The network flow management method according to claim 11, characterized in that: Calculating the stability index according to the average throughput and w historical throughputs includes: according to , calculate the stability index; in, is the stability index, n is the number of all network flows, i is the sequence number of the network flow, w is the number of historical throughputs, j is the sequence number of historical throughputs, t is the tth preset monitoring time period, is the average throughput, is the historical throughput of the i-th network flow in the tj-th preset monitoring time period.

13. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the network flow management method according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the network flow management method according to any one of claims 1 to 12.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the network flow management method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Data stream mapping method and device, equipment, storage medium and product

    CN118890390A

  • Congestion window length adjusting method and device and nonvolatile storage medium

    CN119155254A