Data center network congestion control method based on data model dual drive
By combining deep reinforcement learning and traditional model mapping, the contradiction between data model congestion control in data center networks is solved, and rapid response and flexible adaptation are achieved, and network performance and stability are significantly optimized.
Patent Information
- Application Number
- CN202510173468.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-17
AI Technical Summary
The existing data center network congestion control schemes have contradictions between dynamic response and network adaptability, insufficient response speed and limited adaptability, making it difficult to effectively optimize network resource utilization and performance in high-load scenarios such as AI large-scale model training.
Combining the adaptability of deep reinforcement learning (DRL) technology and the fast response capability of traditional model mapping, a dual-driven congestion control method is designed. This method uses the DRL algorithm to optimize dynamic adjustable parameters by constructing a system framework including optimization engine module, statistical information module and rate adjustment control module, and adjusts the transmission rate in real time through parameterized rate adjustment models to adapt to complex network conditions.
It realizes rapid response and flexible adaptation in a dynamic network environment, significantly optimizes bandwidth utilization, reduces communication delay and queue length, improves network throughput, and enhances the stability of data center networks. It is especially suitable for high-load scenarios such as AI large model training.
Smart Images

Figure CN120017597A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer network technology, and specifically to a network congestion control method and system suitable for data center networks, in particular to a fusion congestion control algorithm that combines traditional model mapping and deep reinforcement learning (DRL) for optimizing network resource utilization, improving throughput, reducing latency, and ensuring network stability and performance for large-scale AI model training. Background Art
[0002] With the rapid development of generative artificial intelligence technology, especially the widespread application of large models such as ChatGPT, the research and development of large AI models has become a hot topic at the forefront of science and technology. This trend has promoted the upgrade of data center computing facilities, but it has also brought new challenges. First, large model training requires high parallel computing and frequent data exchange, which can accelerate computing, but it is easy to cause network congestion, reduce resource utilization and training efficiency; at the same time, training has an urgent need for high bandwidth, but large-scale data transmission will further reduce bandwidth utilization. In addition, high data throughput and complexity pose challenges to network stability. Network congestion or packet loss will prolong task processing time and affect system reliability. Therefore, in order to solve the severe challenges facing data center networks, an efficient congestion control strategy is urgently needed to optimize resource utilization, improve network performance and support large-scale model training.
[0003] Traditional congestion control schemes (such as DCTCP) provide a basis for the research of data center network transmission protocols, but their simple control strategies have significant limitations in optimizing queue length and improving throughput. Improved schemes such as DCQCN, although they have enhanced adaptability to network status, still exhibit slow response speed, high flow delay, and complex parameter adjustment in dynamic environments, which increases the difficulty of actual deployment. Other schemes such as Swift and HPCC, although they have excellent performance, rely on complex hardware configurations and are difficult to promote and apply in ordinary commercial equipment. In addition, these schemes use a fixed event-action mapping mechanism, which lacks flexibility and is difficult to adapt to rapid changes in network conditions.
[0004] In recent years, data-driven congestion control schemes (such as Aurora and AuTO) have shown great potential through DRL technology. These schemes achieve real-time optimization in dynamic environments by learning empirical data in the network, and have excellent generalization and adaptability. Compared with traditional schemes, they reduce the reliance on manual parameter adjustment and can adapt to complex and changing network conditions. However, these schemes may lead to excessive or insufficient utilization of network resources due to improper resource allocation under unknown network conditions. In addition, the DRL algorithm still faces challenges in convergence speed and stability, especially in high-frequency dynamic bandwidth change scenarios, its response speed is slow, which may affect the actual performance.
[0005] In the complex network environment of AI large model training, congestion control of data center networks faces more stringent challenges, especially in scenarios such as dynamic bandwidth changes and frequent flow arrival and departure. Congestion control solutions in AI large model training need to have two capabilities: dynamic response capability and network adaptability. The former requires the solution to quickly adjust the sending rate to cope with frequent changes in network status; the latter requires the solution to flexibly adjust the strategy according to changes in link and traffic patterns. Model mapping-based congestion control solutions have good dynamic response capabilities with a simple signal-action mapping mechanism, but lack adaptability in frequently changing network environments. In contrast, data learning-based solutions, especially those based on DRL technology, show strong network adaptability and can flexibly respond to dynamic changes. However, its high computational overhead and slow response limit its application in high-frequency dynamic scenarios. Therefore, in the environment of AI large model training, how to strike a balance between dynamic response and network adaptation has become a key challenge in designing efficient congestion control.
[0006] The present invention provides a data center network congestion control solution based on dual-drive of data model, which combines the flexible adaptability of DRL and the rapid response capability of traditional model mapping. This method can effectively solve the contradiction between dynamic response and network adaptation in existing congestion control solutions, and provide an efficient and flexible solution for data center networks, which is particularly suitable for high-load scenarios such as AI large model training. Summary of the invention
[0007] Purpose of the invention: The present invention provides a network congestion control method and system based on dual-drive of data models. By combining the adaptability of DRL technology with the rapid response capability of traditional model mapping, the problems of insufficient response speed and limited adaptability of the prior art in a dynamic network environment are solved, thereby optimizing network resource utilization, improving the throughput and stability of data center networks, effectively reducing flow completion time and communication delays, and meeting application scenarios with high bandwidth and high performance requirements such as AI large model training.
[0008] Technical solution: A data center network congestion control method based on dual-drive of data models, the method comprising:
[0009] S1. Construct a system framework including an optimization engine module, a statistical information module and a rate regulation control module; the optimization engine module optimizes the congestion control sub-strategy in real time through an optimization algorithm, the statistical information module is responsible for collecting network operation data, and provides support for the decision-making process and system monitoring of the optimization engine. The optimization engine module and the statistical information module together constitute an RL agent, and the regulation control module is used to execute the sub-strategy generated by the RL agent;
[0010] S2. Construct a network-side rate parameterized rate adjustment model, specifically including:
[0011] S21. Calculate the convergence result through network parameters and queuing delay as a reference for adjusting the rate. The calculation formula is as follows:
[0012]
[0013] d q =rtt new -min{rtt new ,rtt min}
[0014] Where δ is a tuning parameter, and 0.5≤δ≤1.5, which is used to represent the state of the total network bandwidth and transmission demand, d q It is the queuing delay, which is obtained by performing serialization delay and minimum RTT processing on the RTT feedback;
[0015] The rate update cycle needs to be adaptively adjusted according to network demand. The parameter δ reflects the network demand and is determined by the link bandwidth and the number of connections. The update cycle is designed as:
[0016] T a =δ×K
[0017] T++
[0018] Where T a is the adaptive byte counter threshold, K represents the link bandwidth, and T++ represents the update of the byte counter. To improve the convergence speed and fairness, the adaptive byte counter is designed to be independent of the transmission rate, that is:
[0019] BC a =γ×R C
[0020] BC++
[0021] Among them BC a is the adaptive byte counter threshold, γ is the adjustment parameter, and 0.5≤γ≤1.5, R C is the current transmission rate;
[0022] S22. Design a mechanism to adjust the model rate reduction, as follows:
[0023]
[0024] Where R T Indicates the target transmission rate that needs to be adjusted, R CIndicates the current end-side transmission rate. Parameters μ1 and μ2 are adjustment parameters, and 2<μ1≤μ2≤4. These parameters are used to control the aggressiveness of rate reduction when congestion occurs. If CNP is not received within the specified time, α needs to be updated, where α is the adjustment factor and g is a fixed value used to adjust α.
[0025] α=(1-g)α+g
[0026] S23. Design a rate increase mechanism for the adjustment model, where the rate increase includes a rapid recovery phase, an adaptive increase phase, and a super increase phase;
[0027] Fast recovery phase: When both the timer and the byte counter are less than the threshold, that is, T<F and BC<F, a fast recovery operation is performed, where μ3 is an adjustment parameter and 6≤μ3≤10. The design is as follows:
[0028]
[0029] Adaptive increase phase: When there is a timer or byte counter greater than the threshold, that is, T≥F or BC≥F, an adaptive additive speed-up operation is performed; the adaptive growth step size is related to the current transmission rate and link bandwidth, but has nothing to do with the number of carried flows. The design is as follows:
[0030]
[0031] Where R L is the link capacity, β and μ3 are adjustment parameters and 0<β<1;
[0032] Super-increment phase: When both the timer and the byte counter are greater than the threshold, that is, T≥F and BC≥F, a super-increment operation is performed, which is designed as follows:
[0033]
[0034] Where i = min(BC,T-F+1), and R HAI is a fixed value of 100Mbps, and μ3 is an adjustment parameter;
[0035] S24. After determining the overall design of the control model, the corresponding control parameters in the network end-side rate parameterization rate adjustment model are set to be dynamically adjustable, which is expressed as:
[0036] θ:={δ,γ,β,μ1,μ2,μ3}
[0037] S3, based on the DRL algorithm, the dynamically adjustable parameters in S23 are optimized, and the congestion control problem in the network is modeled as a Markov decision model, including setting the state quantity and reward function. The description of the network status includes introducing the switch buffer occupancy rate Occ to describe the congestion degree. This indicator is represented by the proportion of packets marked by ECN;
[0038] S4. Based on the DRL algorithm, the state and reward functions in step S3 are designed and solved. The proximal policy optimization is selected as the RL controller, and the PPO algorithm is used for policy optimization based on the Markov decision process.
[0039] Furthermore, the specific process of step S3 includes:
[0040] S31. Set the state quantity required by the DRL algorithm
[0041] The state quantities include the current network throughput, round-trip delay, and packet loss rate, as well as the measurement of the transmission rate, visualizing the availability of the network and bandwidth, where the transmission rate ΔR t =(R t -R t-1 ) / min(R t-1 ,R t );
[0042] S32. Set the reward function required by the DRL algorithm
[0043] This reward function is used to quantify the performance criteria of the task and guide the agent to improve the generated sub-strategy sequence during the training phase. The RL agent collects rewards at each monitoring interval to interact with the network environment;
[0044] The reward function is designed as follows:
[0045]
[0046] in:
[0047]
[0048] Among them, ω1 represents the penalty for packet loss, and ω2, ω3, and ω4 are weighting coefficients.
[0049] Furthermore, the design of the reward function includes considering power = throughput / rtt. Power maximization can reflect the maximum throughput while minimizing network latency. It also includes considering the ratio change of the sending rate and the occupancy of the switch buffer.
[0050] Further, the objective function of the PPO algorithm described in step S4 includes advantage estimation and calculation probability ratio;
[0051] The advantage function is the difference between the expected value of the state of a given action and the expected value of all possible actions in the same state. In order to estimate the advantage function, PPO uses a neural network model to train the value state function, calculates the total reward from the beginning to the expected realization of a given state, and finally uses the generalized advantage estimate through the value of the critic network to perform advantage estimation;
[0052] The probability ratio is calculated as in is the old strategy, π θ is the updated strategy;
[0053] To avoid large-scale policy updates, the PPO algorithm clips the target within the advantage estimate range, defined as follows:
[0054]
[0055] is the estimated advantage at time t, ∈ is a hyperparameter controlling the scope of cropping;
[0056] The optimization function for training a neural network model includes the squared error loss function:
[0057]
[0058] Combining the actor loss with the critic loss and entropy, we can implement the RL agent objective function:
[0059]
[0060] where c1 and c2 are coefficients.
[0061] Furthermore, in each iteration of the PPO algorithm, participants independently collect observation trajectory data including state, action, and reward sequences;
[0062] For a given trajectory, the GAE algorithm calculates an advantage estimate for each time required for policy update;
[0063] Finally, the proxy function is maximized by stochastic gradient descent (SGD) using the collected observation trajectory data and the estimated value.
[0064] Beneficial effects: The present invention proposes a data model dual-driven congestion control method by combining deep reinforcement learning technology and traditional model mapping mechanism, which has a good balance between dynamic response capability and adaptive capability. Its fast signal-action mapping mechanism can adjust the transmission rate in real time, and DRL technology improves adaptability to complex network conditions through dynamic learning. This method can significantly optimize bandwidth utilization, reduce communication delay and queue length, and improve network throughput. It is suitable for high-performance and dynamic and changeable network environments such as AI large model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 This is a schematic diagram of the idea of dual drive of data model;
[0066] Figure 2 It is a data center network congestion control method framework based on dual-drive of data model;
[0067] Figure 3 The diagram is a flowchart of the data center network algorithm execution based on dual-drive of data models. DETAILED DESCRIPTION
[0068] In order to illustrate the technical solution disclosed by the present invention in detail, the present invention is further explained below in conjunction with the accompanying drawings and embodiments.
[0069] (1) Specific architecture and implementation process of the system
[0070] The data center network congestion control system based on data model dual drive proposed in the present invention mainly consists of three parts: rate adjustment control module, optimization engine module and statistical information module. The functions of each module and their mutual assistance relationship are as follows:
[0071] Statistics module: First, it is responsible for collecting network operation data in real time, including key indicators such as throughput, round-trip delay, packet loss rate, and switch buffer occupancy rate, and then pre-processing the data in a standardized manner. Finally, it is also necessary to regularly transmit the processed network status data to the optimization engine module to provide support for subsequent decision-making optimization.
[0072] Optimization engine module: Use deep reinforcement learning algorithm to analyze the network data provided by the statistical information module. Model the network status through Markov decision process and optimize the congestion control strategy based on PPO. Of course, it is also necessary to combine historical experience data and continuously iterate and optimize the congestion control sub-strategy to adapt to the current network conditions.
[0073] Rate adjustment control module: accepts the congestion control strategy output by the optimization engine module and performs corresponding rate adjustment according to the current network status. The module adopts a parameterized rate adjustment model, including four stages: fast recovery, adaptive adjustment, super increase and speed reduction control. The overall ECN marking and RTT feedback are used to dynamically adjust the data flow transmission rate to ensure efficient use of network resources.
[0074] The system workflow of the present invention adopts a closed-loop feedback control mechanism to achieve efficient network congestion control through a continuously iterative evaluation-execution-strategy generation cycle. First, the statistical information module collects network performance data and transmits it to the optimization engine. The optimization engine analyzes this data, optimizes the congestion control sub-strategies using reinforcement learning, and generates new strategies that adapt to the current network status. Subsequently, the control model is regulated to execute these new strategies, maximizing network performance by adjusting the rate. The system continuously monitors changes in network status. Ensure that each decision is based on the latest data, thereby achieving more efficient network congestion control.
[0075] (2) Design and implementation of network-side rate adjustment model
[0076] The present invention calculates the convergence result through parameters and queuing delay as a reference for adjusting the rate.
[0077]
[0078] d q =rtt new -min{rtt new ,rtt min}
[0079] Where δ is a tuning parameter and 0.5≤δ≤1.5, which is used to represent the state of the total network bandwidth and transmission demand. q It is the queuing delay, which is the result obtained by the algorithm after performing serialization delay and minimum RTT processing on the RTT feedback.
[0080] The timer in the system increases by 1 every 55μs. When the accumulated value exceeds the threshold F (the default value is 5), different speed-up stages are triggered. The rate update cycle needs to be adaptively adjusted according to network requirements. The parameter δ reflects the network requirements and is determined by the link bandwidth and the number of connections. Therefore, the update cycle should be designed as:
[0081] T a =δ×K
[0082] T++
[0083] Where T a is the adaptive byte counter threshold. Similarly, the byte counter is set to 150KB. The counter is incremented by 1 after each 150KB packet is received. When the accumulated value exceeds the threshold, different speed-up stages are also triggered. To improve convergence speed and fairness, the adaptive byte counter should be designed to be independent of the transmission rate:
[0084] BC a =γ×R C
[0085] BC++
[0086] Among them, BC a is the adaptive byte counter threshold, γ is a tuning parameter and 0.5≤γ≤1.5, R C is the current transfer rate.
[0087] Design and adjust the model rate reduction mechanism: In order to speed up the convergence speed, when the reference rate is much lower than the current rate, the model will significantly reduce the speed to avoid severe congestion and quickly adjust to the ideal rate. The design is as follows:
[0088]
[0089] α=(1-g)α+g
[0090] The parameters μ1 and μ2 are adjustment parameters and 2<μ1≤μ2≤4. These parameters can control the aggressiveness of the rate reduction when congestion occurs. If CNP is not received within a period of time, α needs to be updated.
[0091] α=(1-g)α+g
[0092] Design and adjust the model rate increase mechanism: The rate increase is mainly divided into three stages: fast recovery stage, adaptive increase stage, and super increase stage.
[0093] The first is the fast recovery phase. When both the timer and the byte counter are less than the threshold, that is, T<F and BC<F, the fast recovery operation is performed, where μ3 is the adjustment parameter and 6≤μ3≤10. The design is as follows:
[0094]
[0095] The second is the adaptive increase phase. When there is a timer or byte counter greater than the threshold, that is, T≥F or BC≥F, an adaptive additive speed-up operation is performed. The adaptive growth step should be related to the current transmission rate and link bandwidth, but not to the number of carried flows. The design is as follows:
[0096]
[0097] Where R L is the link capacity, β and μ3 are adjustment parameters and 0<β<1.
[0098] Finally, in the super-increment phase, when both the timer and the byte counter are greater than the threshold, that is, T ≥ F and BC ≥ F, the super-increment operation is performed, which is designed as follows:
[0099]
[0100] Where i = min(BC,T-F+1), and R HAI is a fixed value of 100Mbps. μ3 is an adjustable parameter.
[0101] After determining the overall design of the control model, the present invention further sets the corresponding control parameters in the model to be dynamically adjustable, which is recorded as:
[0102] θ:={δ,γ,β,μ1,μ2,μ3}
[0103] Specifically, the parameter δ is used to represent the state of the total network bandwidth and transmission demand, to calculate the reference rate, and to provide the direction and amplitude of the rate adjustment, while ensuring that the timer threshold is reasonable. The parameter γ regulates the byte counter threshold to ensure that the rate increase interval is independent of the transmission rate. The parameter β ensures that the adaptive growth step is related to the current transmission rate and link bandwidth. Parameters μ1 and μ2 timely adjust the aggressiveness of the rate reduction during congestion to accelerate convergence. Parameter μ3 uniformly adjusts the conservativeness of the rate increase to prevent re-congestion and accelerate convergence. Finally, the parameter set {δ, γ, β, μ1, μ2, μ3} determines the sending behavior of the model and is defined as the output of the RL agent. The present invention customizes sub-strategies to adapt to the current network conditions by controlling these parameter settings.
[0104] (3) Design and implementation of optimization engine RL agent
[0105] Setting the state quantity required by the algorithm: The state information collected by the RL agent needs to take into account the current network conditions and flow status. Throughput (thr and thr max ), round-trip delay (rtt and rtt min ) and packet loss rate (loss) reflect the overall performance of the network; combined with the transmission rate (ΔR t ) measurement, which can visualize the network and bandwidth availability, where ΔR t =(R t -R t-1 ) / min(R t-1 ,R t ). In general networks, these three factors can fully describe the network status. However, due to the large amount of burst traffic in the data center network (DCN), we additionally introduce the switch buffer occupancy (Occ) to describe the congestion level. This indicator is represented by the proportion of packets marked by ECN.
[0106] Set the reward function required by the algorithm: The RL agent needs to define a reward function to quantify the performance criteria of the task and guide the agent to improve the sequence of generated sub-strategies during the training phase. The RL agent collects rewards at each monitoring interval to interact with the network environment. Design the reward function as follows:
[0107]
[0108] in:
[0109]
[0110] In the above formula, the first part is designed based on full consideration of power = throughput / rtt. Maximizing power can reflect the maximum throughput while minimizing network delay. ω1 represents the penalty for packet loss. The second part takes into account the ratio change of the sending rate and the occupancy of the switch buffer. ω2, ω3 and ω4 are weighted coefficients.
[0111] (4) Select an optimization algorithm to optimize the parameters.
[0112] The present invention needs to comprehensively consider the response timeliness and convergence speed of the algorithm. Even if the performance is the best, the reinforcement learning algorithm with slow convergence is not applicable. The stability of the RL algorithm is also critical, especially in environmentally sensitive situations. Since a data center may contain thousands of devices, a high-performance, low-resource and easy-to-implement algorithm is required. We chose proximal policy optimization (PPO) as the RL controller because of its fast convergence, low sample complexity and high performance. In addition, PPO is simple to implement, insensitive to hyperparameters, requires few samples, has few convergence steps, and does not require replay buffer memory. Its optimization function design ensures that the policy will not deviate too much after each update, ensuring the stability of the algorithm. These characteristics make PPO the best choice.
[0113] The objective function of PPO contains the advantage estimate and the probability ratio: in is the old strategy, π θ is the updated policy. To avoid large-scale policy updates, PPO clips the target within the advantage estimate. It is defined as follows:
[0114]
[0115] is the estimated advantage at time t, and ∈ is a hyperparameter that controls the scope of clipping. The advantage function is the difference between the expected value of a state for a given action and the expected value of all possible actions for the same state. To estimate the advantage function, PPO uses a neural network model (critic) to train the value state function, calculates the total reward from the beginning to the expected realization of a given state, and finally uses the generalized advantage estimation (GAE) through the value of the critic network for advantage estimation. The optimization function for training the critic model includes the squared error loss function:
[0116]
[0117] Combining the actor loss with the critic loss and entropy, we can achieve the proxy objective function, where c1 and c2 are coefficients.
[0118]
[0119] In each iteration of the PPO algorithm, multiple participants independently collect observation trajectory data for several time steps (state, action, and reward sequence). Given the trajectory, the GAE algorithm calculates the advantage estimate for each time required for policy update. Finally, using the collected data and the estimated value, the proxy function is maximized by stochastic gradient descent SGD (or similar methods).
[0120] Based on the above implementation process, those skilled in the art should be able to know that the present invention combines the adaptive ability of deep reinforcement learning with the rapid response ability of traditional model mapping to solve the problems of slow response speed and insufficient adaptability of the existing technology in dynamic network environments, and can effectively reduce communication delays, reduce queue backlogs, improve throughput, and enhance the stability of data center networks in complex dynamic network environments, and is particularly suitable for high-load scenarios such as AI large model training. The method is flexible in structure, easy to deploy, and can be widely used in various computing network environments that require efficient congestion control.
Claims
1. A data center network congestion control method based on data model dual drive, characterized in that: The method includes: S1. Construct a system framework including an optimization engine module, a statistical information module and a rate regulation control module; the optimization engine module optimizes the congestion control sub-strategy in real time through an optimization algorithm, the statistical information module is responsible for collecting network operation data, and provides support for the decision-making process and system monitoring of the optimization engine. The optimization engine module and the statistical information module together constitute an RL agent, and the regulation control module is used to execute the sub-strategy generated by the RL agent; S2. Construct a network-side rate parameterized rate adjustment model, specifically including: S21. Calculate the convergence result through network parameters and queuing delay as a reference for adjusting the rate. The calculation formula is as follows: d q =rtt new -min{rtt new ,rtt min } Where δ is a tuning parameter, and 0.5≤δ≤1.5, which is used to represent the state of the total network bandwidth and transmission demand, d q It is the queuing delay, which is obtained by performing serialization delay and minimum RTT processing on the RTT feedback; The rate update cycle needs to be adaptively adjusted according to network demand. The parameter δ reflects the network demand and is determined by the link bandwidth and the number of connections. The update cycle is designed as: T a =δ×K T++ Where T a is the adaptive byte counter threshold, K represents the link bandwidth, and T++ represents the update of the byte counter. To improve the convergence speed and fairness, the adaptive byte counter is designed to be independent of the transmission rate, that is: BC a =γ×R C BC++ Among them, BC a is the adaptive byte counter threshold, γ is the adjustment parameter, and 0.5≤γ≤1.5, R C is the current transmission rate; S22. Design a mechanism to adjust the model rate reduction, as follows: Where R T Indicates the target transmission rate that needs to be adjusted, R C Indicates the current end-side transmission rate. Parameters μ1 and μ2 are adjustment parameters, and 2<μ1≤μ2≤4. These parameters are used to control the aggressiveness of rate reduction when congestion occurs. If CNP is not received within the specified time and α needs to be updated, then: α=(1-g)α+g Where α is the adjustment factor, and g is a fixed value used to adjust α; S23. Design a rate increase mechanism for the adjustment model, where the rate increase includes a rapid recovery phase, an adaptive increase phase, and a super increase phase; Fast recovery phase: When both the timer and the byte counter are less than the threshold, that is, T<F and BC<F, a fast recovery operation is performed, where μ3 is an adjustment parameter and 6≤μ3≤10. The design is as follows: Adaptive increase phase: When there is a timer or byte counter greater than the threshold, that is, T≥F or BC≥F, an adaptive additive speed-up operation is performed; the adaptive growth step size is related to the current transmission rate and link bandwidth, but has nothing to do with the number of carried flows. The design is as follows: Where R L is the link capacity, β and μ3 are adjustment parameters and 0<β<1; Super-increment phase: When both the timer and the byte counter are greater than the threshold, that is, T≥F and BC≥F, a super-increment operation is performed, which is designed as follows: Where i = min(BC,T-F+1), and R HAI is a fixed value of 100Mbps, and μ3 is an adjustment parameter; S24. After determining the overall design of the control model, the corresponding control parameters in the network end-side rate parameterization rate regulation model are set to be dynamically adjustable, that is, θ is defined as: θ:={δ,γ,β,μ1,μ2,μ3} S3, based on the DRL algorithm, the dynamically adjustable parameters in S23 are optimized, and the congestion control problem in the network is modeled as a Markov decision model, including setting the state quantity and reward function. The description of the network status includes introducing the switch buffer occupancy rate Occ to describe the congestion degree. This indicator is represented by the proportion of packets marked by ECN; S4. Based on the DRL algorithm, the state and reward functions in step S3 are designed and solved. The proximal policy optimization is selected as the RL controller, and the PPO algorithm is used for policy optimization based on the Markov decision process.
2. The efficient data center network congestion control method based on data model dual drive according to claim 1 is characterized in that: The specific process of step S3 includes: S31. Set the state quantity required by the DRL algorithm The state quantities include the current network throughput, round-trip delay, and packet loss rate, as well as the measurement of the transmission rate, visualizing the availability of the network and bandwidth, where the transmission rate ΔR t =(R t -R t-1 ) / min(R t-1 ,R t ); S32. Set the reward function required by the DRL algorithm This reward function is used to quantify the performance criteria of the task and guide the agent to improve the generated sub-strategy sequence during the training phase. The RL agent collects rewards at each monitoring interval to interact with the network environment; The reward function is designed as follows: in: Among them, ω1 represents the penalty for packet loss, and ω2, ω3, and ω4 are weighting coefficients.
3. The efficient data center network congestion control method based on data model dual drive according to claim 2 is characterized in that: The design of the reward function includes considering power = throughput / rtt. Power maximization can reflect the maximum throughput while minimizing network latency. It also includes considering the ratio change of the sending rate and the occupancy of the switch buffer.
4. The efficient data center network congestion control method based on data model dual drive according to claim 1 is characterized in that: The objective function of the PPO algorithm described in step S4 includes advantage estimation and calculation probability ratio; The advantage function is the difference between the expected value of the state of a given action and the expected value of all possible actions in the same state. In order to estimate the advantage function, PPO uses a neural network model to train the value state function, calculates the total reward from the beginning to the expected realization of a given state, and finally uses the generalized advantage estimate through the value of the critic network to perform advantage estimation; The probability ratio is calculated as in is the old strategy, π θ is the updated strategy; To avoid large-scale policy updates, the PPO algorithm clips the target within the advantage estimate range, defined as follows: is the estimated advantage at time t, ∈ is a hyperparameter controlling the scope of cropping; The optimization function for training a neural network model includes the squared error loss function: Combining the actor loss with the critic loss and entropy, we can implement the RL agent objective function: where c1 and c2 are coefficients.
5. The efficient data center network congestion control method based on data model dual drive according to claim 4 is characterized in that: In each iteration of the PPO algorithm, participants independently collect observation trajectory data including state, action, and reward sequences; For a given trajectory, the GAE algorithm calculates an advantage estimate for each time required for policy update; Finally, the proxy function is maximized by stochastic gradient descent (SGD) using the collected observation trajectory data and the estimated value.
Citation Information
Patent Citations
Congestion control method and system based on deep reinforcement learning
CN110581808A
Training method and device of congestion control model and congestion control method and device
CN112770353A
Network congestion control method and device, storage medium and electronic equipment
CN113872877A
Switch-driven data center congestion control method, system, device and medium
CN118214719A
Internet of vehicles channel congestion control method based on multi-agent deep reinforcement learning
CN118283700A
Cited By
Adaptive congestion control algorithm and system based on large language model
CN120434187A