Adaptive Traffic Control Method for Data Centers Based on Deep Reinforcement Learning

By applying an adaptive flow control method of deep reinforcement learning in the data center, intelligently adjusting the PFC pause threshold, the problem of manual configuration in the prior art being unable to respond to network changes in a timely manner, and more efficient and stable network transmission is achieved.

CN119182723BActive Publication Date: 2025-05-30NANJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411699131.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-05-30
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

The existing data center traffic control solution requires manual configuration of PFC thresholds, and cannot deal with complex and changeable network environments in a timely manner, resulting in network performance degradation and threshold configuration errors.

Method used

Adaptive flow control method based on deep reinforcement learning is adopted, and the PFC pause threshold is dynamically adjusted by the agent, and the traffic control strategy is optimized in real time according to changes in the network environment.

Benefits of technology

It realizes flexible regulation of PFC pause threshold, improves network transmission performance, reduces flow completion time, avoids head of queue blocking and manual configuration errors, and improves the stability and efficiency of the data center network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119182723B_ABST
    Figure CN119182723B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of data center network traffic control, and discloses an adaptive traffic control method for a data center based on deep reinforcement learning, including: S1, selecting action space and state space parameters; S2, designing a reward function according to the action space and the state space; S3, uniformly managing the switch port queues; S4, obtaining the state, calculating the reward, and updating the network to train and optimize the network environment; S5, if no reward is generated for the switch port threshold within multiple time steps, it indicates that this threshold is adapted to the current network environment, and the training is paused; S6, when the accumulated switch port data during the paused training exceeds the limit, it indicates that traffic control may be triggered at this time, and the training of the port threshold is resumed. The present invention enables each port of the switch to achieve adaptive traffic regulation, can effectively improve the network transmission speed, and has the advantages of fast convergence, easy deployment, and strong adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of traffic control in data centers, and specifically relates to an adaptive traffic control method for data centers based on deep reinforcement learning. Background Art

[0002] Current data centers adopt Remote Direct Memory Access (RDMA) technology to ensure high throughput and low latency in data centers. RDMA allows a computer to directly access the memory of a remote system without the participation of the operating system, thus significantly improving the transmission performance of the data center network. The RDMA transmission protocol relies on a lossless network environment, that is, data packets will not be lost due to network congestion, thus avoiding the additional transmission delay caused by data retransmission. To achieve a lossless network, the data center network adopts Priority-based Flow Control (PFC) technology. PFC avoids buffer overflow by pausing the data transmission of the switch port or port priority queue, thus ensuring the lossless transmission of data. Specifically, when the queue length of the switch exceeds the pause threshold of PFC, the switch sends a PFC pause frame to its upstream switch. After receiving the PFC pause frame, the upstream switch pauses the transmission of the data stream. When the queue length of the switch is lower than the recovery threshold of PFC, a PFC recovery frame is sent to the upstream switch. After receiving the PFC recovery frame, the upstream switch resumes the transmission of the data stream. The transmission efficiency of data in the network is closely related to the setting of the PFC pause threshold. If the PFC pause threshold is low, it will cause frequent data pause operations, thus affecting the traffic transmission efficiency; if the PFC pause threshold is high, it will cause the switch to be insensitive to the network state, resulting in serious queue accumulation. Existing traffic control schemes still require manual configuration of the PFC threshold. However, the real network environment is complex and changeable, and manual configuration of the PFC threshold cannot timely respond to emerging network applications and their bursty traffic, resulting in a decline in network performance. Manual configuration is difficult to adapt to the frequent changes in the network environment in a timely manner and may lead to incorrect threshold configuration. Summary of the Invention

[0003] In order to solve the above technical problems, the present invention provides an adaptive traffic control method for data centers based on deep reinforcement learning, designs an adaptive dynamic traffic control threshold configuration scheme, improves network transmission performance, and ensures the reliability of threshold configuration.

[0004] To achieve the above object, the present invention is realized by the following technical solutions:

[0005] The present invention is a data center adaptive traffic control method based on deep reinforcement learning. The implementation of this adaptive traffic control specifically includes the following steps:

[0006] Step 1: Select the threshold for sending pause frames when triggering traffic control as the action space, and select the data traffic received and sent by the switch and the pause time of the upstream switch port as the state space parameters;

[0007] Step 2: Determine the changes in the network environment before and after the action execution according to the action space and the state space, and design an effective reward function;

[0008] Step 3: Design and deploy an agent to uniformly manage the switch port queues;

[0009] Step 4: The agent interacts with the switch, calculates the reward, and updates the network to train and optimize the network environment;

[0010] Step 5: If the thresholds of the trained switch ports do not generate rewards within multiple time steps, it indicates that this threshold is adapted to the current network environment, and the training is paused to reduce the switch performance overhead;

[0011] Step 6: When the accumulated data of the paused training switch exceeds a certain limit, it indicates that traffic control may be triggered at this time, so the training of the port threshold is resumed.

[0012] A further improvement of the present invention lies in: In step 1, the process of selecting the action space and the state space parameters includes:

[0013] Step 1.1: According to the traffic situation flowing through the switch, calculate the remaining space in the switch buffer in real time :

[0014]

[0015] where is the total buffer size, is the total header space, is the total reserved space, is the number of bytes already used in the shared buffer;

[0016] Step 1.2: According to the remaining space in the switch buffer, calculate the threshold for triggering the PFC pause frame

[0017]

[0018] where is a shift factor for adjusting the priority-based flow control (PFC) pause threshold, and its value is ;

[0019] Step 1.3: To ensure throughput and need to recover in a timely manner, the PFC recovery threshold is set to

[0020]

[0021] wherein, is the maximum transmission unit;

[0022] Step 1.4: Select as the action to dynamically adjust the PFC trigger sensitivity, and set its action space to [1, 2, 4, 8, 16, 32, 64]. This action space can be adjusted according to the different switch buffers;

[0023] Step 1.5: Select the data traffic sent and received and the number of received data packets within one time step as state parameters. The data traffic sent and received reflects the traffic transmission situation of the port, and the number of received data packets reflects the traffic density of the port; select the upstream switch pause time within one time step as a state parameter to reflect the impact caused by the current switch PFC trigger. If there is no PFC trigger, the value of this state is 0.

[0024] A further improvement of the present invention is that in step 2, the process of determining the change in the network environment before and after the action execution according to the action space and the state space and designing an effective reward function includes:

[0025] Step 2.1: If the received data volume within one time step after an action execution is equal to the sent data volume or the number of received data packets is 0, and there is no PFC trigger, it means that this action has no impact on the network. Set the total reward to 0 and directly enter the next training;

[0026] Step 2.2: Set the receive reward. When most of the received data packets are large traffic or the number of received data packets is small, the PFC threshold can be appropriately increased to avoid PFC triggering; when there are many received data packets and most of them are small traffic, the PFC threshold can be appropriately decreased to make the PFC trigger in advance to avoid more serious congestion. The above situations can be reflected by the average size of the received data packets reflect, is the received data volume, is the number of received data packets. The average data packet is large, corresponding to the situation of mostly large traffic and few data packets; the average data packet is small, corresponding to the situation of mostly small traffic and many data packets. By comparing the state changes before and after the action execution, design the receive reward function:

[0027]

[0028] wherein, and is the data reception volume before and after the action execution, and is the number of received data packets before and after the action execution, and are the two actions before and after;

[0029] Step 2.3, Set the pause reward. When there are a large number of data packets, appropriately extending the pause time of the upstream switch is beneficial to alleviating network congestion. However, if the pause time is too long, it will also affect the transmission efficiency of the switch port traffic. Therefore, we use to weigh the rationality of this pause duration, and combine the changes in before and after the action execution to reflect the pros and cons of the current action. The pause reward function is designed as follows:

[0030]

[0031] where, and are the pause durations within the time steps before and after the action execution, is the pause duration within the time step;

[0032] Step 2.4, Set the total reward. In order to fully weigh the reception reward and the pause reward, the total reward is set as a weighted function of the reception reward and the pause reward:

[0033]

[0034] where, and are the weights of the reception reward and the pause reward in the total reward respectively.

[0035] A further improvement of the present invention lies in: In step 3, the process of designing and deploying an agent that can obtain the state and calculate the reward by executing actions on the switch to learn and help make the best decision to uniformly manage the switch port queue includes:

[0036] Step 3.1, Select a deep learning algorithm to train the agent, and adaptively adjust the PFC pause threshold through the agent;

[0037] Step 3.2, Initialize a deep neural network as the network, which is used to approximate the value function. During the continuous update and learning process, use the Q-value function to evaluate the value of the current action and guide the action selection to help the agent make the optimal decision when adjusting the PFC pause threshold of the switch port;

[0038] Step 3.3, Initialize a target network, whose structure is the same as that of The networks are the same, but the parameters are updated independently and are used to calculate the target Q value;

[0039] Step 3.4: Create an experience replay buffer to store the experiences obtained during the interaction between the agent and the switch, namely the state, action, reward, and next state;

[0040] Step 3.5: Set the learning rate, discount factor , for the hyperparameters of the random probability epsilon for the policy, batch size, and update frequency.

[0041] A further improvement of the present invention lies in: In step 4, the agent interacts with the switch, calculates the reward, and updates the network. The process of training and optimizing the network environment includes:

[0042] Step 4.1: Select an action according to the policy: With a probability of select a random action to explore a new state, and with a probability of select the action with the largest value in the current network output;

[0043] Step 4.2: The agent determines the action and deploys it to the switch queue, adjusts the sensitivity triggered by priority-based traffic control, then obtains the next time step state from the switch, calculates the reward by combining the action and the state, and stores the current state, action, reward, and next state in the experience replay buffer;

[0044] Step 4.3: Repeat step 4.1 and step 4.2. When the experiences in the experience replay buffer accumulate to a set number, sample a batch of experiences from the experience replay buffer for training the network;

[0045] Step 4.4: Use the target network to calculate the target value, and update the parameters of the network by minimizing the loss function. The loss function is the mean squared error between the target value and the current value:

[0046]

[0047] Every certain number of steps, copy the parameters of the network to the target network to maintain the stability of the target network, where is the target value, is the expected value, representing the average value of all sampled state-action pairs. is the state at the current time step. is the action at the current time step. are the weights of the neural network. is the current estimated value of the network;

[0048] Step 4.5: As the training progresses, will gradually decay to reduce the selection ratio of random actions, and gradually make the agent transition from exploration to exploitation phase.

[0049] A further improvement of the present invention lies in: In step 5, the process of pausing training to reduce the performance overhead of the switch when the specific condition is met includes:

[0050] Step 5.1: Set a counter to track the number of consecutive rounds with a reward of 0, so as to judge whether the current threshold has an impact on the traffic transmission of the switch port;

[0051] Step 5.2: If the reward is 0 within several consecutive rounds, such as 3 - 5 rounds, it means that the current action does not affect the traffic transmission of the switch port. Mark the agent as the idle state and pause the training for this switch port.

[0052] Furthermore, in step 6, the process of resuming the training of the port threshold when traffic control may be triggered includes:

[0053] Step 6.1: In the idle state, check whether the size of the shared buffer exceeds 1 / 2 of the current threshold at each time step, which can be adjusted according to the network bandwidth;

[0054] Step 6.2: If the size of the shared buffer exceeds 1 / 2 of the current threshold, resume the training for this port.

[0055] The beneficial effects achieved by the present invention are: The present invention combines traffic control with deep reinforcement learning, can respond to changes in the network environment in a timely manner, adaptively adjust the PFC pause threshold according to the network environment to reduce the flow completion time, thereby avoiding problems such as head-of-line blocking that may occur, and also avoiding configuration errors caused by manually modifying the PFC threshold, thereby ensuring the high efficiency and stability of data transmission in the data center network. The present invention is an adaptive traffic control method based on deep reinforcement learning, with the advantages of fast convergence, easy deployment, and strong adaptability. Therefore, the adaptive traffic control method is more suitable for the data center network that is always in a changing state. Compared with the prior art, it has the following advantages:

[0056] (1) The present invention can achieve more flexible regulation of the sensitivity of PFC pause threshold triggering to adapt to complex and changing network environments.

[0057] (2) The present invention combines deep reinforcement learning in a programmable switch, which can automatically adapt to the dynamic changes of network traffic according to real-time network load conditions without any device-level modifications and without human intervention, simplifying the network operation of the data center.

[0058] (3) During the online training process, the present invention has a fast convergence speed and can quickly adapt to the current network traffic conditions with only a very short training time to cope with various burst traffic conditions.

[0059] (4) The present invention adopts a distributed agent deployment and controls the pause and recovery of the agent to relieve the huge pressure on the switch caused by excessive training data. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a flowchart of the data center adaptive traffic control method based on deep reinforcement learning according to the present invention.

[0061] Figure 2 It is a schematic diagram of the flow completion time before and after training in an embodiment of the present invention.

[0062] Figure 3 It is a comparison chart of the PFC trigger times of each switch when the present application and the existing algorithm are congested.

[0063] Figure 4 It is a comparison chart of the queue length changes of congested flows when the present invention and the existing algorithm are congested.

[0064] Figure 5 It is a comparison chart of the maximum / average queue lengths during the operation of the present invention and the existing algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] The following will disclose the embodiments of the present invention in a schematic diagram. For the sake of clarity, many practical details will be described together in the following description. However, it should be understood that these practical details should not be used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.

[0066] As Figure 1 shown, the present invention is a data center adaptive traffic control method based on deep reinforcement learning, and the method specifically includes the following steps:

[0067] Step 1, select the threshold for sending pause frames when triggering traffic control as the action space, and select the data traffic received and sent by the switch and the pause time of the upstream switch port as the state space parameters.

[0068] Calculate the remaining space in the switch buffer in real time according to the traffic flowing through the switch :

[0069]

[0070] Among them, is the total buffer size, is the total header space, is the total reserved space, is the number of bytes used in the shared buffer.

[0071] Then, use the remaining space in the switch buffer to calculate the threshold for triggering the PFC pause frame

[0072]

[0073] Among them, is a shift factor for adjusting the PFC pause threshold, and its value is ;

[0074] In order to ensure throughput and need to recover in time, set the PFC recovery threshold to

[0075]

[0076] Among them, is the maximum transmission unit.

[0077] Select as the action for dynamically adjusting the PFC trigger sensitivity, and set its action space to [1, 2, 4, 8, 16, 32, 64]. This action space can be adjusted according to the different switch buffers. Select the data traffic and the number of received data packets in one time step as the state parameters. The data traffic reflects the traffic transmission situation of the port, and the number of received data packets reflects the traffic density of the port; select the upstream switch pause time in one time step as the state parameter to reflect the impact caused by the current switch PFC trigger. If there is no PFC trigger, the value of this state is 0.

[0078] Step 2: Determine the changes in the network environment before and after the action execution according to the action space and the state space, and design an effective reward function.

[0079] If the amount of received data within one time step after a single action execution is equal to the amount of sent data or the number of received data packets is 0, and there is no PFC trigger, it indicates that this action has no impact on the network. Set the total reward to 0 and directly proceed to the next training. Meanwhile, the reward function consists of two parts: the received data packet reward and the pause situation reward. The received reward is set as follows. When the received data packets are of high traffic or the number of received data packets is small, the PFC threshold can be appropriately increased to avoid PFC triggering; when there are numerous received data packets, and most of them are of low traffic, the PFC threshold can be appropriately decreased to make the PFC trigger in advance to avoid more serious congestion. The above situations can be reflected by the average size of the received data packets Reflect is the amount of received data, is the number of received data packets. A larger average data packet corresponds to the situation of mostly high-traffic and fewer data packets; a smaller average data packet corresponds to the situation of mostly low-traffic and more data packets. By comparing the state changes before and after the action execution, design the received reward function:

[0080]

[0081] Among them, and are the amounts of data received before and after the action execution, and are the numbers of received data packets before and after the action execution, and are the two consecutive actions;

[0082] The pause reward is set as follows: When there are numerous data packets, appropriately extending the pause time of the upstream switch is beneficial to alleviating network congestion. However, if the pause time is too long, it will also affect the traffic transmission efficiency of the switch port. Therefore, the present invention uses to weigh the rationality of this pause duration, and combines the changes in before and after the action execution to reflect the pros and cons of the current action. Design the pause reward function as follows:

[0083]

[0084] Among them, and are the pause durations within the time step before the action execution and within the time step after the action execution, is the pause duration within the time step;

[0085] The total reward is jointly composed of the received reward and the pause reward. In order to fully weigh the received reward and the pause reward, the total reward is set as a weighted function of the received reward and the pause reward:

[0086] .

[0087] in, and The weights of receiving rewards and pausing rewards in the total rewards respectively.

[0088] Step 3: Design and deploy an intelligent agent to uniformly manage the switch port queues, and help the switch achieve adaptive traffic control by building an intelligent agent.

[0089] The present invention selects depth The learning algorithm trains the agent and adaptively adjusts the PFC threshold through the agent; initializes a deep neural network as Network, used for approximation The value function initializes a target The network, its structure and The network is the same, but the parameters are updated independently to calculate the target value; create an experience replay buffer to store the agent's experience, namely state, action, reward, and next state; finally set the learning rate for the optimizer's parameter update rate, such as 0.001, discount factor , used to measure the discount rate of future rewards, such as 0.95, The epsilon of the strategy is used to control the trade-off between exploration and exploitation, such as the initial value is 1 and the minimum value is 0.01. The batch size is used to determine the number of samples sampled from the experience replay buffer each time, such as 32. The update frequency is used to specify how many time steps are used to update the target network, such as 10 steps. These hyperparameters help the agent explore and optimize the neural network in the interaction with the switch. Value Function Used to indicate a given state Take an action After that, the expected total return that may be obtained in the future. The goal of learning is to find a policy that maximizes the cumulative reward in the future by taking the action in each state. In the DQN network, deep neural networks are used to approximate Value function. During training, the neural network receives the state As input, and output for each possible action value, and then update the weights of the neural network to approximate This approximation allows the agent to learn effectively in large environments even when the state and action spaces are large.

[0090] Step 4: The agent interacts with the switch, calculates the reward, and updates Network,training optimizes the network environment.

[0091] The agent selects an action according to the policy: with probability it selects a random action to explore new states, and with probability 1 - it selects the current action with the maximum value in the network output; the agent determines the action and deploys it to the switch queue, adjusts the sensitivity of PFC triggering, transfers to the next state, observes the reward, and stores the current state, action, reward, and next state in the experience replay buffer; randomly samples a mini - batch of experiences from the experience replay buffer for training the network; uses the target network to calculate the target value, and updates the parameters of the network by minimizing the loss function, where the loss function is the mean squared error (MSE) between the target value and the current value:

[0092]

[0093] where is the target value, is the expected value, representing the average over all sampled state - action pairs, is the state at the current time step, is the action at the current time step, are the weights of the neural network, is the current estimate of the value by the network;

[0094] Every certain number of steps, copy the parameters of the network to the target network to maintain the stability of the target network; as training progresses, gradually decay epsilon to reduce the proportion of random action selection, and gradually make the agent transition from exploration to exploitation phase. The target value is calculated by the Bellman equation of -learning: , where is the reward at the current time step, is the discount factor, is the state at the next time step, is the action at the next time step, are the old neural network weights, is the maximum value calculated by the target network for the next step. (epsilon) is a value between 0 and 1. At the beginning of training, is larger, such as 0.9, which means the agent has a 90% probability of randomly selecting actions to conduct more exploration. As the training progresses, gradually decreases, and the agent will rely more on existing knowledge to select the optimal actions. The decrease of satisfies the decay function , and this function will be executed before the end of each time step. is the minimum value, for example, 0.01. is the decay rate, for example, 0.995. The initial value can be set by oneself.

[0095] Step 5, if the switch port threshold being trained does not generate rewards within multiple time steps, it indicates that this threshold is adapted to the current network environment, and the training is paused to reduce the switch performance overhead.

[0096] The agent will set a counter to track the number of consecutive rounds with a reward of 0, so as to judge the impact of the current threshold on the network environment: if the reward is 0 within 3 - 5 consecutive rounds, it means that the current action does not affect the switch port traffic transmission, and the agent is marked as the idle state, and the training for this switch port is paused. By pausing through this condition, the training burden of the switch can be reduced, and the switch performance overhead can be decreased.

[0097] Step 6, when the accumulated data of the switch with paused training exceeds a certain limit, it indicates that traffic control may be triggered at this time, so the training for the port threshold is resumed.

[0098] In the idle state, it will check whether the size of the shared buffer exceeds 1 / 2 of the current threshold at each time step, which can be adjusted according to the network bandwidth. If the size of the shared buffer exceeds 1 / 2 of the current threshold, the training for this port is resumed. Through this method, the PFC threshold can be adjusted in time when the network environment changes, improving the network performance.

[0099] Figure 2 is the effect diagram of the present invention, showing that the network data transmission rate designed by the present invention based on the adaptive PFC pause threshold is significantly improved. Specifically, the topology environment used in this method is a fat - tree topology, including 4 core switches, 8 aggregation - layer switches, 8 access - layer switches, and 16 servers. The link bandwidth between switches is 400Gbps, and the link bandwidth between the server and the switch is 100Gbps. The training dataset used is AliStorage, the network load is 90%, and the training traffic size is the traffic transmitted at a host bandwidth of 100Gbps for 0.1 seconds continuously. As Figure 2As shown, compared with the network using the traditional PFC threshold, for the small flow (i.e., <8Kb flow), the completion time is reduced by 83.26%, for the medium flow (i.e., 8Kb - 125Kb flow), the completion time is reduced by 82.11%, and for the large flow (i.e., >125Kb flow), the completion time is reduced by 72.94%. This shows that by using the present invention, the traffic transmission efficiency of the overall network is significantly improved. This indicates that the present invention has obvious advantages in improving network performance.

[0100] Figure 3 It is a comparison chart of the PFC trigger times of each switch when the present application and the existing algorithm are congested. As Figure 3 shown, when using the traditional PFC threshold regulation method, the volatility of the PFC trigger times is relatively large, reaching a maximum of 4000 - 5000 times, showing obvious network congestion problems. Compared with the network using the traditional PFC threshold regulation, the present invention maintains the PFC trigger times of the switch at about 1000 - 2000 times. While effectively reducing the PFC trigger times of the switch, it makes the distribution of the PFC trigger times more uniform, with smaller volatility, reduces congestion events in the network, and improves the overall performance of the switch.

[0101] Figure 4 It is a comparison chart of the queue length changes of the congested flows when the present invention and the existing algorithm are congested. As Figure 4 shown, when using the traditional PFC threshold regulation method, the fluctuation of the queue length is relatively severe. Especially when approaching the end of the event, the queue rises sharply and reaches a peak value of about 2.385MB, indicating that in the traditional PFC threshold regulation method, the network is prone to large congestion, resulting in a large accumulation of data packets in the queue, forming a network bottleneck. Compared with the traditional PFC threshold regulation method, the curve of this method is smoother, and the queue length is maintained at a relatively low level for most of the time, with almost no obvious sharp rise in the queue, indicating that this method can help the switch better predict and allocate network resources, avoid the accumulation of a large number of data packets and network congestion, thereby improving the stability and performance of the switch.

[0102] Figure 5 It is a comparison chart of the maximum / average queue length during the operation of the present invention and the existing algorithm. As Figure 5As shown, in the AliStorage dataset, the average queue length of the traditional PFC threshold regulation method is close to 158 KB, and the maximum queue length is close to 2.385 MB. The average queue length of this method is close to 69 KB, and the maximum queue length is close to 1.051 MB. Compared with the traditional method, the average queue length is reduced by 56%, and the maximum queue length is reduced by 56%. In the WebSearch dataset, the average queue length of the traditional PFC threshold regulation method is close to 641 KB, and the maximum queue length is close to 2.407 MB. The average queue length of this method is close to 60 KB, and the maximum queue length is close to 0.809 MB. Compared with the traditional method, the average queue length is reduced by 90%, and the maximum queue length is reduced by 66%. Whether it is the AliStorage dataset or the WebSearch dataset, this method significantly reduces the average length and the maximum length of the queue, effectively optimizes the queue management of different applications, reduces congestion and improves network performance.

[0103] Through deep reinforcement learning, the present invention adaptively selects a suitable PFC pause threshold interval and reasonably regulates the timing of PFC triggering, thereby ensuring the network throughput, reducing the flow completion time, and eliminating the lag and unnecessary errors caused by cumbersome manual regulation. Compared with manually adjusting the PFC threshold, the adaptive flow control technology used in the present invention can respond to different network traffic and automatically adjust the PFC threshold, thereby ensuring the network throughput and flow completion time. The solution proposed by the present invention may be more effective and practical.

[0104] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiment. Any equivalent modification or change made by those of ordinary skill in the art according to the disclosure of the present invention shall be included in the protection scope recorded in the claims.

Claims

1. A data center adaptive traffic control method based on deep reinforcement learning, characterized by: The data center adaptive flow control method specifically comprises the following steps: Step 1: Select the threshold for sending a pause frame when triggering flow control as the action space, and select the switch data transmission and reception flow and the upstream switch port pause time as the state space parameters; Step 2: According to the action space and state space, determine the changes in the network environment before and after the action is executed, and design the reward function, which specifically includes the following steps: In step 2, the changes in the network environment before and after the action is executed are determined according to the action space and the state space, and a reward function is designed, which specifically includes the following steps: Step 2.1: If the amount of received data in a time step after an action is executed is equal to the amount of sent data or the number of received data packets is 0, and there is no priority-based flow control trigger, it means that this action has no effect on the network, set the total reward to 0, and directly enter the next training; Step 2.2, set the receiving reward: when the received data packets are large and the number of received data packets is small, increase the priority-based flow control threshold to avoid the triggering of priority-based flow control; when the received data packets are large and the flow is small, reduce the priority-based flow threshold to trigger the priority-based flow control in advance to avoid causing more serious congestion. Reflect, D get is the amount of received data, Pa get For the number of received data packets, the average data packet is large, corresponding to the situation of large traffic and small data packets, and the average data packet is small, corresponding to the situation of small traffic and large data packets. Compare the state changes before and after the action is executed, and design the receiving reward function: Among them, D getp and D getc is the amount of data received before and after the action is executed, Pa getp and Pa getc A is the number of packets received before and after the action is executed. p and A c It is two actions, front and back; Step 2.3, set pause reward: When there are many packets, extend the pause time of the upstream switch to help alleviate network congestion. To weigh the duration of this pause, and combine it with the before and after action execution The change of reflects the pros and cons of the current action and designs the pause reward function: Among them, T pausep and T pausec is the pause duration in the time step before and after the action is executed, T pause is the pause duration within the time step; Step 2.

4. Set the total reward: Set the total reward as a weighted function of the received reward and the pause reward: R=ω1×R get +ω2×R pause Among them, ω1 and ω2 are the weights of receiving rewards and pausing rewards in the total rewards respectively; Step 3: Design and deploy an intelligent agent on the switch to obtain the state and calculate the reward by executing actions to learn and help make the best decision to uniformly manage the switch port queue; Step 4: The agent interacts with the switch, calculates rewards, updates the Q network, and trains to optimize the network environment; Step 5: If the trained switch port threshold does not generate rewards in multiple time steps, it means that this threshold is adapted to the current network environment, and the training is suspended to reduce the switch performance overhead; Step 6: When the accumulated data of the switch that has suspended training exceeds the set limit, flow control is triggered and the training of the port threshold is resumed.

2. The data center adaptive traffic control method based on deep reinforcement learning according to claim 1 is characterized in that: In step 1, selecting the threshold for sending a pause frame when triggering flow control as the action space, and selecting the switch data transmission and reception flow and the upstream switch port pause time as the state space parameters specifically includes the following steps: Step 1.1: Calculate the remaining space B in the switch buffer in real time based on the traffic flowing through the switch. remain : B remain =B-B hdrm -B rsrv -B used , Among them, B is the total buffer size, B hdrm is the total headroom, B rsrv is the total reserved space, B used The number of bytes used for the shared buffer; Step 1.2: Calculate the pause frame threshold Th for triggering priority-based flow control based on the remaining space in the switch buffer. p : Where times is a shift factor for adjusting the priority-based flow control pause threshold; Step 1.3: Set the priority-based flow control recovery threshold to Th r Th r =Th p -2MTU, Among them, MTU is the maximum transmission unit; Step 1.4, select times as the action for dynamically adjusting the trigger sensitivity of priority-based flow control; Step 1.5, select the data flow rate of transceiver and sender, the number of received data packets and the pause time of the upstream switch within a time step as the state parameters. The data flow rate of transceiver and sender reflects the traffic transmission status of the port, the number of received data packets reflects the traffic density of the port, and the pause time of the upstream switch reflects the impact of the current switch's priority-based flow control trigger. If there is no priority-based flow control trigger, the value of this state is 0.

3. The data center adaptive traffic control method based on deep reinforcement learning according to claim 1 is characterized in that: In step 3, the intelligent agent designed and deployed on the switch acquires the state and calculates the reward by executing actions to learn, and helps make the best decision to uniformly manage the switch port queue, which specifically includes the following steps: Step 3.1, select the deep Q learning algorithm to train the intelligent agent, and adaptively adjust the priority-based flow control pause threshold through the intelligent agent; Step 3.2, initialize a deep neural network as a Q network to approximate the Q value function. In the continuous updating and learning process, the approximate Q value function is used to evaluate the value of the current action and guide the selection of actions to help the intelligent agent make the best decision when adjusting the priority-based flow control pause threshold of the switch port; Step 3.3: Initialize a target Q network. The structure of the target Q network is the same as that of the Q network, but the parameters are updated independently to calculate the target Q value. Step 3.4: Create an experience replay buffer to store the experience gained during the interaction between the agent and the switch, namely the state, action, reward, and next state; Step 3.

5. Set the hyperparameters of learning rate, discount factor γ, random probability for ε-greedy strategy, batch size, and update frequency.

4. The data center adaptive traffic control method based on deep reinforcement learning according to claim 1 is characterized in that: In step 4, the agent interacts with the switch, calculates rewards, updates the Q network, and trains and optimizes the network environment, specifically including the following steps: Step 4.1, select actions according to the ε-greedy strategy: select random actions with probability ε, explore new states, and select the action with the largest Q value in the current Q network output with probability 1-ε; Step 4.2: The agent determines the action and deploys it to the switch queue, adjusts the sensitivity of the priority-based flow control trigger, then obtains the next time step state from the switch, calculates the reward based on the action and state, and stores the current state, action, reward, and next state in the experience replay buffer; Step 4.3: Repeat steps 4.1 and 4.

2. When the experience in the experience replay buffer reaches the set number, a batch of experience is sampled from the experience replay buffer for training the Q network. Step 4.4, use the target Q network to calculate the target Q value, and update the parameters of the Q network by minimizing the loss function. The loss function is the mean square error between the target Q value and the current Q value: L(θ)=E[(yQ(s,a;θ)) 2 ] Every fixed number of steps, the parameters of the Q network are copied to the target Q network to maintain the stability of the target Q network, where y is the target Q value, E is the expected value, which represents the average value of all sampled state and action pairs, s is the current time step state, a is the current time step action, θ is the weight of the neural network, and Q(s,a;θ) is the estimated Q value of the current Q network; Step 4.5: As the training progresses, the parameter ε is gradually decayed to reduce the selection ratio of random actions and gradually make the agent transition from the exploration stage to the utilization stage.

5. The data center adaptive traffic control method based on deep reinforcement learning according to claim 1 is characterized in that: In step 5, if the trained switch port threshold does not generate rewards in multiple time steps, it means that the threshold is adapted to the current network environment, and the training is suspended to reduce the switch performance overhead, which specifically includes the following steps: Step 5.1, set a counter to track the number of consecutive rounds with a reward of 0, so as to determine whether the current threshold has an impact on the traffic transmission of the switch port; Step 5.2: If the reward is 0 for multiple consecutive rounds, it means that the current action does not affect the traffic transmission of the switch port. The agent is marked as idle and the training of this switch port is suspended.

6. The data center adaptive traffic control method based on deep reinforcement learning according to claim 1 is characterized in that: The step 6 specifically includes the following steps: Step 6.1: In the idle state, the shared buffer size is checked at each time step to see if it exceeds 1 / 2 of the current threshold. Step 6.2: If the shared buffer size exceeds 1 / 2 of the current threshold, the training of this port is resumed.

Citation Information

Patent Citations

  • Data center network ECN automatic regulation and control method based on multi-agent reinforcement learning

    CN115529278A