Network congestion control method and device based on deep learning, and readable storage medium
Through the deep learning-based network congestion control method, the congestion flag threshold is dynamically adjusted, which solves the problems of high system costs and difficulty in adapting to dynamic network changes in the prior art, and achieves efficient network congestion control.
Patent Information
- Application Number
- CN202311564630.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-30
AI Technical Summary
Existing model-based display congestion notification methods are small in the data center, increasing system costs and difficult to adapt to dynamic changes in network traffic and complex network states.
The network congestion control method based on deep learning is adopted to obtain the network status of the current congestion queue, determine the optimal congestion control strategy, and adjust the congestion flag threshold in real time to dynamically control network traffic.
It reduces the adjustment time of ECN parameters, improves the congestion adjustment efficiency and congestion control accuracy of network traffic, and reduces system costs.
Smart Images

Figure CN120075137A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a network congestion control method, device, and computer-readable storage medium based on deep learning. Background Art
[0002] In traditional network traffic control, when network congestion occurs, the router will discard the data packet or put the data packet into the queue, which easily leads to delays and packet losses and affects network performance. Therefore, a congestion control method is needed to control the total amount of data entering the network so that the network traffic remains at an acceptable level. And Explicit Congestion Notification (ECN) provides a more intelligent congestion control method. When the buffer of the router is close to overflow or congestion, this method will change the ECN flag bit of the data packet and then send the data packet to the destination. After receiving the data packet with the ECN flag bit, the receiving end will send a message reply to the sending end to notify the sending end of the network congestion situation. After receiving the ECN reply, the sending end takes appropriate measures according to the ECN flag bit to relieve network congestion, reduce the data traffic in the case of network congestion, thereby avoiding packet loss and delay and improving network performance. Since the traffic situation and congestion degree of the network will change with the dynamic network environment. However, the current model-based explicit congestion notification method needs to deploy a regulation model on each switch. This deployment method increases the system cost when the scale of the data center is small. Summary of the Invention
[0003] The main purpose of this application is to provide a network congestion control method, device, and storage medium based on deep learning, aiming at the technical problem of how to reduce the system cost of the current network congestion control method.
[0004] To achieve the above object, an embodiment of this application provides a network congestion control method based on deep learning. The network congestion control method based on deep learning includes:
[0005] Obtain the identifier of the target data center to which the current congestion queue belongs, and determine the model deployment method of the target data center based on the identifier of the target data center;
[0006] Determine the target node based on the model deployment method;
[0007] Obtain the current network state of the current congestion queue, and determine the optimal congestion control policy corresponding to the current network state based on the congestion control model in the target node;
[0008] Based on the optimal congestion control strategy, adjust the congestion flag threshold of the current congestion queue to control the network traffic of the current congestion queue based on the congestion flag threshold.
[0009] In addition, to achieve the above object, an embodiment of the present application further provides a network congestion control device based on deep learning. The network congestion control device based on deep learning includes a processor, a memory, and a network congestion control program based on deep learning stored on the memory and executable by the processor. When the network congestion control program based on deep learning is executed by the processor, the steps of the network congestion control method based on deep learning as described above are implemented.
[0010] In addition, to achieve the above object, an embodiment of the present application further provides a computer-readable storage medium. A network congestion control program based on deep learning is stored on the computer-readable storage medium. When the network congestion control program based on deep learning is executed by a processor, the steps of the network congestion control method based on deep learning as described above are implemented.
[0011] One of the above technical solutions has the following beneficial effects: Obtain the identifier of the target data center to which the current congestion queue belongs, and determine the model deployment method of the target data center based on the identifier of the target data center; based on the model deployment method, determine the target node; obtain the current network state of the current congestion queue, and determine the optimal congestion control strategy corresponding to the current network state based on the congestion control model in the target node; based on the optimal congestion control strategy, adjust the congestion flag threshold of the current congestion queue to control the network traffic of the current congestion queue based on the congestion flag threshold. In this way, the embodiment of the present application determines the model deployment method according to the scale of the data center in advance, and then determines the node where the model is located according to the deployment method. Then, according to the real-time situation of the current congestion queue (i.e., the current network state) and the pre-trained congestion control model, the optimal congestion control strategy suitable for the current congestion queue is determined. And based on the optimal congestion control strategy, the congestion flag threshold of the fixed current congestion queue is adjusted in real time. Thus, this embodiment provides a congestion control strategy that can adjust the congestion flag threshold in real time based on a dynamically changing network environment, reduces the adjustment time of ECN parameters, improves the congestion adjustment efficiency and congestion regulation accuracy of network traffic. In addition, further determine the model deployment method according to the scale of the data center, avoid improper model deployment, and reduce system costs. Description of the Drawings
[0012] Figure 1 It is a schematic hardware structure diagram of a network congestion control device based on deep learning involved in the solution of the embodiment of the present application;
[0013] Figure 2 This is a schematic flowchart of the first embodiment of the network congestion control method based on deep learning in this application;
[0014] Figure 3 This is a schematic flowchart of the first embodiment of the network congestion control method based on deep learning in this application;
[0015] Figure 4 It is a block diagram of the overall structure of centralized deployment;
[0016] Figure 5 It is a communication schematic diagram between the client and the server;
[0017] Figure 6 It is an interaction schematic diagram between the client and the server during separate training;
[0018] Figure 7 It is a schematic diagram of the reinforcement learning algorithm logic.
[0019] The realization, functional features and advantages of the purpose of this application will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments
[0020] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0021] The network congestion control method based on deep learning involved in the embodiments of this application is mainly applied to a network congestion control device based on deep learning. The network congestion control device based on deep learning can be a device with display and processing functions such as a PC, a portable computer, a mobile terminal, etc.
[0022] Refer to Figure 1 , Figure 1 This is a schematic hardware structure diagram of the network congestion control device based on deep learning involved in the solution of the embodiments of this application. In the embodiments of this application, the network congestion control device based on deep learning may include a processor 1001 (such as a CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components; the user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard); the network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface); the memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0023] Those skilled in the art can understand that Figure 1 the hardware structure shown in Figure 1 does not constitute a limitation on the network congestion control device based on deep learning, and may include more or fewer components than shown, or combine some components, or have different component arrangements.
[0024] Continuing to refer to Figure 1 , Figure 1 in Figure 1 , the memory 1005 as a computer-readable storage medium may include an operating system, a network communication module, and a network congestion control program based on deep learning.
[0025] In Figure 1 Figure 1 , the network communication module is mainly used to connect to the server and communicate with the server for data; while the processor 1001 can call the network congestion control program stored in the memory 1005 and execute the network congestion control method provided by the embodiments of the present application.
[0026] Congestion control refers to controlling the total amount of data entering the network when the network is congested, so that the network traffic is maintained at an acceptable level, in order to reduce the data traffic in the case of network congestion, thereby avoiding packet loss and delay and improving network performance. However, the current static fixed ECN threshold adjustment strategy of ECN is difficult to adapt to the dynamic changes of network traffic and complex network states, and the ECN tuning configuration requires a lot of time to determine the adjustment strategy corresponding to different network environments. In addition, for small-scale production data centers, deploying an adjustment model for each switch will additionally increase the system cost.
[0027] To solve the above problems, the embodiments of the present application provide a network congestion control method based on deep learning applied to an ECN congestion control system, and the congestion control system includes:
[0028] A simulation network module for constructing a simulation network model;
[0029] An initialization module for initializing the congestion control model;
[0030] A separation training module for training the congestion control model using a reinforcement learning algorithm;
[0031] An online deployment module for deploying the congestion control model to the switch and retraining the congestion control model using actual data;
[0032] A congestion control module for adjusting the congestion flag threshold of the queue in the network using the congestion control model to achieve network congestion control.
[0033] The simulation network module first models the environment of the network congestion control model, that is, collects historical data of the sender and receiver in different network environments, such as network topology data, transmission delay data, data traffic, and link capacity. Normalize the values of the historical data, and establish a simulation network environment required for model training based on the normalized historical data.
[0034] The initialization module initializes the congestion control model. Specifically, the initialization module first defines the state space, action space, agent, and memory bank of the congestion control model, and defines the reward function and loss function based on the characteristics of the traffic network. Among them, the state space includes the current queue length, the output data rate of each link, the high marking threshold, the low marking threshold, and the marking probability. The action space includes high marking threshold adjustment actions, low marking threshold adjustment actions, and marking probability adjustment actions. That is, by adjusting one or more parameters among the high marking threshold, the low marking threshold, and the marking probability, the network traffic is controlled to avoid congestion.
[0035] The separate training module uses the reinforcement learning algorithm DQN (Deep Q Network, a neural network algorithm based on deep learning) to train the agent in the congestion control model in the separate training environment.
[0036] The online deployment module deploys the trained congestion control model to the switch of the corresponding node network according to the scale of the data center to which the congestion queue belongs. Each node network determines the congestion control policy corresponding to the current network state based on the congestion control model in the switch, so as to adjust each ECN threshold based on the congestion control policy, thereby solving the problem of continuously changing network congestion.
[0037] Based on the above congestion control system, an embodiment of the present application provides a network congestion control method based on deep learning.
[0038] Refer to Figure 2 , Figure 2 which is a schematic flowchart of the first embodiment of the network congestion control method based on deep learning of the present application.
[0039] In this embodiment, the network congestion control method based on deep learning includes the following steps:
[0040] Step S10, obtain the identifier of the target data center to which the current congestion queue belongs, and determine the model deployment method of the target data center based on the identifier of the target data center;
[0041] In this embodiment, the congestion control module obtains any queue in the currently congested queue as the current congested queue. The online deployment module pre-deploys the model according to the scale of the data center to which each congested queue belongs. After completing the model deployment, it stores each data center, its model deployment method, and the node identifier of the deployed congestion control model in a storage manner such as a data table or a database. The congestion control module first obtains the identifier of the data center to which the current congested queue belongs (i.e., the target data center), and then determines the deployment method of the target data center in the data table based on the identifier of the target data center.
[0042] It can be understood that the processing flow for each congested queue is the same as that of the current congested queue, and multiple congested queues can also be processed in parallel.
[0043] Step S20: Determine the target node based on the model deployment method;
[0044] In this embodiment, the congestion control module determines the model deployment node (i.e., the target node) based on the model deployment method. The model deployment method includes distributed deployment and centralized deployment. The target node can be the node to which the current congestion model belongs or a preset node.
[0045] Step S30: Obtain the current network state of the current congested queue, and determine the optimal congestion control policy corresponding to the current network state based on the congestion control model in the target node;
[0046] In this embodiment, the current network state includes the current queue length of the current congested queue, the output data rate of each link, the high marking threshold, the low marking threshold, the marking probability parameter, etc. The congestion control policy is a state adjustment action, such as adjusting the high marking threshold and / or the low marking threshold of a certain link from the current value to a preset value. The congestion control module determines the congestion control model from the target node. At least one congestion control policy corresponding to the current network state is determined through the congestion control model. When the congestion control policy is unique, it is used as the optimal congestion control policy; when the congestion control policy is not unique, the congestion control model scores each congestion control policy, and the congestion control policy with the highest score is the optimal congestion control policy.
[0047] Step S40: Adjust the congestion flag threshold of the current congested queue based on the optimal congestion control policy to control the network traffic of the current congested queue based on the congestion flag threshold.
[0048] In this embodiment, the congestion control module adjusts the congestion flag threshold of the current congestion queue based on the optimal congestion control strategy, such as lowering or raising the high marking threshold, or raising or lowering the low marking threshold, etc. Thus, by adjusting the high marking threshold and / or the low marking threshold, the network traffic in the current congestion queue is correspondingly adjusted to avoid congestion.
[0049] The above embodiment provides a network congestion control method based on deep learning. The method obtains the identifier of the target data center to which the current congestion queue belongs, and determines the model deployment method of the target data center based on the identifier of the target data center; determines the target node based on the model deployment method; obtains the current network state of the current congestion queue, and determines the optimal congestion control strategy corresponding to the current network state based on the congestion control model in the target node; adjusts the congestion flag threshold of the current congestion queue based on the optimal congestion control strategy to control the network traffic of the current congestion queue based on the congestion flag threshold. By the above method, the embodiment of the present application determines the model deployment method according to the scale of the data center in advance, and then determines the node where the model is located according to the deployment method. Then, according to the real-time situation of the current congestion queue (i.e., the current network state) and the pre-trained congestion control model, the optimal congestion control strategy suitable for the current congestion queue is determined. And based on the optimal congestion control strategy, the congestion flag threshold of the fixed current congestion queue is adjusted in real time. Thus, this embodiment provides a congestion control strategy that can adjust the congestion flag threshold in real time based on a dynamically changing network environment, reduces the adjustment time of ECN parameters, improves the congestion adjustment efficiency and congestion regulation accuracy of network traffic. In addition, further determining the model deployment method according to the scale of the data center avoids improper model deployment and reduces system costs.
[0050] Refer to Figure 3 , Figure 3 which is a schematic diagram of the functional modules of the first embodiment of the network congestion control device based on deep learning of the present application.
[0051] The current congestion control method generally adopts the method of deploying a distributed deployment model to control the data center traffic. For small-scale data centers, a model is deployed on each switch, which increases the system cost.
[0052] In this embodiment, to solve the above problems, step S20 specifically includes:
[0053] Step S21, when the model deployment method is centralized deployment, obtain the master node as the target node.
[0054] Step S22, when the model deployment method is distributed deployment, determine the node to which the current congestion queue belongs as the target node.
[0055] Specifically, if the scale of the data center is larger than a preset scale, such as the data volume is larger than a preset value, in order to improve the congestion control efficiency of the large-scale data center, the congestion control model of the data center is deployed in a distributed manner. When performing distributed deployment, the memory bank information is saved in the cache, the model is trained through the memory bank information, and a congestion control model is deployed in the switch corresponding to each node of the data center. The switch of each node can make decisions independently, and each node controls its own traffic based on the congestion control model stored by itself; if the scale of the data center is not larger than the preset scale, in order to save system costs, the congestion control model of the data center is deployed in a centralized manner. When performing centralized deployment, data is collected through the switches of each node and transmitted to the central controller, and the model is trained through the central controller. And a congestion control model is deployed in the switch corresponding to the main node of the data center or an additional single node set, and all traffic control of the entire network is completed through the congestion control model in this node.
[0056] Specifically, when performing deployment, centralized deployment or distributed deployment is selected according to the scale of the data center to which the congestion queue belongs. And the trained congestion control model is deployed to the switch of the main node network of the centralized deployment or the switches of the networks of each node of the distributed deployment. Based on the congestion control model in the switch, the congestion control strategy corresponding to the current network state is determined, so as to adjust each ECN threshold based on the congestion control strategy, thereby solving the problem of continuously changing network congestion.
[0057] As Figure 4 shown, when performing centralized deployment, network data is collected online through the control interface of each switch, and based on the network data collected by each switch, the memory bank of the central controller is filled. The central controller retrains the model based on the network data in the memory bank to perform the reinforcement learning agent. Then, the congestion control strategy is determined based on the agent, and the ECN threshold is adjusted to perform queue management based on the ECN threshold, that is, to control the network traffic of the congestion queue.
[0058] Through the above method, after training the congestion control model, the congestion control model is deployed into the switch of the main node or each sub-node, and then the congestion control threshold is adjusted according to the real-time traffic state of the network. It not only realizes traffic regulation and avoids congestion; but also flexibly deploys the congestion control model according to the scale of the data center, adopts centralized deployment for the model in small-scale data centers to save costs, and adopts distributed deployment for the model in large-scale data centers to improve efficiency. Thus, while improving the congestion control efficiency, the congestion control cost is reduced.
[0059] Furthermore, when the current control model is trained offline, it is necessary to first build a simulation network locally and then train the model. Not only does it require storing the model and training data in the same environment, reducing the flexibility of the training data, but also it does not use the data of the actual environment to which the model belongs as training data, resulting in the problem of underfitting in the trained model.
[0060] In this embodiment, to solve the above problems, the method is applied to a congestion control system, which includes a client and a server. Data is transmitted between the client and the server through a network. Before determining the optimal congestion control strategy corresponding to the current network state based on the congestion control model in the target node, it further includes:
[0061] Receiving the initial network state data sent by the client, where the initial network state data is collected from the simulation network environment and the actual network environment of the client;
[0062] Determining the congestion control strategy corresponding to the initial network state and sending the congestion control strategy to the client, so that the client performs traffic control based on the congestion control strategy;
[0063] Obtaining the adjusted network state, and using the initial network state data, the congestion control strategy, and the adjusted network state as a set of training data;
[0064] Training the initial congestion control model based on the training data to generate the congestion control model, and storing the congestion control model in the target node.
[0065] Specifically, a dual - end interaction strategy between the client and the server is adopted, that is, the training environment of the congestion control model is divided into the client and the server, and the client and the server are connected through network communication,
[0066] and data interaction and transmission are carried out through the network. The initial network state data is collected in the simulation network environment and the actual network environment in the client, and the collected data is sent to the server; the server feeds back the corresponding congestion control strategy to the client according to the network state data. The client performs traffic control according to the congestion control strategy sent by the server and records the adjusted network state data.
[0067] Taking the received initial network state data, congestion control strategy, and adjusted network state data as training data, and storing the training data in the local memory bank. Thus, the training data in the memory bank is continuously updated, and then the congestion control model is trained in real - time based on the continuously updated training data, and the model after each update is deployed to the target node.
[0068] As shown Figure 5 in the figure, a Server and a Client are constructed. The Server and the Client communicate through a gRPC network, and the data format of the communication content is defined by protobuf. The Client sends a connection request ProtoRequest to the Server, and the Server sends a connection response ProtoResponse to the Client to establish a connection, and then the two ends interact based on the connection. Specifically, the information on the change of the network status collected is sent to the Server, the ECN threshold is adjusted according to the adjustment policy information returned by the Server, and the traffic control of the congestion queue is performed based on the adjusted ECN threshold; the Server is used to send the adjusted ECN marking threshold, and store the received network status and its adjustment policy in the memory bank for congestion control model training. Thus, before training, a simulation network environment is established and a communication network is built.
[0069] Furthermore, taking the initial network status data, the congestion control policy, and the adjusted network status as a set of training data includes:
[0070] Generating a network status vector based on the initial network status data corresponding to each congestion queue;
[0071] Generating the training data based on the network status vector, the congestion control policy corresponding to each group of network status data in the network status vector, and the adjusted network status.
[0072] In this embodiment, multiple sets of training data are input in the form of vectors for unified training. That is, the multiple network status data corresponding to each congestion queue are converted into a network status vector, and based on the congestion control policy corresponding to each network status data, it is converted into a corresponding control policy vector. Based on the network status vector and its corresponding control policy vector, the congestion control model is trained, that is, the status data of each queue is uniformly input in the form of vectors, and the sending rate of data packets is dynamically adjusted according to the congestion degree of each queue during training. Thus, the input operation of training data is simplified, and the input efficiency of training data is improved.
[0073] In the above manner, the training environment is divided into a client and a server. There is no need to set the training data and the model on the same side, which not only can update the training data of the server at any time and improve the flexibility of the training data; but also based on the real data collected by the client in the actual environment required by the model as training data, avoiding the problem of underfitting of the model after training.
[0074] Further, determining the optimal congestion control policy corresponding to the current network state based on the congestion control model includes:
[0075] Obtain the network state parameters in the current network state, where the network state parameters include the current queue length, the output data rate of each link, the initial high marking threshold, the initial low marking threshold, and the initial marking probability;
[0076] Based on the congestion control model and the network state parameters, determine the target adjustment actions for the state parameters to be adjusted, as the optimal congestion control policy, where the state parameters to be adjusted include the initial high marking threshold, the initial low marking threshold, and the initial marking probability.
[0077] In this embodiment, before training the congestion control model based on the reinforcement learning algorithm, first define the network state parameters in the current network state:
[0078] S t ={qlen,txRate,[Kmax 0 ,Kmin 0 ,Pmax 0}
[0079] Where S t is the network state parameter, qlen is the current queue length, txRat is the output data rate of each link, Kmax 0 is the initial high marking threshold, Kmin 0 is the initial low marking threshold, and Pmax 0 is the initial marking probability.
[0080] Input the network state parameters into the congestion control model, that is, based on the initial high marking threshold, the initial low marking threshold, and the initial marking probability, perform network control on the current congestion queue of the current queue length and the output data rate of each link, and network congestion occurs. Therefore, adjust the traffic of the current congestion queue to avoid continued congestion of the current congestion queue. That is, based on the congestion control model, adjust the state parameters to be adjusted in the network state parameters to determine the optimal congestion control policy, that is, the best threshold after adjusting the state parameters to be adjusted:
[0081] a t ={Kmax 1 ,Kmin 1 ,Pmax 1}
[0082] Where a t is the state parameter to be adjusted after executing the target adjustment action, Kmax 1 is the adjusted high marking threshold, and Kmin 1is the adjusted low marking threshold, Pmax 1 is the adjusted marking probability.
[0083] It can be understood that the adjusted optimal threshold can be any one of the initial high marking threshold, the initial low marking threshold, and the initial marking probability, or two of them, or all of them.
[0084] Based on the difference between the initial high marking threshold and the adjusted high marking threshold, the difference between the initial low marking threshold and the adjusted low marking threshold, and the difference between the initial marking probability and the adjusted probability threshold, determine the target adjustment action corresponding to the state parameter to be adjusted.
[0085] In the above manner, in this embodiment, the network state parameters are defined based on the necessary parameters, unnecessary parameters are removed, the complexity of the parameter space during network training is reduced, and the overhead of the training model is reduced.
[0086] Further, determining the target adjustment actions of the initial high marking threshold, the initial low marking threshold, and the initial marking probability based on the congestion control model and the network state parameters includes:
[0087] Based on the congestion control model, determine at least one adjustment action corresponding to the state parameter to be adjusted and its corresponding reward value;
[0088] Among the adjustment actions, determine the adjustment action corresponding to the maximum reward value as the target adjustment action.
[0089] In this embodiment, while the congestion control model determines each adjustment action corresponding to the state parameter to be adjusted, it generates the reward value of each adjustment action. That is, the client adjusts the network according to each adjustment action and feeds back the congestion situation of the adjustment to the server, and the model scores the adjustment action according to the adjusted congestion situation. And determine the adjustment action corresponding to the maximum reward value as the target adjustment action.
[0090] The reward value of the adjustment action is the reward value of the adjusted network state, and the reward value of the adjusted network state is calculated according to the reward function. Among them, the calculation formula of the reward function is:
[0091] R = ε 1 × L + ε 2 × T
[0092]
[0093]
[0094] Among them, ε 1 and ε 2ε is a weighted parameter for balancing throughput and latency parameters 1 and ε 2 The sum of them is 1; L is a latency parameter; T is the link utilization; BW is the total link bandwidth.
[0095] Furthermore, before determining the optimal congestion control strategy corresponding to the current network state based on the congestion control model in the target node, it further includes:[[]]
[0096] Establish a simulation network environment based on historical environment parameters in the actual network environment;
[0097] Determine a reward function based on the network characteristic parameters of the simulation network environment, and calculate the reward value of the model to be trained based on the reward function;
[0098] Determine the loss function of the model to be trained based on the reward value;
[0099] Train the model to be trained based on the training samples in the sample library until the loss function value of the model to be trained is less than a preset value to obtain the congestion control model.
[0100] In this embodiment, first, the environment of the network congestion control model is modeled, and historical data of the client and the server in different network environments are collected, such as network topology data, transmission delay data, data traffic, and link capacity. Normalize the values of the historical data, and establish a simulation network environment required for model training based on the normalized historical data.
[0101] Then initialize the congestion control model, and define the network state parameters and the state parameters to be adjusted of the congestion control model:
[0102] S t ={qlen, txRate, [Kmax 0 , Kmin 0 , Pmax 0}
[0103] a t ={Kmax 1 , Kmin 1 , Pmax 1}
[0104] Among them, S t is the network state parameter, and a t is the state parameter after performing the adjustment action.
[0105] Next, a four-layer fully connected neural network is constructed, which includes an evaluation network and a target network. The network structures of the evaluation network and the target network are the same. Initialize the parameters such as the network structure and learning rate of the two networks.
[0106] Use the reinforcement learning algorithm DQN (Deep Q Network, a neural network algorithm based on deep learning) to train znengti in the congestion control model;
[0107] Input each network state parameter S t into the initialized neural network to obtain the state parameter a to be adjusted output by the neural network t and its vector of evaluation values. Use the DQN algorithm to train the two networks to continuously update the weights of the evaluation network, and keep the target network unchanged within a preset time period to reduce the correlation between the evaluation network value and the target network value and improve the stability of the training algorithm. Stop training until the mean square error between the predicted value obtained by the evaluation network and the target value obtained by the target network reaches the minimum, and use the trained neural network as the congestion control model.
[0108] Specifically, in the separate training environment, adopt the dual-end interaction strategy of the client and the server. The client and the server communicate through the network, and define the communication content data in the protobuf (Protocol Buffers, a structured data storage format). The client sends a message including initial state data, adjusted state data, and a flag bit indicating whether the model training is completed, etc. to the server through protobuf, and the server sends a message including action state data corresponding to the initial state data, the reward value of the action state data, etc. to the client.
[0109] During separate training, the simulation network environment and the actual network environment are located on the client, and the training data is stored on the server, and the model training is carried out on the server.
[0110] Such as Figure 6As shown, the client Client first initializes the environment (including the current queue length, the output data rate of each link, the initial high marking threshold, the initial low marking threshold, and the initial marking probability). That is, the client sends the current network status to the server Server through the communication connection with the server. After receiving the current network status, the server outputs an action to the client through the reinforcement learning agent, that is, based on the agent trained by reinforcement learning, it outputs the threshold adjustment action corresponding to the current network status to the client. The client receives the action information and interacts with the environment based on this action information, that is, the client adjusts the threshold based on the action information, such as increasing the high marking threshold by a preset value and decreasing the low marking threshold by a preset value, etc. The client sends the adjusted network status, the current network status, and its corresponding adjustment action information to the server.
[0111] Store the current network status state, its corresponding adjustment action information action, the adjusted network status obs_state, the reward value rewards, and the completion flag obs_done as a set of training data in the memory bank of the server. Among them, obs_done is the flag bit indicating whether the training is completed. When this set of data is the last set of training data, it is the completion identifier, and the flags of other sets of training are all uncompleted. The server stores each set of training data in the memory bank in the format of state, action, obs_state, rewards, obs_done.
[0112] The client sends the actual network status to the server in real time as training data. The server regularly trains the agent in the congestion control model according to the training data in the memory bank. Thus, the client and the server continuously interact to obtain the best reinforcement learning agent through training with real data.
[0113] Thus, run the environment program, the agent interacts with the environment, collects sample data and saves it to the memory bank. When the amount of sample data in the memory bank reaches the preset value, start training. During training, randomly sample a preset amount of data from the memory bank. For example, the number of samples drawn is 32, that is, randomly select 32 sets of state data in the memory bank each time as the input of the first layer of the neural network. The output after passing through four fully connected network layers is the Q value corresponding to different adjustment actions. The Q value is the expected benefit that can be obtained by executing a certain adjustment action at a certain moment.
[0114] Use the reward function and the loss function to adjust and update the parameters of the neural network for training, that is, update and train the model according to the value of the loss function until the value of the loss function is minimized to complete the update. Among them, the calculation formula of the loss function is:
[0115]
[0116] Among them, E is the expectation, and r t is the reward value of the congestion control model at time t; γ is the return discount factor. The closer the discount factor is to 0, the more the optimization goal focuses on short-term benefits. The closer it is to 1, the more the optimization goal focuses on long-term cumulative benefits. Users can set the coefficient according to the actual model training time requirements; Q represents the function corresponding to each group of network state parameters; s t+1 is the network state parameter at time t + 1; a represents the set of adjustment actions; θ and θ′ are the weight parameters of the target network and the evaluation network respectively. That is, at time t, an action a t is performed on the parameter to be adjusted, and the reward value r t and the next state s t+1 are obtained.
[0117] The reward value of the congestion control model is the reward value of the adjusted network state.
[0118] After calculating the value function, DQN adopts a greedy strategy to accelerate the algorithm convergence. Taking the probability □ as the criterion, a random adjustment action is taken; or an adjustment action with the largest Q value is selected with a probability of 1 - □.
[0119] Specifically, as Figure 7 shown, the state (i.e., the network state parameter) is obtained from the simulation network environment, and each action (i.e., the adjustment action) is performed on the network state parameter to obtain the adjusted state parameter, and the Q value of all adjusted state parameters is calculated by the agent. Then, the adjustment action is selected according to the greedy strategy to make a decision. The ECN marking threshold is adjusted according to the adjustment action, and the reward value and the next state value are calculated according to the adjusted state parameter to complete a training step. The model training parameters are updated according to the reward value, and then the above steps are repeated for training until the optimal reinforcement learning agent is trained to complete the congestion control model training. Thus, through the congestion control model, multiple queues can be controlled in real time. When a congested queue is detected, its bandwidth can be reduced or its sending rate can be restricted, thereby alleviating the congestion phenomenon. At the same time, non-congested queues can continue to send data packets to maintain their normal sending rate.
[0120] Furthermore, after completing the congestion control model training, a traffic congestion control model verified on historical data is obtained. Then, when the data center scale is small, the model is deployed to the central controller (i.e., the switch where the master node belongs). After each switch obtains the network state information, it is transmitted to the central controller as sample data and saved in the memory bank. Sample data is regularly taken from the memory bank to train the congestion control model in the central controller.
[0121] In the above manner, this embodiment adopts a method based on reinforcement learning to implement ECN traffic control. The sender sends data packets, and the reinforcement learning model controls the ECN configuration policy according to the current network state, controls the network traffic in real time, and avoids congestion. The congestion control model is separately trained in a simulation environment, and a deep reinforcement learning algorithm is used to optimize the traffic control policy. Then, the trained model is applied to the actual network environment, and online training is carried out, and the policy is continuously optimized according to the feedback information of the actual network. In this way, it is possible to dynamically adjust the traffic control policy in a real network to cope with different network conditions and congestion levels, and improve network performance and stability.
[0122] In addition, an embodiment of the present application also provides a computer-readable storage medium.
[0123] A network congestion control program based on deep learning is stored on the computer-readable storage medium of the present application. When the network congestion control program based on deep learning is executed by a processor, the steps of the network congestion control method based on deep learning as described above are implemented.
[0124] Among them, the method implemented when the network congestion control program based on deep learning is executed can refer to each embodiment of the network congestion control method based on deep learning of the present application, and will not be elaborated here.
[0125] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.
[0126] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0127] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0128] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of this application.
[0129] The above are only the preferred embodiments of this application, and do not limit the patent scope of this application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.
Claims
1. A network congestion control method based on deep learning, characterized in that, the network congestion control method includes: obtaining the identifier of the target data center to which the current congestion queue belongs, and determining the model deployment method of the target data center based on the identifier of the target data center; determining a target node based on the model deployment method; obtaining the current network state of the current congestion queue, and determining the optimal congestion control policy corresponding to the current network state based on the congestion control model in the target node; adjusting the congestion flag threshold of the current congestion queue based on the optimal congestion control policy, so as to control the network traffic of the current congestion queue based on the congestion flag threshold.
2. The network congestion control method according to claim 1, characterized in that, the determining a target node based on the model deployment method includes: when the model deployment method is centralized deployment, obtaining the master node as the target node.
3. The network congestion control method according to claim 1, characterized in that, the determining a target node based on the model deployment method includes: when the model deployment method is distributed deployment, determining the node to which the current congestion queue belongs as the target node.
4. The network congestion control method according to claim 1, characterized in that, the method is applied to a congestion control system, the congestion control system includes a client and a server, data is transmitted between the client and the server through a network, and before the step of determining the optimal congestion control policy corresponding to the current network state based on the congestion control model in the target node, it further includes: receiving the initial network state data sent by the client, wherein the initial network state data is collected from the simulation network environment and the actual network environment of the client; determining the congestion control policy corresponding to the initial network state, and sending the congestion control policy to the client, so that the client performs traffic control based on the congestion control policy; obtaining the adjusted network state, and using the initial network state data, the congestion control policy, and the adjusted network state as a set of training data; training the initial congestion control model based on the training data to generate the congestion control model, and storing the congestion control model in the target node.
5. The network congestion control method according to claim 4, characterized in that, using the initial network state data, the congestion control policy, and the adjusted network state as a set of training data includes: generating a network state vector based on the initial network state data corresponding to each congestion queue; generating the training data based on the network state vector, the congestion control policy corresponding to each group of network state data in the network state vector, and the adjusted network state.
6. The network congestion control method according to claim 1, characterized in that, the determining the optimal congestion control policy corresponding to the current network state based on the congestion control model includes: Obtain the network state parameters in the current network state, where the network state parameters include the current queue length, the output data rate of each link, the initial high marking threshold, the initial low marking threshold, and the initial marking probability; Based on the congestion control model and the network state parameters, determine the target adjustment actions of the state parameters to be adjusted as the optimal congestion control strategy, where the state parameters to be adjusted include the initial high marking threshold, the initial low marking threshold, and the initial marking probability.
7. The network congestion control method according to claim 6, characterized in that The determining the target adjustment actions of the initial high marking threshold, the initial low marking threshold, and the initial marking probability based on the congestion control model and the network state parameters includes: Based on the congestion control model, determine at least one adjustment action corresponding to the state parameter to be adjusted and its corresponding reward value; Among the respective adjustment actions, determine the adjustment action corresponding to the maximum reward value as the target adjustment action.
8. The network congestion control method according to any one of claims 1-7, characterized in that Before determining the optimal congestion control strategy corresponding to the current network state based on the congestion control model in the target node, it further includes: Based on the historical environmental parameters in the actual network environment, establish a simulation network environment; Based on the network characteristic parameters of the simulation network environment, determine a reward function, and based on the reward function, calculate the reward value of the model to be trained; Based on the reward value, determine the loss function of the model to be trained; Based on the training samples in the sample library, train the model to be trained until the loss function value of the model to be trained is less than a preset value to obtain the congestion control model.
9. A network congestion control device based on deep learning, characterized in that The network congestion control device based on deep learning includes a processor, a memory, and a network congestion control program based on deep learning stored on the memory and executable by the processor. When the network congestion control program based on deep learning is executed by the processor, the steps of the network congestion control method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium, characterized in that A network congestion control program based on deep learning is stored on the computer-readable storage medium. When the network congestion control program based on deep learning is executed by a processor, the steps of the network congestion control method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Congestion control method and device of network equipment, computer equipment and medium
CN120416166A