Network traffic optimization scheduling method based on deep reinforcement learning and suitable for periodic traffic characteristics
By introducing deep reinforcement learning technology in network traffic scheduling, using the LSTM model to predict periodic traffic and combining the DQN model for path selection, the lag problem of periodic traffic scheduling in the existing technology is solved, and more efficient resource utilization and network performance are achieved.
Patent Information
- Application Number
- CN202510276281.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-03
AI Technical Summary
In the face of dynamic and complex network environments, it is difficult for the prior art to achieve efficient resource utilization and load balancing, especially in periodic traffic scenarios, which lacks the ability to predict traffic changes, resulting in lag in scheduling strategies.
The network traffic optimization scheduling method based on deep reinforcement learning is adopted. By collecting historical periodic traffic data, an LSTM model is built for traffic prediction, and the network status is monitored in real time with SDN technology, and the DQN model is used to select the optimal path for scheduling.
It realizes prediction and real-time response to periodic traffic changes, improves resource utilization and network performance, and enhances the flexibility and adaptability of scheduling strategies.
Smart Images

Figure CN120090989A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communication technologies, and particularly relates to a network traffic optimization and scheduling method based on deep reinforcement learning applicable to periodic traffic characteristics. Background Art
[0002] With the continuous expansion of the scale of data center networks and the rapid growth of network traffic, traditional traffic scheduling methods face many challenges in dealing with dynamic and complex network environments. Existing traffic scheduling technologies mainly include rule-based methods (such as Equal-Cost Multi-Path Routing - ECMP) and heuristic algorithm-based methods (such as Shortest Path First - SPF). Although these methods perform well in simple network environments, they often struggle to achieve efficient resource utilization and load balancing when faced with dynamically changing network traffic. Specifically, the existing technologies have the following deficiencies: Limitations of ECMP: ECMP evenly distributes traffic to multiple paths through a hash algorithm, but it cannot distinguish between large flows and small flows, resulting in large flows possibly occupying too much bandwidth and affecting the transmission efficiency of small flows. In addition, ECMP lacks the ability to perceive the network state in real time and cannot dynamically adjust path selection. Lack of prediction ability: Existing methods fail to fully utilize the periodic characteristics of historical traffic data and cannot predict traffic changes in advance, resulting in the lag of scheduling strategies. For example, when faced with periodic traffic, existing methods cannot optimize path selection in advance, leading to low resource utilization.
[0003] In recent years, the rapid development of software-defined network (SDN) and deep reinforcement learning (DRL) technologies has provided new solutions for network traffic scheduling. SDN realizes real-time monitoring and flexible scheduling of network states through centralized control, while DRL can adapt to complex network environments through intelligent decision-making capabilities. However, existing traffic scheduling methods based on SDN and DRL still have the following problems: Poor dynamic adaptability: Existing methods fail to adjust scheduling strategies in real time when faced with dynamic changes in the network environment, resulting in low resource utilization and degraded network performance. Summary of the Invention
[0004] To overcome the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to provide a network traffic optimization and scheduling method based on deep reinforcement learning applicable to periodic traffic characteristics.
[0005] The technical solution of the present invention is as follows:
[0006] A network traffic optimization and scheduling method based on deep reinforcement learning applicable to periodic traffic characteristics, characterized by including the following stages:
[0007] I. Preparation stage:
[0008] Step 1: Collect historical periodic traffic data; monitor the network status in real time through the SDN controller to obtain historical periodic traffic data, including delay, bandwidth utilization, and load information;
[0009] Step 2: Data preprocessing, clean, normalize, and perform time series segmentation on the collected historical traffic data;
[0010] Step 3: Build an LSTM model, use the preprocessed historical data to build an LSTM model, and define input features and output labels;
[0011] Step 4: Train the LSTM model, use the training set data to train the LSTM model, and generate traffic prediction data for the next period;
[0012] Step 5: Build a state space, fuse the predicted traffic data with the real-time traffic data to build a state space;
[0013] Step 6: Train the DQN model, use the constructed state space to train the Deep Q-Network (DQN) model;
[0014] II. Scheduling stage:
[0015] Step 1: Obtain network topology and link status information, obtain the current network topology through the topology discovery module of the SDN controller, and collect link bandwidth utilization, delay, and packet loss rate information in real time through the traffic monitoring module;
[0016] Step 2: Distinguish between large flows and small flows;
[0017] Step 3: Use ECMP for scheduling of small flows;
[0018] Step 4: Use the DQN model for scheduling of large flows.
[0019] 2. According to the network traffic optimization scheduling method based on deep reinforcement learning applicable to periodic traffic characteristics described in claim 1, wherein the preparation stage:
[0020] (1) Collect historical periodic traffic data;
[0021] Monitor the network status in real time through the SDN controller to obtain historical periodic traffic data, including delay, bandwidth utilization, and load information. The data collection frequency is once per second to ensure the timeliness and integrity of the data; the storage format of historical data is:
[0022] (1);
[0023] wherein represents the timestamp, represents the delay, Indicates the bandwidth utilization rate, Indicates the load information;
[0024] (2)Data preprocessing:
[0025] Preprocess the collected historical traffic data, specifically including:
[0026] Data cleaning: Remove outliers and missing values to ensure data quality;
[0027] Data normalization: Scale the data to the range [0, 1]. The normalization formula is:
[0028] (2);
[0029] Where, is the original data, and are the minimum and maximum values of the data respectively;
[0030] Time series segmentation: Segment the normalized data into a training set and a validation set according to the time series. The training set accounts for 80% and the validation set accounts for 20%;
[0031] (3)Build an LSTM model;
[0032] Build an LSTM model using the preprocessed historical data, specifically including:
[0033] Define the input features and output labels: The input features are historical latency, bandwidth utilization rate, and load information, and the output label is the traffic prediction value for the next period; The input features and the output label Y are in the format of:
[0034] (3);
[0035] (4);
[0036] Initialize the model parameters: Initialize the parameters of the LSTM model, including the number of hidden units , time step and learning rate ; The calculation formula for the number of hidden layer units is:
[0037] (5);
[0038] Where: is the number of training set samples; Indicates rounding down;
[0039] The purpose of this formula is to dynamically adjust the number of hidden layer units according to the data scale to avoid the model being too complex or too simple;
[0040] Time step : Time step is the length of the input sequence of the LSTM model and is usually determined according to the periodicity or temporal correlation of the data; if the data has an obvious periodicity (such as daily or weekly traffic patterns), the time step can be set to the length of the period; for example:
[0041] If the data is collected hourly and has a daily periodicity, then ;
[0042] If the data is collected every minute and has an hourly periodicity, then ;
[0043] If there is no obvious periodicity, the appropriate time step can be selected through experiments, and the common value range is ;
[0044] Learning rate : Learning rate is the step size for parameter update during model training. Usually, the initial value is selected based on experience and dynamically adjusted during training. A common method for selecting the initial learning rate is:
[0045] (6);
[0046] where: is the dimension of the input features (i.e., the number of features); its purpose is to avoid the learning rate being too large or too small;
[0047] (4) Train the LSTM model;
[0048] Use the training set data to train the LSTM model, which specifically includes:
[0049] Forward propagation: Calculate the output of the LSTM model, and the formula is:
[0050] (7);
[0051] (8);
[0052] (9);
[0053] (10);
[0054] (11);
[0055] (12);
[0056] where , , represent the forget gate, input gate, and output gate respectively, represents the cell state, represents the hidden state, and are the weight matrix and bias vector respectively, represents the Sigmoid activation function;
[0057] Backpropagation: Optimize the model parameters through the backpropagation algorithm. The loss function is the mean squared error (MSE), and the formula is:
[0058] (13);
[0059] where is the true value, is the predicted value, is the number of samples;
[0060] Model evaluation: Use the validation set to evaluate the prediction accuracy of the LSTM model. The accuracy calculation formula is:
[0061] (14);
[0062] where is the number of samples in the validation set;
[0063] (5) Generate prediction data;
[0064] Use the trained LSTM model to generate traffic prediction data for the next period. The format of the prediction data is:
[0065] (15);
[0066] where represent the predicted traffic values for the next period respectively;
[0067] Dynamic weight adjustment mechanism: Integrate the predicted traffic data with the real-time traffic data to construct the state space , and the formula for constructing the state space is:
[0068] (16);
[0069] (17);
[0070] where represents the weight of the prediction data; represents the weight of the real-time data. The prediction accuracy is calculated through the validation set, and the formula is:
[0071] (18);
[0072] Wherein, is the actual traffic value, is the predicted traffic value, is the number of samples in the validation set;
[0073] The real-time traffic fluctuation coefficient is obtained by calculating the standard deviation of the real-time traffic, and the formula is:
[0074] (19);
[0075] Wherein, is the mean value of the real-time traffic, is the number of real-time traffic samples;
[0076] (6) Train the DQN model;
[0077] Use the constructed state space to train the Deep Q-Network (DQN) model, which specifically includes:
[0078] Define the state space and action space: The state space includes the bandwidth utilization rate, delay, packet loss rate of the current link, and predicted traffic data; The action space includes the optional paths from the source node to the destination node;
[0079] Reward function design: The reward function Comprehensively consider the delay, bandwidth utilization rate, and packet loss rate. The formula of the reward function is:
[0080] (20);
[0081] Wherein, , , are weight parameters;
[0082] Experience replay mechanism: Store the historical state, action, reward, and next state in the replay buffer, and randomly sample data from the buffer during training. The target value calculation formula is:
[0083] (21);
[0084] Wherein, represents the immediate reward, represents the discount factor, represents the parameters of the target network.
[0085] 3. A network traffic optimization scheduling method based on deep reinforcement learning applicable to periodic traffic characteristics according to claim 1, wherein, in the said scheduling phase:
[0086] (1) Obtain network topology and link state information;
[0087] Obtain the current network topology through the topology discovery module of the SDN controller, including the connection relationship and link state information between switches; collect the bandwidth utilization, delay, and packet loss rate information of the link in real time through the traffic monitoring module, and store this information as a state data set;
[0088] (2) Distinguish large flows and small flows;
[0089] According to the traffic volume Distinguish large flows and small flows, and set the threshold of the traffic volume When it is determined as a large flow; otherwise it is determined as a small flow; the threshold can be dynamically adjusted according to the network environment and service requirements;
[0090] (3) Schedule small flows using ECMP;
[0091] For the traffic determined to be small flows, use Equal-Cost Multi-Path Routing (ECMP) for scheduling, specifically including:
[0092] Calculate multiple equivalent paths from the source node to the destination node;
[0093] Use the hash algorithm to evenly distribute small flows to multiple paths, and the hash function is:
[0094] (22);
[0095] Among them, is the unique identifier of the flow, is the number of equivalent paths;
[0096] (4) Schedule large flows using the DQN model;
[0097] For the traffic determined to be large flows, use the trained Deep Q-Network (DQN) model for scheduling, specifically including:
[0098] Construct the state space : The state space includes the bandwidth utilization, delay, packet loss rate of the current link, and the predicted traffic data generated by the LSTM model; the construction formula of the state space is:
[0099] (23);
[0100] Among them, represents the weight of the predicted data; represents the weight of the real-time data;
[0101] Define the state space : Action space Includes optional paths from the source node to the destination node, and each path corresponds to an action ;
[0102] Select the optimal path: Use the trained DQN model to select the optimal action according to the current state space That is, the optimal path; The action selection formula is: That is, the optimal path; The action selection formula is:
[0103] (24);
[0104] Wherein, Represents the parameters of the DQN model, Represents at state Select action Of Value;
[0105] Reward function design: The reward function Comprehensively consider the delay, bandwidth utilization rate and packet loss rate. The reward function formula is:
[0106] (25);
[0107] Wherein, , , Are weight parameters used to balance the importance of different optimization objectives;
[0108] Experience replay mechanism: Store the historical state, action, reward and next state in the replay buffer. During training, randomly sample data from the buffer. The target Value calculation formula is:
[0109] (26);
[0110] Wherein, Represents the immediate reward, Represents the discount factor, Represents the parameters of the target network;
[0111] (5) Real-time update the network state information;
[0112] Regularly (such as every few seconds) obtain the latest link state information through the SDN controller, and dynamically adjust the scheduling strategy to adapt to the changes in the network environment; The update process includes:
[0113] Recalculate the bandwidth utilization rate, delay and packet loss rate of the link;
[0114] According to the latest state space Re - select the optimal path.
[0115] Compared with the prior art, the present invention has the following advantages:
[0116] 1. Starting from historical data, the present invention utilizes the periodic characteristics of historical traffic data to predict the changing trend of future traffic, and to a certain extent, solves the lag in the collection of link traffic information by the SDN monitoring module.
[0117] 2. The present invention introduces a dynamic weight adjustment mechanism, which dynamically adjusts the weights of predicted data and real - time data according to the prediction accuracy, improving the flexibility and adaptability of the scheduling strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0118] Figure 1 It is a schematic architecture diagram of a network traffic optimization scheduling method based on deep reinforcement learning applicable to the periodic traffic characteristic scenario provided by an embodiment of the present invention;
[0119] Figure 2 It is a flowchart of the preparation stage of a network traffic optimization scheduling method based on deep reinforcement learning applicable to the periodic traffic characteristic scenario provided by an embodiment of the present invention;
[0120] Figure 3 It is a flowchart of the scheduling stage of a network traffic optimization scheduling method based on deep reinforcement learning applicable to the periodic traffic characteristic scenario provided by an embodiment of the present invention;
[0121] Figure 4 It is a network topology diagram used in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0122] The present invention will be further described below in conjunction with the drawings and embodiments.
[0123] As Figure 1 shown, a network traffic optimization scheduling method based on deep reinforcement learning applicable to the periodic traffic characteristic scenario includes the following steps:
[0124] Preparation stage:
[0125] Step 1: Collect historical periodic traffic data. The SDN controller is used to monitor the network status in real - time to obtain historical periodic traffic data, including delay, bandwidth utilization rate, and load information.
[0126] Step 2: Data pre - processing. The collected historical traffic data is cleaned, normalized, and segmented into time series.
[0127] Step 3: Build an LSTM model. The pre - processed historical data is used to build an LSTM model, and the input features and output labels are defined.
[0128] Step 4: Train the LSTM model. Use the training set data to train the LSTM model to generate traffic prediction data for the next period.
[0129] Step 5: Construct the state space. Integrate the predicted traffic data with the real-time traffic data to construct the state space.
[0130] Step 6: Train the DQN model. Use the constructed state space to train the Deep Q-Network (DQN) model.
[0131] Furthermore, the specific steps of Step 1 include:
[0132] Collect historical periodic traffic data;
[0133] Monitor the network status in real time through the SDN controller to obtain historical periodic traffic data, including delay, bandwidth utilization, and load information. The data collection frequency is once per second to ensure the timeliness and integrity of the data. The storage format of historical data is:
[0134] (1);
[0135] where represents the timestamp, represents the delay, represents the bandwidth utilization, represents the load information
[0136] Furthermore, the specific steps of Step 2 include:
[0137] Data preprocessing
[0138] Preprocess the collected historical traffic data, which specifically includes:
[0139] Data cleaning: Remove outliers and missing values to ensure data quality
[0140] Data normalization: Scale the data to the range of [0, 1]. The normalization formula is:
[0141] (2);
[0142] where, is the original data, and are the minimum and maximum values of the data respectively;
[0143] Time series segmentation: Segment the normalized data into a training set and a validation set according to the time series. The training set accounts for 80% and the validation set accounts for 20%.
[0144] Furthermore, the specific steps of Step 3 include:
[0145] Build an LSTM model using preprocessed historical data, specifically including:
[0146] Define the input features and output labels: The input features are historical latency, bandwidth utilization, and load information, and the output label is the traffic prediction value for the next period. The input features and the output label Y are in the format of:
[0147] (3);
[0148] (4);
[0149] Initialize the model parameters: Initialize the parameters of the LSTM model, including the hidden units , the time step and the learning rate . The calculation formula for the number of hidden layer units is:
[0150] (5);
[0151] Where: is the number of samples in the training set; represents rounding down.
[0152] The purpose of this formula is to dynamically adjust the number of hidden layer units according to the data scale to avoid the model being too complex or too simple.
[0153] The time step : The time step is the length of the input sequence of the LSTM model, usually determined according to the periodicity or time correlation of the data. If the data has obvious periodicity (such as daily or weekly traffic patterns), the time step can be set to the length of the period. For example:
[0154] If the data is collected once per hour and has a daily periodicity, then ;
[0155] If the data is collected once per minute and has an hourly periodicity, then .
[0156] If there is no obvious periodicity, the appropriate time step can be selected through experiments, usually in the range of ;
[0157] The learning rate : The learning rate is the step size for parameter update during model training. Usually, the initial value is selected based on experience and dynamically adjusted during training. A common method for selecting the initial learning rate is:
[0158] (6);
[0159] Wherein: is the dimension of the input features (i.e., the number of features). Its purpose is to avoid the learning rate being too large or too small.
[0160] Further, the specific steps of Step 4 include:
[0161] Using the training set data to train the LSTM model, specifically including:
[0162] Forward propagation: Calculate the output of the LSTM model, and the formula is:
[0163] (7);
[0164] (8);
[0165] (9);
[0166] (10);
[0167] (11);
[0168] (12);
[0169] Where , , represent the forget gate, input gate, and output gate respectively, represents the cell state, represents the hidden state, and are the weight matrix and bias vector respectively, represents the Sigmoid activation function.
[0170] Backward propagation: Optimize the model parameters through the backpropagation algorithm, and the loss function is the mean squared error (MSE), and the formula is:
[0171] (13);
[0172] Wherein, is the true value, is the predicted value, is the number of samples.
[0173] Model evaluation: Use the validation set to evaluate the prediction accuracy of the LSTM model, and the accuracy calculation formula is:
[0174] (14);
[0175] Wherein, is the number of samples in the validation set.
[0176] Further, the specific steps of Step Five include:
[0177] Use the trained LSTM model to generate traffic prediction data for the next period. The format of the prediction data is:
[0178] (15);
[0179] Where respectively represent the predicted traffic values for the next period.
[0180] Dynamic weight adjustment mechanism: Integrate the predicted traffic data with the real-time traffic data to construct a state space , and the construction formula of the state space is:
[0181] (16);
[0182] (17);
[0183] Where, represents the weight of the prediction data; represents the weight of the real-time data. The prediction accuracy is calculated through the validation set, and the formula is:
[0184] (18)
[0185] Where, is the real traffic value, is the predicted traffic value, is the number of samples in the validation set;
[0186] The real-time traffic fluctuation coefficient is obtained by calculating the standard deviation of the real-time traffic, and the formula is:
[0187] (19);
[0188] Where, is the mean value of the real-time traffic, is the number of real-time traffic samples.
[0189] Further, the specific steps of Step Six include:
[0190] Use the constructed state space to train a Deep Q-Network (DQN) model, specifically including:
[0191] Define the state space and action space: The state space includes the bandwidth utilization rate, delay, packet loss rate of the current link, and the predicted traffic data; the action space includes the optional paths from the source node to the destination node.
[0192] Reward Function Design: Reward Function Taking into account latency, bandwidth utilization, and packet loss rate comprehensively, the reward function formula is as follows:
[0193] (20);
[0194] Among them, , , are weight parameters.
[0195] Experience Replay Mechanism: Store historical states, actions, rewards, and next states in a replay buffer. During training, randomly sample data from the buffer. The target value calculation formula is as follows:
[0196] (21);
[0197] Among them, represents the immediate reward, represents the discount factor, represents the parameters of the target network,
[0198] Scheduling Phase:
[0199] Step 1: Obtain network topology and link state information. Obtain the current network topology through the topology discovery module of the SDN controller, and collect real-time bandwidth utilization, latency, and packet loss rate information of the links through the traffic monitoring module.
[0200] Step 2: Distinguish between large flows and small flows.
[0201] Step 3: Schedule small flows using ECMP.
[0202] Step 4: Schedule large flows using the DQN model.
[0203] Furthermore, the specific content of Step 1 includes:
[0204] Obtain the current network topology through the topology discovery module of the SDN controller, including the connection relationship between switches and link state information; collect real-time bandwidth utilization, latency, and packet loss rate information of the links through the traffic monitoring module, and store this information as a state dataset.
[0205] Furthermore, the specific content of Step 2 includes:
[0206] Distinguish between large flows and small flows according to the traffic volume Set a threshold for the traffic volume , when , it is determined as a large flow; otherwise, it is determined as a small flow. The threshold can be dynamically adjusted according to the network environment and business requirements.
[0207] Further, step three specifically includes:
[0208] For the traffic determined to be small flows, use Equal-Cost Multi-Path Routing (ECMP) for scheduling, which specifically includes:
[0209] Calculate multiple equivalent paths from the source node to the destination node;
[0210] Use the hash algorithm to evenly distribute small flows onto multiple paths. The hash function is:
[0211] (22);
[0212] Wherein, is the unique identifier of the flow, is the number of equivalent paths.
[0213] Further, step four specifically includes:
[0214] For the traffic determined to be large flows, use the trained Deep Q-Network (DQN) model for scheduling, which specifically includes:
[0215] Construct the state space : The state space includes the bandwidth utilization, delay, packet loss rate of the current link, and the predicted traffic data generated by the LSTM model. The construction formula of the state space is:
[0216] (23);
[0217] Wherein, represents the weight of the predicted data; represents the weight of the real-time data.
[0218] Define the state space : The action space includes the optional paths from the source node to the destination node, and each path corresponds to an action .
[0219] Select the optimal path: Use the trained DQN model to select the optimal action according to the current state space i.e., the optimal path. The action selection formula is:
[0220] (24);
[0221] Wherein, represents the parameters of the DQN model, represents the probability of selecting action in state of Value
[0222] Reward function design: Reward function Taking into account latency, bandwidth utilization, and packet loss rate comprehensively, the reward function formula is:
[0223] (25);
[0224] Wherein, , , are weight parameters used to balance the importance of different optimization objectives.
[0225] Experience replay mechanism: Store historical states, actions, rewards, and next states in a replay buffer. During training, randomly sample data from the buffer. The target value calculation formula is:
[0226] (26);
[0227] Wherein, represents the immediate reward, represents the discount factor, represents the parameters of the target network,
[0228] Further, the specific steps of step five include:
[0229] Regularly (such as every few seconds) obtain the latest link state information through the SDN controller, and dynamically adjust the scheduling strategy to adapt to changes in the network environment. The update process includes:
[0230] Recalculate the bandwidth utilization, latency, and packet loss rate of the link;
[0231] According to the latest state space Re-select the optimal path.
[0232] Such as Figure 1As shown in the figure. The data exchange layer, as the underlying infrastructure, is responsible for the actual transmission and exchange of network traffic. It realizes the efficient forwarding of data packets through physical network devices (such as switches and routers), and real-time collects information such as link bandwidth utilization, latency, and packet loss rate. The control layer, as the core scheduling center, realizes centralized network control based on software-defined network (SDN) technology. It is responsible for topology discovery, traffic monitoring, path calculation, and the issuance of scheduling strategies, and dynamically adjusts network configurations to ensure the real-time and effectiveness of scheduling strategies. The agent layer, as the decision-making brain, realizes intelligent traffic scheduling based on deep reinforcement learning technology. It predicts traffic changes through a long short-term memory (LSTM) model, constructs a state space, and uses a deep Q-network (DQN) model to select the optimal path. At the same time, it continuously optimizes the scheduling strategy through real-time interaction with the network environment to ensure the optimality of network performance. These three planes work together to jointly achieve the efficient scheduling of network traffic and resource optimization.
[0233] The present invention will be further described in detail below in combination with the system structure and specific implementation steps:
[0234] Preparation stage:
[0235] (1) Collect historical periodic traffic data;
[0236] In a Fat-Tree structure with K = 4, the network is divided into a core layer, an aggregation layer, and an edge layer. The SDN controller monitors the network status of each layer in real time to obtain historical periodic traffic data, including latency, bandwidth utilization, and load information.
[0237] The data collection frequency is once per second to ensure the timeliness and integrity of the data. The storage format of historical data is:
[0238] (1)
[0239] (2) Data preprocessing
[0240] Remove outliers and missing values from the collected historical traffic data,
[0241] Normalize the cleaned data: Scale the data to the range [0, 1]. The normalization formula is:
[0242] (2)
[0243] Time series segmentation: Segment the normalized data into a training set and a validation set according to the time series. The training set accounts for 80% and the validation set accounts for 20%.
[0244] The pseudo code for data preprocessing is as follows:
[0245]
[0246] (3) Build the LSTM model
[0247] Define the input features and output labels: The input features are historical latency, bandwidth utilization, and load information, and the output label is the predicted value of the traffic in the next period. The input features and the output label Y are in the format of:
[0248] (3)
[0249] (4)
[0250] Initialize the model parameters: Initialize the parameters of the LSTM model, including the hidden units , time step and learning rate . The calculation formula for the number of hidden layer units is:
[0251] (5)
[0252] (4) Train the LSTM model
[0253] Use the training set data to train the LSTM model, optimize the model parameters through the backpropagation algorithm, and the loss function is the mean squared error (MSE), and the formula is:
[0254] (6)
[0255] Use the validation set to evaluate the prediction accuracy of the LSTM model, and the accuracy calculation formula is:
[0256] (7)
[0257] (5) Generate prediction data
[0258] Use the trained LSTM model to generate the predicted traffic data for the next period, and the format of the prediction data is:
[0259] (8)
[0260] (6) Build the state space
[0261] Fuse the predicted traffic data with the real-time traffic data to build the state space , and the construction formula of the state space is:
[0262] (9)
[0263] Balance the weights of the real-time data and the predicted data through the dynamic weight adjustment mechanism, and the weight calculation formula is:
[0264] (10)
[0265] (11)
[0266] (7) Train the DQN model
[0267] Define the state space and the action space : where the state space includes the real-time traffic and the predicted traffic, and the action space includes the optional paths from the source node to the destination node.
[0268] Use the historical network state data for training, optimize the model parameters through the experience replay mechanism, and the formula for the target Q value is:[[]]
[0269] (12)
[0270] Scheduling phase:
[0271] (1) Obtain the network topology and link state information
[0272] In the Fat-Tree structure with K = 4, obtain the current network topology through the topology discovery module of the SDN controller, including the connection relationships of the core layer, aggregation layer, and edge layer.
[0273] Collect the bandwidth utilization, delay, and packet loss rate information of each layer link in real time through the traffic monitoring module.
[0274] (2) Distinguish large flows and small flows
[0275] According to the traffic volume distinguish large flows and small flows, and set the threshold of the traffic volume , when , it is determined as a large flow; otherwise, it is determined as a small flow. The threshold can be dynamically adjusted according to the network environment and service requirements.
[0276] (3) Schedule small flows using ECMP
[0277] In the Fat-Tree structure, calculate multiple equivalent paths from the source node to the destination node. For example, in the Fat-Tree with K = 4, there are 2 equivalent paths from each edge switch to the core switch.
[0278] Use the hash algorithm to evenly distribute small flows to multiple paths, and the hash function is:[[]]
[0279] (1)
[0280] The pseudo-code for scheduling small flows using ECMP is as follows:
[0281]
[0282] (4)Schedule the large flow using the DQN model
[0283] Use the trained DQN model and select the optimal action according to the current state space That is, the optimal path. The action selection formula is as follows: That is, the optimal path. The action selection formula is as follows:
[0284] (2)
[0285] The pseudo-code for scheduling the large flow using DQN is as follows:
[0286]
Claims
1. A network traffic optimization scheduling method based on deep reinforcement learning suitable for periodic traffic characteristics, characterized in that: The following phases are included:
1. Preparation stage: Step 1: Collect historical periodic traffic data; Monitor network status in real time through the SDN controller to obtain historical periodic traffic data, including latency, bandwidth utilization, and load information; Step 2: Data preprocessing: cleaning, normalizing and time series segmentation of the collected historical traffic data; Step 3: Build an LSTM model using preprocessed historical data and define input features and output labels. Step 4: Train the LSTM model. Use the training set data to train the LSTM model and generate traffic forecast data for the next cycle. Step 5: Construct state space, fuse predicted traffic data with real-time traffic data to construct state space; Step 6: Train the DQN model and use the constructed state space to train the deep Q network (DQN) model; 2. Scheduling stage: Step 1: Obtain network topology and link status information. Obtain the current network topology through the topology discovery module of the SDN controller, and collect link bandwidth utilization, latency, and packet loss rate information in real time through the traffic monitoring module. Step 2: Distinguish between large and small flows; Step 3: Use ECMP to schedule small flows; Step 4: Use the DQN model to schedule large flows.
2. According to claim 1, a network traffic optimization scheduling method based on deep reinforcement learning applicable to periodic traffic characteristics is characterized in that: The preparation phase described: (1) Collect historical periodic traffic data; The SDN controller monitors the network status in real time and obtains historical periodic traffic data, including latency, bandwidth utilization, and load information. The data collection frequency is once per second to ensure the timeliness and integrity of the data. The historical data storage format is: (1); in Indicates the timestamp, Indicates delay, Indicates bandwidth utilization. Indicates load information; (2) Data preprocessing: Preprocess the collected historical traffic data, including: Data cleaning: remove outliers and missing values to ensure data quality; Data normalization: Scale the data to the range of [0,1]. The normalization formula is: (2); in, is the original data, and are the minimum and maximum values of the data respectively; Time series segmentation: The normalized data is divided into a training set and a validation set according to the time series, with the training set accounting for 80% and the validation set accounting for 20%; (3) Build an LSTM model; Use the preprocessed historical data to build an LSTM model, including: Define input features and output labels: Input features are historical latency, bandwidth utilization, and load information, and output labels are traffic prediction values for the next cycle; Input features and the output label Y is in the format: (3); (4); Initialize model parameters: Initialize the parameters of the LSTM model, including hidden units , time step and learning rate ; The calculation formula for the number of hidden layer units is: (5); in: is the number of samples in the training set; Indicates rounding down; The purpose of this formula is to dynamically adjust the number of hidden layer units according to the data scale to avoid the model being too complex or too simple; Time step :Time step is the length of the input sequence to the LSTM model, which is usually determined based on the periodicity or time correlation of the data; if the data has obvious periodicity (such as daily or weekly traffic patterns), the time step can be set to the length of the period; for example: If data is collected every hour and has a daily periodicity, then ; If the data is collected once per packet = minute and has an hourly periodicity, then ; If there is no obvious periodicity, the appropriate time step can be selected through experiments, usually in the range of ; Learning Rate : learning rate It is the step size of parameter update during model training. The initial value is usually selected based on experience and dynamically adjusted during training. A commonly used method for selecting the initial learning rate is: (6); in: is the dimension of the input features (i.e. the number of features); its purpose is to avoid the learning rate being too large or too small; (4) Training the LSTM model; Use the training set data to train the LSTM model, including: Forward propagation: Calculate the output of the LSTM model. The formula is: (7); (8); (9); (10); (11); (12); in , , Represent the forget gate, input gate and output gate respectively. Indicates the cell state, Indicates the hidden state, and are the weight matrix and bias vector respectively, Represents the Sigmoid activation function; Back propagation: The model parameters are optimized through the back propagation algorithm. The loss function is the mean square error (MSE), and the formula is: (13); in, is the true value, is the predicted value, is the number of samples; Model evaluation: Use the validation set to evaluate the prediction accuracy of the LSTM model. The accuracy calculation formula is: (14); in, is the number of samples in the validation set; (5) Generate prediction data; Use the trained LSTM model to generate traffic forecast data for the next cycle. The forecast data format is: (15); in They represent the predicted flow values for the next cycle respectively; Dynamic weight adjustment mechanism: Fusion of predicted traffic data with real-time traffic data to construct state space , the construction formula of the state space is: (16); (17); in, Represents the weight of the predicted data; Represents the weight of real-time data. The prediction accuracy is calculated through the validation set. The formula is: (18); in, is the actual flow value, To predict the flow value, is the number of samples in the validation set; The real-time traffic fluctuation coefficient is obtained by calculating the standard deviation of the real-time traffic. The formula is: (19); in, is the mean of real-time traffic, is the number of real-time traffic samples; (6) Training the DQN model; Use the constructed state space to train the Deep Q Network (DQN) model, including: Define state space and action space: state space Including the bandwidth utilization, delay, packet loss rate and predicted traffic data of the current link; action space Includes optional paths from the source node to the destination node; Reward Function Design: Reward Function Taking latency, bandwidth utilization and packet loss rate into consideration, the reward function formula is: (20); in, , , is the weight parameter; Experience replay mechanism: store historical states, actions, rewards, and next states in a replay buffer, and randomly sample data from the buffer during training. The value calculation formula is: (21); in, Indicates immediate reward, represents the discount factor, Represents the parameters of the target network.
3. According to claim 1, a network traffic optimization scheduling method based on deep reinforcement learning applicable to periodic traffic characteristics is characterized in that: The scheduling phase described: (1) Obtain network topology and link status information; The topology discovery module of the SDN controller is used to obtain the current network topology, including the connection relationship and link status information between switches. The traffic monitoring module collects link bandwidth utilization, latency, and packet loss rate information in real time and stores this information as a status data set. (2) Distinguish between large flows and small flows; According to the flow size Distinguish between large and small flows and set the flow size threshold ,when When , it is judged as a large flow; otherwise, it is judged as a small flow; the threshold Can be dynamically adjusted according to network environment and business needs; (3) Use ECMP to schedule small flows; For traffic that is determined to be a small flow, equal-cost multi-path routing (ECMP) is used for scheduling, including: Calculate multiple equivalent paths from the source node to the destination node; Use a hash algorithm to evenly distribute small flows to multiple paths. The hash function is: (22); in, is the unique identifier of the stream, is the number of equivalent cost paths; (4) Use the DQN model to schedule large flows; For traffic that is determined to be a large flow, the trained deep Q network (DQN) model is used for scheduling, including: Constructing the state space : State space Including the bandwidth utilization, delay, packet loss rate of the current link and the predicted traffic data generated by the LSTM model; the construction formula of the state space is: (23); in, Represents the weight of the predicted data; Indicates the weight of real-time data; Defining the state space : Action Space Includes optional paths from the source node to the destination node, each path corresponds to an action ; Select the optimal path: Use the trained DQN model to select the optimal path based on the current state space. Choose the best action That is the optimal path; the action selection formula is: (24); in, represents the parameters of the DQN model, Indicates in status Next select action of value; Reward Function Design: Reward Function Taking latency, bandwidth utilization and packet loss rate into consideration, the reward function formula is: (25); in, , , is a weight parameter used to balance the importance of different optimization objectives; Experience replay mechanism: store historical states, actions, rewards, and next states in a replay buffer, and randomly sample data from the buffer during training. The value calculation formula is: (26); in, Indicates immediate reward, represents the discount factor, Represents the parameters of the target network; (5) Update network status information in real time; The latest link status information is obtained periodically (e.g., every few seconds) through the SDN controller, and the scheduling strategy is dynamically adjusted to adapt to changes in the network environment. The update process includes: Recalculate the link's bandwidth utilization, latency, and packet loss rate; According to the latest state space Reselect the optimal path.
Citation Information
Cited By
Network traffic scheduling optimization method based on deep learning
CN120281665A
Deep Learning-Based Network Traffic Scheduling Optimization Method
CN120281665B
Periodic traffic transmission optimization for model distributed training
CN120750862A
Optimization of periodic traffic transmission for distributed model training
CN120750862B
Token distribution method and device based on traffic prediction, equipment and storage medium
CN120835037A