A Method for Adjusting Dynamic Buffer Size Based on Deep Reinforcement Learning
By applying the Dueling DQN model of deep reinforcement learning in the network, combining the self-attention mechanism and noise layer, dynamically adjusting the buffer size, the problem that the existing technology is difficult to adapt to complex and dynamically changing network traffic is solved, and the network transmission performance and service quality are improved.
Patent Information
- Application Number
- CN202510081574.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The prior art is difficult to adapt to complex and dynamically changing network traffic, resulting in buffer overflow, network delay and packet loss rate problems.
The dynamic buffer size adjustment method based on deep reinforcement learning is adopted, and the buffer size is dynamically adjusted to cope with changes in network traffic through the Dueling DQN model combined with the self-attention mechanism and noise layer.
It effectively reduces network jitter and delay, improves network transmission performance and service quality, and is suitable for complex and dynamically changing network environments.
Smart Images

Figure CN119520450B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information engineering, and particularly to a method for dynamically adjusting buffer size based on deep reinforcement learning. Background Art
[0002] With the development of network technology, traffic management in complex network environments has become increasingly important. Traditional fixed buffer management strategies are unable to flexibly cope with the dynamic changes of network traffic. Especially in scenarios with high traffic fluctuations, buffer overflow, network latency, and packet loss rate become significant problems. In response to these problems, dynamic buffer management strategies have emerged, aiming to adjust the buffer size in real time according to the network state, thereby improving the transmission performance of the network, reducing latency and packet loss.
[0003] With the rapid growth of the scale and capacity of the Internet, the traditional bandwidth-delay product (BDP) rule is no longer applicable to the future Internet. At the same time, projects such as the Global Environment for Network Innovations (GENI) and the Future Internet Design (FIND) have promoted the innovation of Internet technology, bringing challenges and opportunities to the buffer management of next-generation routers. The research community's interest in the buffer size problem has been increasing, and a large number of studies have emerged in the past few years. However, these studies are based on certain assumptions about Internet traffic and may be limited in application in other traffic models, showing inconsistent or even contradictory results. For example, existing fixed buffer management methods or rule-based dynamic management methods often have difficulty coping with complex network changes. Especially when facing burst traffic or high bandwidth requirements, the adaptability of traditional methods is weak.
[0004] With the development of deep learning (DL) in recent years, deep reinforcement learning (DRL), which combines deep learning and reinforcement learning (RL), provides a new idea for buffer capacity management. In particular, DQN (Deep QNetwork) has demonstrated its superior performance in complex decision-making scenarios by combining Q-learning with neural networks. However, simply using DQN has limitations in some complex scenarios. For example, when facing a continuous state space, its performance may be limited. To further optimize the performance of buffer management, Dueling DQN proposes an idea of separately modeling state value and action advantage. This method makes the impact of actions that are not significant in most states smaller, while better decision-making can be made in critical states.
[0005] Based on this, the present invention constructs a dynamic adjustment model for the router buffer size through Dueling DQN, and enables it to make more accurate decisions in a dynamically changing network environment by introducing a self-attention mechanism. At the same time, a noise layer is introduced to avoid falling into local optimal solutions and accelerate the convergence process of the algorithm, aiming to dynamically adjust the buffer size to cope with traffic changes in complex networks. Summary of the Invention
[0006] Objective of the present invention: Aiming at the problem that the buffer management strategy in the current network is difficult to adapt to complex and dynamically changing network traffic, the present invention proposes a method for dynamically adjusting the buffer size based on deep reinforcement learning. This method aims to dynamically adjust the buffer size through a reinforcement learning algorithm to reduce jitter and delay under different traffic loads and network states, thereby effectively improving the transmission performance of the network and enhancing the quality of service (QoS).
[0007] To achieve the above functions, the present invention designs a method for dynamically adjusting the buffer size based on deep reinforcement learning. For a network with an agent and a buffer, the following steps S1 - S7 are executed to complete the automatic adjustment of the buffer size for the network:
[0008] Step S1: Initialize the initial value and upper and lower bounds of the buffer;
[0009] Step S2: Construct a Dueling DQN model, including network state, action, and reward function, as well as a state value network and an action advantage network; where the network state is composed of the characteristics of the buffer, the action is defined as the adjustment strategy of the agent for the buffer, and the reward function gives rewards based on the actions of the agent; the state value network and the action advantage network evaluate the network state and action respectively;
[0010] Step S3: Construct an experience replay pool for storing network environment samples, including experience data of the current network state, action, reward, next network state, and whether it ends;
[0011] Step S4: For the network state at the current moment, extract features through two fully connected layers in sequence, perform attention weight allocation on the features extracted by the two fully connected layers through the self-attention mechanism, combine the attention weights with the extracted features, and then introduce noise perturbations through the noise layer;
[0012] Step S5: Use the action advantage network of the Dueling DQN model to evaluate the actions that the agent can take, use the state value network to evaluate the network state at the current moment, obtain the Q value of the network at the current moment, select and execute actions according to the Q value, and form a network environment sample based on the experience data of the current network state, action, reward, next network state, and whether it ends, and store it in the experience replay pool;
[0013] Step S6: Extract network environment samples from the experience replay pool, use the mean squared error as the loss function, and train the Dueling DQN model through the backpropagation algorithm. After iterative training, obtain the trained Dueling DQN model;
[0014] Step S7: Apply the trained Dueling DQN model to complete the automatic adjustment of the buffer size for the network.
[0015] Beneficial effects: Compared with the prior art, the advantages of the present invention include:
[0016] The present invention designs a dynamic buffer size adjustment method based on deep reinforcement learning, constructs an intelligent agent using the Dueling DQN architecture, effectively realizes the dual reduction of jitter and delay in the network by introducing a noise layer and a self-attention mechanism, combined with the intelligent dynamic adjustment of the buffer capacity, especially in the scenario of high traffic load. Compared with the traditional static buffer management method, the present invention can more flexibly adapt to the changes of network traffic, automatically adjust the buffer size, and improve the overall stability of network transmission. By reducing the average transmission delay and jitter, the present invention not only directly optimizes the network performance, but also indirectly improves the reliability of data transmission, especially in the data transmission scenarios requiring high reliability and low delay, ensuring the continuous and stable operation of the network. In addition, the present invention avoids the waste of resources by dynamically managing the utilization of buffer resources, improving the resource utilization efficiency of the network. The present invention is applicable to a variety of complex dynamic network environments, especially applicable to high-performance network scenarios requiring low delay and low jitter, and has a wide application prospect. Description of the Drawings
[0017] Figure 1 is a flowchart of a dynamic buffer size adjustment method based on deep reinforcement learning provided by an embodiment of the present invention;
[0018] Figure 2 is a framework diagram of the automatic adjustment of the network buffer size provided by an embodiment of the present invention;
[0019] Figure 3 is a neural network structure diagram of the Dueling DQN model provided by an embodiment of the present invention;
[0020] Figure 4 is a training flowchart of the Dueling DQN model provided by an embodiment of the present invention;
[0021] Figure 5 is a structure diagram of the dumbbell topology provided by an embodiment of the present invention;
[0022] Figure 6 is a comparison chart of jitter results provided by an embodiment of the present invention;
[0023] Figure 7 is a comparison chart of latency results provided by an embodiment of the present invention;
[0024] Figure 8 is a comparison chart of average throughput provided by an embodiment of the present invention;
[0025] Figure 9 is a comparison chart of packet loss rate provided by an embodiment of the present invention. Detailed implementation manners
[0026] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and cannot be used to limit the protection scope of the present invention.
[0027] A method for dynamically adjusting buffer size based on deep reinforcement learning provided by an embodiment of the present invention, with reference to Figure 1 , for a network with an agent and a buffer, the following steps S1 - S7 are executed to complete the automatic adjustment of the buffer size of the network; the framework diagram of the automatic adjustment of the buffer size is shown in Figure 2 :
[0028] Step S1: Initialize the initial value and upper and lower bounds of the buffer;
[0029] Step S2: Construct a Dueling DQN model, including network state, action, and reward function, as well as a state value network and an action advantage network; where the network state is composed of the characteristics of the buffer, the action is defined as the adjustment strategy of the agent for the buffer, and the reward function gives rewards according to the actions of the agent; the state value network and the action advantage network evaluate the network state and action respectively;
[0030] The neural network structure diagram of the Dueling DQN model is as shown in Figure 3 ; the network state, action, and reward function in step S2 are specifically as follows:
[0031] Define the network state as the following formula:
[0032] ;
[0033] In the formula, is the network state at time represents the current buffer size, defined as the actual cache amount allocated to the router; represents the current dequeue rate, which measures the speed at which data packets flow out of the buffer to the link and affects the timeliness of data transmission; Indicates the current queue delay, which is the average waiting time of packets in the buffer queue. Indicates the number of packets queued in the current queue, which is used to measure the utilization degree of the current buffer.
[0034] Define the action space as follows:
[0035] ;
[0036] In the formula, represents the action space, represents the action, which is defined that the action space of the agent is the adjustment strategy for the buffer. , where represents increasing the buffer size, represents decreasing the buffer size, represents keeping the buffer size unchanged;
[0037] According to the changes in network performance, design the reward function. The reward value will consider the packet loss rate, latency, and throughput. When the buffer adjustment effectively reduces the packet loss rate or decreases the latency, a positive reward is given; otherwise, a negative reward is given. The agent is guided by the reward feedback to continuously optimize the strategy. Define the reward function as follows:
[0038] ;
[0039] In the formula, represents the reward function, is the queuing delay reward, is the packet loss rate reward, is the throughput reward; , , are the weights of the queuing delay reward, packet loss rate reward, and throughput reward respectively. Each weight satisfies , , .
[0040] Step S3: Construct an experience replay pool for storing network environment samples, including the network state, action, reward, network state at the next moment, and experience data indicating whether it is over at the current moment; the role of the experience replay pool is to store past experiences so that the agent can randomly sample from them for training, thereby improving the stability of training and avoiding overfitting.
[0041] The capacity of the experience replay pool in Step S3 is set to N In this embodiment N = 5000. When the capacity reaches the upper limit, the oldest network environment sample will be removed;
[0042] The storage form of the experience replay pool is as follows:
[0043] ;
[0044] In the formula, represents the network environment sample, represents the network state at time represents the action, represents the reward, represents the network state at the next time, represents the experience data indicating whether the episode ends.
[0045] Step S4: For the network state at the current time, extract features through two layers of fully connected layers in sequence, perform attention weight allocation on the features extracted by the two layers of fully connected layers through the self-attention mechanism, combine the attention weights with the extracted features, and then introduce noise perturbation through the noise layer;
[0046] In step S4, the self-attention mechanism is used to improve the feature modeling ability, and the exploration ability and training stability are enhanced through the noise layer. The specific steps are as follows:
[0047] Step S4.1: Input the network state at the current time into the first fully connected layer. Let the input network state be , and the first fully connected layer maps the input network state to the first hidden layer to obtain the output of the first hidden layer, and process it using the ReLU activation function. The formula is as follows:
[0048] ;
[0049] In the formula, represents the weight matrix of the first hidden layer, which is used to linearly transform the input network state , represents the bias vector of the first hidden layer, which is added to the result of the linear transformation to increase the expressive ability of the model. Its dimension is the same as the number of neurons in the first hidden layer;
[0050] Step S4.2: Input the output of the first hidden layer into the second fully connected layer to obtain the output of the second hidden layer:
[0051] ;
[0052] In the formula, represents the weight matrix connecting the first hidden layer to the second hidden layer, which is used to linearly transform the output of the first hidden layer, The bias vector of the second hidden layer is added to the result of the linear transformation to increase the expressiveness of the model;
[0053] Step S4.3: The output of the second hidden layer is processed by the self-attention layer Dynamic weight allocation is performed. The self-attention mechanism can capture the interdependence of different features in the input network state. The query vector is first calculated , key vector Sum value vector , the calculation formula is as follows:
[0054] ;
[0055] In the formula, , , is a trainable weight matrix;
[0056] Then, using the query vector and key vector The inner product of is used to calculate the correlation between the network state at each historical moment and the network state at the current moment to obtain the weight. The weight calculation formula is as follows:
[0057] ;
[0058] In the formula, is the attention weight, is the query vector, is the key matrix, is the dimension of the key vector; the softmax operation ensures that the weight values are normalized so that the sum of the weights of the network state at all historical moments is 1;
[0059] Step S4.4: Combine the attention weights and input features, and the weighted output is:
[0060] ;
[0061] In the formula, is the attention weight, is a value vector, is the weighted output;
[0062] Step S4.5: The output of the self-attention layer is passed to the noise layer. The noise layer enhances the model's exploration ability by introducing noise disturbance, so that the model can not only rely on historical data for training, but also randomly explore different actions. The calculation formula of the noise layer is as follows:
[0063] ;
[0064] In the formula, is the output of the self-attention layer, is the noise layer, which is a random perturbation sampled from the noise distribution.
[0065] Step S5: Use the action advantage network of the Dueling DQN model to evaluate the actions that the agent can take, use the state value network to evaluate the network state at the current moment, obtain the Q value of the network at the current moment, select and execute an action according to the Q value, and form a network environment sample based on the experience data of the current moment network state, action, reward, next moment network state, and whether it ends, and store it in the experience replay pool;
[0066] The specific steps of Step S5 are as follows:
[0067] Step S5.1: Evaluate the action through the advantage function of the action advantage network and the value function of the state value network of the Dueling DQN model. The advantage function is used to measure the pros and cons of a certain action relative to other actions, while the value function evaluates the value of the current network state;
[0068] The value function is as follows:
[0069] ;
[0070] where, represents the function from the network state to its corresponding value, is the value of the network state ;
[0071] The advantage function is as follows:
[0072] ;
[0073] where, represents the function from the network state and the action to its corresponding advantage value, is the advantage value of the action ;
[0074] Step S5.2: Calculate the Q value as follows:
[0075] ;
[0076] where, is the Q value of taking the action in the network state , represents the number of actions, is the advantage value of the action , is the network state Value represents all possible actions in the action space represents the advantage values of all possible actions in the action space;
[0077] Step S5.3: Select an action using the Boltzmann Softmax strategy , the Boltzmann Softmax strategy calculates the selection probability of each action through Q values, making actions with higher Q values more likely to be selected. The calculation formula for the selection probability of an action is:
[0078] ;
[0079] where is the temperature parameter, used to control the balance between exploration and exploitation represents all possible actions in the action space; represents the current network state under which all possible actions the expected future cumulative return, that is, the Q values of all possible actions ;
[0080] Step S5.4: Execute the action , and obtain the feedback reward and the network state at the next moment , and at the same time store the network environment samples in the experience replay pool.
[0081] Step S6: Extract network environment samples from the experience replay pool, use the mean squared error (MSE) as the loss function, and train the Dueling DQN model through the backpropagation algorithm. After iterative training, obtain the trained Dueling DQN model;
[0082] Refer to Figure 4 , the specific steps of Step S6 are as follows:
[0083] Step S6.1: Randomly extract a batch of network environment samples from the experience replay pool and calculate the target Q value. The formula is:
[0084] ;
[0085] In the formula, represents the target network Q value, represents the immediate reward, represents the discount factor, represents the advantage value of the action, represents the value of the network state, is the normalization operation for the action, represents the target network parameters represents the number of actions; is the advantage function value of the target network estimating to execute an action under the network state ; represents all possible actions in the action space, i.e., increasing, decreasing, and maintaining the buffer size;
[0086] Step S6.2: Update the main network parameters by minimizing the mean square error as the loss function. The loss function is as follows:
[0087] ;
[0088] where is the loss function, are the main network parameters; represents the expectation of a batch of data sampled from the experience replay pool, i.e., taking the average of the losses of each sample, represents the target network Q value, is the main network Q value;
[0089] Step S6.3: Update the network parameters using the soft update strategy. By using the soft update strategy, the balance between the target network and the current network can be maintained more effectively, making the training more stable and progressive, and avoiding large parameter jumps in the model. The formula is as follows:
[0090] ;
[0091] where is the weight of the soft update, represents the target network parameters, are the main network parameters;
[0092] Step S6.4: Through multiple iterations of training the Dueling DQN model, obtain the trained Dueling DQN model, and finally enable the model to accurately predict and dynamically adjust the network buffer size, effectively reducing jitter and latency.
[0093] After training is completed, the Dueling DQN model receives the current network state and outputs the optimal buffer adjustment action in real time. When the network traffic changes, the model can quickly adjust the buffer size to avoid jitter and excessive latency.
[0094] Step S7: Apply the trained Dueling DQN model to complete the automatic adjustment of the buffer size for the network.
[0095] The following is an application embodiment for verifying a dynamic buffer size adjustment method based on deep reinforcement learning designed by the present invention:
[0096] This embodiment adopts a dumbbell topology to simulate the dynamic cache management performance in a high-load environment. The structure of the dumbbell topology is as Figure 5 shown. It includes 50 source nodes on the left and 50 destination nodes on the right. All source nodes are connected to the destination nodes through a central bottleneck link. The bandwidth of this bottleneck link is set to 10 kbps, and the delay is 5 ms; the cache capacity is initially 5 packets, the upper bound is 20 packets, and the lower bound is set to 2 packets.
[0097] To test the processing ability of the system and the optimization effect of the algorithm, this embodiment adopts a traffic flow based on the ON / OFF traffic model to simulate three typical network services: fax, video, and voice. In the sending rule of the traffic flow, numbers from 1 to 50 send data to nodes from 51 to 100 respectively. The simulation duration is 40 seconds. Voice service starts to be sent during the period from 1 to 20 seconds, video service starts to be sent during the period from 10 to 30 seconds, and fax service starts to be sent during the period from 20 to 40 seconds.
[0098] The experimental results show a comparison with the static cache management algorithm BDP and the dynamic cache capacity management algorithm DQN. In terms of jitter, the present invention reduces by 40.85% and 15% respectively compared with the static cache management algorithm BDP and the dynamic cache capacity management algorithm DQN under high load. The results are as Figure 6 seen. The present invention shows higher stability. In terms of delay, the present invention reduces by 20.95% and 10.76% respectively compared with the BDP algorithm and the DQN algorithm when the transmission rate is continuously increasing. The results are as Figure 7 shown, effectively meeting the low-delay requirement and significantly improving the transmission efficiency of delay-sensitive services. The comparison of the average throughput is as Figure 8 shown. The throughputs of the three algorithms are 225.555, 225.342, and 225.542 respectively. The present invention is slightly higher than the DQN algorithm and slightly lower than the BDP algorithm. This is because the cache size setting is smaller than that of the BDP algorithm. Therefore, the present invention can obtain a relatively objective throughput under the condition of a lower buffer capacity. The comparison of the packet loss rate is as Figure 9 shown. Their values are: 0.87558, 0.875708, and 0.879057 respectively. It can be seen that the packet loss rate of the present invention is slightly higher than the other two. This is because the present algorithm makes a trade-off between the packet loss rate and the buffer capacity. Although the packet loss rate is slightly higher than the other two algorithms, this result is acceptable.
[0099] In summary, the technical effects achieved by the method designed in the present invention include:
[0100] 1. Improve the adaptive adjustment ability of the buffer in the network, and reduce the average delay jitter and the average queuing delay;
[0101] 2. Dynamically optimize buffer allocation in a complex network environment to adapt to traffic fluctuations;
[0102] 3. Utilize the learning ability of the Dueling DQN model to reduce the need for manual adjustment of network parameters and improve the overall network efficiency;
[0103] 4. Reduce the buffer layout cost.
[0104] The present invention aims to solve the deficiencies in the existing buffer capacity management methods and provide a solution with intelligent adjustment capabilities to meet the performance requirements in future highly dynamic network environments.
[0105] The above has described in detail the embodiments of the present invention in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A dynamic buffer size adjustment method based on deep reinforcement learning, characterized in that: For a network with an agent and a buffer, the following steps S1 to S7 are performed to complete automatic adjustment of the buffer size for the network: Step S1: Initialize the initial value and upper and lower bounds of the buffer; Step S2: Construct the Dueling DQN model, including network state, action and reward function, as well as state value network and action advantage network; The network state consists of the characteristics of the buffer, the action is defined as the agent's adjustment strategy for the buffer, and the reward function feedbacks the reward according to the agent's action; the state value network and the action advantage network evaluate the network state and action respectively; Step S3: Construct an experience replay pool to store network environment samples, including the current network state, action, reward, the next network state, and experience data on whether the process is over; Step S4: for the current network state, extract features through two fully connected layers in sequence, assign attention weights to the features extracted by the two fully connected layers through the self-attention mechanism, combine the attention weights with the extracted features, and then introduce noise disturbance through the noise layer; The specific steps of step S4 are as follows: Step S4.1: Input the current network state into the first fully connected layer, assuming that the input network state is , the first fully connected layer maps the input network state to the first hidden layer, and obtains the output of the first hidden layer , and processed using the ReLU activation function, the formula is as follows: ; In the formula, represents the weight matrix of the first hidden layer, represents the bias vector of the first hidden layer; Step S4.2: The output of the first hidden layer Input the second fully connected layer to get the output of the second hidden layer : ; In the formula, represents the weight matrix connecting the first hidden layer to the second hidden layer, represents the bias vector of the second hidden layer; Step S4.3: The output of the second hidden layer is processed by the self-attention layer To perform dynamic weight allocation, first calculate the query vector , key vector Sum value vector , the calculation formula is as follows: ; In the formula, , , is a trainable weight matrix; Then, using the query vector and key vector The inner product of is used to calculate the correlation between the network state at each historical moment and the network state at the current moment to obtain the weight. The weight calculation formula is as follows: ; In the formula, is the attention weight, is the query vector, is the key matrix, is the dimension of the key vector; Step S4.4: Combine the attention weights and input features, and the weighted output is: ; In the formula, is the attention weight, is a value vector, is the weighted output; Step S4.5: Pass the output of the self-attention layer to the noise layer. The calculation formula of the noise layer is as follows: ; In the formula, is the output of the self-attention layer, is the noise floor, is a random perturbation sampled from a noise distribution; Step S5: Use the action advantage network of the Dueling DQN model to evaluate the actions that the agent can take, use the state value network to evaluate the network state at the current moment, obtain the Q value of the network at the current moment, select and execute actions based on the Q value, and construct a network environment sample based on the current network state, action, reward, next moment network state, and experience data of whether the process is over, and store it in the experience replay pool; Step S6: extract network environment samples from the experience replay pool, use mean square error as the loss function, train the Dueling DQN model through the back propagation algorithm, and obtain the trained Dueling DQN model after iterative training; Step S7: Apply the trained Dueling DQN model to complete automatic adjustment of the network buffer size.
2. According to the method of dynamic buffer size adjustment based on deep reinforcement learning in claim 1, it is characterized in that: The network state, action, and reward function in step S2 are as follows: The network status is defined as follows: ; In the formula, for The network status at any time, Indicates the current buffer size, Indicates the current dequeue rate, Indicates the current queue delay, Indicates the number of queued packets in the current queue; The action space is defined as follows: ; In the formula, represents the action space, Indicates action, ,in Indicates increasing the buffer size. Indicates reducing the buffer size, Indicates keeping the buffer size unchanged; The reward function is defined as follows: ; In the formula, represents the reward function, It is a queue delay reward. is the packet loss rate reward, is the throughput reward; , , They are the weights of queuing delay reward, packet loss rate reward, and throughput reward. Each weight satisfies , , .
3. The method for dynamic buffer size adjustment based on deep reinforcement learning according to claim 1, characterized in that: The capacity of the experience replay pool in step S3 is set to N , which is stored in the following format: ; In the formula, Indicates a sample network environment. express The network status at any time, Indicates action, Indicates reward, Indicates the network status at the next moment. Experience data indicating whether the round is over.
4. The method for dynamic buffer size adjustment based on deep reinforcement learning according to claim 1, characterized in that: The specific steps of step S5 are as follows: Step S5.1: Evaluate the action using the advantage function of the action advantage network and the value function of the state value network of the Dueling DQN model. The advantage function is used to measure the advantages and disadvantages of an action relative to other actions, while the value function evaluates the value of the current network state. The value function is as follows: ; in, Indicates network status to its corresponding value, Network status the value of The advantage function is as follows: ; in, Indicates network status and actions to its corresponding advantage value, For Action The advantage value of Step S5.2: Calculate the Q value as follows: ; in, In network status Take Action The Q value, Indicates the number of actions, For Action The advantage value of Network status The value of represents all possible actions in the action space, Represents the advantage value of all possible actions in the action space; Step S5.3: Select actions using the Boltzmann Softmax strategy , the calculation formula for the probability of selecting an action is: ; in, is the temperature parameter, represents all possible actions in the action space, Indicates the current network status All possible actions Q value; Step S5.4: Execute action , and get rewards for feedback And the network status at the next moment , and store the network environment samples in the experience replay pool.
5. The method for dynamic buffer size adjustment based on deep reinforcement learning according to claim 1, characterized in that: The specific steps of step S6 are as follows: Step S6.1: Randomly extract a batch of network environment samples from the experience replay pool and calculate the target Q value. The formula is: ; In the formula, represents the target network Q value, Indicates immediate reward, represents the discount factor, represents the advantage value of the action, A value representing the state of the network, is the normalization operation of the action, represents the target network parameters, Indicates the number of actions; is the target network estimated in the network state Next action The advantage function value of Represents all possible actions in the action space; Step S6.2: Update the main network parameters by minimizing the mean square error as the loss function. The loss function is as follows: ; in, is the loss function, Main network parameters; It means to average the loss of each sample. represents the target network Q value, is the Q value of the main network; Step S6.3: Update network parameters using the soft update strategy, the formula is as follows: ; in, is the weight of soft update, represents the target network parameters, is the main network parameter; Step S6.4: Train the Dueling DQN model through multiple iterations to obtain a trained Dueling DQN model.
Citation Information
Patent Citations
Active queue management method based on dynamic cache
CN118400336A
Intelligent path optimization method and system based on link state perception enhancement
CN119011463A