A Resource Allocation Method for Delay-Sensitive Networks Based on Burst Awareness
By reserving time slot resources for burst streams in a delay-sensitive network and using multi-queue loop forwarding and deep reinforcement learning models, the problem of unreliable burst stream transmission in a delay-sensitive network is solved, and efficient resource allocation and service flow scheduling is achieved.
Patent Information
- Application Number
- CN202410758521.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-06-13
AI Technical Summary
When existing delay-sensitive networks deal with unknown burst streams, there are packet loss dead zones, and the reliability of transmission cannot be guaranteed.
Time slot resources are reserved for burst streams in each time slot, the queue offset value of each switch is determined through multi-queue cyclic forwarding, and the trained deep reinforcement learning model is used to allocate resources based on network state information.
It improves the reliability of burst streaming and the network's scheduling ability of heterogeneous service traffic, avoids packet loss dead zones, and ensures successful scheduling of real-time service flows.
Smart Images

Figure CN118631761B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of delay-sensitive networks, and in particular to a method for allocating delay-sensitive network resources based on burst perception. Background Art
[0002] With more and more applications with real-time requirements, the end-to-end delay and reliability requirements for real-time service flow transmission are becoming more and more stringent. Therefore, solving the heterogeneous service flow scheduling problem in delay-sensitive networks (Temporal Segment Networks, TSN) while ensuring the needs of real-time application service flows is an urgent problem to be solved. Most of the existing research work is based on the prerequisite that the characteristic information of real-time service flows is known in advance, and the processing methods of unknown burst flows in actual network scenarios are not fully considered. The existing queue forwarding model that guarantees service flow delay often has packet loss dead zones in the receiving time slot, and cannot guarantee the reliability of burst flow transmission. Summary of the invention
[0003] In view of the above analysis, an embodiment of the present invention aims to provide a delay-sensitive network resource allocation method based on burst perception, so as to solve the problem that the reliability of burst flow transmission cannot be guaranteed in the prior art.
[0004] On the one hand, an embodiment of the present invention provides a method for allocating delay-sensitive network resources based on burst perception, comprising the following steps:
[0005] Reserve time slot resources for burst flow in each time slot;
[0006] Determine an offset value for each queue of each switch in each time slot based on multi-queue round-robin forwarding;
[0007] According to the current network status information after reserving time slot resources, the offset value of each queue of each switch in each time slot, and the flow information of each real-time business flow, the time slot resource allocation plan for each real-time business flow is obtained based on the trained deep reinforcement learning model.
[0008] Based on a further improvement of the above method, the offset value of each queue of each switch in each time slot is determined based on multi-queue cyclic forwarding, including:
[0009] Determine the sending queue of each switch in each time slot according to the time slot sequence number and set the offset value of the sending queue;
[0010] The other queues are determined as receiving queues, and the offset values of the other queues are calculated based on the distances from the other queues to the sending queue and the offset value of the sending queue.
[0011] For further improvement based on the above method, the offset value of other queues is calculated using the following formula:
[0012]
[0013] where ψ base represents the offset value of the sending queue, ψ j represents the offset value of the j-th queue, K represents the total number of queues, and ξ represents the slot number.
[0014] For further improvement based on the above method, the deep reinforcement learning model includes a prediction network and a target network;
[0015] The trained deep reinforcement learning model is obtained in the following manner:
[0016] S31. For each round of training, use the first real-time traffic flow as the current real-time traffic flow;
[0017] S32. Construct the current state and input it into the prediction network to predict the Q values corresponding to the current real-time traffic flow when allocated to different time slots and offset values, and determine the current action according to the prediction result; the current state includes the offset value of each queue of each switch in each time slot, the current network state information, and the flow information of the current real-time traffic flow; the current action includes the time slot and offset value corresponding to the current real-time traffic flow;
[0018] S33. Determine whether the constraints are satisfied according to the time slot and offset value corresponding to the current real-time traffic flow. If not, the scheduling of the current real-time traffic flow fails; if satisfied, the scheduling of the current real-time traffic flow succeeds; generate an experience sample based on the current state and the current action and store it in the experience sample library; the experience sample library is used to update the parameters of the prediction network and the target network;
[0019] S34. If the scheduling of the current real-time traffic flow fails, update the time slot resource utilization rate to the initial resource utilization rate, end the current round of training, and execute step S35; if the scheduling of the current real-time traffic flow succeeds, use the next real-time traffic flow as the current real-time traffic flow and return to step S32; if there is no next real-time traffic flow, execute step S35;
[0020] S35. If the number of training rounds reaches the preset number of rounds, end the training to obtain the trained deep reinforcement learning model; otherwise, perform the next round of training.
[0021] For further improvement based on the above method, generating an experience sample based on the current state and the current action and storing it in the experience sample library includes:
[0022] Calculate the reward value according to the time slot and offset value corresponding to the current real-time traffic flow;
[0023] Store the current state, current action, reward value, next state, and the scheduling result of the current real-time service flow as an experience sample in the experience sample library.
[0024] Based on a further improvement of the above method, the following formula is used to calculate the reward value according to the time slot and offset value corresponding to the current real-time service flow:
[0025]
[0026] where γ a and γ b represent weight coefficients, represents the difference in the number of service flow schedules in two consecutive steps, represents the difference in the load balancing ability of service flow schedules in two consecutive steps.
[0027] Based on a further improvement of the above method, the following formula is used to calculate the difference in the number of service flow schedules in two consecutive steps:
[0028]
[0029] where χ t represents the number of real-time service flows scheduled in the t-th step, represents the number of real-time service flows within the scheduling period, χ t-1 represents the number of real-time service flows scheduled in the (t - 1)-th step, and μ represents the weight coefficient.
[0030] Based on a further improvement of the above method, the following formula is used to calculate the difference in the load balancing ability of service flow schedules in two consecutive steps:
[0031]
[0032] where, represents the maximum time slot resource utilization rate in the t-th step, Ξ represents the ideal time slot resource utilization rate, represents the maximum time slot resource utilization rate in the (t - 1)-th step, and θ represents the weight coefficient.
[0033] Based on a further improvement of the above method, the constraint conditions include: the burst flow reception time slot length constraint, the time slot capacity constraint, and the end-to-end bounded low latency, jitter, and zero packet loss constraints for each real-time service flow in the network transmission.
[0034] Based on a further improvement of the above method, the burst flow reception time slot length constraint is: the product of the queue offset value of the burst flow and the time slot length is greater than or equal to the sum of the time slot length, the transmission delay of the burst flow on the switch interface, the transmission delay of the burst flow on the link, and the clock deviation.
[0035] Compared with the prior art, the present invention reserves time-slot resources for bursty flows within each time slot, thereby ensuring the reliability of bursty flow transmission. By determining the offset value of each queue of each switch within each time slot based on multi-queue cyclic forwarding, the mapping relationship between the queue type and the time slot is determined. According to the current network state information after reserving time-slot resources, the offset value of each queue of each switch within each time slot, and the flow information of each real-time traffic flow, based on the trained deep reinforcement learning model, a time-slot resource allocation scheme for each real-time traffic flow can be efficiently obtained while ensuring the reliability of bursty flows, improving the network's scheduling ability for heterogeneous traffic flows.
[0036] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combined solutions. Other features and advantages of the present invention will be described in the following specification, and some advantages can be made obvious from the specification, or understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the content specifically pointed out in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings are only for the purpose of showing specific embodiments and are not considered as a limitation to the present invention. Throughout the drawings, the same reference signs denote the same components;
[0038] Figure 1 It is a flowchart of a method for allocating resources in a burst-aware delay-sensitive network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The following will specifically describe the preferred embodiments of the present invention in conjunction with the drawings, where the drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.
[0040] A specific embodiment of the present invention discloses a method for allocating resources in a burst-aware delay-sensitive network, as Figure 1 shown, including the following steps:
[0041] S1. Reserve time-slot resources for bursty flows within each time slot;
[0042] S2. Determine the offset value of each queue of each switch within each time slot based on multi-queue cyclic forwarding;
[0043] S3. Based on the current network state information after reserving time-slot resources, the offset value of each queue of each switch within each time slot, and the flow information of each real-time traffic flow, obtain a time-slot resource allocation scheme for each real-time traffic flow based on the trained deep reinforcement learning model.
[0044] In implementation, the time-sensitive network adopts a fully centralized network control architecture. Its control plane includes a global controller, in which a centralized user configuration (CUC), a centralized network configuration (CNC), and a database are deployed. CUC provides a Web interface to interact with users, collect information about registered terminals, TSN switches, and real-time application traffic flows. CNC is responsible for centrally controlling the traffic flow transmission scheduling of the data plane, sending configuration information, and performing the calculations necessary for planning traffic flows. The database is used to connect CUC and CNC, and record and store process data. The data plane consists of terminals and TSN switches.
[0045] In implementation, 8 queues are evenly arranged on each switch node for sending and receiving traffic flows. Among them, the first K queues are used to circularly forward bursty flows and real-time application traffic flows, and the other queues are used to store best-effort flows. The number of switch nodes in the network is not limited in the present invention. 3 ≤ K ≤ 8.
[0046] Compared with the prior art, the method for allocating resources of a delay-sensitive network based on burst awareness provided in this embodiment ensures the reliability of bursty flow transmission by reserving time slot resources for bursty flows in each time slot. By determining the offset value of each queue of each switch in each time slot based on multi-queue circular forwarding, the mapping relationship between the queue type and the time slot is determined. Based on the current network state information after reserving time slot resources, the offset value of each queue of each switch in each time slot, and the flow information of each real-time traffic flow, a time slot resource allocation scheme for each real-time traffic flow can be efficiently obtained while ensuring the reliability of bursty flows based on a trained deep reinforcement learning model, improving the network's scheduling ability for heterogeneous traffic flows.
[0047] In implementation, according to the distribution state of historical bursty flows, time slot resources can be reserved for bursty flows in each time slot of the scheduling period in advance to accommodate the sending of bursty flows in this time slot segment.
[0048] In implementation, in order to avoid packet loss dead zones in the receiving time slot and ensure that the traffic flows of different queues do not conflict on the same sending interface, a multi-queue circular forwarding mechanism is adopted to determine the offset value of each queue in each time slot.
[0049] Specifically, determining the offset value of each queue of each switch in each time slot based on multi-queue circular forwarding in step S2 includes:
[0050] S201. Determine the sending queue of each switch in each time slot according to the time slot serial number and set the offset value of the sending queue;
[0051] S202. Determine the other queues as receiving queues, and calculate the offset values of the other queues based on the distance from the other queues to the sending queue and the offset value of the sending queue.
[0052] During implementation, first determine the transmission queue of each switch within each time slot. For example, for the ξ-th time slot, the (ξ % K)-th queue on the switch is used as the transmission queue for the burst stream and the real-time traffic stream, and the other queues are used as the receiving queues for the burst stream and the real-time traffic stream, ensuring that within any time slot, there is exactly one queue on a switch node for transmitting traffic streams, and the other queues are all used for receiving traffic streams, thereby avoiding conflicts of traffic streams in different queues on the same transmission interface. ξ = 0, …, λ - 1, where λ represents the number of time slots in the scheduling period.
[0053] Meanwhile, the queues on the switch node cycle as the transmission queue and the receiving queue, and the queue type is not specified and changes with the time slot, thus avoiding the emergence of dead zones and affecting the reliability of traffic streams.
[0054] During implementation, set the offset value ψ of the transmission queue base = 0.
[0055] Use the following formula to calculate the offset values of other queues:
[0056]
[0057] where ψ base represents the offset value of the transmission queue, ψ j represents the offset value of the j-th queue, K represents the total number of queues, and ξ represents the time slot number.
[0058] During implementation, based on the offset value of the transmission queue, use Equation (1) to calculate the offset values of other queues, that is, calculate the offset values of the receiving queues. The offset value of other queues is the distance from this queue to the transmission queue, and the offset values are incremented by 1 in sequence. The offset values of each queue are different in different time slots. By calculating the offset values of the queues, the corresponding relationship between the queues and the traffic streams can be established by allocating queue offset values for each traffic stream. For the burst stream, as long as the allocated offset value is greater than 1, the dead zone introduced by the traditional two-queue cyclic forwarding mechanism in each time slot segment can be eliminated, and the legal time slot receiving range is expanded; for each real-time application traffic stream, it is relatively easy to be managed by the central controller, and it only needs to ensure that the allocated offset value is greater than or equal to 0.
[0059] Then, based on the current network state information after reserving time slot resources, the offset value of each queue of each switch within each time slot, and the flow information of each real-time traffic stream, obtain the time slot resource allocation scheme for each real-time traffic stream based on the trained deep reinforcement learning model.
[0060] Among them, the network state information includes the network topology structure, link rate, and resource utilization rate of each time slot.
[0061] Specifically, the flow information of each real-time service flow includes an ID number, a flow period, a frame size, a source address, a destination address, an end-to-end delay requirement, a jitter requirement, and a reliability requirement. It should be noted that for a fully centralized network control architecture, the real-time service flow information within each scheduling period is known.
[0062] For the real-time service flows within a scheduling period, the current network state information, the offset value of each queue of each switch within each time slot, and the flow information of each real-time service flow are used as the state input of the reinforcement learning to train the well-trained deep reinforcement learning model, and then the resource allocation result of the real-time service flow can be obtained, that is, the time slot and offset value allocated to the real-time service flow.
[0063] It should be noted that the jump path of the real-time service flow on the switch is determined. After allocating a real-time service flow, it is necessary to update the time slot resource utilization rate of each current queue.
[0064] The state space of the deep reinforcement learning model includes the network state information, the offset value of each queue of each switch within each time slot, and the flow information of each real-time service flow, which is the input of the deep reinforcement learning model. The action space of the deep reinforcement learning model includes the time slots and offset values within the scheduling period, which is the output of the model.
[0065] The deep reinforcement learning model includes a prediction network and a target network, and the well-trained deep reinforcement learning model is obtained after multiple rounds of training reach the number of training rounds.
[0066] Specifically, the well-trained deep reinforcement learning model is obtained in the following way:
[0067] S31. For each round of training, use the first real-time service flow as the current real-time service flow;
[0068] S32. Construct the current state and input it into the prediction network to predict the Q values corresponding to the current real-time service flow when allocated to different time slots and offset values, and determine the current action according to the prediction result; the current state includes the offset value of each queue of each switch within each time slot, the current network state information, and the flow information of the current real-time service flow; the current action includes the time slot and offset value corresponding to the current real-time service flow;
[0069] S33. Determine whether the time slot and offset value corresponding to the current real-time service flow meet the constraint conditions. If not, the scheduling of the current real-time service flow fails; if so, the scheduling of the current real-time service flow succeeds. Generate an experience sample based on the current state and the current action and store it in the experience sample library; the experience sample library is used to update the parameters of the prediction network and the target network.
[0070] S34. If the current real-time traffic flow scheduling fails, update the time slot resource utilization rate to the initial resource utilization rate, end the current round of training, and execute step S35; if the current real-time traffic flow scheduling is successful, use the next real-time traffic flow as the current real-time traffic flow, return to step S32; if there is no next real-time traffic flow, execute step S35;
[0071] S35. If the number of training rounds reaches the preset number of rounds, end the training to obtain the trained deep reinforcement learning model; otherwise, perform the next round of training.
[0072] For example, in the first round of training, for the t-th step of training, the offset value of each queue of each switch in each time slot, the current network state information, and the flow information of the t-th real-time traffic flow are used as the state of the t-th step, and the Q values when the t-th real-time traffic flow is assigned to different time slots and queues are predicted by inputting into the prediction network. The ε-greedy strategy is used to select the action of the t-th step, that is, the time slot and queue assigned to the t-th real-time traffic flow.
[0073] Judge whether the allocation result predicted by the prediction network can successfully schedule the t-th real-time traffic flow, that is, judge whether the constraint conditions are met. If not, the t-th real-time traffic flow scheduling fails; if so, the t-th real-time traffic flow scheduling is successful.
[0074] If the t-th real-time traffic flow scheduling is successful, continue to use the remaining information of the (t + 1)-th real-time traffic flow, the offset value of each queue of each switch in each time slot, and the current network state information as the current state, that is, the state of the (t + 1)-th step of training, and input it into the prediction network for the (t + 1)-th step of training. It should be noted that at this time, the time slot resource utilization rate of each current queue is the time slot resource utilization rate of each queue after the t-th real-time traffic flow is assigned. If all real-time traffic flows are assigned, end the current round of training and perform the next round of training.
[0075] If the t-th real-time traffic flow scheduling fails, update the time slot resource utilization rate of each time slot to the initial resource utilization rate, end the current round of training, and perform the next round of training.
[0076] It should be noted that before the next round of training, clear the allocation result of the previous round and re-allocate from the first real-time traffic flow. The initial resource utilization rate is the time slot resource utilization rate of each queue when no real-time traffic flow has been assigned yet after reserving time slot resources for the burst flow in each time slot.
[0077] When the number of training rounds reaches the preset number of rounds, stop the training to obtain the trained deep reinforcement learning model, and input the state into the prediction network of the trained deep reinforcement learning model to obtain the time slot and queue assigned to the real-time traffic flow.
[0078] Specifically, the constraint conditions include: the constraint on the length of the receiving time slot for bursty traffic flows, the constraint on the slot capacity, and the end-to-end bounded low latency, jitter, and zero packet loss constraints for each real-time traffic flow during network transmission.
[0079] Specifically, the constraint on the length of the receiving time slot for bursty traffic flows is that the product of the queue offset value of the bursty traffic flow and the time slot length is greater than or equal to the sum of the time slot length, the transmission delay of the bursty traffic flow on the switch interface, the transmission delay of the bursty traffic flow on the link, and the clock deviation.
[0080] It should be noted that the queue offset value of the bursty traffic flow is determined in advance. The transmission delay of the bursty traffic flow on the switch interface and the transmission delay of the bursty traffic flow on the link can be measured.
[0081] The slot capacity constraint requires that the total sum of the data packets transmitted in each time slot does not exceed the maximum capacity of that time slot.
[0082] Specifically, the slot capacity constraint is expressed as:
[0083]
[0084] Among them, represents the number of real-time traffic flows within the scheduling period, represents the total number of data packets sent by the i-th real-time traffic flow within a scheduling period, size i represents the size of the data packets of the i-th real-time traffic flow, represents the indicator function. If the m-th data packet of the i-th traffic flow is successfully scheduled within the ξ-th time slot, otherwise it is C ξ represents the maximum capacity of the ξ-th time slot excluding the reserved time slot resources.
[0085] The end-to-end bounded low latency, jitter, and zero packet loss constraints for each real-time traffic flow during network transmission are expressed as:
[0086]
[0087] Among them, represents the queue offset value allocated to the i-th real-time traffic flow on the n-th switch. T represents the time slot length, and N represents the number of switches. represents the end-to-end latency requirement of the i-th real-time traffic flow, represents the latency required for the j-th data packet of the i-th real-time traffic flow to be sent from the source end to the destination end, represents the jitter requirement of the i-th real-time traffic flow, represents the total number of data packets sent by the i-th real-time traffic flow within a scheduling period, Indicates the indicator function, and λ represents the number of time slots in the scheduling period.
[0088] In order to enable the deep reinforcement learning model to allocate reasonable resources for real-time service flows and enable real-time service flows to be successfully scheduled, therefore, during the training process, experience samples are generated and stored in the experience sample library to facilitate updating the prediction network, the target network, and their parameters based on the experience sample library.
[0089] Specifically, generating experience samples based on the current state and the current action and storing them in the experience sample library includes:
[0090] Calculating the reward value according to the time slot and offset value corresponding to the current real-time service flow;
[0091] Taking the current state, the current action, the reward value, the next state, and the scheduling result of the current real-time service flow as experience samples and storing them in the experience sample library.
[0092] The experience sample is represented as (s t , a t , r t ,, s t+1 , d t ). s t represents the state at the t-th step, a t represents the action at the t-th step, r t represents the reward at the t-th step, d t represents the flag indicating whether the real-time service flow at the t-th step is successfully scheduled. If it is successfully scheduled, it is 1; otherwise, it is 0., s t+1 represents the state at the (t + 1)-th step. Specifically, the goal of resource allocation in the present invention is to maximize the number of schedulable service flows and the load balancing degree of time slot resources. The goal can be expressed as:
[0093]
[0094] Among them, represents the revenue obtained by allocating the i-th real-time service flow to the ξ-th time slot, and Φ ξ (f i ) represents the indicator function indicating whether the i-th real-time service flow is successfully scheduled. If it is successfully scheduled, it is 1; otherwise, it is 0.
[0095] Revenue Among them, α and β are weights, and A(f i , λ) represents the scheduling quantity revenue of the i-th real-time service flow, represents the load balancing degree function.
[0096]
[0097] Among them, κ ξ represents the resource utilization rate of the ξ-th time slot, Indicates the average resource occupancy rate of all time slots.
[0098] The reward of deep reinforcement learning is related to the allocation goal. Based on the above analysis, the reward includes two parts: the number of service flow scheduling and load balancing. Specifically, the following formula is used to calculate the reward value according to the time slot and offset value corresponding to the current real-time service flow:
[0099]
[0100] Among them, γ a and γ b represent weight coefficients, represents the difference in the number of service flow scheduling between two consecutive steps, represents the difference in the load balancing ability of service flow scheduling between two consecutive steps.
[0101] To accelerate the algorithm convergence speed, the reward for the number of service flow scheduling is the difference in the number of service flow scheduling between the previous and the next step. The following formula is used to calculate the difference in the number of service flow scheduling between two consecutive steps:
[0102]
[0103] Among them, χ t represents the number of real-time service flows scheduled at the t-th step, represents the number of real-time service flows within the scheduling period, χ t-1 represents the number of real-time service flows scheduled at the (t - 1)-th step, and μ represents the weight coefficient.
[0104] To accelerate the algorithm convergence speed, the load balancing reward is the difference in the load balancing ability of service flow scheduling between two consecutive steps. The following formula is used to calculate the difference in the load balancing ability of service flow scheduling between two consecutive steps:
[0105]
[0106] Among them, represents the maximum time slot resource utilization rate at the t-th step, Ξ represents the ideal time slot resource utilization rate, represents the maximum time slot resource utilization rate at the (t - 1)-th step, and θ represents the weight coefficient.
[0107] During implementation, the resource utilization rate of each time slot is obtained according to the ratio of the total size of the service flows allocated to each time slot to the capacity of the time slot, and the maximum resource utilization rate is taken as
[0108] During implementation, the prediction network and the target network have the same structure, but different parameters.
[0109] Update the parameters of the prediction network and the target network based on the experience sample library, including:
[0110] After generating an experience sample based on the state and the current action and storing it in the experience sample library, a random experience sample is drawn from the experience samples. The current state of the drawn experience sample is input into the experience network to obtain a predicted Q value, and the next state is input into the target network to obtain a target Q value;
[0111] Calculate the loss based on the predicted Q value, the target Q value, the reward in the experience sample, and the flag indicating whether scheduling is successful, and update the parameters of the prediction network based on the calculated loss;
[0112] If the current training step meets the target network update condition, update the parameters of the target network to the parameters of the prediction network.
[0113] Specifically, the following formula is used to calculate the loss:
[0114] L = E[(y i - Q fc ) 2 (10)
[0115]
[0116] where Q fc represents the predicted Q value, Q T represents the target Q value, r i represents the reward in the i-th experience sample, d i represents the flag indicating whether the service flow in the i-th experience sample is successfully scheduled, γ represents a parameter, E[·] represents the expectation, and L represents the loss. In implementation, the target network update condition can be, for example, to update the target network every 100 training steps.
[0117] To enhance the ability to extract network state features and speed up the inference of the mapping relationship between real-time service flows and different time slots, the prediction network and the target network adopt a residual neural network structure.
[0118] The residual neural network includes a plurality of residual blocks connected in sequence. The residual blocks include dimensionality-increasing residual blocks and basic residual blocks, and the dimensionality-increasing residual blocks and the basic residual blocks are alternately arranged. In implementation, the first residual block of the residual neural network is a dimensionality-increasing residual block.
[0119] In implementation, the first two layers of the dimensionality-increasing residual block use a convolutional kernel with a 3*3 structure, the padding is 1 for both layers, the convolutional kernel sliding step of the first layer network is 2, the convolutional kernel sliding step of the second layer network is 1, the third layer network uses a 1*1 convolutional kernel, the default padding is 0, and the sliding step is 2.
[0120] The basic residual block uses two layers of network, the convolutional kernel size is 3*3, the padding is 1, and the sliding step is 1.
[0121] The residual neural network structure is adopted to effectively avoid the problems of gradient explosion and gradient disappearance during the training process, and accelerate the algorithm convergence speed.
[0122] Those skilled in the art can understand that all or part of the processes of implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory or a random access memory, etc.
[0123] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for resource allocation in a latency-sensitive network based on burst awareness, characterized in that It includes the following steps: Reserve time slot resources for the burst stream within each time slot; Determine the offset value of each queue of each switch within each time slot based on multi-queue cyclic forwarding; Based on the current network state information after reserving time slot resources, the offset value of each queue of each switch within each time slot, and the flow information of each real-time service flow, obtain the time slot resource allocation scheme for each real-time service flow based on the trained deep reinforcement learning model; The deep reinforcement learning model includes a prediction network and a target network; Obtain the trained deep reinforcement learning model in the following way: S31. For each round of training, use the first real-time service flow as the current real-time service flow; S32. Construct the current state and input it into the prediction network to predict the Q values corresponding to the current real-time service flow allocated at different time slots and offset values, and determine the current action according to the prediction result; the current state includes the offset value of each queue of each switch within each time slot, the current network state information, and the flow information of the current real-time service flow; the current action includes the time slot and offset value corresponding to the current real-time service flow; S33. Judge whether the constraint conditions are satisfied according to the time slot and offset value corresponding to the current real-time service flow. If not, the current real-time service flow scheduling fails. If satisfied, the current real-time service flow scheduling succeeds; generate an experience sample based on the current state and the current action and store it in the experience sample library; the experience sample library is used to update the parameters of the prediction network and the target network; S34. If the current real-time service flow scheduling fails, update the time slot resource utilization rate to the initial resource utilization rate, end the current round of training, and execute step S35; if the current real-time service flow scheduling succeeds, use the next real-time service flow as the current real-time service flow and return to step S32; if there is no next real-time service flow, execute step S35; S35. If the number of training rounds reaches the preset number of rounds, end the training and obtain the trained deep reinforcement learning model. Otherwise, perform the next round of training; Generating an experience sample based on the current state and the current action and storing it in the experience sample library includes: Calculate the reward value according to the time slot and offset value corresponding to the current real-time service flow; Take the current state, the current action, the reward value, the next state, and the scheduling result of the current real-time service flow as an experience sample and store it in the experience sample library; Updating the parameters of the prediction network and the target network based on the experience sample library includes: After generating an experience sample based on the current state and the current action and storing it in the experience sample library, randomly extract an experience sample from the experience sample library, input the current state of the extracted experience sample into the prediction network to obtain the predicted Q value, and input the next state into the target network to obtain the target Q value; Calculate the loss based on the predicted Q value, the target Q value, the reward in the experience sample, and the success flag of the scheduling, and update the parameters of the prediction network based on the calculated loss; If the current training step satisfies the target network update condition, update the parameters of the target network to the parameters of the prediction network; The constraint conditions include: the burst stream receiving time slot length constraint, the time slot capacity constraint, and the end-to-end bounded low latency, jitter, and zero packet loss constraints for each real-time service flow in the network; The length constraint of the received time slot for the burst stream is that the product of the queue offset value of the burst stream and the time slot length is greater than or equal to the sum of the time slot length, the transmission delay of the burst stream on the switch interface, the transmission delay of the burst stream on the link, and the clock deviation.
2. The method for allocating resources in a latency-sensitive network based on burst perception according to claim 1, wherein Determining the offset value of each queue of each switch in each time slot based on multi-queue cyclic forwarding includes: Determining the sending queue of each switch in each time slot according to the time slot sequence number and setting the offset value of the sending queue; Determining the other queues as receiving queues, and calculating the offset values of the other queues based on the distance from the other queues to the sending queue and the offset value of the sending queue.
3. The method for allocating time-delay sensitive network resources based on burst perception according to claim 2, wherein Calculating the offset values of the other queues using the following formula: where, ψ base represents the offset value of the transmission queue, ψ j represents the offset value of the j-th queue, K represents the total number of queues, and ξ represents the time slot number.
4. The method for allocating time-delay sensitive network resources based on burst perception according to claim 1, wherein Calculating the reward value according to the time slot and offset value corresponding to the current real-time traffic flow using the following formula: Among them, γ a and γ b represent weight coefficients, represents the difference in the number of service flow schedules in two consecutive steps, represents the difference in the load balancing ability of service flow schedules in two consecutive steps.
5. The method for allocating time-delay sensitive network resources based on burst perception according to claim 4, wherein Calculating the difference in the number of traffic flow schedules for two consecutive steps using the following formula: Among them, χ t represents the number of real-time traffic flows in the t-th step of scheduling, represents the number of real-time traffic flows within the scheduling period, and χ t-1 represents the number of real-time traffic flows in the (t - 1)-th step of scheduling, and μ represents the weight coefficient.
6. The method for allocating time-delay sensitive network resources based on burst perception according to claim 4, characterized in that Calculating the difference in the load balancing ability of the traffic flow schedules for two consecutive steps using the following formula: Among them, represents the maximum slot resource utilization rate at the t-th step, represents the ideal slot resource utilization rate, represents the maximum slot resource utilization rate at the (t - 1)-th step, and θ represents the weight coefficient.