A scheduling method, device, equipment and storage medium
By generating target vectors and time slot queue packets, combined with the Dueling DQN algorithm to optimize the scheduling strategy, the problem of coexistence of service flows in a large-scale deterministic network is solved, and ultra-low delay transmission is achieved.
Patent Information
- Application Number
- CN202211356883.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-01
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-11-01
AI Technical Summary
The existing large-scale deterministic network technology cannot support multiple service scenarios with service quality requirements at the same time, and cannot achieve ultra-low delay transmission, resulting in the inability to meet strict delay requirements such as AR services and remote industrial control.
By receiving the current time slot load, connection information and transmission information sent by the data plane device, the target vector is generated, the data transmission path and time slot queue grouping are determined, and the scheduling strategy is optimized by the Dueling DQN algorithm to realize online real-time scheduling of service flows.
It realizes that in the coexistence of service flows with multiple service quality requirements, it can be scheduled online in real time to achieve ultra-low delay transmission, solving the problems of complex transmission and difficult scheduling.
Smart Images

Figure CN115914128B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a scheduling method, apparatus, device, and storage medium. Background Art
[0002] Existing large-scale deterministic network technologies achieve end-to-end deterministic quality of service guarantee through a loop forwarding mechanism. In a deterministic network, all nodes divide time into equal-length time slots, and the length is denoted as Δ. dip . By measuring the maximum transmission delay between neighbor nodes, a time slot mapping relationship is established, and then the forwarding time slot of the data packet is determined. Through the time slot mapping and forwarding relationship, deterministic end-to-end transmission is guaranteed.
[0003] However, the quality of service of large-scale deterministic network technologies strongly depends on the system parameter Δ. dip . Its end-to-end delay is positively correlated with Δ. dip , and the upper bound of the delay jitter is 2Δ. dip . Therefore, in the actual application process, the existing large-scale deterministic network technologies have the following two deficiencies: (1) They cannot support service scenarios with multiple quality of service requirements simultaneously. For example, AR services require a delay within 2 milliseconds, while remote industrial control requires a stricter delay within 500 microseconds. The existing deterministic network technologies cannot meet the quality of service requirements of both; (2) They cannot achieve ultra-low delay transmission: The existing deterministic networks require (where Len is the maximum length of data packets in the network, and BW is the minimum bandwidth of the links in the network), and the end-to-end delay is positively correlated with Δ. dip . When the required transmission delay is less than , the existing deterministic networks will not be able to meet the requirement. Summary of the Invention
[0004] Embodiments of the present invention provide a scheduling method, apparatus, device, and storage medium, which solve the problem of being unable to support service scenarios with multiple quality of service requirements simultaneously and can achieve ultra-low delay transmission.
[0005] According to one aspect of the present invention, a scheduling method is provided, including:
[0006] Receiving the current time slot load, connection information, and transmission information sent by a data plane device;
[0007] Generating a target vector according to the current time slot load, the connection information, and the transmission information;
[0008] Determining a data transmission path and a time slot queue grouping according to the target vector, and sending the data transmission path and the time slot queue grouping to the data plane device.
[0009] According to another aspect of the present invention, a scheduling device is provided, and the scheduling device includes:
[0010] a receiving module, configured to receive the current time slot load, connection information, and transmission information sent by a data plane device;
[0011] a target vector generation module, configured to generate a target vector according to the current time slot load, the connection information, and the transmission information;
[0012] a determination module, configured to determine a data transmission path and a time slot queue grouping according to the target vector, and send the data transmission path and the time slot queue grouping to the data plane device.
[0013] According to another aspect of the present invention, an electronic device is provided, and the electronic device includes:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the scheduling method according to any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the scheduling method according to any embodiment of the present invention when executed by a processor.
[0018] In the embodiments of the present invention, by receiving the current time slot load, connection information, and transmission information sent by a data plane device; generating a target vector according to the current time slot load, the connection information, and the transmission information; determining a data transmission path and a time slot queue grouping according to the target vector, and sending the data transmission path and the time slot queue grouping to the data plane device, since the quality of service of large-scale deterministic network technology highly depends on the system parameter Δ dip (time slot length), its end-to-end delay is positively correlated with Δ dip , and the upper bound of delay jitter is 2Δ dip . When the current large-scale deterministic network transmission technology only supports configuring a single Δ dip , there are both problems that it cannot support service scenarios with multiple quality of service requirements at the same time and that it cannot achieve ultra-low delay transmission. The embodiments of the present invention support the coexistence of multiple Δ dip parameters, and have the ability to allocate different traffic flows to different Δ dipThe conditions for transmission under the parameters are presented, and an online scheduling method is proposed to solve the transmission scheduling problem of service flows under this system. It can achieve online and immediate scheduling, and thus achieve ultra-low latency transmission.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Brief Description of the Drawings
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.
[0021] Figure 1 is a flowchart of a scheduling method in an embodiment of the present invention;
[0022] Figure 2 is a diagram of a scheduling system in an embodiment of the present invention;
[0023] Figure 3 is a schematic structural diagram of a scheduling device in an embodiment of the present invention;
[0024] Figure 4 is a schematic structural diagram of an electronic device in an embodiment of the present invention. Detailed Embodiments
[0025] In order to enable those skilled in the art of this technology to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts should fall within the scope of protection of the present invention.
[0026] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] Embodiment 1
[0028] Figure 1 As shown in the flowchart of a scheduling method provided by an embodiment of the present invention, this embodiment is applicable to the case of large-scale deterministic network scheduling. This method can be executed by the scheduling device in the embodiment of the present invention, and the device can be implemented in a software and / or hardware manner, such as Figure 1 shown, the method specifically includes the following steps:
[0029] S110, receive the current time slot load, connection information and transmission information sent by the data plane device.
[0030] Among them, a time slot refers to a time slice obtained by dividing time. There is a certain traffic flow to be sent within each time slot, and the amount of traffic flow data transmitted in each time slot is the result of traffic flow scheduling and is a known quantity. The current time slot load is: at the current moment, the amount of traffic flow data to be transmitted in each time slot divided by the maximum data amount that can be transmitted in this time slot.
[0031] It should be noted that the scheduling in the present invention refers to arranging the traffic flow to be transmitted on a specific time slot. Therefore, the amount of data in each time slot is related to the scheduling result.
[0032] Among them, the data plane device can be a deterministic network device or a user terminal, etc., and the embodiments of the present invention do not limit this.
[0033] Specifically, the data plane device regularly reports the load information of each time slot of the link, where the load information of each time slot is the load information of the device itself, that is, the time slot load information on each sending port of the device.
[0034] Among them, the connection information is the connection information between data plane devices. For example, the user terminal can be connected to a deterministic network device, or two deterministic network devices can be connected, or the user terminal can be connected to another user terminal through a deterministic network device.
[0035] Among them, the transmission information includes at least one of service data, the sending end of the service data, the receiving end of the service data, the time slot index corresponding to the service data, the generation period of the service data, the service data packet size, and the end-to-end delay corresponding to the service data.
[0036] Among them, the time slot index corresponding to the service data is the time slot corresponding to the generation of the service data, indicating the time slot where the sending end is located when the service data is generated.
[0037] S120. Generate a target vector according to the current time slot load, the connection information, and the transmission information.
[0038] Among them, the target vector is a state vector.
[0039] Specifically, the method of generating a target vector according to the current time slot load, the connection information, and the transmission information can be: determining state information according to the current time slot load, the connection information, and the transmission information, and determining the target vector according to the state information. For example, it can be to determine the remaining time slot load corresponding to each time slot according to the current time slot load, the connection information, and the transmission information, and generate a target vector according to the remaining time slot load.
[0040] In a specific example, determine the target vector S according to the sending end of the service data, the receiving end of the service data, the time slot index corresponding to the service data, the generation period of the service data, the service data packet size, the end-to-end delay corresponding to the service data, and the remaining time slot load corresponding to each time slot. The target vector:
[0041] S=(src,dest,arr_time,duration,payload,ddl,w1,w2,…,w C ), where src is the sending end of the service data, dest is the receiving end of the service data, arr_time is the time slot index corresponding to the service data, duration is the generation period of the service data, payload is the service data packet size, ddl is the end-to-end delay corresponding to the service data, and w c is the remaining time slot load corresponding to time slot c, where c∈{1,2,…,C}.
[0042] S130. Determine the data transmission path and time slot queue grouping according to the target vector, and send the data transmission path and time slot queue grouping to the data plane device.
[0043] It should be noted that, due to the use of time-division multiplexing scheduling, network resources are uniformly identified by paths and time slots on the paths. Therefore, when allocating resources, it is necessary to clarify the corresponding path and the time slots occupied on the path during the transmission of the traffic flow.
[0044] Specifically, the method for determining the data transmission path and the time slot queue grouping according to the target vector may be: inputting the target vector into the Q-network model to obtain the target action value; obtaining the data transmission path and the time slot queue grouping corresponding to the target action value.
[0045] Specifically, the centralized scheduling controller sends the data transmission path and the time slot queue grouping to the data plane device, so that the data plane device places the data packets into the corresponding time slot queues for forwarding according to the data transmission path and the time slot queue grouping.
[0046] Among them, the time slot queue grouping is a part of the scheduling result. The scheduling algorithm outputs: the data transmission path and the time slot queue grouping. Among them, the data transmission path and the time slot queue grouping are the scheduling result. The time slot queue grouping refers to which egress forwarding queue group the data packet enters. Different queue groups correspond to different time slot lengths, but share the underlying link resources. The scheduling algorithm ensures that the occupancy of the underlying resources does not overflow.
[0047] In a specific example, the deterministic network device obtains the data transmission path and the time slot queue grouping from the centralized scheduling controller, and can determine the forwarding port of the traffic flow in this node, the time slot queue grouping of the traffic flow in the forwarding port, and the time slots occupied by the traffic flow in the time slot queue grouping according to the data transmission path and the time slot queue grouping. The deterministic network device places the data packets into the corresponding queues for forwarding according to the correspondence between the time slots and the time slot queues configured locally. It should be noted that the queues in the device output port are divided into multiple cyclic queue groups, and different groups are forwarded based on different time slot lengths.
[0048] Optionally, the transmission information includes at least one of: service data, the sending end of the service data, the receiving end of the service data, the time slot index corresponding to the service data, the generation period of the service data, the service data packet size, and the end-to-end delay corresponding to the service data.
[0049] Among them, the end-to-end delay corresponding to the service data is: the maximum acceptable end-to-end delay corresponding to the service data.
[0050] Optionally, determining the data transmission path and the time slot queue grouping according to the target vector and sending the data transmission path and the time slot queue grouping to the data plane device includes:
[0051] Input the target vector into the Q-network model to obtain the target action value;
[0052] Obtain the data transmission path and time slot queue grouping corresponding to the target action value;
[0053] Send the data transmission path and the time slot queue grouping to the data plane device.
[0054] Specifically, the Q-network model can be Dueling DQN. Dueling DQN contains two neural networks with the same structure. One is called the Q-network, and the other is called the target Q-network. The Q-network can be represented by the function Q(S, A; φ), where S represents the input state, A represents the action space, and φ represents the parameters of the Q-network; the target Q-network is represented by the function Q t (S, A; φ t ) represents, where φ t represents the parameters of the target Q-network. The outputs of both the Q-network and the target Q-network are the Q-values of each action in A.
[0055] Dueling DQN inserts the advantage function Adv(S, A; θ, α) and the value function V(S; θ, β) between the hidden layer and the output layer in the Q-network structure. Here, θ is the parameter between the input layer and the hidden layer, α is the parameter of the advantage function, and β is the parameter of the value function. The advantage function is a fully connected layer that outputs a vector of the same size as the action space; the value function is also a fully connected layer that outputs a scalar. The value finally output by the entire Q-network is:
[0056]
[0057] The output is a vector of the same size as the action space, and the value of each dimension represents the Q-value of each action. Among them, the role of the Q-network is to output the Q-values corresponding to different actions and select the action with the largest Q-value as the target action. The target actions include: data transmission path and time slot queue grouping.
[0058] Specifically, the method of inputting the target vector into the Q-network model to obtain the target action value can be: input the target vector into the Q-network model, obtain the Q-values of each action output by the Q-network model, and determine the action value with the largest Q-value as the target action value.
[0059] Specifically, the method of obtaining the data transmission path and time slot queue grouping corresponding to the target action value can be: pre-establish a database on the correspondence between action values and data transmission paths and time slot queue groupings, and obtain the data transmission path and time slot queue grouping corresponding to the target action value by querying the database.
[0060] In a specific example, the embodiment of the present invention selects the K - shortest paths algorithm (KSP, k - shortest paths), where K is a preset parameter related to the number of neurons in the outermost layer of the Q - network. Before the scheduling starts, calculate the first K shortest paths between every two client - sides. The selection principle of the shortest path is the number of hops. Therefore, the size of the action space is |A| = K * M, which is the product of the number of selectable paths and the number of selectable time - slot queue groups M. The Q - network only contains one hidden layer with the number of neurons Neu, and the activation function selects ReLU. The output dimension of the advantage function is |A|, and the output dimension of the output layer of the final Q - network is also |A|. The value of each dimension represents the Q - value of the corresponding action. Select the action value with the largest Q - value according to the epsilon - greedy algorithm.
[0061] Optionally, after determining the data transmission path and the time - slot queue grouping according to the target vector, the following steps are further included:
[0062] Determine the target transmission delay according to the data transmission path, the time - slot queue grouping, and the transmission information;
[0063] If the target transmission delay is greater than the end - to - end delay corresponding to the service data, determine that the scheduling fails;
[0064] If the target transmission delay is less than or equal to the end - to - end delay corresponding to the service data, determine the target time - slot load corresponding to each along - the - way time - slot of the data transmission path according to the data transmission path, the time - slot queue grouping, and the transmission information;
[0065] Allocate the target time - slot load corresponding to each along - the - way time - slot to the corresponding along - the - way time - slot in sequence.
[0066] Among them, the along - the - way time - slot is the time - slot occupied along the path.
[0067] Specifically, the method for determining the target transmission delay according to the data transmission path, the time - slot queue grouping, and the transmission information can be as follows: Let p=(v1, v2,…, v |p| )(where v1 is the sending end of the service data, v |p| is the receiving end of the service data, and the number of nodes included in the path p is |p|), which is the transmission path of the service data, and v i is a node in the path p. The delay from v1 to v i is defined as d i , and d i is determined based on the following formula:
[0068]
[0069] According to the above formula, the end-to-end delay d corresponding to the service data can be obtained through recursive calculation |p| . Among them, c i-1 is the number of the transmission cycle on node i - 1, and Δ m is the slot length corresponding to the slot queue packet m, is defined as follows:
[0070]
[0071] Among them, mod(x, 1) refers to the fractional part of x, is the propagation delay of the link (v i-1 , v i ), is the difference between the start time of the supercycle on v i-1 and the start time of the supercycle on v i , and Δ m satisfies the following relationship:
[0072] Δ m = k m Δ m-1 , m ∈ {2, 3,..., M}
[0073] Among them, k m is a positive integer. It should be noted that different nodes have multiple egress ports, and the same egress port has multiple queue packets. The queue packet m is forwarded based on the slot length Δ m , and one port of a node contains M slot queue packets (referred to as packets). In addition, the start times of the transmission cycles of all packets in the same port are aligned, that is, precise clock synchronization is achieved.
[0074] Specifically, the method for determining the target slot load corresponding to each slot along the data transmission path according to the data transmission path, the slot queue packet, and the transmission information may be: first, determine the transmission cycle corresponding to each node in the path according to the data transmission path, the slot queue packet, and the generation cycle of the service data, and then determine the target slot load corresponding to each slot along the path according to the transmission cycle corresponding to each node in the path. Among them, the target slot load is the load on the transmission cycle. The transmission cycle and the slot are synonymous.
[0075] In a specific example, the transmission cycle corresponding to each node in the path is determined based on the following process:
[0076] Step 1: When the cycle γ of the packet m of the service data arrives at the sending end v1, it is arranged for transmission in the cycle . Among them, is the number of transmission cycles included in a supercycle;
[0077] Step 2: When node v i receives the service data sent at period c of packet m from the upstream node v i-1 and needs to forward the service data to the downstream node v i+1 , the transmission period number on the queue packet m of v i can be calculated by the following formula:
[0078]
[0079] where, is the propagation delay of the link (v i-1 , v i ), is the difference between the start time of the supercycle on v i-1 and the start time of the supercycle on v i , let b be the transmission period number. Then forward the service data in the transmission period b on the queue packet m of v i . It should be noted that different queue packets are forwarded based on different time slot lengths, where the transmission period b is the transmission period with the transmission period number b.
[0080] Step 3: Repeat Step 2 until i = |p| - 1. The last node is the receiving end and does not need to calculate the transmission period. In this way, the transmission periods along the data transmission path can be obtained.
[0081] It should be noted that the length Δ hc of a supercycle is an integer multiple of the transmission period lengths of all packets in the port and cannot be equal to the transmission period of any one packet (i.e., it must contain 2 or more periods). The start time of the supercycle is the same as the start time of the transmission period of the m = argmax m (Δ m )th packet. The transmission periods within a supercycle are numbered starting from 0.
[0082] The transmission period mapping function for different packets in the same port is defined as: There exist packets m i and m j , and m i < m j , and the mapping function is:
[0083]
[0084] where γ is the transmission period number on packet m i , and the value range of γ is indicating packet mi The transmission period γ of and packet m j The transmission period t of overlaps in time.
[0085] In a specific example, the action value is converted into the corresponding path and queue packet, and then the time slots occupied along the path are calculated according to the time slot index arr_time corresponding to the service data, as well as the upper bound of the end-to-end delay transmitted by this scheme. If the end-to-end delay ddl corresponding to the service data is exceeded, the data flow scheduling fails; if it is less than or equal to ddl, the load of this service flow is added to the original load of these time slots. Then, the new time slot load information is transmitted to the data plane abstraction layer. The data plane abstraction layer determines whether there is a situation where the time slot load capacity is exceeded according to the new time slot load information. If it exists, the service flow scheduling fails; if it does not exist, the service flow scheduling is successful, and the time slot load information of the network node is updated. According to whether the scheduling is successful and the new time slot load information, the reward value R of this round of scheduling is calculated. Among them, the time slots occupied along the path are calculated based on the following formula:
[0086]
[0087] The upper bound of the end-to-end delay is calculated based on the following formula:
[0088]
[0089] d when i = |p| i Is the upper bound of the end-to-end delay.
[0090] Optionally, it further includes:
[0091] Receiving the maximum load of the time slot sent by the data plane device;
[0092] If the target time slot load corresponding to any time slot along the path is greater than the maximum load of the time slot, it is determined that the scheduling fails;
[0093] If the target time slot loads corresponding to the time slots along the path are all less than or equal to the maximum load of the time slot, the current time slot load is updated according to the target time slot loads corresponding to the time slots along the path.
[0094] Among them, the maximum load of the time slot is the maximum load that the time slot can bear. If the time slot load exceeds the maximum load of the time slot, a situation of time slot transmission overflow will occur.
[0095] Among them, comparing the target time slot load corresponding to any along-the-way time slot with the maximum time slot load. If the target time slot load corresponding to any along-the-way time slot is greater than the maximum time slot load, it is determined that the scheduling fails. If the target time slot loads corresponding to the along-the-way time slots are all less than or equal to the maximum time slot load, the current time slot load is updated according to the target time slot load corresponding to the along-the-way time slot, in order to determine whether there is a situation of time slot transmission overflow, and thus effectively prevent the occurrence of time slot transmission overflow.
[0096] Optionally, after updating the current time slot load according to the target time slot load corresponding to the along-the-way time slot when the target time slot loads corresponding to the along-the-way time slots are all less than or equal to the maximum time slot load, it further includes:
[0097] Determining the remaining time slot load according to the maximum time slot load and the target time slot load;
[0098] Determining the reward value according to the remaining time slot load;
[0099] Optimizing the Q-network model according to the reward value to obtain an optimized Q-network model.
[0100] Specifically, if the target time slot loads corresponding to the along-the-way time slots are all less than or equal to the maximum time slot load, the remaining time slot load is determined according to the maximum time slot load and the target time slot load. For example, the data plane abstraction layer abstracts the network topology and time slot resources of the data plane. The set of network nodes is V, and the set of links is E (the links are all abstracted as unidirectional links). The number of queue groups for each port in the network is M, where the time slot length of the m-th group is Δ m , it should be noted that the time slot length is the transmission cycle length. Define the supercycle, and the length is denoted as Δ hc . The length of the supercycle is a multiple of the time slot lengths of all packets, that is, for Δ hc = N m Δ m , N m is a positive integer. For a link e = (v i , v j ) ∈ E, v i , v j ∈ V, the number of time slots is Further, the number of time slots in the entire network can be obtained as C = ∑ e∈E C e . The remaining time slot load of a time slot c is w c ∈ [0, 1], where c ∈ {1, 2,..., C}. Determine the remaining time slot load of time slot c according to the maximum time slot load and the target time slot load of time slot c.
[0101] Specifically, the method for determining the reward value according to the remaining time slot load may be: calculating the reward value based on the following formula:
[0102]
[0103] where R is the reward value, max(·) is the function for finding the maximum value, and avg(·) is the function for finding the average value.
[0104] Specifically, the training process of Dueling DQN is as follows:
[0105] 1. Randomly initialize the parameters φ of the Q network and assign them to the target Q network φ t = φ.
[0106] 2. In each training step, for the current state S, randomly select an action a from A with probability ∈, and select the action with the largest Q value with probability 1 - ∈, that is, select the action with the largest Q value through the Q network:
[0107]
[0108] 3. Execute the action a to obtain the reward value R and the next state S'.
[0109] 4. Store the experience (S, a, R, S') in the experience replay buffer.
[0110] 5. Randomly extract an experience sample (S f , A f , R f , S' f ) with a sample size of β from the experience replay buffer, f ∈ [1, β].
[0111] 6. If S f ' is the end state, then the target y f of the value function (i.e., the Q network) is R f . Otherwise, the value of y f is:
[0112]
[0113] y f = R f + γQ t (S' f , a max ; φ t )
[0114] where a max is the maximum action value, γ is the decay factor, γ ∈ (0, 1]; Q t (S' f , amax ; φ t ) is the target Q network.
[0115] 7. Update the parameters of the Q network by performing a one-step minimization of the loss function L on the experience samples. The loss function is expressed as:
[0116]
[0117] 8. Every certain number of training steps U, assign the parameters of the Q network to the target Q network. For example, if U = 200, then perform φ every 200 steps t = φ.
[0118] 9. Decay the probability ∈, and denote the decay coefficient as ∈_decay. Then the update expression is:
[0119] ∈ = ∈ * (1 - ∈_decay).
[0120] Repeat steps 1 - 9 until ∈ is very small, such as ∈ < 0.00001 (parameters set by the user in advance), to obtain the over-trained Q network model.
[0121] In a specific example, as Figure 2 shown, a scheduling system is provided. The scheduling system includes three main entities, namely: the user side, the deterministic network device, and the centralized scheduling controller;
[0122] User side: After the deterministic traffic flow accesses the user side, the user side applies for network resources from the centralized scheduling controller for the transmission of the deterministic traffic flow. The user side needs to send the transmission information of the deterministic traffic flow to the centralized scheduling controller. The sending port of the user side can implement a data transmission structure based on multi-cycle queue grouping. The user side regularly sends the current time slot load to the centralized scheduling controller to synchronize the link load status information with the centralized scheduling controller.
[0123] Deterministic network device: The deterministic network device needs to regularly send the load situation in the queue to the centralized scheduling controller to synchronize the link load information with the centralized scheduling controller.
[0124] Centralized scheduling controller: The centralized scheduling controller includes: a data plane abstraction module, a status construction module, a deep reinforcement learning module, and an action conversion module.
[0125] Among them, the data plane abstraction module abstracts the queue grouping in the network node and the data load situation in the queue into a data structure. In addition, the data plane abstraction layer also needs to update the current time slot load of each node in the network according to the time slot load information output by the action conversion module, calculate the reward value of the scheduling result and return it to the deep reinforcement learning module, and send the generated scheduling policy (data transmission path and time slot queue grouping) to the data plane device. Specifically,
[0126] The state construction module is responsible for generating a target vector according to the current time slot load, connection information and transmission information sent by the data plane device, and inputting the target vector into the deep reinforcement learning module.
[0127] The deep reinforcement learning module includes components such as the neural network and experience replay cache in the Dueling DQN algorithm, as well as the neural network update and optimization algorithm. The deep reinforcement learning module selects the action value with the largest Q value according to the Q value of each action output by the Q network and outputs it to the action conversion module.
[0128] The action conversion module converts the input action value into resources in the data plane, that is, time slot resources. Then, the time slot load generated by the action is input into the data plane abstraction layer to update the time slot load state of the data plane.
[0129] The technical solution of this embodiment solves the problems of complex transmission and difficult scheduling in the case of coexistence of deterministic traffic flows with multiple quality of service requirements by receiving the current time slot load, connection information and transmission information sent by the data plane device; generating a target vector according to the current time slot load, the connection information and the transmission information; determining the data transmission path and time slot queue grouping according to the target vector, and sending the data transmission path and time slot queue grouping to the data plane device, and can achieve online real-time scheduling, and then achieve ultra-low latency transmission.
[0130] Embodiment 2
[0131] Figure 3 It is a schematic structural diagram of a scheduling device provided by an embodiment of the present invention. This embodiment is applicable to the situation of scheduling. The device can be implemented in software and / or hardware, and the device can be integrated in any device providing scheduling functions, such as Figure 3 As shown, the scheduling device specifically includes: a receiving module 210, a target vector generation module 220, and a determination module 230.
[0132] Among them, the receiving module is used to receive the current time slot load, connection information and transmission information sent by the data plane device;
[0133] The target vector generation module is used to generate a target vector according to the current time slot load, the connection information and the transmission information;
[0134] A determination module, configured to determine a data transmission path and a time slot queue grouping according to the target vector, and send the data transmission path and the time slot queue grouping to the data plane device.
[0135] Optionally, the transmission information includes: service data, a sending end of the service data, a receiving end of the service data, a time slot index corresponding to the service data, a generation period of the service data, a service data packet size, and an end-to-end delay corresponding to the service data.
[0136] The above product can execute the method provided in any embodiment of the present invention, and has corresponding function modules and beneficial effects for executing the method.
[0137] The technical solution of this embodiment is to receive the current time slot load, connection information, and transmission information sent by the data plane device; generate a target vector according to the current time slot load, the connection information, and the transmission information; determine a data transmission path and a time slot queue grouping according to the target vector, and send the data transmission path and the time slot queue grouping to the data plane device, which solves the problems of complex transmission and difficult scheduling in the case of coexistence of deterministic traffic flows with multiple quality of service requirements, can achieve online instant scheduling, and further achieve ultra-low latency transmission.
[0138] Embodiment III
[0139] Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0140] As Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as read-only memory (ROM) 12, random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0141] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0142] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the scheduling method.
[0143] In some embodiments, the scheduling method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the scheduling method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the scheduling method in any other appropriate manner (e.g., by means of firmware).
[0144] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0145] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0146] In the context of the present invention, a computer-readable storage medium may be a tangible medium that can contain, or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0147] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0148] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0149] A computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0150] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0151] The above specific implementation manners do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A scheduling method, characterized in that, including: Receiving the current time slot load, connection information, and transmission information sent by the data plane device; Generating a target vector based on the current time slot load, the connection information, and the transmission information; Determining a data transmission path and a time slot queue grouping according to the target vector, and sending the data transmission path and the time slot queue grouping to the data plane device; Determining a data transmission path and a time slot queue grouping according to the target vector, and sending the data transmission path and the time slot queue grouping to the data plane device, including: Inputting the target vector into a Q-network model to obtain a target action value; Obtaining the data transmission path and the time slot queue grouping corresponding to the target action value; Sending the data transmission path and the time slot queue grouping to the data plane device; Inputting the target vector into a Q-network model to obtain a target action value, including: Inputting the target vector into a Q-network model; Obtaining the Q value of each action output by the Q-network model, where the actions include: data transmission path and time slot queue grouping; Determining the action value with the largest Q value as the target action value.
2. The method according to claim 1, characterized in that, The transmission information includes at least one of: service data, the sending end of the service data, the receiving end of the service data, the time slot index corresponding to the service data, the generation period of the service data, the service data packet size, and the end-to-end delay corresponding to the service data.
3. The method according to claim 1, characterized in that, After determining the data transmission path and the time slot queue grouping according to the target vector, it further includes: Determining a target transmission delay according to the data transmission path, the time slot queue grouping, and the transmission information; If the target transmission delay is greater than the end-to-end delay corresponding to the service data, determining that the scheduling fails; If the target transmission delay is less than or equal to the end-to-end delay corresponding to the service data, determining the target time slot load corresponding to each along-the-way time slot of the data transmission path according to the data transmission path, the time slot queue grouping, and the transmission information; Sequentially allocating the target time slot load corresponding to each along-the-way time slot to the corresponding along-the-way time slot.
4. The method according to claim 3, wherein It further includes: Receiving the maximum time slot load sent by the data plane device; If the target time slot load corresponding to any along-the-way time slot is greater than the maximum time slot load, determining that the scheduling fails; If the target time slot loads corresponding to the along-the-way time slots are all less than or equal to the maximum time slot load, updating the current time slot load according to the target time slot loads corresponding to the along-the-way time slots.
5. The method according to claim 4, wherein After, if the target time slot loads corresponding to the along-the-way time slots are all less than or equal to the maximum time slot load, updating the current time slot load according to the target time slot loads corresponding to the along-the-way time slots, it further includes: Determining the remaining time slot load according to the maximum time slot load and the target time slot load; Determining a reward value according to the remaining time slot load; Optimizing the Q-network model according to the reward value to obtain an optimized Q-network model.
6. A scheduling device, characterized in that, including: A receiving module, configured to receive the current time slot load, connection information, and transmission information sent by the data plane device; A target vector generation module, configured to generate a target vector based on the current time slot load, the connection information, and the transmission information; A determination module, configured to determine a data transmission path and a time slot queue grouping according to the target vector, and send the data transmission path and the time slot queue grouping to the data plane device; Specifically, the determination module is configured to: Input the target vector into a Q-network model to obtain a target action value; Obtain the data transmission path and the time slot queue grouping corresponding to the target action value; Send the data transmission path and the time slot queue grouping to the data plane device; The determination module is further configured to: Input the target vector into the Q-network model; Obtain the Q value of each action output by the Q-network model, where the actions include: data transmission path and time slot queue grouping; Determine the action value with the largest Q value as the target action value.
7. The device according to claim 6, characterized in that, The transmission information includes: service data, the sending end of the service data, the receiving end of the service data, the time slot index corresponding to the service data, the generation period of the service data, the service data packet size, and the end-to-end delay corresponding to the service data.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor can execute the scheduling method according to any one of claims 1-5.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the scheduling method according to any one of claims 1-5 is implemented.