Network scheduling method and apparatus

The network scheduling method and apparatus efficiently allocate network resources by applying multiple stages of arbitration and circulating failed requests, resulting in higher throughput and deterministic latency, addressing the challenges of exponential data growth and nanosecond timescale requirements.

WO2025120314A1PCT designated stage expired Publication Date: 2025-06-12UCL BUSINESS LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
PCT/GB2024/053031
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2024-12-04
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing network scheduling technologies face challenges in efficiently allocating network resources, such as path, timeslot, and wavelength, on nanosecond timescales to handle exponential growth in on-demand data while managing cost, power consumption, and latency.

Method used

A network scheduling method and apparatus that receive multiple requests for network resources, apply multiple stages of arbitration in different resource domains to identify contentious requests, and circulate failed requests to retry in other timeslots, ultimately granting surviving requests and buffering failed requests for reconsideration in the next epoch.

Benefits of technology

The solution achieves higher throughput, more deterministic latency, and reduces the number of clock cycles per epoch, outperforming previous network schedule processing units by increasing spatial parallelism and optimizing resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2024053031_12062025_PF_FP_ABST
    Figure GB2024053031_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A network scheduling method, for allocating network resources for communications between a plurality of nodes connected in a network, is performed as follows. Multiple requests for network resources are received, each request being for connection from one node among the plurality of nodes, acting as a source node, to another node among the plurality of nodes, acting as a destination node. A forthcoming epoch of time is considered, the epoch being the network reconfiguration cycle time, the epoch being divided into T timeslots, a timeslot being the shortest period of time that a communication circuit in the network remains active for communication. A plurality of windows of requests, from said received requests, are selected for processing; each window comprising up to T requests for each source node. Multiple stages of arbitration are applied sequentially to the requests selected for processing in different re source domains to identify contentious requests in different domains. For each stage of arbitration, (i) where a request is uncontentious, that request is indicated for potential grant, and (ii) where there is a contention between requests, one request is picked for potential grant and the other contentious requests are indicated as failed. Requests are circulated within each window to retry failed requests in other timeslots. After circulation and repeated arbitration: all surviving requests indicated for potential grant are granted, which comprises outputting grant information to source nodes at each timeslot iteration, and outputting network resource configuration information at each epoch for the granted requests. Any remaining requests that are indicated as 'failed requests' are buffered in a buffer for reconsideration in another epoch.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] NETWORK SCHEDULING METHOD AND APPARATUS

[0002] FIELD OF THE INVENTION

[0003] The present invention relates to a method and apparatus for network scheduling, in particular to allocate network resources for communication between nodes connected in the network.

[0004] BACKGROUND OF THE INVENTION

[0005] Recent growth in volume of on-demand data has been exponential. Much of the data is stored / processed in and accessed from data centers in which large numbers of servers are connected in a network. It is a continual challenge to scale these and other networks and still handle the network traffic, while managing cost, power consumption, latency and so forth. A fundamental challenge is how to schedule and re -configure network resources (switches, buffers, paths, etc.) in nanoseconds in order to transfer information data from any to any network termination point in a deterministic way. Processing the traffic demands in order to allocate network resources, such as path, timeslot and wavelength, is a computationally hard problem, and needs to be handled on nanosecond timescales.

[0006] The present invention has been devised in view of the above problems.

[0007] SUMMARY OF THE INVENTION

[0008] According to a first aspect of the invention there is provided a network scheduling method for allocating network resources for communications between a plurality of nodes connected in a network, the method comprising: receiving multiple requests for network resources, each request being for connection from one node among the plurality of nodes, acting as a source node, to another node among the plurality of nodes, acting as a destination node; considering a forthcoming epoch of time, the epoch being the network reconfiguration cycle time, the epoch being divided into T timeslots, a timeslot being the shortest period of time that a communication circuit in the network remains active for communication, and selecting for processing, from said received requests, a plurality of windows of requests, each window comprising up to T requests for each source node; applying, to the requests selected for processing, multiple stages of arbitration sequentially in different resource domains to identify contentious requests in different domains, and, for each stage of arbitration, (i) where a request is uncontentious, indicating that request for potential grant, and (ii) where there is a contention between requests, picking one request for potential grant and indicating the other contentious requests as failed; circulating requests within each window to retry failed requests in other timeslots; and after circulation and repeated arbitration: granting all surviving requests indicated for potential grant, comprising outputting grant information to source nodes at each timeslot iteration, and outputting network resource configuration information at each epoch for the granted requests; and buffering the requests that remain indicated as failed requests in a buffer for reconsideration in another epoch.

[0009] Another aspect of the invention provides a network scheduling apparatus for allocating network resources for communications between a plurality of nodes connected in a network, the apparatus being arranged to: receive multiple requests for network resources, each request being for connection from one node among the plurality of nodes, acting as a source node, to another node among the plurality of nodes, acting as a destination node; consider a forthcoming epoch of time, the epoch being the network reconfiguration cycle time, the epoch being divided into T timeslots, a timeslot being the shortest period of time that a communication circuit in the network remains active for communication, and select for processing, from said received requests, a plurality of windows of requests, each window comprising up to T requests for each source node; apply, to the requests selected for processing, multiple stages of arbitration sequentially in different resource domains to identify contentious requests in different domains, and, for each stage of arbitration, (i) where a request is uncontentious, indicate that request for potential grant, and (ii) where there is a contention between requests, pick one request for potential grant and indicate the other contentious requests as failed; circulate requests within each window to retry failed requests in other timeslots; and after circulation and repeated arbitration: grant all surviving requests indicated for potential grant, comprising outputting grant information to source nodes at each timeslot iteration, and outputting network resource configuration information at each epoch for the granted requests; and buffer the requests that remain indicated as failed requests in a buffer for reconsideration in another epoch. The network scheduling method and apparatus, with a novel scheduling architecture, outperforms previous network schedule processing units. In examples, spatial parallelism is carefully increased in the architecture to increase decisions per cycle and reduce the required number of cycles per epoch, enabling a reduced execution time. An example implementing the invention is able to achieve a higher throughput, more deterministic latency and a smaller number of clock cycles per epoch.

[0010] Further optional features of the invention are defined in the dependent claims .

[0011] DESCRIPTION OF THE DRAWINGS

[0012] Embodiments of the invention will now be described, by way of non -limiting example, with reference to the accompanying drawings. The invention may further comprise, in any combination, any features of the embodiments which will now be described.

[0013] Fig. 1 illustrates an example of a data centre optical circuit switched (OCS) network architecture;

[0014] Fig. 2 is a schematic flow chart showing an outline of a network scheduling method;

[0015] Fig. 3 shows schematically a hardware architecture of a network scheduling apparatus also known as a data scheduling unit (DSU) ;

[0016] Fig. 4A illustrates schematically features of the initialization block of Fig.3 ;

[0017] Fig. 4B shows an example of load balancing performed by the load balancer of Fig.4A;

[0018] Fig. 5 illustrates the operational principle of circulation and arbitration performed by the DSU compute engine in the hardware example of Fig. 3 ;

[0019] Fig. 6 illustrates a worked example of circulation and arbitration for network scheduling; and

[0020] Fig. 7 schematically illustrates an example of the resource allocation process performed by the resource allocation unit of Fig. 3.

[0021] DETAILED DESCRIPTION OF THE INVENTION

[0022] Fig. 1 shows a portion of an example of an OCS network. The complete network comprises a plurality of communication groups (CGs or clusters). Each CG comprises a plurality of racks, and each rack comprises a plurality of nodes (i.e. end nodes, e.g. servers). Each node has multiple wavelength-tunable transceivers, each connected to a passive or active all-to-all core network that can be realised by a range of network architectures. In the illustration of Fig. 1, each port of the 1 :X splitter (coupler) has a semiconductor optical amplifier (SOA) to switchably and selectively amplify an optical signal for broadcast into waveguides. Modular subnets interconnect the racks and CGs. The optical ‘connections’ , defining paths between transmitting transceivers and receiving transceivers, are reconfigured (switched) every cycle of operation, and the time duration of each cycle is referred to as an epoch. The task of a network scheduler is to allocate the network resources, such as path, wavelength and timeslot, such that the network can be efficiently configured to meet the traffic demands.

[0023] This is purely one example of a network; the network scheduling described herein can be applied to any suitable circuit-based switching network architecture, which can be optical or electrical.

[0024] Fig. 2 shows a high-level flow chart of the operation of a network scheduling method for allocating network resources for communications between a plurality of nodes connected in a network. Box S10 comprises receiving multiple requests for network resources, each request being for connection from one node among the plurality of nodes, acting as a source node, to another mode among the plurality of nodes, acting as a destination node. Box S12 comprises considering a forthcoming epoch of time, divided into T timeslots, a timeslot being the shortest period of time to be scheduled for communication, and selecting for processing, from said received requests, a window of requests comprising up to T requests for each source node. Box S14 comprises: applying, to the requests selected for processing, multiple stages of arbitration sequentially in different resource domains to identify contentious requests in different domains, and, for each stage of arbitration, (i) where a request is uncontentious, indicating that request for potential grant, and (ii) where there is a contention between requests, picking one request for potential grant and indicating the other contentious requests as failed; and circulating requests within each window to retry failed requests in other timeslots. Box S16 comprises, after circulation and repeated arbitration: granting all surviving requests indicated for potential grant and outputting information for respective network resource allocation for the relevant timeslots for the granted requests; and buffering the requests that remain indicated as failed requests in a buffer for reconsideration in another epoch. More specific details of the network scheduling method and hardware architecture to implement it will be explained below.

[0025] DATA SCHEDULING UNIT (DSU)

[0026] A. Working Principle

[0027] Fig. 3 shows a hardware architecture of a network scheduling apparatus also known as a data scheduling unit (DSU) or scheduler. The DSU works across x2subnets and, therefore, resolves the contention at x destination receivers. The DSU works across x sets of source groups and x sets of destination racks i.e., a total of xN nodes, where N is the number of nodes per rack.

[0028] The role of the DSU is to: (1) collect requests from xN sources and build the demand; (2) allocate path (space division multiplexing, SDM), wavelength (wavelength division multiplexing, WDM) and timeslot (time division multiplexing, TDM) resources to match the demand for every reconfiguration cycle (epoch) ; and (3) respond the grants to the sources along with the corresponding request ID back to the sources. The resource allocation matrix (final target matrix of the scheduler) has an epoch worth of X transceiver configurations, W wavelengths and T timeslots for each sub-net. The goal of the DSU is to maximise the usage of resources in this matrix for every reconfiguration cycle.

[0029] At any given point in time, the DSU deals with a single epoch. For the scheduling to be efficient, it has three primary goals: (1) the scheduling time must be < the epoch size; (2) the scheduling should achieve close to maximal matching with respect to the demand; and (3) failed requests should not live for a long time in the buffer to avoid long tail latencies. The behavioural model of the DSU architecture can be divided into 7 functional blocks: initialization, circulation, arbitration, resource allocation, structural integrity manager, buffer manager and demand / grant manager as shown in Figure 3. Figure 3 shows the architecture and the building blocks of the hardware-based register transfer level (RTL) design (which can be implemented, for example, on ASIC / FPGA) of the DSU. The number shown in the circle is the order in which the modules are used whilst scheduling. Each of these independent blocks and their working principle will be explained in detail in the following sub-sections. B. Overview

[0030] Step 1 : The fresh requests from the Demand Manager and retry requests from Buffer Manager are sent to the Initialize block.

[0031] Step 2: The initialize block has: (i) an iteration manager and (ii) a load balancer to ensure that the load is equally distributed across the multi-stage arbiters. It also validates the requests every iteration. The initialized requests (both buffered requests B and new requests R) are stored on the Request BRAM.

[0032] Step 3: There are two stages of circulation here: (1) Request window selection and (2) Request window circulation. The objective of Request window selection is to select a window of requests N x T requests out of either the N x R or B requests depending on iteration number, where T is the number of timeslots in an epoch. This is sequentially selected from the Request BRAM. The first T requests are selected first, then the subsequent T requests and so on. Each T requests are given a block of Ititerations. The objective of Request window circulation is to retry failed (leftover) requests in other timeslots (in other arbiter domains).

[0033] Step 4: The multi-stage arbiter aims to segregate and isolate different contention planes: sources, wavelengths, path, timeslots, subnets and destination. To achieve this, each of these arbitration planes are either spatially (replicated) or temporally parallelised (staged).

[0034] Note: Step 3 and Step 4 work in concert with each other to circulate failures across multiple different timeslots and perform parallel retries within each window.

[0035] Step 5: The circulated window of requests that were sent to the arbiter in Step 3 are now granted by the arbiters and grants are generated (same layout as request but failed requests are dropped). The grants are reshaped and fed back to the circulator to update the valid bits and request sizes. The structural integrity manager re-shifts the circulator by the same amount, re-attaches the head and tail for the window and sends it to resource allocation to update the valid bits. Step 6: The arbitration stage in Step 4 allocates the wavelength, timeslot, path and subnet and updates the resource map. On the other hand, the de-circulated grants from Step 5 are used to update valid bits - signals used for Step 2 in the next iteration.

[0036] Step 7: Failed requests in all the epochs are dumped in the buffer and the grants are sent to the grant manager for sending it out.

[0037] C. Initialization

[0038] In the DSU, requests are always arranged per source and an arbitrary number of requests can exist per source as the requests build up over time. Each request specifies the destination, the flow size (number of bytes or timeslots required to be communicated), the timestamp, epoch when it was first requested and a valid bit (1 - valid request, 0 - invalid request). The requests that belong to the most recent epoch can be found in the Demand Manager and the requests that have failed in the previous epochs and are waiting for retries dwell in the Buffer Manager - none of the requests are dropped.

[0039] At every epoch, the DSU starts with the initialization (Initialize) functional block. The initialization block begins every epoch by emptying resource allocation map (wavelength, port and timeslot). Once the scheduling starts, a total of I iterations are performed per epoch, which is equal to the number of boot cycles and the number of clock cycles used.

[0040] As mentioned in Step 2 of the overview and shown in Figure 4A, the initialize block has: (i) an iteration manager (controlled based on Equation 1 below) that controls which iterations are given to the buffer / demand manager, (ii) a load balancer based on fixed rotation barrel shifters to ensure requests are evenly spread for parallel multistage arbiter (only once every epoch) , and (iii) a valid updater signal from resource allocation invalidates contradictory requests every iteration. The initialization block either handles (i) failed requests (from previous epochs) that are re-attempting from the buffer or (ii) current requests from the source nodes in any given iteration. An iteration is shown as a schedule process that goes through Steps 2-7 in Figure 3. So, the DSU either works on (i) N x R requests or (ii) N x B requests, if the current iteration, i, is < or > IB, respectively, which is calculated using Eq. 1. As the first few iterations inherently have priority (as they have an empty allocation map available), buffer requests are given priority. The valid bits in the buffer manager are connected to an OR Gate together for identifying the depth of the buffer (Simple buffer size flag). Based on equation (1), the number of buffer iterations IB is optimised with a weight, WB, and bias, BB, depending on the current size of the buffer with valid requests, SB-

[0041] The Load Balancer in the ‘Initialize’ functional block (Fig. 4A) works both on the N x R requests from the demand manager and the N x B requests from the buffer manager once per epoch to circulate the requests across each of the N nodes to distribute the requests evenly. As shown by the example in Figure 4B, where each matrix element is a destination number, the request on each source is circulated such that the variance in the number of valid requests per column is minimised (on the left of Fig. 4B the number of valid requests per column varies from zero to six; on the right of Fig. 4B, after load balancing, the number of valid requests per column varies from two to five) . The circulators use a barrel shifter with static configurations to ensure circulation of all requests on all nodes by a unique shift value based on a look-up table (LUT). Majority of the time, the buffer is full, so IB > I and hence, all the iterations are given to the buffer.

[0042] At each iteration, a valid signal from the resource manager block informs the initialize functional block of the updated contradictory (invalid) requests every iteration, which updates the Request RAM. If the request is invalid across all retries, the request is dropped form the RAM to the Buffer Manager for retry in next epoch (shown by the arrow in Figure 3).

[0043] D. Circulators and Multi-stage Arbiters

[0044] The requests that are received by the DSUs have contention at multiple levels: sources, wavelengths, path, timeslots, subnets (if X > 1 , where X is the number of racks per communication group) and destinations. The numerous layers of contention make the computation of scheduling an extremely complex problem that is impossible to solve perfectly. However, in this stage (highlighted as step 4 in the overview), the convoluted connection requests are dissected into multiple orthogonal planes of contention and each of them is dealt with separately. [In an example: D is the number of DSUs; P is the number of transceivers; X is the number of racks; in this case X is also the number of splits, but in general it can be x, where x is any integer between 1 and X].

[0045] In Figure 3, the two functional blocks ‘circulators’ and ‘multistage arbiters’ are highlighted (along with ‘Structure Integrity Manager’) as these blocks form the ‘engine’ of the DSU that perform the computation of the schedule (the ‘DSU Compute Engine’). In the overview, they are described in step 3 -6 in section B above.

[0046] D.l. Working Principle:

[0047] This block employs T x N-port arbiters, one for each timeslot. Each of the T requests per source from the request shaper are mapped to a different timeslot space (1 , 2, ...T) and hence, do not contend (in the R or B dimension). However, within a given timeslot, requests of different sources to the same destination is where the arbiter comes into action. The priority encoder within these arbiters can be set to simple round-robin (RR) (which is the preferred embodiment) or, for example, can be oldest first. The critical path length of the DSU lives in the arbiters due to the long carry chains of the priority encoder. In one example, a highly-parallel 1000 x 64-port block on 45 nm CMOS ASIC can achieve 435 MHz clock speed with the use of the correct carry look- ahead chain. The present method is original in the way the arbitration is employed i.e., the co-ordination of distributing and circulating the requests.

[0048] As mentioned before, T parallel arbiters are spatially parallelised to allocate timeslot resources. For other contending domains, a temporally parallelised arbitration method is employed. So, for every timeslot, when one request from a single source has been assigned for a particular destination:

[0049] 1. Source contention: the source cannot be activated for a path or wavelength corresponding to a different destination.

[0050] 2. Destination contention: the destination cannot be activated for a request corresponding to a different source. To resolve these contentions, a multi-staged arbitration engine is used as shown in Figure 5. As deciding the destination is assumed to decide the destination node, the three contention stages (wavelength, path, source and destination) are resolved by pipelining multiple arbiters back-to-back of different dimensions and reshaping (wiring) the inputs accordingly. The output of these arbiters is a contention-free grant matrix that is similar to the request matrix but has no contentions.

[0051] As mentioned before, an epoch is the computational / reconfiguration cycle time that is composed of T timeslots. The DSU engine begins by performing request window selection. The window selection stage is where a window of requests is selected and priority is managed; up to T sets of requests (destination node #, destination rack #, size) from NP sources are selected every iteration. The DSU has up to I iterations and the order of priority is sequential (first gets priority over second and so on). Hence, the long-dwelling buffer requests (i.e., demands of the oldest epoch, which is sorted in the buffer manager) are chosen first to avoid long latency and choose the windows sequentially. To parallelise the compute of requests from all NP nodes in a group across all timeslots, T sets of hardware round-robin arbiters are replicated i.e., one for each timeslot. The arbiters that work on each of the T timeslots are fixed with a number (Ti, T2. . . TT).

[0052] D.2. Visual Example:

[0053] In this section, further explanation of an example is given by delving into the intricate details of step 3-6 of the overview in sub-section B by breaking it down into vii (seven) stages. Figure 6 serves as a visual aid to show how each of the functional blocks works together for 1 iteration. Some of these stages are commonly highlighted in Fig. 5. Figure 6 shows an example of how a request is processed by the DSU engine. In this example, 4 source nodes from 2 racks are shown to request 4 destinations each. The top half shows the destination nodes requested and the bottom half in blue shows the destination rack.

[0054] Stage (i): In the window selection stage, 1 window of T timeslots is selected; hence, T = 2 parallel requests out of B = 4 parallel requests across all source nodes are shown to be selected and held for up to Ii = 3 scheduler iterations. In the first stage of the DSU engine, window selection is shown by the highlighted two requests from each source. As shown, the system example has a total of 6 available iterations, 3 of which will be used for the first time window and the remaining 3 for the second. The total number of iterations for which the window is held is based on a sum of natural numbers formula, implemented on a LUT.

[0055] Stage (ii): For every iteration that the window selection is held, the circulator re-attempts the selected requests across up to T sets of arbiters. The circulator does a barrel shift on every element of the NPT matrix by one step per iteration. This operation is based on the iteration number as shown in Fig. 5 in the black box under the arbiter elements. In the example in Fig. 6, the requests, i.e. destination node # and rack # can be seen shifted by 1 in stage (ii).

[0056] Stage (iii): After circulation, the requests go through the first set of arbitration where the arbiters ensure that a unique destination per rack number is requested by each source transmitter. Hence, this is a source contention resolver. Requests are granted provided requested destination nodes are unique or the requested rack numbers for same destination node numbers are unique; the rest fail and are gathered as invalids. In Fig. 5, this stage is highlighted as Arbitration Engine 1. In Fig. 6, stage (iii) shows examples of how some requests fail in arbitration and how only one destination node # per rack# is selected. For example, in Fig. 6, both node 1 and 2 of rack 1 request destination node 2, but as they request to different rack #, they do not fail in arbitration. However, when node 3 and 4 request destination node 2 of rack 2, one of them is selected by the arbiters.

[0057] Stage (iv): The matrix is reshaped in this stage by stitching together requests of the same source node number accounting previously arbitrated information. In other words, the grant matrix of the first arbitration engine is reshaped into the request matrix of the second arbitration engine, which arbitrates across a different plane. In the visual example, requests from source node n of rack 1 and rack 2 form n different matrices. In summary, Stage (iii) has N x T x P requests being processed in parallel with N-port destination arbiters; Stage (iv) has N x N x T requests being processed with P-port subnet arbiters. So, the reshape in Stage (iv) re-arranges a N x T x P x N matrix to a N x N x T x P matrix to prepare the requests for the next arbitration. This second set of arbitration is denoted by Arbitration Engine 2 in Fig. 5. Each destination receiver in each node shares a path across multiple sub -nets and hence, multiple sources. Hence, the role of these arbiters is to ensure that only one source node is given access to a single destination node per rack. Hence, this is a destination contention resolver. A similar layout to arbitration engine 1 , this stage is composed of smaller P-port arbiters.

[0058] Stage (v): The final grant matrix is rearranged and the invalid grants are separated from the valid grants. The re -mapped output (only 1 -valid bit for each request) is stored in another set of registers. Invalidation is the final stage where a request can drop, even after making it through the final grant matrix. Here, each resource (wavelength, timeslot and path) has an availability flag (bit) to ensure that only one valid resource is granted per destination node per rack.

[0059] Stage (vi): After this, the requests are re-arranged to the original order by a symmetrical barrel-shifter-based de-circulator, as shown by the box under arbitration engine 2 in Fig. 5.

[0060] E. Resource Allocation and Register Updates:

[0061] As shown in Fig. 3, the resource allocation functional block has a block RAM, two sets of inputs and three outputs. For inputs, (i) the block takes wavelength, timeslots, subnet, and path allocations from the arbiter and updates the resource map and (ii) updates the valid and invalid bits of the current requests (from the circulator). For outputs, (i) the buffer manager is updated with the failed requests from the previous buffer and the demand manager, (ii) the grant manager is updated with grant ID, network resource allocation from the resource map and (iii) the valid flags are sent to the initialize module for requesting only valid requests every iteration.

[0062] This functional block is responsible for the following actions: (i) At the end of every iteration, this block writes the corresponding control registers in the initialize (request / buffer manager) to co-ordinate the invalidation of already granted requests to stop / cancel multiple grants requesting the same timeslot, (ii) The resource map (wavelength / timeslot / path for every sub -net) is stored in the BRAM. (iii) The grants generated are sent to the Grant manager. See Fig. 7 for further information on the resource manager block operation.

[0063] Stage (vii): In the example in Fig. 6, resource allocation as to how wavelength and timeslots are allocated when the destination number is the wavelength number in the system, thus simplifying the receiver design. The entire process (all vii stages) is repeated by the scheduler in multiple iterations, as much as the clock speed and the epoch allows. Multiple different windows are selected as iterations progress to maximise network utilisation in order of priority. At every iteration the valid requests and invalid / failed requests are stored. The valid requests make it to the grant manager, while invalid / failed requests are queued for the buffer.

[0064] F. Grant Shaper / Manager:

[0065] The Grant Shaper block is represented by the structural integrity manager inside Fig. 3. The grants are organized and arranged with the corresponding IDs, wavelength, timeslot and epoch number for communication to the nodes. The functional block reshapes the circulated and arbitrated requests to the original shape to make invalidation and grant mapping easier. The N x T grants from the arbiters are stitched together and two actions are performed by the structural integrity manager: (i) the requests are decirculated back to the original order and (ii) the truncated requests from the current iteration are re-appended or if zeros were padded, they are removed. A finite state - machine inside the structural integrity manager decides the order in which these functions is executed.

[0066] Once the grants are organized and arranged, they are stored in the BRAM within the Grant Manager. At the end of every epoch, the Grant Manager takes the responsibility of sending the grants and wavelength control information back to the nodes.

[0067] G. Buffer Manager:

[0068] At the end of every I iterations, the failed requests are all queued per source and stored in the buffer for retry in the future epochs. Sorting is simplified by the careful co-ordination of de-circulating the load balancers at the end of every epoch. This ensures that the oldest (longest living) requests in the buffer are always at the head of the line. There are two types of requests that are buffered: (i) failed requests from buffers and (ii) failed requests from servers. The buffer manager arranges the buffer in the same order. H. Pipelining:

[0069] The DSU uses both spatial and temporal parallelism to achieve maximum network utilization and minimise the clock rate. As requests never re-attempt in the same timeslot because of the barrel shifter (circulation before every arbitration), there is no possibility of requests repeatedly requesting the same timeslot. Hence, the pipelining does not enforce cancellation, ensuring a sustained throughput.

[0070] Implementation

[0071] One example has been implemented as an ASIC-based fast-OCS scheduler. The DSU computes a schedule in 10s of nanoseconds at 90% network load, scales to support 8 clusters of 64-nodes (512-node network) and achieves a throughput as high as >98%, low average and median latency <10 ps, and tail latency of <100 ps with tolerance to skewed traffic, in merely 66 clock cycles, greatly lowering the conventional execution time. The resource allocation is different from conventional network schedulers in the way it prioritises, retries and arbitrates to achieve this performance. This technology can support a larger high-radix optical network and greatly simplify the control plane, eliminating convergence and shared-memory problems.

[0072] It is possible to implement embodiments of apparatus of the invention as one or more hard-wired electronic circuits, some or all of which can be integrated onto a single electronic chip or plurality of chips, such as dedicated FPGAs or ASICs.

[0073] It is understood that the term ‘network’ used herein encompasses not only a ‘single’ network, but also a plurality of nodes interconnected with a plurality of sub networks (or subnets) forming a giant network.

Claims

CLAIMS1. A network scheduling method for allocating network resources for communications between a plurality of nodes connected in a network, the method comprising: receiving multiple requests for network resources, each request being for connection from one node among the plurality of nodes, acting as a source node, to another node among the plurality of nodes, acting as a destination node; considering a forthcoming epoch of time, the epoch being the network reconfiguration cycle time, the epoch being divided into T timeslots, a timeslot being the shortest period of time that a communication circuit in the network remains active for communication, and selecting for processing, from said received requests, a plurality of windows of requests, each window comprising up to T requests for each source node; applying, to the requests selected for processing, multiple stages of arbitration sequentially in different resource domains to identify contentious requests in different domains, and, for each stage of arbitration, (i) where a request is uncontentious, indicating that request for potential grant, and (ii) where there is a contention between requests, picking one request for potential grant and indicating the other contentious requests as failed; circulating requests within each window to retry failed requests in other timeslots; and after circulation and repeated arbitration: granting all surviving requests indicated for potential grant, comprising outputting grant information to source nodes at each timeslot iteration, and outputting network resource configuration information at each epoch for the granted requests; and buffering the requests that remain indicated as failed requests in a buffer for reconsideration in another epoch.

2. A method according to claim 1 , wherein receiving multiple requests for network resources comprises one of: receiving requests not previously considered; retrieving previously failed requests that have been buffered; a mix of receiving requests not previously considered and retrieving previously failed requests that have been buffered.

3. A method according to claiml or 2, wherein priority is given to previously failed requests that have been buffered.

4. A method according to claim 1 or 2, further comprising an initialization stage to load balance requests across the timeslots in an epoch.

5. A method according to any preceding claim, wherein the requests selected for processing are arranged in a matrix according to source nodes and requests for epoch, and circulation is performed using a barrel shift on every element in the matrix.

6. A method according to any preceding claim, wherein the network comprises a plurality of subnets, and wherein the resource domains of contention include at least two of: source node, destination node, and subnet.

7. A method according to any preceding claim, further comprising pre -computing the number of iterations allocated for circulation and arbitration for each epoch.

8. A method according to any preceding claim, further comprising selecting a window using a barrel shifter, shifting the selected window to the head of the line for processing, and saving the rest of the window as a tail to be processed in subsequent iterations.

9. A method according to any preceding claim, further comprising performing circulation of requests across different timeslots or arbitration planes within each window to re-attempt and generate more grants with incremental iterations.

10. A method according to any preceding claim, wherein the grant information output to source nodes comprises identification information for timeslot, destination and epoch; and wherein the output network resource configuration information comprises identification information for transmission wavelength, path, timeslot, and epoch for each network interface card.

11. A network scheduling apparatus for allocating network resources for communications between a plurality of nodes connected in a network, the apparatus being arranged to: receive multiple requests for network resources, each request being for connection from one node among the plurality of nodes, acting as a source node, to another node among the plurality of nodes, acting as a destination node;consider a forthcoming epoch of time, the epoch being the network reconfiguration cycle time, the epoch being divided into T timeslots, a timeslot being the shortest period of time that a communication circuit in the network remains active for communication, and select for processing, from said received requests, a plurality of windows of requests, each window comprising up to T requests for each source node; apply, to the requests selected for processing, multiple stages of arbitration sequentially in different resource domains to identify contentious requests in different domains, and, for each stage of arbitration, (i) where a request is uncontentious, indicate that request for potential grant, and (ii) where there is a contention between requests, pick one request for potential grant and indicate the other contentious requests as failed; circulate requests within each window to retry failed requests in other timeslots; and after circulation and repeated arbitration: grant all surviving requests indicated for potential grant, comprising outputting grant information to source nodes at each timeslot iteration, and outputting network resource configuration information at each epoch for the granted requests; and buffer the requests that remain indicated as failed requests in a buffer for reconsideration in another epoch.

12. An apparatus according to claim 11 , wherein the network is an optical circuit switched network supporting communication over multiple wavelength channels .

13. An apparatus according to claim 11 or 12, wherein the network comprises a plurality of communication groups, each communication group comprising a plurality of racks, each rack comprising a plurality of nodes, and the network further comprising a plurality of subnets interconnecting the communication groups and racks.

14. An apparatus according to claim 11 , 12 or 13, comprising a barrel shifter for selecting a window, shifting the selected window to the head of the line for processing, and saving the rest of the window as a tail to be processed in subsequent iterations.

Citation Information

Cited By

  • Multi-service broadcast light receiving scheduling method and system based on wavelength allocation

    CN121442228A