System and methods for fast-switched optical data center networks

HK40135167APending Publication Date: 2026-07-17MAX PLANCK GESELLSCHAFT ZUR FOERDERUNG DER WISSENSCHAFTEN EV +1

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Applications
Current Assignee / Owner
MAX PLANCK GESELLSCHAFT ZUR FOERDERUNG DER WISSENSCHAFTEN EV
Filing Date
2026-05-19
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing fast-switching fiber data center network (DCN) architectures face challenges in time synchronization and device scalability, making large-scale deployment difficult, and existing systems fail to provide a practical end-to-end system implementation.

Method used

The system employs an in-band synchronization method to synchronize ToR switches, combined with Hop-On-Hop-Off (HOHO) routing technology. Through a unified routing approach, it achieves network device synchronization and routing optimization, and uses commercial equipment to build an easily scalable fiber optic data center network architecture that supports plug-in integration of different optical hardware.

Benefits of technology

It achieves an in-band synchronization error of less than 15ns, a network utilization rate of 99.93%, and 99.4% of data packets are sent within the predetermined time slice. It supports flexible integration and upgrades of various optical DCN architectures, reducing deployment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to a method and apparatus for communicating data packets in a data center network (DCN), the data center comprising a plurality of host servers, a plurality of ToR switches connected to the host servers, and an optical network fabric connected to the plurality of ToR switches, where the optical network fabric operates according to a given schedule, wherein a given schedule defines, for each time slice in a sequence of time slices, which top-of-rack switch pairs are connected via a dedicated optical circuit established for the time slice by an optical controller of an optical network structure, where a top-of-rack switch: is synchronized with another top-of-rack switch; receiving a data packet at the ingress port; and sending the data packet to the egress port. According to the invention, the synchronization comprises sending a synchronization message to the other top-of-rack switch within the band.
Need to check novelty before this filing date? Find Prior Art

Description

(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202380099932.3 (22) Application Date 2023.06.29 (85) PCT International Application Entering National Phase Date 2025.12.29 (86) PCT International Application Application Data PCT / EP2023 / 000041 2023.06.29 (87) PCT International Application Publication Data WO2025 / 002526 EN 2025.01.02 (71) Applicant Max Planck Society for the Advancement of Science Address Germany Applicant Free University Foundation (72) Inventors Xia Yiting Li Jialong Lei Yiming F. de Marki R. Josh B. Chandra Sekaran (74) Patent Agency China Council for the Promotion of International Trade Patent & Trademark Office Co., Ltd. 11038 Patent Attorney Zhou Hengwei (51) Int.Cl. G06F 1 / 04(2006.01) H04J 3 / 06(2006.01) H04Q 11 / 00(2006.01) (54) Title of Invention: System and Method for Fast Switching Optical Data Center Network (57) Abstract: This invention relates to a method and apparatus for transmitting data packets in a data center network (DCN), the data center including multiple host servers, multiple top-of-rack (ToR) switches connected to the host servers, and an optical network structure connected to the multiple ToR switches, wherein the optical network structure, according to a given scheduling operation, wherein the given scheduling, for each time slice in a time-slice sequence, defines which ToR switches are connected to a dedicated optical circuit established for the time slice by an optical controller of the optical network structure, wherein the ToR switches: synchronize with another ToR switch; receive data packets at an ingress port; and send data packets to an egress port. According to the invention, synchronization includes sending a synchronization message to the other ToR switch within a band.Claims 2 pages, Description 25 pages, Drawings 9 pages, CN 121420263 A 2026.01.27 CN 1 21 42 02 63 A 1. A method for transmitting data packets in a data center network (DCN), the data center comprising a plurality of host servers, a plurality of top-of-rack (ToR) switches connected to the host servers, and an optical network structure connected to the plurality of ToR switches, wherein the optical network structure, according to a given scheduling operation, wherein the given scheduling, for each time slice in a time-slice sequence, defines which ToR switches are connected to a dedicated optical circuit established for the time slice by an optical controller of the optical network structure, wherein the ToR switches: - synchronize with another ToR switch; - receive data packets at an ingress port; and - send the data packets to an egress port, characterized in that synchronization comprises: sending a synchronization message to the other ToR switch in-band via a circuit of the optical network structure. 2. The method of claim 1, wherein the synchronization message is sent based on the given scheduling. 3. The method of claim 2, wherein the given schedule is an initial synchronization schedule for coarsely accurate synchronization prior to the data center network being put into operation. 4. The method of claim 3, wherein the given schedule is a schedule for operation after the initial synchronization has terminated. 5. The method of claim 4, wherein synchronization includes resynchronizing the top-of-rack switch. 6. The method of claim 5, wherein resynchronization is performed when the synchronized clock of the top-of-rack switch has drifted from the master clock by a current amount exceeding a predefined threshold T. 7. The method of claim 6, wherein the drift amount corresponds to the sum of the synchronization error and the drift error of the top-of-rack switch. 8. The method of claim 7, wherein the synchronization error and the drift error are estimated based on empirical aggregation or statistics of the synchronization error and the drift error of the top-of-rack switch. 9. The method of claim 8, wherein the predefined threshold T is determined based on the overhead and accuracy of resynchronization. 10. The method of claim 9, wherein the predefined threshold T is also determined based on the relative importance of the synchronization error and the drift error. 11. The method of claim 10, wherein if the synchronization error is dominant, the value of T is increased to reduce the resynchronization phases, such that the top-of-rack switch is resynchronized directly with the dominant top-of-rack switch once per cycle of the scheduling. 12. The method of claim 11, wherein if the drift error is dominant, T is decreased to resynchronize each ToR multiple times per cycle via an intermediate reference ToR before the clock drift of the ToR.13. The method of claim 12, wherein the ToR with the smallest drift is selected as the other ToR for resynchronization. 14. The method of claim 13, wherein time slices and ports connected to the other top-of-rack switch are stored in a lookup table on the top-of-rack switch, the lookup table being used to direct synchronization messages in different time slices to specific ports connected to the top-of-rack switch to be synchronized with. 15. The method of claim 14, wherein the lookup table is preloaded into the control plane of the top-of-rack switch. 16. The method of claim 15, wherein batches of the lookup table are periodically injected into the data plane of the top-of-rack switch. 17. The method of claim 16, wherein synchronization returns an offset for adjusting the local clock of the top-of-rack switch. 18. The method of claim 17, wherein the synchronization message is a DPTP message. 19. A top-of-rack switch implementing the method steps of any one of claims 1 to 18. Claims 2 / 2 Page 3 CN 121420263 A System and Method for High-Speed ​​Switching Optical Data Center Networks

[0001] This invention relates to systems and methods for transmitting data in a data center network (DCN). Background Art

[0002] The design of data center networks has largely benefited from Moore's law for networks—the bandwidth of electrical switches doubles every two years at the same cost and power. As such bandwidth expansion slows, a series of optical DCN architectures have been proposed to take advantage of the bandwidth, power, and cost advantages of optical interconnects [13-16, 20, 23, 27, 31-33, 36, 37, 42-16]. Compared to electrical interconnects in conventional DCNs, optical interconnects use circuit switching to establish dedicated optical circuits between endpoints and switch circuits across “time slices” to create time-shared networks.

[0003] So-called slow-switching optical DCNs have switching delays of tens of milliseconds [20, 37, 42, 44-46]. Due to the limitation of switching speed, this type of optical network must work in conjunction with electrical networks to avoid network fragmentation, for example, by offloading high traffic through on-demand circuitry enhancement of the electrical DCN [20, 42], or by acting as a “patch panel” for the electrical switch and reconfiguring the network topology at a granularity of seconds to hours [14, 37, 44-46].For example, after deploying slow-switched optical interconnects in the network core, Google’s Jupiter DCN architecture has achieved a 5x increase in capacity, a 41% reduction in power consumption, and a 30% reduction in cost

[37] . These optical interconnects provide a large port count to interconnect electrical switches and reconfigure the DCN topology when needed, such as during equipment upgrades and failures, or every few hours as DCN traffic evolves.

[0004] On the other hand, Figure 1 shows an example of a typical general-purpose fast-switched optical DCN, which has been increasingly recognized in recent years as an alternative to slow-switched DCNs. Fast-switched optical DCNs consist of optical switches that interconnect top-of-rack (ToR) switches and end servers [31-33, 36]. The architecture uses circuit switching to establish dedicated optical circuits that are time-divided between different ToR pairs for high-speed transmission of aggregated traffic. After establishment, the circuits are reserved for fixed time intervals (called time slices) during which the connected ToR (Fibre Channel module) has exclusive use of them, i.e., without contention with other ToRs. An optical controller (e.g., an FPGA board [13, 27, 32, and 33]) controls a circuit switch to continuously change the circuitry, once per time slice, to route traffic in the optical domain over an all-optical network architecture. The sequence of ToR aspect connections associated with its time slice constitutes the optical schedule. Typically, the schedule is predefined and repeats every optical cycle. There is at least one circuit between each ToR per cycle. The removal of the electrical network also reduces costs compared to slow-switched optical DCNs, but also deviates from the full connectivity assumed by conventional DCN designs. The switching delay of fast-switched optical DCNs varies between a few nanoseconds

[13] and tens of microseconds [32, 33, 36].

[0005] Prior Art

[0006] Table 1 summarizes the limitations of the fast-switched DCN architectures proposed to date:

[0007] Specification 1 / 25 pages 4 CN 121420263 A

[0008] Table 1: Implementation limitations of prior work.

[0009] As shown in Table 1, existing architectures have not yet gone beyond the proof-of-concept prototype of the proposed optical network structure, and their systems, which only have minimal test capabilities, do not suggest practical end-to-end system implementations.

[0010] A common problem with ToRs [33, 36] simulated using Linux servers and hosts [13, 27] simulated using FPGA boards is time synchronization between the ToR and the host and the optical structure. Some architectures have optical controllers [27, 36] that notify the ToR of upcoming circuits, but it is difficult to scale such designs to large DCNs. Some others rely on link up / down events on the ToR / host ports to detect circuit on / off [32, 33], but resetting the port introduces millisecond-level latency—insufficient for the response of a fast-switching optical DCN.

[0011] Mordia

[36] and RotorNet

[33] apply multi-hop routing to long or heavy (“elephant”) flows and short or light (“mouse”) flows (latency-sensitive traffic), which requires large buffers on the ToRs, e.g., about 70MB per switch port in Mordia on a 100Gbps DCN. Mordia and SiP-Ring use dynamic optical scheduling [27, 36] that utilizes real-time traffic estimation calculations, which is difficult in fast-switched optical DCNs, especially for bursty traffic, and neither implemented traffic estimation in their prototypes. Opera imposes a rigid relationship between the number of uplinks and ToRs to ensure the existence of multi-hop paths for each ToR at any given time

[32] , making deployment and scaling challenging. Sirius and SiP-Ring require custom hardware [13, 27], e.g., custom optical modules and optical interfaces on GPUs, which are not yet in production.

[0012] Therefore, several challenges remain for the implementation and eventual deployment of fast-switched optical DCNs. First, network devices need to be synchronized across the network with sub-microsecond or even nanosecond accuracy to maintain traffic and circuit synchronization for rapid reconfiguration. Second, as circuit durations decrease to the same scale as latency on DCN RTTs and host stacks, top-of-rack switches (ToRs) and host systems need to maintain good performance. Third, even if achieved, each optical architecture is a closed ecosystem with highly coupled optical hardware and networking systems, requiring upgrades from one architecture to another after deployment.

[0013] Objective of the Invention

[0014] Therefore, the object of the present invention is to provide methods and apparatus for transmitting data packets in data center networks that address the aspects discussed above. Summary of the Invention

[0015] This object is achieved by the methods and apparatus according to the independent claims. Advantageous embodiments are defined in the dependent claims.

[0016] According to a first aspect, the present invention includes a method for transmitting data packets in a data center network (DCN), the data center including a plurality of host servers, a plurality of top-of-rack (ToR) switches connected to the host servers, and an optical network structure connected to the plurality of ToR switches, wherein the optical network structure, according to a given scheduling operation, wherein the given scheduling, for each time slice in a time-slice sequence, defines which ToR switches are connected to a dedicated optical circuit established for the time slice by an optical controller of the optical network structure, wherein the ToR switches: synchronize with another ToR switch; receive data packets at an ingress port; and send data packets to an egress port, wherein synchronization includes sending a synchronization message to the other ToR switch in-band (i.e., via the circuitry of the optical network structure). In-band synchronization allows the elimination of additional out-of-band networks (such as electrical networks) for synchronization purposes.

[0017] Synchronization messages can be sent based on a given schedule. The given schedule can be an initial synchronization schedule used to synchronize with coarse accuracy before the data center network is put into operation. The given schedule can be a schedule for operation after the initial synchronization has terminated. Synchronization can include resynchronizing the top-of-rack switch. Resynchronization can be performed when the synchronized clock of the top-of-rack switch has drifted from the master clock by a current amount exceeding a predetermined threshold T. The drift amount can correspond to the sum of synchronization and drift errors of the top-of-rack switch. Synchronization errors and drift errors can be estimated based on an empirical profile or statistics of the synchronization and drift errors of the top-of-rack switch. The predefined threshold T can also be determined based on the overhead and accuracy of resynchronization. The predefined threshold T can also be determined based on the relative importance of synchronization errors and drift errors. If synchronization errors are dominant, the value of T can be increased to reduce the resynchronization phase, such that the top-of-rack switch is directly resynchronized with the dominant top-of-rack switch once per cycle of the schedule. If the drift error is dominant, T can be made smaller to resynchronize each ToR multiple times in a cycle via an intermediate reference ToR before its clock drifts. The ToR with the smallest drift can be selected as the other ToR for resynchronization. Time slices and ports connected to the other top-mount switch can be stored in a lookup table on the top-mount switch, which is used to direct synchronization messages in different time slices to specific ports connected to the top-mount switch to be synchronized. The lookup table can be preloaded into the control plane of the top-mount switch. Batch lookup tables can be periodically injected into the data plane of the top-mount switch. Synchronization can return an offset for adjusting the local clock of the top-mount switch. The synchronization message can be a DPTP message.

[0018] According to a second aspect, the invention also includes a top-mount switch that implements one or more of the steps in the method described above.

[0019] According to a third aspect, the present invention includes a method for transmitting data packets in a data center network (DCN), the data center including multiple host servers, multiple top-of-rack (ToR) switches connected to the host servers, and an optical network structure connected to the multiple top-of-rack switches, wherein the optical network structure defines, according to a given scheduling operation, which, for each time slice in a time-slice sequence, which top-of-rack switch pairs are connected via dedicated optical circuits established for said time slice by an optical controller of the optical network structure, wherein the top-of-rack switch: synchronizes with another top-of-rack switch; receives data packets at an ingress port; and sends data packets to an egress port, wherein mouse flows are routed on the fastest (multi-hop) path according to a technique described below in conjunction with HOHO routing. This allows for minimizing their latency.

[0020] According to a fourth aspect, the present invention also includes a top-of-rack switch that implements one or more steps of the HOHO routing method.

[0021] Hereinafter, the methods, apparatus, and systems according to the invention are also referred to by the general term OpenOptics. OpenOptics provides a general framework for easy end-to-end implementation of fast-switching optical DCNs. Its efficient design overcomes the system limitations of existing work. It uses commercial equipment in an architecture that is easy to implement. However, by decoupling the system from the optical hardware, OpenOptics can adapt to new optical hardware if it becomes commercialized in the future.

[0022] OpenOptics does the same thing for fast-switching optical DCNs as OpenFlow[9] does for traditional networks. It allows specific optical hardware to be integrated into a general framework in a plug-and-play manner to have working end-to-end systems, and cloud applications can run without modification, just like on a traditional DCN. As optical technology advances, different optical architectures can be implemented directly on top of OpenOptics, and the system can remain intact when the DCN architecture is upgraded to newer optical hardware. By decoupling the software system from the optical hardware, the niche area of ​​optical DCNs becomes more accessible to network researchers.

[0023] The general enabler in OpenOptics is a unified routing method, subsequently known as Hop-On-Hop-Off (HOHO) routing, which can be applied to different Fast Switched Optical DCN architectures. Most Fast Switched Optical DCN architectures use predefined optical scheduling, i.e., a repeating sequence of circuit connections on time slices, to avoid costly real-time traffic estimation and circuit planning over short time slice durations. HOHO routing leverages this fact, abstracting each architecture through its optical scheduling. It takes the optical scheduling as input and computes the shortest latency path (which has been proven optimal) offline for mouse streams.Each architecture-specific routing algorithm is replaced by HOHO routing, which generates better paths for mouse flows (see specification 3 / 25, page 6, CN 121420263 A) and preserves a direct path between the source and destination ToRs for elephant flows.

[0024] Through unified HOHO routing, ToRs and host systems can also be unified across architectures. The runtime system of the offline HOHO algorithm can be implemented on ToRs and hosts, using P4 on Intel Tofino2 switches and libvma on Mellanox NICs. Considering the aforementioned challenges, the system design revolves around testing the boundaries of these commercial tools used for fast switching DCNs. Specifically, network-wide in-band ToR synchronization is achieved based on aggregated synchronization errors between Tofino2 switches; HOHO routing on ToRs is achieved through careful measurement of system latency on Tofino2 switches at each critical step; and application-agnostic host networking is constructed through a fair assessment of kernel overhead and kernel bypass options.

[0025] Micro-benchmark evaluation of OpenOptics performance shows that the in-band ToR synchronization according to the invention can keep the synchronization error below 15 ns, the ToR system of the invention achieves zero packet loss with 99.93% achievable network utilization, and the host system of the invention sends 99.4% of the packets within the scheduled time slice. The versatility of OpenOptics is demonstrated by implementing Mordia

[36] , RotorNet

[33] and Opera

[32] (three fast-switched optical DCN architectures) on top of it. Case studies of running Memcached[6] and Gloo[2] applications on top of them show that the tail flow completion time of mouse flow in OpenOptics is comparable to that of electric DCN. Brief Description of the Drawings

[0026] Figure 1 shows an example of a fast-switched optical data center network (DCN).

[0027] Figure 2 shows an OpenOptics system according to an embodiment of the invention.

[0028] Figure 3 shows a diagram of a backtracking algorithm for HOHO routing according to an embodiment of the invention.

[0029] Figure 4 illustrates a schematic setup for pooling synchronization errors according to an embodiment of the present invention.

[0030] Figure 5 shows (a) multi-hop synchronization errors and (b) drift errors over different durations at the 99.9 percentile across switch pairs.

[0031] Figure 6 illustrates the initial synchronization phase of the ToR synchronization algorithm according to an embodiment of the present invention. Hk is the pooled k-hop synchronization error (Figure 5a). Di is the pooled drift error of ToRj relative to ToRi in a time slice (Figure 5b).

[0032] Figure 7 shows an example of a routing process according to an embodiment of the present invention.

[0033] Figure 8 shows the queue (a) latency estimation error and (b) rotation offset according to an embodiment of the present invention.

[0034] Figure 9 shows the kernel and libvma latency differences according to an embodiment of the present invention. (a) Signal reception distance. (b) Signal response turnaround time.

[0035] Figure 10 shows the ToR synchronization evaluation according to an embodiment of the present invention—(a) setup and (b) error; the subplot in (b) shows the error of synchronizing each ToR with the dominant ToR once per cycle (top part) and the error of frequently resynchronizing using a threshold T = 15ns (bottom part).

[0036] Figure 11 shows (a) total queue occupancy per port and (b) count of consecutive queues in use according to an embodiment of the present invention.

[0037] Figure 12 shows the delivery time of the last packet of each time slice relative to the end of the time slice according to an embodiment of the present invention.

[0038] Figure 13 shows the flow completion time of (a) Memcached SET and (b) Gloo allreduce migration according to an embodiment of the present invention. Specification 4 / 25 Page 7 CN 121420263 A Detailed Implementation

[0039] Figure 2 illustrates an OpenOptics system according to an embodiment of the present invention, including an offline part and an online part. The core of the OpenOptics framework is HOHO routing, which takes a uniform input of optical scheduling as input for offline path computation, regardless of the specific optical architecture. It generates routing tables for each ToR, which route mouse streams on the fastest (multi-hop) path to minimize their latency. The routing table gives provisional egress ports and queues mapped to the routed paths. Packet admission control checks whether the provisional path is feasible. If the given route is feasible, the packet is enqueued; otherwise, the packet is re-looped to check the next optimal route, with an incrementing re-loop count. The subsystem that implements it on the ToR is described below. ToR synchronization is also described below as a prerequisite for HOHO routing to correctly time time slices. In offline computation, the synchronization scheduler receives the aggregated synchronization and drift errors between switches and the optical schedule to generate a synchronization schedule. Runtime synchronization processes enforce scheduling to achieve the desired synchronization performance and cost.

[0040] Due to memory limitations, the ToR system only stores mouse streams. Elephant streams, which require throughput but are not sensitive to latency, are routed on a direct circuit between ToRs, rather than via a multi-hop path through intermediate ToRs, to save bandwidth. As will be explained below, the host system sends mouse streams to ToRs for processing by HOHO routing and suspends elephant streams until arrival at the destination ToR is signaled to the direct circuit via ToR.

[0041] This design is driven by the short time slice duration of fast-switching optical DCNs and the versatility requirements of OpenOptics. Programmable switches reduce the complexity of ToR systems, and the latency and system overhead on the switch data plane are orders of magnitude lower than those on the host. The ToR-centric design is also less likely to require modifications to the application, for example, the host system is completely transparent to TCP applications.

[0042] Routing

[0043] As illustrated in Figure 1, existing optical routing algorithms route mouse streams over multi-hop optical paths through intermediate ToRs to avoid long wait times for direct connections. Such paths are “uninterrupted,” in the sense that packets must persist until they reach the destination ToR after hopping onto the path. In contrast, HOHO routing seeks the “fastest” path. With the help of new features of programmable switches, HOHO routing allows packets to hop off the original path at an intermediate ToR and onto another optical path, which is later supplied with an earlier arrival time at the destination ToR.

[0044] By decoupling route computation from the runtime system, HOHO routing becomes generalizable. Assuming that packets traverse ToRs (with empty ToR queues) in zero time, the path can be computed offline, taking optical scheduling as input. HOHO generates the fastest path for each source-destination ToR pair for each time slice in the optical scheduling where packets may arrive. The path includes all traversed ToRs, each associated with a time slice leaving the ToR, for example, the departure time slice in Figure 1 (denoted as "t="). <n>HOHO routing has the property that a complete path can be decomposed into a hop-by-hop lookup on each ToR. Therefore, the output path is converted into a static route for the next-hop lookup on each ToR. Before the lookup, the runtime system on each ToR predicts whether the actual queuing delay will cause the packet to miss the scheduled departure time slice, and the ToR reroutes the packet to the next fastest path as needed.

[0045] The goal of HOHO routing is to forward packets from the source to the destination ToR via the fastest path. The fastest path in the optical DCN is defined as the path that requires the minimum number of time slices. Depending on when the packet arrives at the ToR (i.e., its arrival time slice) and the optical scheduling, the fastest path can "hop" through intermediate ToRs.

[0046] HOHO routing consists of an offline routing algorithm that computes the fastest path and a runtime system that orchestrates the forwarding of packets along these paths.

[0047] The HOHO routing algorithm is agnostic to the optical DCN architecture and is general for a wide range of time slice durations. Taking into account cyclic optical scheduling, the offline routing algorithm first computes all corresponding arrival time slices. 121420263 A The fastest path for a source-destination ToR pair. The complete path is then converted into a next-hop lookup table for implementation in the switch's data plane. Logically, for each ToR, there is a next-hop lookup table for each arrival time slice. During runtime, when a ToR receives a packet within a specific time slice, that ToR looks up the next-hop lookup table corresponding to that (arrival) time slice to obtain the egress port and the sending time slice in which the packet is to be transmitted. If the sending time slice is later than the arrival time slice, the switch temporarily buffers the packet until the sending time slice, i.e., the time slice in which the packet should be transmitted. When designing HOHO routing, we assume that packets always arrive at the beginning of a time slice and that there is no queuing delay at the ToR. These assumptions effectively reduce static... The design of the HOHO routing algorithm is decoupled from the runtime system on the switch. In practice, if breaking these assumptions causes a packet to miss its planned transmission time slice to reach the next hop, we perform runtime adjustments to match it with the next time slice. This mechanism finds the next optimal path. The runtime system on the ToR detects upcoming slice misses based on a fair estimate of the packet's arrival time and queuing delay relative to the packet's egress time. The cost function in the HOHO routing algorithm at this point only considers the transmission delay as a cost. To reduce slice misses, we can modify this cost function, for example, to include queuing delay. Then, we need a reasonable measurement or estimate of the queuing delay across different paths across the network scope, which can be collected via network telemetry. Due to performance limitations...The trade-off between improvement and system complexity is controversial and unclear, so we leave this discussion to future work.

[0048] Figure 3 illustrates the backtracking algorithm for HOHO routing. The time slices of the optical circuits are represented as absolute values, where the arrival time slice is t = 0. The destination calls Routing to find the earliest last hop (A and B in the case of t = 5), and then it calls Subpath to find the shortest feasible path from the source through them (S→B→D). Paths that violate various constraints (see the explanation in red) are pruned from this backtracking search.

[0049] The (optical) circuits in the optical DCN are analogous to "buses": the circuits (buses) transport data packets (people) from the source ToR to the destination ToR. The time slice of the circuit is the "deadline" for boarding the "bus". The earliest time to reach the destination ToR (as illustrated in Figure 3) is determined entirely by the earliest time slice of the last-hop ToR (i.e., when the "bus" departs from the last "stop" to the destination (e.g., t=5 in Figure 3)). To find the fastest path, the earliest "bus" to the last hop to the destination must be found (step 1). The "journey" from the source ToR to the last-hop ToR is then planned, satisfying the "deadline" for making all "transfers" (step 2): arriving at the next-hop ToR earlier than the time slice of the previous hop (e.g., waiting for the next "bus"), or within the same time slice (e.g., catching the next "bus" on time). In Figure 3, for example, a packet arriving at D from S→B in slice t=3 must wait 2 time units at B, while a packet arriving at B from G→B in slice t=5 "on time" uses the B→D circuit. Once the last-hop ToR is selected in step 1, the arrival time does not change regardless of how complex the "journey" is in step 2, although shorter paths with fewer "transfers" are preferred.

[0050] Based on this intuition, a backtracking algorithm (Algorithm 1) for HOHO routing can be designed. This algorithm includes two processes: Routing and Subpath, corresponding to steps 1 and 2 respectively:

[0051] Specification 6 / 25 page 9 CN 121420263 A

[0052]

[0053] For a data packet arriving at the source ToR in a specific time slice, the routing process finds the fastest optical path to the destination ToR. It finds the last-hop ToR that provides the earliest arrival at the destination ToR by sorting the time slices of all candidate ToRs connected to that destination (line 2). For each candidate ToR, it calls the Subpath process to find the earliest available optical path from the source ToR.The subpath of the row (line 20). The process exits when the first valid path is found (line 6) or when the search ends (line 11). As per specification 7 / 25 page 10 CN 121420263 A, each ToR pair is guaranteed circuitry in optical scheduling, Routing will always find a path—in the worst case, a direct path (S→D in Figure 3) if no faster path exists. When multiple fastest paths (via A and B in Figure 3) exist, the shortest path is selected (lines 9-10).

[0054] The Subpath process recursively finds a feasible subpath from the source ToR to an intermediate ToR. The process can terminate in two ways: (i) when it cannot find a path with a length of at most the maximum hop count (lines 13-14), for example, S→H→E→A→D in Figure 3, or (ii) when it finds a connection from the source ToR and its time slice can catch up with the "deadline" for the next hop transmission (lines 15-16). For example, in Figure 3, S→B→D satisfies this condition because the time slice t=3 of S→B is earlier than the time slice t=5 of B→D, while S→A→D violates this condition. Regardless of the number of hops traversed, the path must start from the source ToR. Therefore, a path is found if and only if the source ToR is directly connected to the current intermediate ToR. Otherwise, Subpath calls itself to search forward for other intermediate ToRs that are not yet in the subpath and finds a feasible subpath that always satisfies the "deadline" (lines 18-22). If Subpath finds multiple feasible subpaths, we select the shortest one (line 23). In Figure 3, A→F→A→D is filtered out because of the repeated "A", and S→B→D is ultimately selected because it is shorter than another feasible path S→G→B→D;

[0055] The HOHO routing algorithm is optimal: the selected path is the shortest, which brings the least delay. Proof. Let p be the selected path, its last hop ToR to dst be r, and the path length be l. The time slice of the optical connection between r and dst is s. If there exists a better path ˆ from src to dst, which has a last hop ToR ˆ at slice ˆ and a path length of lˆ, then s > ˆ, or s = ˆ and l > lˆ. This can be proven by contradiction: Case I: s > ˆ. In routing, the last hop ToR is traversed in ascending order of its time slice to dst. Therefore, ˆ must be found earlier than p, which is a contradiction. Case II: s = ˆ and l > lˆ. When a tie is broken at the same time slice during routing, ˆ will overwrite p and be selected (lines 9-10), which is a contradiction.

[0056] HOHO routing produces a complete path, including every hop along the way, but only the route lookup on each intermediate ToR is performed.Based on the next hop immediately following. This implementation preserves the optimal path. In other words, each hop lookup produces the optimal path. Proof. Let p be the chosen path, with the first hop ToR from src being r, the last hop ToR to dst being r', the optical connection between src and r at time slice s, the connection between r' and dst at time slice s', and the path length being l. The residual path from r to dst is p' = p - src, the arrival time at r is s, and the path length is l' = l - 1. It can be proven that p' is the optimal path for Routing(r, dst, s). If there exists a better path pˆ' from r to dst at slice s than p', with the last hop ToR to dst being rˆ', then the optical connection between rˆ' and dst is sˆ', and the path length is lˆ'. Therefore, s' > sˆ' or s' = sˆ' and l' > lˆ'. For either case, since pˆ' starts at slice s where src and r are connected, there must be a path pˆ' = src + pˆ' from src to dst, reaching dst at slice sˆ', with a path length of l̂ = lˆ' + 1. Comparing pˆ to p, we can have s' > sˆ' or s' = sˆ' and l > lˆ'. Therefore, pˆ is better than p, which contradicts property 1, which states that the selected path p is optimal. Since p' is optimal, because Routing selects a single path from the feasible paths, Routing(r, dst, s) can return a different optimal path pˆ' equivalent to p', i.e., s' = sˆ' and l' > lˆ'. Then for the complete paths p and pˆ from src, s' = sˆ' and l = lˆ. Therefore, pˆ is also optimal. Repeating the above proof hop-by-hop until dst, we make hop-by-hop lookup produce the optimal path.

[0057] If a packet misses its planned time slice, the switching system adjusts at runtime to reroute the packet through the next available time slice. According to the invention, the runtime adjustment is robust to find the next optimal path starting from the current ToR. In other words, rerouting after missing a planned transmission time slice gives the next optimal path. Proof. Let p be the optimal path from src to dst, R be the set of intermediate ToRs along p, and S be the set of time slices for the ToR connections of adjacent hops. Suppose that Si is missed at Ri and the current time slice is Sc (Sc > Si). By hop-by-hop search, we find the path p' from Ri to dst at slice Sc. According to property 2, p' is optimal with respect to the current time slice Sc.

[0058] For a given optical schedule, the HOHO routing algorithm runs offline per source-destination ToR pair per time slice, and the output is a complete path containing all traversed ToRs, each associated with a time slice leaving the ToR. HOHO routing has the property of reproducing the complete path with a hop-by-hop lookup, so the output path is encoded as a static route for next-hop lookup in the HOHO routing table on each ToR.

[0059] ToR synchronization is a prerequisite for HOHO routing for correctly timing time slices. Due to memory limitations, the ToR system only stores mouse streams. Elephant streams, which require throughput but are not sensitive to latency, are routed on the direct circuit between ToRs instead of multi-hop paths via intermediate ToRs to save bandwidth. The host system sends mouse streams to the ToRs for processing by HOHO routing and suspends elephant streams until arrival at the destination ToR is signaled to the direct circuit via the ToR.

[0060] TOR Synchronization

[0061] Current fast-switching optical DCN architectures propose synchronizing both the ToRs and the hosts with the optical schedule. They perform such time synchronization by detecting link up / down events [32, 33] or via out-of-band signaling over the management network [27, 36]. Using any of these methods in OpenOptics violates its principles of feasibility and universality. Setting up ports takes up to several milliseconds after a link up / down event is detected on commercial switches and hosts, and frequent link up / down events in fast-switching optical DCNs can lead to port flapping. While DCNs typically have out-of-band management networks, it is difficult to always have a high-speed management network connecting all devices for accurate time synchronization.

[0062] Instead, the present invention relies on in-band time synchronization, which simply works on fast-switching optical DCNs without any additional assumptions beyond the optical network architecture. The ToR (rather than the host) is synchronized according to the optical schedule for better scalability. DPTP is tailored to be used as a time synchronization protocol because it supports both switch-to-switch synchronization and switch-to-host synchronization in the data plane, with accuracy on the order of tens of nanoseconds

[26] .

[0063] However, applying DPTP in-band to optical DCNs is challenging. DPTP assumes a fully connected network for continuous-time synchronization, even though the optical DCN only connects ToR pairs on discrete time slices. The transient nature of optical circuits fundamentally limits when and where synchronization messages can be sent. Furthermore, the ToR is unaware of the time slice prior to synchronization, making reliable message exchange difficult.

[0064] Intuitively, the opto-router is reconfigured with the optical controller (e.g., an FPGA board). The ToR can synchronize with such a controller via PTP, which is natively supported by many FPGA boards

[21] and is compatible with DPTP running on the ToR. It can then...Used as a master ToR for synchronizing other ToRs.

[0065] Although ToRs can be preloaded with optical scheduling and time slice durations, they do not know the start point of the optical scheduling before synchronization; therefore, they cannot send synchronization messages at the correct time. To solve this problem, initial synchronization is performed before the network is put into operation. This process is guided by a custom optical scheduling that is used only for synchronization and is distinct from the operational optical scheduling. Initially, the master ToR has a clock. Circuitry is then added between the master ToR and another ToR in the synchronization scheduling so that the master ToR can "push" the clock to that ToR. The ToR that now receives the clock can propagate it to yet another ToR in the next time slice of the synchronization scheduling. This synchronization scheduling and time slice duration are preloaded onto each ToR. After a ToR learns the clock, it knows where and when to propagate it. After the ToRs are synchronized, the network moves to the operational optical scheduling and begins carrying traffic. The ToRs then use periodic resynchronization to recalibrate their clocks.

[0066] Initial time synchronization can be performed naively by pushing the clock from the dominant ToR one by one to each other ToR. Each individual synchronization requires different optical circuits to last for a time slice. Therefore, this process takes as many time slices as the number of ToRs. In the case of high clock drift or a large number of ToRs, the clock of an early synchronized ToR may have drifted by the time the remaining ToRs are synchronized by the specification page 9 / 25 12 CN 121420263 A. Therefore, fast convergence of the global clock requires ToRs to be synchronized in a chain through reference ToRs that have already been synchronized.

[0067] Synchronization accuracy is affected by the interaction of synchronization error (i.e., artifacts of the synchronization protocol) and drift error (i.e., clock drift between two synchronizations of a ToR)

[26] . Here, synchronization error determines the length of the synchronization chain, and drift error determines the frequency of synchronization. These two key parameters are specific to the physical characteristics of the switch chip, and therefore they are combined to design a general synchronization algorithm for different types of ToRs.

[0068] Both synchronization and drift errors need to be measured on the same physical clock. According to an embodiment of the invention, they are estimated by pooling values ​​in an offline synchronization algorithm. The algorithm statically calculates how the estimated errors evolve and how they should be reduced by ToR aspect synchronization. The algorithm outputs optical scheduling for initial synchronization, and synchronization steps for both initial synchronization and resynchronization.

[0069] Parameter Pooling

[0070] In this embodiment, synchronization and drift errors are pooled on three EdgeCore DCS810 Intel Tofino2 switches (denoted as SW1, SW2, SW3). The pooling results are hardware-specific, but general insights can be derived from them.The design of the synchronization algorithm is explained.

[0071] Figure 4 shows a schematic setup for pooling synchronization errors according to an embodiment of the invention. The ToRs are synchronized with each other on a direct optical path, so synchronization and drift errors are pooled between the switch pairs. The three switches are arranged such that for each pair (in order), one switch is the master switch and the other is the secondary switch. They are virtualized as logical ToRs, where the numbers represent logical ToR IDs. Each logical ToR propagates the clock to the next downstream ToR. ToR0 is the master ToR, and the ToR ID in Figure 4 also represents the synchronization hop count. Logical ToRs on the same physical switch share the clock, so the 2, 4, 6, 8, and 10-hop synchronization errors can be measured by calculating the difference between the synchronized clock and the ground truth. Since in-band time synchronization is interleaved with operational traffic, the experiment was run at a 100Gbps line rate using 64B and 1500B packets with and without background traffic.

[0072] Drift errors are independent of synchronization hops and background traffic. Following the methodology in the DPTP paper

[26] , the drift error between each pair of switches (also in sequence) is measured. Specifically, one switch is directly synchronized with another switch (1-hop synchronization), and for consecutive synchronization requests, the drift error is calculated as the elapsed time between two requests based on the local clock minus the elapsed time based on the master clock. The drift error is measured over various time durations.

[0073] Observations

[0074] The 99.9 percentile synchronization error between each pair of the three switches is summarized in Figure 5a as a stable upper limit of the synchronization error that needs to be considered in the synchronization algorithm. It is clear that the 99.9 percentile of the synchronization error is independent of the switch and increases linearly with the synchronization hop count. There are significant differences in the synchronization error across the switch pairs, but the 99.9 percentile number covers a very small range—10 hops within 6 ns and 2 hops within 2 ns. Observations are made by fitting a linear function to estimate the synchronization error for each synchronization hop count, including odd hop counts that cannot be measured. The multi-hop synchronization error (used for the synchronization algorithm) derived from this fitting function is not affected by a specific ToR.

[0075] Figure 5b summarizes the 99.9 percentile drift error between each pair of the three switches as a stable upper limit for the drift error. The 99.9 percentile of the drift error is switch-specific and increases linearly with duration. Clock drift is inherent to the physical characteristics of the switch chip, so the drift error is switch-specific. The drift error between each pair of switches is fitted with a separate linear function used to predict the drift error across different time durations. OpenOptics users must measure the drift error between each pair of switches before running the synchronization algorithm; they canUse our measurement tools to automate this process. Specification 10 / 25 pages 13 CN 121420263 A

[0076] Synchronization Algorithm

[0077] Algorithm 2 presents a ToR synchronization algorithm according to an embodiment of the present invention:

[0078] Specification 11 / 25 pages 14 CN 121420263 A

[0079]

[0080]

[0081] Figure 6 illustrates the key steps of the synchronization algorithm. In the initial synchronization phase, each ToR is quickly synchronized (once) with coarse accuracy using a custom optical schedule. The fastest way to propagate the master clock from the dominant ToR to the remaining ToRs is to construct a binary tree, as shown in Figure 6. The level t of the binary tree consists of ToRs that have already been synchronized after t time slices. These ToRs are used as intermediaries to synchronize other ToRs in the chain. In the next time slice, each synchronized ToR pushes its clock to the remaining ToR with the minimum drift error. Specification 12 / 25 pages 15 CN 121420263 A The idea is that for each ToR, it is synchronized with similar ToRs, which will further propagate the clock to other ToRs.

[0082] The estimate of how much a synchronized clock has drifted from the master clock is based on the pooled results. The estimation error of a ToR is the sum of its pooled synchronization and drift errors, and changes in the estimation error are tracked at each synchronization step, as shown in Figure 6. The estimation error of ToR0 is always 0 because it has the master clock. At the end of time slice t=1, ToR1 is synchronized with ToR0, so the estimation error of ToR1 is H1+D0,1: it experiences a drift error D0,1 relative to ToR0 during one time slice (inferred from Figure 5b) and a 1-hop synchronization error H1 (derived from Figure 5a). In the next time slice, the estimation error of ToR1 becomes H1+2D0,1 because the drift error has increased by D0,1 after one time slice. The newly synchronized ToR5 inherits the estimation error from its reference clock ToR1, but since R5 has already undergone 2-hop synchronization, we need to change H1 to H2 and add a drift error D1,5 relative to ToR1 for this time slice. This process continues over time for both already synchronized and newly synchronized ToRs.

[0083] The resynchronization phase uses operational optical scheduling and synchronizes ToRs to each other to improve synchronization accuracy. Typically, depending on how many optical uplinks a ToR has and the optical scheduling, each time slice that ToR is connected to many other ToRs. The ToR with the smallest estimation error (if less than its own error) is a candidate for resynchronization.

[0084] Resynchronization is only performed when the current estimation error of the ToR exceeds a predefined threshold T that balances the overhead and accuracy of resynchronization.Resynchronization is then performed. It is determined by the duration of the time slice (and tolerance for asynchrony) and the relative importance of synchronization error versus drift error. If synchronization error is dominant, T can be made larger to reduce the resynchronization phase, such that each ToR is resynchronized directly with the dominant ToR once per cycle of the schedule. However, if drift error is dominant, T can be made smaller to resynchronize each ToR multiple times per cycle via an intermediate reference ToR before its clock drifts.

[0085] The algorithm generates a synchronous optical schedule for initial synchronization, and synchronization steps for both initial synchronization and resynchronization. The synchronization schedule, similar to the operation of the schedule, is enforced by the optical controller as a circuit connection on the optical network structure. Synchronization steps can be implemented via lookup tables on the ToRs to direct synchronization messages in different time slices to specific ports connected to the ToRs to be synchronized with. If there are a large number of synchronization steps, they can be preloaded into the control plane of the programmable ToRs and later injected into the data plane in batches at regular intervals. In addition to these explicit outputs, each DPTP synchronization returns an offset to adjust the local clock. The implementation of time slices utilizing clock offsets will be described below.

[0086] Routing Runtime System

[0087] The system design will now be outlined to demonstrate the feasibility of implementing HOHO routing in practice. Each ToR requires three key functions. First, for an arriving data packet, the next hop (egress port) determined by HOHO routing depends on the arrival time slice of the data packet. Therefore, each ToR switch will need to keep track of the current time slice by synchronizing with the optical network infrastructure time. Second, each ToR needs to implement a routing table that can match the arrival time slice of the data packet with the destination ToR and look up the egress port (next hop) and the transmission slice. Third, the ToR also needs to implement time-scheduled data packet delivery so that each data packet can be delivered precisely within the transmission slice determined by HOHO routing.

[0088] Since HOHO routing requires data packets to be routed precisely based on time slices, there are several challenges to its implementation. First, since the optimal path of HOHO routing may cause data packets to be sent out in a later time slice, a mechanism is needed to temporarily buffer data packets until the exact time slice in the future. Secondly, since DPTP in ToR synchronization provides a clock offset that cannot be directly used for queue management in time scheduling, a method is needed to convert the clock offset into a time trigger that causes packets to be dequeued. Finally, since HOHO routing calculates static paths assuming empty switch queues (page 13 / 25, CN 121420263 A), although in practice queuing delays may cause packets to miss their allocated transmission time slices, a method is needed...Queuing delays need to be measured, and admission control needs to be performed before packets are enqueued to prevent slices from being missed. These challenges are addressed below.

[0089] To enable queue rotation, modern programmable switches (e.g., Intel Tofmo2) support per-packet queue selection and pause / resume of target queues triggered by ingress packets in the data plane

[28] . These features can be used to enqueue packets intended to be sent out in the same time slice into a designated queue that is paused until the start of that sending time slice. After being resumed, the queue remains active for exactly one time slice. The pause / resume of the queue and its duration of activity can be controlled by ingress packets from an on-chip packet generator

[25] . The packet generator reliably sends one packet into the ingress pipeline every configured time interval, which in our case is the duration of the time slice. This design can be built on top of a calendar queue framework

[40] . A calendar queue is a priority queue where each priority is associated with a "calendar day". For future calendar days, packets can be enqueued according to priority (i.e., "rank"). In this example, calendar days are time slices, and each egress port has a set of calendar queues because packets are matched with ports on a per-time-slice basis. Each time slice is sequentially assigned to a queue in each egress port until the queue is exhausted for wrap-around. Queue rotation occurs whenever the system advances to a new time slice, i.e., the current queue is paused and the next queue is resumed, triggered by two consecutive packets from the packet generator.

[0090] The rank of an incoming packet is the difference between the sending time slice and the current time slice, i.e., how far into the future the packet's delivery is scheduled.

[0091] Figure 7 illustrates an example of the routing process. An incoming packet at time slice 2, which initially matches the first entry in the lookup table to obtain the optimal path (1), should be enqueued into queue 0 of egress port 5 (for the current time slice t2), which is full, and then matched with the second entry to obtain the secondary path (2), successfully enqueuing into queue 2 of port 1 (for t4, which is two time slices away). In Figure 7, by looking up the third entry in the table, packets arriving at ToR in time slice t=3 will be scheduled to be transmitted from egress port=1 in t=4. The active queue for the current time slice is queue=0 in port=1, and packets are enqueued into queue=1 to be released after one time slice.

[0092] For time slice enforcement, the packet generator for queue rotation needs to start after the initial time synchronization and is updated with global synchronization with the optical scheduler every time it is resynchronized. This task can utilize clock offsets from DPTP synchronization.The clock offset is stored in an SRAM-based register on the programmable ToR. The kickoff packet generator starts with the minimum packet interval, which is consistently within 10 ns for each test performed by the inventors. The absolute time of the first time slice, which is globally agreed upon, is a parameter that can be set by the control plane program. Each packet from the kickoff packet generator reads the offset value and checks whether the first time slice has arrived by adding the offset to the ingress timestamp as the local data plane clock time. When the start of the first time slice is detected, the P4 program signal notifies the control plane to shut down the kickoff packet generator to save packet generation and processing resources on the switch. Intel Tofino2 allows the packet generator to be configured by the control plane and then started in the data plane later via the triggering of the ingress packet

[28] . A pair of identical rotating packet generators are defined for queue rotation, with a packet interval equal to the duration of the time slice. One such generator is started by the last packet from the kickoff packet generator, which marks the start of the first time slice with a maximum delay of 10 ns (based on the minimum packet interval). Thus, the queue rotation process can be bootstrapping.

[0093] Queue rotation is adjusted only after the clock drift has exceeded a predefined threshold, unlike (the above) T: the resynchronization decision is based on an estimated (aggregated) error, while the offset change in DPTP comes from real-world measurements. The clock offset used for the current queue rotation is stored in another register. Each DPTP packet reads this register to check for offset changes. When the threshold is reached, the DPTP packet restarts the startup packet generator, which now monitors the new clock offset to determine the start of the next time slice derived from the start time of the first time slice, the duration of the time slice, and the number of time slices elapsed. Based on the new clock offset, packets from the startup packet generator that have detected the start of the next time slice are used to update the queue rotation. It simultaneously disables the current rotation packet generator and allows another to follow the updated clock. The P4 program recalculates the current time slice number with the new clock to start a new queue rotation from the correct queue level. Again, the startup packet generator is shut down immediately after use. This process continues, and the two rotating packet generators constantly switch, one in use and one idle, to enable real-time clock adjustment on the data plane.

[0094] (As previously mentioned) queue delays need to be estimated in order to assess whether packets can be successfully delivered within the scheduled time slice. However, they are difficult to predict on commercial switches: before packets enter the queuing system, the actual...Queue information is not accessible in the ingress pipeline.

[0095] To circumvent this limitation, the queuing delay estimation method according to this embodiment of the invention utilizes a register array to track the occupancy of each queue. Packets can be queued into any queue, but only the active queue in the current time slice of each port can discharge traffic. If a packet is queued, the queue occupancy increases, while each update interval (via the packet generator) reduces the occupancy of the active queue by the volume of traffic discharged during that time period (line rate multiplied by the update interval). Therefore, the queuing delay of incoming packets is the current queue occupancy plus the packet size (assuming the packet will be queued) divided by the line transmission rate. The choice of update interval is influenced by a trade-off between estimation accuracy and pipeline processing resource consumption. The inventors found that a 50ns interval is reasonable in practice: the packet generator sends a packet every 50ns or at 20Mpps, which creates only 1.3% pipeline forwarding overhead on a Tofino2 switch with a processing capacity of 1.5Bpps. The evaluation in Figure 8a shows that the estimated queuing delay with a 50ns update interval is only 58ns.

[0096] According to this embodiment of the invention, a recursive lookup is introduced to enhance the static lookup using dynamic packet admission control. For each inbound packet, the admission control compares its estimated queuing delay with the allowed transmission time of the assigned queue, which is the duration of the full time slice of the inactive (paused) queue and the remaining time slice duration of the active (exited) queue. The packet is enqueued only if the admission control determines that the packet can be sent out in the scheduled time slice; otherwise, the packet is re-circulated to re-make a lookup for the next optimal path. The lookup table is implemented as a matching action table on the Tofino2 switch. The matching fields are the packet's arrival time slice, destination ToR, and the re-circulation count that the packet has already undergone. The returned lookup (action) data consists of the egress port and the transmission time slice when the packet should be sent to the next hop.

[0097] Figure 7 illustrates how the lookup table is generated. Assume that the packet arrives in time slice 2, with no packet re-circulation in the initial lookup (e.g., the first entry in the lookup table), corresponding to the optimal path, e.g., path (1). Each re-loop will look for a less desirable path, such as a second entry with one re-loop for the suboptimal path (2). A packet after n re-loops has the same effect as making a new packet arrive after n time slices. This is why the second and third entries in Figure 7 both point to path (2).

[0098] Although packet re-looping is a laborious operation, it prevents slices from being missed and simplifies the design and operation of the HOHO routing algorithm. According to an embodiment of the invention, a limit is set on the re-loop count, and packets that have been re-looped too many times are discarded.

[0099] The evaluation in Table 3 (to be continued below) shows that, in a very short 1μs time slice, less than 0.5% of the production DCN traffic packets experienced recirculation, and packets were recirculated at most once. By eliminating data packet loss caused by slice misses, recirculation prevents packet reordering caused by sporadic packet retransmissions. Packet reordering still occurs, but as will be discussed later, it is rare.

[0100] In addition to the dynamic queuing delay that can be estimated, there are fixed pipeline processing, packet serialization, and on-wire propagation delays in the system. These fixed delays are offset, i.e., by sending traffic earlier to maximize circuit utilization. The offset is the delay from queue rotation to the delivery of the first packet at the destination ToR. The offset is measured using source and destination ToRs virtualized on the Intel Tofmo2 switch, and 100Gbps traffic is sent between them with different packet sizes. As shown in Figure 8b, the minimum delay is 1287 ns. This offset ensures that the packet with the minimum delay can catch the upcoming time slice. The maximum delay is 1324 ns, resulting in a narrow delay range of 34 ns. This indicates that, in the worst case, the packet with the maximum delay arrives 34 ns after the circuit is established, wasting the minimum amount of slice time.

[0101] Other unavoidable system overheads, including errors from ToR synchronization, time slice enforcement, and queuing delay estimation, must be protected by guards between consecutive optical time slices during which data transmission is not permitted. According to Figure 10b, the error of ToR synchronization in the 512-ToR DCN is less than 20 ns, and a large T is set to ensure that each ToR has a worst-case error of less than 40 ns in synchronizing with the dominant ToR. As explained, queue rotation begins within a 10 ns delay after the start of the time slice. In Figure 8a, the queuing delay estimation with a 50 ns update interval shows a tail error of 58 ns. A guard band is introduced for reconfiguring optical circuitry [13, 20, 33, 36], which can overlap with slice updates in the system. The overall system overhead of 126 ns is easily covered by today's mainstream microsecond-level optical switching technologies, therefore OpenOptics does not require an additional guard band.

[0102] Host System

[0103] As already explained, the host normally sends mouse streams to be processed by the ToR, but releases elephant streams when direct circuitry to the destination ToR is available. For versatility, the host system is designed to be transparent to TCP applications.

[0104] A recent work on slow-switched optical DCNs utilizes signaling messages from the ToR to notify the host of an impending arrival.The optical circuit

[17] , but it is unknown whether this approach fits well into fast-switching optical DCNs, given the signal transmission delay and other overhead from the host stack. The feasibility of circuit notification via signaling was tested using a small measurement study. Circuit signaling was implemented in the kernel and libvma kernel bypass versions, where synchronized ToR signals notify the start of each time slice and the circuit connection to the host to which it is connected. To minimize the number of signal messages, the effective time slice duration (considering guard bands) was programmed to the hosts, allowing them to infer the end of the time slice using the CPU clock. Signals travel through a high-priority path: from a high-priority queue on the ToR to a dedicated ring in libvma or a dedicated RX queue in the kernel module. Moving average techniques were used to filter signal jitter: each host tracked the signal arrival interval and used the average value to replace abnormally early and late signals that deviated from the average by a considerable amount. Signaling packets were continuously sent from the ToR to the connected hosts at 100 μs intervals. Upon each signal arrival, the host replied with a 1500B packet to simulate the first data packet.

[0105] Figure 9a shows the original signal distance before moving average correction. For 95% of the data points, the libvma implementation has a consistent signal distance within ±0.25 μs of the expected 100 μs.

[0106] Figure 9b plots the time duration from transmitting a signal to receiving a response as measured on the ToR. With the libvma implementation, such turnaround times mostly vary within a small range of 0.75 μs. Here, a moving average is applied, so the variance of the turnaround time is dominated by the variance of the delay on the return path.

[0107] These results demonstrate the usability of the signaling method. After offsetting the minimum turnaround delay in Figure 8b, the sub-microsecond variance is shorter than the microsecond-level switching delay of most fast switching techniques [27, 32, 33, 36], so the host system of the present invention can work well on time slices longer than tens of microseconds. Instructions 16 / 25 pages 19 CN 121420263 A

[0108] To achieve flow pausing, flow aging

[12] is employed to distinguish between mouse and elephant flows in the absence of flow size information, which is essentially transparent to the application. A flow is considered a mouse flow until the accumulated flow volume exceeds a step threshold and is then downgraded to an elephant flow. The step threshold is set to 10KB based on the flow size distribution from DCN flow measurements [11, 34, 39].

[0109] Kernel module with libvma. As a strawman solution, we implemented flow pausing in a kernel module on the netfilter framework [8]. We allocated dynamic storage for each destination ToR.Outgoing packets are stored in their corresponding buffers until they are released to the destination ToR by the incoming circuitry. The inventors also implemented a libvma[5] kernel bypass alternative and compared their performance. Libvma links the TCP socket to the user-space IwIP TCP stack, thus requiring zero changes to the application. We maintain normal socket behavior by allowing the application to write to the segment queue normally. New segments can be written whenever they are appropriate, and send attempts are postponed when the queue is full. The inventors compromised the TCP implementation to prevent outgoing segments from being sent when the stream is paused. The segment queue is flushed after a stream resumption signal is received until the send window is exhausted. In application experiments conducted by the inventors, the throughput of the libvma implementation was proportional to the throughput of the vanilla libvma (without pauses) as a percentage of the active send time slice.

[0110] As can be observed from Figure 9, libvma has much higher signaling consistency (Figure 9a) and lower and more stable signal response turnaround latency (Figure 9b). In addition, libvma has a cleaner design. Implementing the pause mechanism directly within the user-space TCP stack ensures that the TCP state machine matches the actual network stack state, unlike kernel modules which operate at different layers. Netfilter pauses packets that TCP considers to have been sent, unaware of the sending window. The inventors used libvma for the final host system implementation, but still made the kernel module exposed to users of NICs without libvma capabilities.

[0111] Elephant streams to the same destination ToR fairly share circuitry. Regarding host offset and margin, signals should be sent in advance to offset the delay between the ToR and the host. A minimum turnaround delay of 3 μs from Figure 9b is used as the offset value, ensuring that the earliest arriving packets catch the start of the upcoming time slice. Packets with larger delays can be sent out of the slice, i.e., sent to the ToR after the time slice ends. The ToR system will treat them as mouse streams and handle them using HOHO routing, thus expecting no packet loss. 99.94% of the delay measurements fell within 1 μs, so setting a margin of 1 μs before the end of the time slice can avoid the burden on the ToR system regarding out-of-slice packets.

[0112] Evaluation

[0113] In this section, we evaluate the performance of the OpenOptics framework through a micro-benchmark study. Regarding the optical topology and scheduling in the experimental setup, Opera provides the most efficient topology routing cooperative design to date for fast-switching optical DCNs

[32] . The inventors simulated an Opera network with 108 ToRs, six 100 Gbps optical uplinks per ToR, andEach ToR has six hosts, also at 100Gbps. Opera optical scheduling was used, but native Opera routes were replaced with HOHO routes to integrate Opera into the OpenOptics framework. ToR, hosts and optical architecture. The entire Opera network could not fit in the test bench, so representative sender ToRs, representative receiver ToRs and simulated optical network architectures were each implemented on three converged switches. Three servers (each with a Mellanox ConnectX-5 100Gbps dual-port NIC) were connected to the sender ToR to work as six separate hosts.

[0114] The inventors ran Facebook traces

[39] collected from databases, Hadoop and Web services, which are now widely used in DCN research [19, 24, 30, 34, 47]. The traces were scaled to the current topology size and replayed on the server to feed incoming traffic to the sender ToR. Indirect traffic that the sender ToR should forward on behalf of other ToRs was also injected from the server. Traffic sent to unrealized ToRs (other than the sender and receiver ToRs) is discarded by the simulated optical structure (see page 17 / 25 of the specification, 20 CN 121420263 A). For worst-case analysis of the ToRs and host system, the sender ToR and server are made the heaviest speakers in the trace. The inventors also designed resynthesized traffic, such as line rate traffic and burst traffic, to further stress test certain aspects of the system. Measurement of synchronization accuracy requires the observed ToRs to use the same physical clock, so the setup in Figure 10a is designed to simulate a large-scale DCN. Three converged switches are reused and virtualized as ten logical ToRs. They are chained together in the direction of clock propagation to represent the rightmost node of the synchronization tree (Figure 6) in the initial synchronization phase of the synchronization algorithm. These nodes are the worst synchronized at each tree level. In this way, if the tree is fully populated, they act as a subset of the relatively inaccurately synchronized ToRs in the 512-ToR DCN. After initial synchronization in chain order, ToRs resynchronize with each other after the second phase of the algorithm. The round-robin operation of optical scheduling between ToRs is assumed to be a reasonable subset of abundant connections in a large DCN. Synchronization errors of ToR3, ToR6, and ToR9, which are located on the same switch as the dominant ToR ToR0, were measured. Since there is no drift for ToRs on the same switch, drift errors were added to them from the aggregated numbers of the other two switches.

[0115] Figure 10b compares the synchronization error of the algorithm of the present invention with a scarecrow solution that synchronizes individual ToRs with the dominant ToR once per optical cycle. Time slice 0 shows the results of initial synchronization. The tree-building algorithm of the present invention is effective: withCompared to the 44ns error in the Scarecrow solution that had to wait for the direct connection to the master ToR, the synchronization error after a short initial phase was only 10ns. The resynchronization algorithm of the present invention keeps the error below 20ns, compared to the Scarecrow solution drifting up to 58ns per optical cycle. The threshold T = 15ns is set based on the estimated error from the aggregated results, i.e., resynchronization is performed whenever the estimated error exceeds 15ns. Considering the unavoidable one-hop synchronization error of up to 15ns in Figure 5a, this very close actual error of 20ns verifies the accuracy of the error aggregation and estimation scheme of the present invention. When a large T is set, the resynchronization phase step regresses to the Scarecrow solution.

[0116] Regarding the accuracy of queuing delay estimation, Figure 8a plots the queuing delay estimation error under stress testing, where we combine line rate flow and burst flow to periodically fill and remove queues. The error is the difference between our estimated queuing delay and the measured basic true queuing delay. The results validate the method of the present invention, as the estimation accuracy is determined by the interval of the update queue occupying the registers: the estimation accuracy improves with more frequent register updates, and the 99.9 percentile estimation error across the curve falls within the accuracy margin of the update interval.

[0117] Table 2 below shows the relative resource utilization of the ToR with OpenOptics enabled.

[0118]

[0119] As can be seen from Table 2, the resource utilization in the fairly large 108-ToR DCN is a small percentage relative to switch.p4 (the baseline P4 program that implements the core L2 / L3 switching function). The low utilization of SRAM, VLIW Actions, and TCAM indicates efficient implementation of registers and lookup tables in OpenOptics. The Stateful ALU and Ternary Xbar have higher utilization due to arithmetic calculations and branch operations used for packet admission control. All resources are less than 15% of the Intel Tofmo2 switch capacity (pages 18 / 25, 21 CN 121420263 A), leaving enough room for OpenOptics to scale to even larger DCNs.

[0120] Figure 11a shows the total queue occupancy per port, i.e., the sum of individual queue occupancy across calendar queues. For heavy Hadoop traffic, the tail occupancy is as low as 5.704KB. For the high-end 128-port ToR where half of the ports are used as optical uplinks, we only need a total buffer size of 365KB. Figure 11b shows the number of consecutive queues used per port, or the minimum number of queues required by the calendar queues. At any given time, the calendar queues occupy a maximum of 6 queues per port. Our queue usage is significantly lower than the capacity limitations of commercial switch ASICs

[38] , and this justifies the queue wheel of the present invention.The effectiveness of packet recirculation in timely traffic clearing is demonstrated.

[0121] Regarding packet recirculation, it has been explained that recirculation will cause packets to miss scheduled transmission time slices. Table 3 shows the rarity of packet recirculation under production DCN traffic.

[0122]

[0123] At the extremely low time slice duration of 1 μs, less than 0.5% of packets are recirculated, which is almost unattainable considering the microsecond-level switching latency of mainstream optical technologies. For a 10 μs time slice, the number is further reduced to less than 0.05% because it is easier to adapt packets to longer time slices. Hadoop traffic experienced slightly more recirculation because it mainly contains large packets for bulk data migration. Packets are recirculated at most once. These observations indicate that packet recirculation adds very little overhead to the ToR system of the present invention.

[0124] Regarding packet loss and circuit utilization, OpenOptics prevents packet loss through a combined effort of packet admission control, delay offset, and guard band. No packet loss has been observed from trace and even line rate traffic, which proves the effectiveness of these mechanisms. The results are verified by measuring when the last packet of each time slice is received by the destination ToR, as there is a trade-off between packet loss and circuit utilization. The queuing delay estimation of the present invention has an error (Figure 8a). If the admission control conservatively queues packets, packet loss can be avoided, but at the cost of underutilization of the time slice. As shown in Figure 12, at the end of the time slice, at most one packet is rejected. In the worst case of rejecting 1500B packets, the leftmost point in Figure 12 results in an expected waste of 145ns × 50% = 72.5ns, which translates to 99.93% utilization on a 100μs time slice and 99.28% utilization on a 10μs time slice.

[0125] Furthermore, OpenOptics theoretically has packet reordering due to packet recirculation and the resulting dynamics on different paths to the same destination ToR (on different time slices). The "forwarding ToR" is virtualized on the receiver ToR only to forward traffic from the sender ToR to the receiver ToR. With the new path having an added hop and specification page 19 / 25 22 CN 121420263 A, the original direct path is set up for use in the alternative time slice. No packet reordering was observed from the operation of a large number of traces, which is consistent with the rarity of packet recirculation.

[0126] Through flow pausing, the host system of the present invention only releases long flows in the scheduled time slice, but due to host-ToR latency variations, some packets may still arrive outside the expected time slice. Host offset and margin mitigateThis issue was addressed, and zero out-of-slice packets were observed when running traffic traces. As a stress test, iperf was run at full speed on the host. Very few out-of-slice packets were observed, and the margin almost eliminated them. Out-of-slice packets were not lost, but stored on the ToR and disposed of as if from a mouse stream.

[0127] Case Study

[0128] As a case study, the inventors also implemented three fast-switching optical DCN architectures on top of OpenOptics. Real DCN applications ran on these optical architectures and demonstrated end-to-end application performance.

[0129] For the cluster setup, three converged Tofino2 switches (SW1-SW3) were used to implement eight logical ToRs (four on SW1 and four on SW2) and a simulated optical structure (on SW3). Each logical ToR was connected to the simulated optical structure via four 100Gbps optical uplinks. Four servers, each with a Mellanox ConnectX-5 100Gbps dual-port NIC, were connected to the eight ToRs (each via a 100Gbps NIC port) to operate as eight separate hosts.

[0130] Opera

[32] , RotorNet

[33] and Mordia

[36] are implemented on OpenOptics, and the "pass-through" rule is set on a simulated optical structure to have a baseline electrical DCN. Opera supports multi-hop routing of mouse flows (through intermediate ToRs) and ensures the existence of an optical path (usually multi-hop) between any ToR pairs at any given time. It limits the queue size on the ToR to 8 maximum packets to estimate a fixed queuing delay and requires the time slice to be longer than the maximum RTT in the network under the estimated queuing delay. Its optical scheduling for fine-grained routing design has rigid requirements on network structure and time slice duration. This cluster topology is suitable for a minimal Opera network, and as required by the Opera size, we set the time slice duration to 50 μs for all architectures. Opera routes elephant flows from source to destination ToRs via direct circuits to save bandwidth. RotorNet uses round-robin optical scheduling and routes traffic on direct circuits. Mordia proposes on-demand optical scheduling based on real-time traffic generation. This is difficult to implement in practice, so we approximate on-demand scheduling with efficient scheduling for traffic patterns known in advance in the application. These architectures run HOHO routing on OpenOptics. Queue rotation was disabled for comparison with their native routing schemes, and a lookup table for each specific architecture was loaded.

[0131] Latency-sensitive and throughput-intensive applications were run to verify the performance of mouse and elephant streams on OpenOptics. The inventors used a Memcached [6] key-value store for the latency-sensitive application, with two Memcached servers andSix Memslap[7] benchmark clients each ran on a host, and the client wrote 4.2KB of data to the server in each SET operation. The Gloo collective communication library[2] was selected for throughput-intensive applications, and the ring allreduce ran on eight hosts, with data sizes varying from 800KB to 20MB.

[0132] As can be seen from Figure 13a, the multi-hop “always-on” path in Opera effectively reduced the flow completion time (FCT), while RotorNet suffered long delays while waiting for direct circuits to the destination ToR. As expected, the “on-demand” scheduling approximation for Mordia was superior to the round-robin scheduling in RotorNet. Since each server talks to six clients, the optimal scheduling is to iterate the four optical uplinks on the server ToR through the six client ToRs, which is more efficient than global round-robin on the seven other client ToRs. HOHO routing improves routing for each architecture by providing the fastest path that can span multiple hops.

[0133] In particular, HOHO's tail FCT is 69% lower than RotorNet's tail FCT and 85% lower than Mordia's tail FCT because HOHO enables "uninterrupted" paths to unblock packets waiting for direct paths. It is 16% lower than Opera's tail FCT because HOHO achieves a shorter average path length by allowing packets to pause at intermediate ToRs. Furthermore, the fixed queuing delay in Opera conservatively allows packets only to time slices, and our accurate queuing delay estimate eliminates false negatives. Therefore, Opera's tail FCT using HOHO is only 30% longer than the tail FCT of the electrical DCN baseline. Note that our electrical DCN setting sets a highly idealized, loose upper bound where ToRs are directly connected, whereas in practice, multi-layer Clos networks have longer path lengths, more congestion, and higher FCTs.

[0134] In Figure 13b, the median allreduce FCT of Opera and RotorNet is similar because the elephant stream can only pass through direct circuits. However, the direct circuits in Opera scheduling are more diverse than pure round-robin, resulting in a slightly lower FCT. In this case, Mordia has enough ToR uplinks to form a static eight-ToR ring, so the performance is the same as the electric DCN baseline. HOHO routing only handles mouse streams, so it is irrelevant here. From these results, we demonstrate the versatility and simplicity of OpenOptics for implementing different fast-switching DCN architectures with comparable performance. OpenOptics achieves similar performance to electric DCN: for mouse streams utilizing HOHO, and for streams where...The required flow of the circuitry is determined. Its correctness is also verified by the expected behavior of these architectures.

[0135] References

[0136] [1] [n .d .] . All‑reduce collective communication pattern, https: / / mpitutorial.com / tutorials / mpi‑reduce‑and‑allreduce / . ([nd]) .

[0137] [2] [n .d .] .Gloo .https: / / github.com / facebookincubator / gloo . ([n. d.]) .

[0138] [3] [n . d .] . How to achieve low latency with lOGbps Ethernet . https: / / blog.cloudflare.com / how‑to‑achieve‑low‑latency / . ([nd]) .

[0139] [4] [nd]. linuxptp. https: / / linuxptp.sourceforge.net / . ([nd]) .

[0140] [5] [n . d .] . Mellanox Messaging Accelerator . https: / / github .com / Mellanox / libvma / blob / master / README. ([nd]) .

[0141] [6] [nd]. Memchached. https: / / memcached.org / . ([nd]) .

[0142] [7] [n. .org / bin / memslap.html. ([nd]) .

[0143] [8] [nd]. Netfilter. https: / / www.netfilter.org / . ([nd]) .

[0144] [9] [n . d .] . .] . Scaling in the LinuxNetworking Stack. https: / / www.kemel.org / doc / Documentation / networking / scaling.txt. ([n. d.]).

[0146]

[11] Berk Atikoglu, Yuehai Xu, Eitan Frachtenberg, Song Jiang, and Mike Paleczny. 2012. Workload analysis of a large‑scale key‑value store. In Proceedings of the 12th ACM SIGMETRICS / PERFORMANCE joint international conference on Measurement and Modeling of ComputerSystems. 53‑64.

[0147]

[12] Wei Bai, Li Chen, Kai Chen, Dongsu Han, Chen Tian, and Hao Wang. 2017. PIAS: Practical information‑agnostic flow scheduling for commodity data centers. IEEE / ACM Transactions on Networking 25, 4 (2017), 1954—1967.

[0148]

[13] Hitesh Ballani, Paolo Costa, Raphael Behrendt, Daniel Cletheroe, Istvan Haller, Krzysztof Jozwik, Fotini Karinou, Sophie Lange, Kai Shi, Benn Thomsen, et al. 2020. Sirius: A flat datacenter network with nanosecond specification 21 / 25 pages 24 CN 121420263 A optical switching. In Proceedings of the Annual conference of the ACM Special Interest Group on Data Communication on theapplications , technologies , architectures, and protocols for computer communication. 782‑797.

[0149]

[14] Kai Chen , Ankit Singla , Atul Singh , Kishore Ramachandran , Lei Xu, Yueping Zhang , Xitao Wen, and Yan Chen. 2013. OSA: An optical switching architecture for data center networks with unprecedented flexibility. IEEE / ACM Transactions on Networking 22, 2 (2013) , 498‑511.

[0150]

[15] Kai Chen , Xitao Wen , Xingyu Ma , Yan Chen , Yong Xia , Chengchen Hu , and Qunfeng Dong . 2015 . WaveCube: A scalable , fault‑tolerant , high performance optical data center architecture . In 2015 IEEE Conference on Computer Communications (INFOCOM) . IEEE, 1903‑1911.

[0151]

[16] Li Chen , Kai Chen , Zhonghua Zhu , Minlan Yu , George Porter , Chunming Qiao, and Shan Zhong. 2017. Enabling {Wide‑Spread} Communications on Optical Fabric with {MegaSwitch} . In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) . 577‑593.

[0152]

[17] Shawn Shuoshuo Chen, Weiyang Wang,Christopher Canel, Srinivasan Seshan , Alex C Snoeren , and Peter Steenkiste . 2022. Time‑division TCP for reconfigurable data center networks. In Proceedings of the ACM SIGCOMM 2022 Conference. 19‑35.

[0153]

[18] Dah‑Ming Chiu and Raj Jain. 1989. Analysis of the increase and decrease algorithms for congestion avoidance in computer networks. Computer Networks and ISDN systems 17, 1 (1989) , 1‑14.

[0154]

[19] Inho Cho , Keon Jang , and Dongsu Han . 2017 . Credit‑scheduled delay bounded congestion control for datacenters . In Proceedings of the Conference of the ACM Special Interest Group on Data Communication.239‑252.

[0155]

[20] Nathan Farrington , George Porter , Sivasankar Radhakrishnan , Hamid Hajabdolali Bazzaz, Vikram Subramanya, Yeshaiahu Fainman, George Papen, and Amin Vahdat . 2010 . Helios: a hybrid electrical / optical switch architecture for modular data centers. In Proceedings of the ACM SIGCOMM 2010 Conference. 339‑350.

[0156]

[21] Alex Forencich, Alex C Snoeren, GeorgePorter, and George Papen. 2020. Corundum: An open‑source 100‑gbps nic. In 2020 IEEE 28th Annual International Symposium on Field‑Programmable Custom Computing Machines (FCCM). IEEE, 38‑46.

[0157]

[22] Yilong Geng, Shiyu Liu, Zi Yin, Ashish Naik, Balaji Prabhakar, Mendel Rosenblum, and Amin Vahdat. 2018. Exploiting a natural network effect for scalable, fine‑grained clock synchronization. In 15th {USENIX} Symposium on Networked Systems Design and Implementation ({NSDI} 18). 81‑94.

[0158]

[23] Monia Ghobadi, Ratul Mahajan, Amar Phanishayee, Nikhil Devanur, 说明书 22 / 25 页 25 CN 121420263 A Janardhan Kulkami, Gireeja Ranade, Pierre‑Alexandre Blanche, Houman Rastegarfar, Madeleine Glick, and Daniel Kilper. 2016. Projector: Agile reconfigurable data center interconnect. In Proceedings of the 2016 ACM SIGCOMM Conference. 216‑229.

[0159]

[24] Shuihai Hu, Wei Bai, Gaoxiong Zeng, Zilong Wang, Baochen Qiao, Kai Chen, Kun Tan, and Yi Wang. 2020. Aeolus: A building block forproactive transport in datacenters. In Proceedings of the Annual conference of the ACM Special Interest Group on Data Communication on the applications , technologies, architectures, and protocols for computer communication. 422‑ 434.

[0160]

[25] Raj Joshi , Ben Leong , and Mun Choon Chan . 2019 . Timertasks: Towards time‑driven execution in programmable dataplanes. In Proceedings of the ACM SIGCOMM 2019 Conference Posters and Demos. 69‑71.

[0161]

[26] Pravein Govindan Kannan , Raj Joshi , and Mun Choon Chan . 2019. Precise time‑synchronization in the data‑plane using programmable switching asics. In Proceedings of the 2019 ACM Symposium on SDN Research. 8‑20.

[0162]

[27] Mehrdad Khani , Manya Ghobadi , Mohammad Alizadeh , Ziyi Zhu , Madeleine Glick , Keren Bergman , Amin Vahdat , Benjamin Klenk , and Eiman Ebrahimi . 2021 . SiP‑ML: high‑bandwidth optical network interconnects for machine learning training . In Proceedings of the 2021 ACM SIGCOMM 2021 Conference. 657‑675.

[0163]

[28] Jeongkeun Lee . 2020 . Advanced congestion & flow control with programmable switches. In P4 Expert Roundtable Series. https: / / bit.lyZ3J8x7fw

[0164]

[29] Ki Suh Lee, Han Wang, Vishal Shrivastav, and Hakim Weatherspoon. 2016. Globally synchronized time via datacenter networks. In Proceedings of the 2016 ACM SIGCOMM. Conference. 454‑467.

[0165]

[30] Yuliang Li, Rui Miao, Hongqiang Harry Liu, Yan Zhuang, Fei Feng, Lingbo Tang , Zheng Cao , Ming Zhang , Frank Kelly , Mohammad Alizadeh, et al. 2019 . HPCC: High precision congestion control . In Proceedings of the ACM Special Interest Group on Data Communication. 44‑58.

[0166]

[31] Yunpeng James Liu, Peter Xiang Gao, Bernard Wong, and Srinivasan Keshav. 2014. Quartz: a new design element for low‑latency DCNs. ACM SIGCOMM Computer Communication Review 44, 4 (2014) , 283‑294.

[0167]

[32] William M Mellette, Rajdeep Das, Yibo Guo, Rob McGuinness, Alex C Snoeren , and George Porter . 2020 . Expanding across time to deliver bandwidth efficiency以及低延迟。在第17届USENIX网络系统设计与实现研讨会(NSDI 20)上。第1 - 18页。

[0168]

[33] 威廉·M·梅利特、罗布·麦吉尼斯、阿尔琼·罗伊、亚历克斯·福伦西奇、乔治·帕彭、亚历克斯·C·斯诺伦和乔治·波特。2017年。Rotomet:一种可扩展的、低复杂度的光学数据中心网络。在ACM数据通信特别兴趣小组会议论文集。第267 - 280页。

[0169]

[34] 贝纳姆·蒙塔泽里、李一龙、穆罕默德·阿里扎德和约翰·奥斯特霍特。2018年。Homa:一种使用网络优先级的接收方驱动的低延迟传输协议。在2018年ACM数据通信特别兴趣小组会议论文集。第221 - 235页。

[0170]

[35] 马修·K·穆克吉、克里斯托弗·卡内尔、魏洋·王、金大赫、斯里尼瓦桑·塞山和亚历克斯·C·斯诺伦。2020年。为可重构数据中心网络适配TCP。在NSDI。第651 - 666页。

[0171]

[36] 乔治·波特、理查德·斯特朗、内森·法林顿、亚历克斯·福伦西奇,Pang Chen‑Sun , Taj ana Rosing , Yeshaiahu Fainman , George Papen , and Amin Vahdat . 2013 . Integrating microsecond circuit switching into the data center. ACM SIGCOMM Computer Communication Review 43, 4 (2013) , 447‑458.

[0172]

[37] Leon Poutievski , Omid Mashayekhi , Joon Ong , Arjun Singh , Mukarram Tariq, Rui Wang, Jianan Zhang, Virginia Beauregard , Patrick Conner, Steve Gribble , et al . 2022 . Jupiter evolving: transforming google’s datacenter network via optical circuit switches and software‑defined networking. In Proceedings of the ACM SIGCOMM 2022 Conference.66‑85.

[0173]

[38] Ting Qu , Raj Joshi , Mun Choon Chan , Ben Leong , Deke Guo , and Zhong Liu. 2019. SQR: In‑network packet loss recovery from link failures for highly reliable datacenter networks. In Proceedings of ICNP.

[0174]

[39] Arjun Roy, Hongyi Zeng, Jasmeet Bagga, George Porter, and Alex C Snoeren . 2015 . Inside the social network’s (datacenter) network . In Proceedings of the 2015 ACM Conference on Special数据通信兴趣小组。123 - 137。

[0175]

[40] Naveen Kr Sharma、Chenxingyu Zhao、Ming Liu、Pravein G Kannan、Changhoon Kim、Arvind Krishnamurthy和Anirudh Sivaraman。2020年。用于高速分组调度的可编程日历队列。发表于NSDI会议论文集。

[0176]

[41] Arjun Singhvi、Aditya Akella、Dan Gibson、Thomas F Wenisch、Monica Wong - Chan、Sean Clark、Milo MK Martin、Moray McLaren、Prashant Chandra、Rob Cauble等人。2020年。Irma:重新构想多租户数据中心的远程内存访问。发表于美国计算机协会数据通信特别兴趣小组年度会议关于计算机通信的应用、技术、架构和协议的论文集。708 - 721。

[0177]

[42] Guohui Wang、David G Andersen、Michael Kaminsky、Konstantina Papagiannaki、TS Eugene Ng、Michael Kozuch和Michael Ryan。2010年。c - Through:数据中心的兼职光学。发表于ACM SIGCOMM会议论文集 说明书24 / 25页 27 CN 121420263 A2010 Conference. 327‑338.

[0178]

[43] Weiyang Wang , Moein Khazraee , Zhizhen Zhong , Zhijao Jia , Dheevatsa Mudigere , Ying Zhang , Anthony Kewitsch, and Manya Ghobadi. 2022. TopoOpt: Optimizing the Network Topology for Distributed DNN Training. arXiv preprint arXiv:2202.00433 (2022) .

[0179]

[44] Dingming Wu , Yiting Xia , Xiaoye Steven Sun , Xin Sunny Huang , Simbarashe Dzi‑namarira , and TS Eugene Ng . 2018 . Masking failures from application performance in data center networks with shareable backup. In Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communication. 176‑190.

[0180]

[45] Yiting Xia , Mike Schlansker, TS Eugene Ng , and Jean Tourrilhes. 2015. Enabling Topological Flexibility for Data Centers Using OmniSwitch. In HotCloud.

[0181]

[46] Yiting Xia , Xiaoye Steven Sun, Simbarashe Dzinamarira , Dingming Wu , Xin Sunny Huang , and TS Eugene Ng . 2017 . A tale of two topologies: Exploring convertible data center network architectures withflat-tree. In Proceedings of the Conference of the ACM Special Interest Group on Data Communication. 295-308.

[0182]

[47] Qiao Zhang, Vincent Liu, Hongyi Zeng, and Arvind Krishnamurthy. 2017. High-resolution measurement of data center microbursts. In Proceedings of the 2017 Internet Measurement Conference. 78-85. Instruction manual 25 / 25 pages 28 CN 121420263 A Figure 1 Figure 2 Instruction manual Figure 1 / 9 pages 29 CN 121420263 A Figure 3 Figure 4 Figure 5a Instruction manual Figure 2 / 9 pages 30 CN 121420263 A Figure 5b Figure 6 Instruction manual Figure 3 / 9 pages 31 CN 121420263 A Figure 7 Figure 8a Figure 8b Instruction manual Figure 4 / 9 pages 32 CN 121420263 A Figure 9a Figure 9b Instruction manual figures 5 / 9, page 33, CN 121420263 A, Figure 10a, Figure 10b; Instruction manual figures 6 / 9, page 34, CN 121420263 A, Figure 11a, Figure 11b; Instruction manual figures 7 / 9, page 35, CN 121420263 A, Figure 12, Figure 13a; Instruction manual figures 8 / 9, page 36, CN 121420263 A, Figure 13b; Instruction manual figures 9 / 9, page 37, CN 121420263 A< / n>

Claims

1. A method for transmitting data packets in a data center network (DCN), the data center including a plurality of host servers, a plurality of top-of-rack (ToR) switches connected to the host servers, and an optical network structure connected to the plurality of ToR switches, wherein the optical network structure defines, according to a given scheduling operation, which ToR switch pairs are connected via dedicated optical circuits established for the time slice by an optical controller of the optical network structure for each time slice in a time-slice sequence, wherein the ToR switches: - Synchronize with another top switch; - Receive data packets at the ingress port; as well as - Send the data packet to the exit port. Its characteristic is that synchronization includes: Within the band, a synchronization message is sent to the other top switch via the circuitry of the optical network structure.

2. The method of claim 1, wherein the synchronization message is sent based on the given schedule.

3. The method of claim 2, wherein the given schedule is an initial synchronization schedule, the initial synchronization schedule being used to synchronize with coarse accuracy before the data center network is put into operation.

4. The method of claim 3, wherein the given schedule is a schedule of operations after the initial synchronization has been terminated.

5. The method of claim 4, wherein synchronization includes resynchronizing the top-of-rack switch.

6. The method of claim 5, wherein resynchronization is performed when the synchronized clock of the top-of-rack switch has drifted from the master clock by a current amount exceeding a predefined threshold T.

7. The method of claim 6, wherein the drift amount corresponds to the sum of the synchronization error and the drift error of the top-of-rack switch.

8. The method of claim 7, wherein the synchronization error and the drift error are estimated based on empirical aggregation or statistics of the synchronization error and the drift error of the top-of-rack switch.

9. The method of claim 8, wherein the predefined threshold T is determined based on the overhead and accuracy of resynchronization.

10. The method of claim 9, wherein the predefined threshold T is further determined based on the relative importance of the synchronization error and the drift error.

11. The method of claim 10, wherein, If the synchronization error is dominant, the T value is increased to reduce the resynchronization phase, such that the top-of-rack switch is resynchronized directly with the dominant top-of-rack switch once in each cycle of the scheduling.

12. The method of claim 11, wherein, If the drift error is dominant, then T is made smaller so that each ToR is resynchronized multiple times in the cycle via an intermediate reference ToR before the clock of the ToR drifts.

13. The method of claim 12, wherein the ToR having the minimum drift is selected as the other ToR for resynchronization.

14. The method of claim 13, wherein time slices and ports connected to the other top-mount switch are stored in a lookup table on the top-mount switch, the lookup table being used to direct synchronization messages in different time slices to specific ports connected to the top-mount switch to be synchronized.

15. The method of claim 14, wherein the lookup table is preloaded into the control plane of the top-of-rack switch.

16. The method of claim 15, wherein batches of the lookup table are periodically injected into the data plane of the top-of-rack switch.

17. The method of claim 16, wherein the synchronous return is used to adjust the offset of the local clock of the top-of-rack switch.

18. The method of claim 17, wherein the synchronization message is a DPTP message.

19. A top-of-rack switch, wherein the top-of-rack switch implements the method steps of any one of claims 1 to 18.