Techniques for Exchanging Network Traffic in a Data Center

By adopting multi-mode optical switching infrastructure in the data center, the problem of limited practicality of data centers for specific workload types is solved, and multi-protocol support with high bandwidth and low latency and dynamic resource allocation is achieved.

CN115695337BActive Publication Date: 2025-06-27INTEL CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211194590.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-12-30
Filing Date
2017-06-21
Publication Date
2025-06-27
Estimated Expiration
2037-06-21

AI Technical Summary

Technical Problem

A typical data center has limited utility for specific workload types because network components are not equipped to manage high-performance computing network services and other types of network services.

Method used

Adopting a multi-mode optical switching infrastructure, coupled to the switch via fiber, provides high bandwidth and low latency connectivity, supporting a variety of network protocols, including Ethernet and high-performance computing protocols.

Benefits of technology

Support for multiple network protocols is realized, resource utilization efficiency of data centers is improved, allowing pooling resources such as accelerators, memory and storage to dynamically redistribute resources to adapt to different types of workloads.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115695337B_ABST
    Figure CN115695337B_ABST
Patent Text Reader

Abstract

Techniques for switching network traffic include a network switch. The network switch includes one or more processors and communication circuitry coupled to the one or more processors. The communication circuitry is capable of switching network traffic for multiple link layer protocols. Additionally, the network switch includes one or more memory devices storing instructions that, when executed, cause the network switch to receive, via an optical connection, network traffic to be forwarded using the communication circuitry and to determine the link layer protocol of the received network traffic. The instructions further cause the network switch to forward the network traffic based on the determined link layer protocol. Other embodiments are also described and claimed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Utility Patent Application Ser. No. 15 / 395,203, filed Dec. 30, 2016, titled "TECHNOLOGIES FOR SWITCHING NETWORK TRAFFIC IN A DATA CENTER", and claims priority to the following U.S. Provisional Patent Applications: U.S. Provisional Patent Application No. 62 / 365,969, filed Jul. 22, 2016; U.S. Provisional Patent Application No. 62 / 376,859, filed Aug. 18, 2016; and U.S. Provisional Patent Application No. 62 / 427,268, filed Nov. 29, 2016. Background of the Invention

[0003] In a typical data center that provides computing services such as cloud services, workloads can be assigned to multiple computing devices to provide the requested services to customers. Given the latency and bandwidth limitations of twisted - pair copper cables and corresponding networking components (e.g., switches) in such data centers, the physical hardware resources that can be utilized to execute any given workload (including processors, volatile and non - volatile memories, accelerator devices (e.g., coprocessors, field - programmable gate arrays (FPGAs), digital signal processors (DSPs), application - specific integrated circuits (ASICs), etc.)) and data storage devices are typically contained locally within each computing device rather than being spread throughout the data center. As such, depending on the type of workload assigned (e.g., processor - intensive but light on memory usage, memory - intensive but light on processor usage, etc.), the data center can contain many unused physical hardware resources and still be unable to take on additional work (without overloading the computing devices).

[0004] In addition, some typical data centers are designed to operate as high - performance computing (HPC) clusters, using a proprietary network protocol (e.g., Intel OmniPath) to coordinate the communication and processing of workloads, while other data centers are designed to use other communication protocols (such as Ethernet) for communication. The networking components in typical data centers are not equipped to manage both HPC network traffic and other types of network traffic, thus limiting their utility for specific workload types. Brief Description of the Drawings

[0005] The concepts described herein are illustrated by way of example and not limitation in the accompanying drawings. For the sake of brevity and clarity of illustration, the elements illustrated in the drawings are not necessarily drawn to scale. Where considered appropriate, reference numerals have been repeated among the drawings to indicate corresponding or analogous elements.

[0006] Figure 1 is a diagram of a conceptual overview of a data center according to various embodiments, in which one or more of the techniques described herein may be implemented;

[0007] Figure 2 is Figure 1 a diagram of an example embodiment of a logical configuration of a rack of a data center;

[0008] Figure 3 is a diagram of an example embodiment of another data center according to various embodiments, in which one or more of the techniques described herein may be implemented;

[0009] Figure 4 is a diagram of an example embodiment of another data center according to various embodiments, in which one or more of the techniques described herein may be implemented;

[0010] Figure 5 is a diagram representing the link - layer connectivity that can be established between various sled in a Figure 1 , Figure 3 and Figure 4 data center;

[0011] Figure 6 is a diagram of a rack architecture that can represent, according to some embodiments, the architecture of any particular rack among the racks depicted in Figures 1 - 4 ;

[0012] Figure 7 is a diagram of an example embodiment of a sled that can be used with the Figure 6 rack architecture;

[0013] Figure 8 is a diagram of an example embodiment of a rack architecture that provides support for sleds characterized by expansion capabilities;

[0014] Figure 9 is a diagram of an example embodiment of a rack implemented according to the Figure 8 rack architecture;

[0015] Figure 10 is a diagram of an example embodiment of a sled designed to be used in conjunction with a Figure 9 rack;

[0016] Figure 11is a diagram of an example embodiment of a data center according to various embodiments, in which one or more of the techniques described herein may be implemented;

[0017] Figure 12 is a simplified block diagram of at least one embodiment of a switch that may be used in Figure 5 a connectivity solution;

[0018] Figure 13 is a simplified block diagram of at least one embodiment of an environment that may be established by the switches of Figure 5 and 12 ; and

[0019] Figures 14 - 15 is a simplified flowchart of at least one embodiment of a method for switching network traffic that may be performed by the switches of Figure 5 , 12 and Figure 13 . DETAILED DESCRIPTION

[0020] While the concepts of the present disclosure may admit of various modifications and alternative forms, specific embodiments thereof have been shown by way of example in the drawings and will be described in detail herein. It should be understood, however, that the intention is not to limit the concepts of the present disclosure to the particular forms disclosed, but on the contrary, the invention will cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure and the appended claims.

[0021] References in the specification to "one embodiment", "an embodiment", "an illustrative embodiment", etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but each embodiment may or may not necessarily include that particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Additionally, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is considered within the knowledge of one of ordinary skill in the art to implement such feature, structure, or characteristic in connection with other embodiments (whether or not explicitly described). Further, it should be recognized that items included in a list in the form of "at least one of A, B, and C" can mean (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). Similarly, items listed in the form of "at least one of A, B, or C" can mean (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).

[0022] The disclosed embodiments may in some cases be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on a transient or non-transient machine-readable (e.g., computer-readable) storage medium, which may be read and executed by one or more processors. A machine-readable storage medium may be implemented as any storage device, mechanism, or other physical structure for storing or transmitting information in a form readable by a machine (e.g., volatile or non-volatile memory, media disk, or other media device).

[0023] In the figures, some structural or method features may be shown in a particular arrangement and / or order. However, it should be recognized that such particular arrangements and / or orders may not be required. Rather, in some embodiments, such features may be arranged in a manner and / or order different from that shown in the illustrative figures. Additionally, structural or method features included in a particular figure do not mean that such features are required in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.

[0024] Figure 1 A conceptual overview of a data center 100 is illustrated according to various embodiments, which generally may represent a data center or other type of computing network in which / for which one or more of the techniques described herein may be implemented. As Figure 1 shown, the data center 100 generally may include a plurality of racks, each of which may house computing devices including a corresponding set of physical resources. In Figure 1 the specific non-limiting example depicted, the data center 100 includes four racks 102A through 102D, which house computing devices including corresponding sets of physical resources (PCRs) 105A through 105D. According to this example, the common set of physical resources 106 of the data center 100 includes the various sets of physical resources 105A through 105D distributed among the racks 102A through 102D. The physical resources 106 may include various types of resources such as - for example - processors, coprocessors, accelerators, field programmable gate arrays (FPGAs), memories, and storage devices. Embodiments are not limited to these examples.

[0025] The illustrative data center 100 differs from a typical data center in many ways. For example, in the illustrative embodiments, the circuit boards ("skids") on which components such as CPUs, memories, and other components are placed are designed for increased thermal performance. In particular, in the illustrative embodiments, the skids are shallower and thinner than typical boards. In other words, the skids are shorter from front to back (where the cooling fans are located). This reduces the length of the path that air must travel across the components on the board. Additionally, the components on the skids are spaced farther apart than in a typical circuit board, and the components are arranged to reduce or eliminate shadowing (i.e., one component in the air flow path of another component). In the illustrative embodiments, processing components such as processors are located on the top side of the skid, while near-memory or other memory modules or stacks such as dual in-line memory modules (DIMMs) are located on the bottom side of the skid. As a result of the enhanced air flow provided by this design, the components can operate at higher frequencies and power levels than in a typical system, thereby increasing performance. Additionally, the skids are configured to blind mate with the power and data communication interfaces (such as cables, bus bars, optical interfaces, etc.) in each of the racks 102A, 102B, 102C, 102D, thus enhancing their ability to be quickly removed, upgraded, reinstalled, and / or replaced. Similarly, the individual components located on the skids (such as processors, accelerators, memories, and data storage drives) are configured to be easily upgraded (due to their increased spacing from each other). In the illustrative embodiments, the components additionally include hardware attestation features to confirm their reliability.

[0026] Furthermore, in the illustrative embodiments, the data center 100 utilizes a single network architecture ("fabric") that supports multiple other network architectures including Ethernet and Omni-Path. In the illustrative embodiments, the skids are coupled to switches via optical fibers, which provide higher bandwidth and lower latency than typical twisted pair cables (such as Category 5, Category 5e, Category 6, etc.). Due to the high-bandwidth, low-latency interconnects and network architecture, the data center 100 can use physically disaggregated pool resources (such as memories, accelerators (such as graphics accelerators, FPGAs, application specific integrated circuits (ASICs), etc.), and data storage drives), and provide them to compute resources (such as processors) on a demand basis, enabling the compute resources to access the pooled resources as if they were local. The illustrative data center 100 additionally receives utilization information for various resources, predicts the resource utilization for different types of workloads based on past resource utilization, and dynamically reallocates resources based on this information.

[0027] The racks 102A, 102B, 102C, 102D of the data center 100 can include physical design features that facilitate the automation of a wide variety of types of maintenance tasks. For example, the data center 100 can be implemented using racks designed for robotic access and accepting and accommodating robotically manipulable resource skids. Additionally, in an illustrative embodiment, the racks 102A, 102B, 102C, 102D include integrated power sources that receive a voltage greater than that typical for a power source. The increased voltage enables the power source to provide additional power to components on each skid, enabling the components to operate at a frequency higher than typical. In an illustrative embodiment, the power source includes a 277 VAC input to a power supply unit (PSU) to reduce the input current and reduce losses that might occur if a higher input current were used to compensate for a lower input voltage. Additionally, in an illustrative embodiment, more current inputs are provided to each skid to allow each skid to reach a higher power operating point.

[0028] Figure 2 A exemplary logical configuration of a rack 202 of the data center 100 is illustrated. As Figure 2 shown, the rack 202 can generally accommodate a plurality of skids, each of which can include a corresponding set of physical resources. In Figure 2 a specific non-limiting example depicted, the rack 202 accommodates skids 204-1 to 204-4 that include corresponding sets of physical resources 205-1 to 205-4, each of which forms part of a common set of physical resources 206 included in the rack 202. With respect to Figure 1 , if the rack 202 represents - for example - the rack 102A, the physical resources 206 can correspond to the physical resources 105A included in the rack 102A. In the context of this example, the physical resources 105A can thus consist of corresponding sets of physical resources, including physical storage resources 205-1, physical accelerator resources 205-2, physical memory resources 204-3, and physical computing resources 205-5 included in the skids 204-1 to 204-4 of the rack 202. The embodiments are not limited to this example. Each skid can contain a pool of each of the various types of physical resources (e.g., computing, memory, accelerator, storage). By having robotically accessible and robotically manipulable skids that include disaggregated resources, each type of resource can be upgraded independently of one another and at its own optimized refresh rate. In an illustrative embodiment, "robot accessible" and "robot manipulable" mean easily accessible and easily manipulable such that a robot or a person can perform the operations.

[0029] Figure 3Illustrated is an example of a data center 300 according to various embodiments, which may generally represent a data center in / for which one or more of the techniques described herein may be implemented. In Figure 3 the specific non-limiting example depicted, the data center 300 includes racks 302-1 through 302-32. In various embodiments, the racks of the data center 300 may be arranged in such a manner as to define and / or accommodate various access paths. For example, as Figure 3 shown, the racks of the data center 300 may be arranged in such a manner as to define and / or accommodate access paths 311A, 311B, 311C, and 311D. In some embodiments, the presence of such access paths may generally enable automated maintenance equipment (e.g., robotic maintenance equipment) to physically access computing devices housed in the various racks of the data center 300 and perform automated maintenance tasks (e.g., replacing a faulty sled, upgrading a sled). In various embodiments, the dimensions of the access paths 311A, 311B, 311C, and 311D, the dimensions of the racks 302-1 through 302-32, and / or one or more other aspects of the physical layout of the data center 300 may be selected to facilitate such automated operations. Embodiments are not limited to this context.

[0030] Figure 4 Illustrated is an example of a data center 400 according to various embodiments, which may generally represent a data center in / for which one or more of the techniques described herein may be implemented. As Figure 4 shown, the data center 400 may be characterized by an optical fabric 412. The optical fabric 412 may generally include a combination of optical signaling media (e.g., optical cables, also referred to herein as optical fibers or fiber bundles) and an optical switching infrastructure that includes one or more switches 515 (also referred to herein as network switches) that may be included in a central location 406 through which any particular sled in the data center 400 may send signals to and receive signals from every other sled in the data center 400. The signaling connectivity provided by the optical fabric 412 to any given sled may include connectivity to other sleds in the same rack as well as sleds in other racks. In Figure 4In the specific non-limiting example depicted, data center 400 includes four racks 402A through 402D. Racks 402A through 402D house respective pairs of skids 404A-1 and 404A-2, 404B-1 and 404B-2, 404C-1 and 404C-2, and 404D-1 and 404D-2. Thus, in this example, data center 400 includes a total of eight skids. Via optical fabric 412, each such skid can have signaling connectivity with each of the other seven skids in data center 400. For example, via optical fabric 412, skid 404A-1 in rack 402A can have signaling connectivity with skid 404A-2 in rack 402A, as well as with the other six skids 404B-1, 404B-2, 404C-1, 404C-2, 404D-1, and 404D-2 distributed among the other racks 402B, 402C, and 402D of data center 400. The embodiments are not limited to this example.

[0031] Figure 5 FIG. illustrates an overview of connectivity scheme 500, which generally may represent link layer connectivity that can be established among various skids in a data center (such as Figure 1 , 3 and any of the example data centers 100, 300, and 400 of FIGS. 1, 3, and 4). Connectivity scheme 500 can be implemented using an optical fabric characterized by a multimode optical switching infrastructure 514. Multimode optical switching infrastructure 514 generally may include a switching infrastructure that is capable of receiving communications via the same unified set of optical signaling media according to multiple link layer protocols and appropriately switching such communications. In various embodiments, multimode optical switching infrastructure 514 can be implemented using one or more multimode optical switches 515. In various embodiments, multimode optical switches 515 generally may include high-radix switches. In some embodiments, multimode optical switches 515 can include multilayer switches, such as four-layer switches. In various embodiments, multimode optical switches 515 can be characterized by integrated silicon photonics (enabling them to switch communications with significantly reduced latency compared to conventional switching devices). In some embodiments, multimode optical switches 515 can form leaf switches 530 in a leaf-spine architecture, which additionally includes one or more multimode optical spine switches 520.

[0032] In various embodiments, the multimode optical switch 515 may be capable of receiving Ethernet protocol communications carrying Internet Protocol (IP packets) and communications according to a second high-performance computing (HPC) link layer protocol (e.g., InfiniBand of Intel's Omni-Path architecture) via an optically configured optical signaling medium. Other native protocols may be included, such as raw accelerated intercommunication protocols, storage protocols, or even application-specific protocols not embedded or tunneled within existing IP or Omni-Path fabric protocols. As Figure 5 reflected therein, for any particular pair of skids 504A and 504B having optical signaling connectivity to the optical fabric, the connectivity scheme 500 can thus provide support for link layer connectivity via Ethernet links and HPC links. Thus, both Ethernet and HPC communications can be supported by a single high-bandwidth, low-latency switching fabric. Embodiments are not limited to this example. In some embodiments, the switch 515 is in a central location (e.g., central location 406) in the data center 100 rather than being located within a rack, thereby enabling easy access to the switch 515. Additionally, each skid 504 can be coupled to four switches 515, each switch 515 providing one-quarter of the total bandwidth, such as a 50 gigabit per second upstream and 50 gigabit per second downstream fiber optic connection of a total bandwidth of 200 gigabits per second upstream and 200 gigabits per second downstream. In other embodiments, the total bandwidth can be an amount different from 200 gigabits per second upstream and 200 gigabits per second downstream. For example, in other embodiments, each fiber can provide greater than 50 gigabits per second upstream and greater than 50 gigabits per second downstream. As a result, each skid 504 can receive a total upstream bandwidth of 200 gigabits per second and a total downstream bandwidth of 200 gigabits per second. By spreading the bandwidth among multiple switches 515, if any one switch 515 becomes inoperable, the other three-quarters of the bandwidth (e.g., 150 gigabits) remains available. As such, the connectivity scheme in such embodiments is more fault-tolerant than embodiments in which all available bandwidth is integrated in a single switch.

[0033] Figure 6 Illustrated is a general overview of a rack architecture 600 according to some embodiments, which may represent Figures 1 to 4 the architecture of any particular rack depicted in Figure 6 reflected therein, the rack architecture 600 can generally be characterized by a plurality of skid spaces into which skids can be inserted, each skid space being robotically or manually accessible via a rack access area 601. In Figure 6In the specific non - limiting example depicted, the rack architecture 600 features five sled spaces 603 - 1 through 603 - 5. The sled spaces 603 - 1 through 603 - 5 feature corresponding Multi - Purpose Connector Modules (MPCMs) 616 - 1 through 616 - 5. When a sled is inserted into any given one of the sled spaces 603 - 1 through 603 - 5, the corresponding MPCM (e.g., MPCM 616 - 3) can be coupled to the mating MPCM of the inserted sled. This coupling can provide the inserted sled with connectivity to the signaling infrastructure and power infrastructure of the rack in which it is housed. Among the types of sleds accommodated by the rack architecture 600 can be one or more types of sleds characterized by expandability.

[0034] Figure 7 An example of a sled 704 that can represent such types of sleds is illustrated. As Figure 7 shown, the sled 704 can include a set of physical resources 705, as well as an MPCM 716, which is designed to couple to the mating MPCM when the sled 704 is inserted into a sled space (e.g., Figure 6 any of the sled spaces 603 - 1 through 603 - 5). The sled 704 can also feature an expansion connector 717. The expansion connector 717 can generally include sockets, slots, or other types of connection elements (which can accept one or more types of expansion modules, such as an expansion sled 718). By coupling to the mating connector on the expansion sled 718, the expansion connector 717 can provide the physical resources 705 with access to complementary computing resources 705B residing on the expansion sled 718. The embodiments are not limited to this context.

[0035] Figure 8 An example of a rack architecture 800 that can represent a rack architecture is illustrated, which can be implemented to support sleds (e.g., Figure 7 the sled 704) characterized by expandability. In Figure 8 the specific non - limiting example depicted, the rack architecture 800 includes seven sled spaces 803 - 1 through 803 - 7, which feature corresponding MPCMs 816 - 1 through 816 - 7. The sled spaces 803 - 1 through 803 - 7 include corresponding main regions 803 - 1A through 803 - 7A and corresponding expansion regions 803 - 1B through 803 - 7B. For each such sled space, when the corresponding MPCM is coupled to the mating MPCM of the inserted sled, the main region can generally constitute the region of the sled space that can physically accommodate the inserted sled. The expansion region can generally constitute the region of the sled space that can physically accommodate expansion modules, such as Figure 7 the expansion sled 718 (in the case where the inserted sled is configured with such a module).

[0036] Figure 9 illustrates an example of a rack 902 according to some embodiments, which may represent a rack implemented according to the Figure 8 rack architecture 800. In the specific non-limiting example depicted in Figure 9 , the rack 902 is characterized by seven sled spaces 903-1 through 903-7, which include corresponding primary regions 903-1A through 903-7A and corresponding expansion regions 903-1B through 903-7B. In various embodiments, an air cooling system may be used to implement temperature control in the rack 902. For example, as reflected in Figure 9 , the rack 902 may be characterized by a plurality of fans 919, which are typically arranged to provide air cooling within the various sled spaces 903-1 through 903-7. In some embodiments, the height of the sled space is greater than the conventional "1U" server height. In such embodiments, the fans 919 may typically include relatively slow large-diameter cooling fans as compared to the fans used in conventional rack configurations. Operating a larger-diameter cooling fan at a lower speed can increase fan life while still providing the same amount of cooling as compared to a smaller-diameter cooling fan operating at a higher speed. The sled is physically shallower than conventional rack dimensions. Additionally, components are arranged on each sled to reduce thermal shadowing (i.e., not arranged in series in the air flow direction). Thus, the wider and shallower sled allows for an increase in device performance because, due to improved cooling (i.e., no thermal shadowing, more space between devices, more space for larger heatsinks, etc.), the devices can operate at a higher thermal envelope (e.g., 250W).

[0037] The MPCMs 916-1 through 916-7 may be configured to provide power usage to the inserted sleds from the power supplied by the respective power modules 920-1 through 920-7, and each power module may draw power from an external power source 921. In various embodiments, the external power source 921 may deliver alternating current (AC) power to the rack 902, and the power modules 920-1 through 920-7 may be configured to convert such AC power to direct current (DC) power to be supplied to the inserted sleds. In some embodiments, for example, the power modules 920-1 through 920-7 may be configured to convert 277 volts AC power to 12 volts DC power to be provided to the inserted sleds via the respective MPCMs 916-1 through 916-7. The embodiments are not limited to this example.

[0038] The MPCMs 916-1 through 916-7 may also be arranged to provide optical signaling connectivity to the inserted sleds to a multimode optical switching infrastructure 914, and the multimode optical switching infrastructure 914 may be associated with Figure 5is the same as or similar to the multi-mode optical switching infrastructure 514. In various embodiments, the optical connectors included in the MPCMs 916-1 through 916-7 can be designed to couple with corresponding optical connectors included in the MPCM of the inserted sled to provide optical signaling connectivity to the multi-mode optical switching infrastructure 914 for such sled via respective lengths of optical cables 922-1 through 922-7 (also referred to herein as optical fibers). In some embodiments, each such length of optical cable can extend from its corresponding MPCM to an optical interconnection loom 923 outside the sled space of the rack 902. In various embodiments, the optical interconnection loom 923 can be arranged to pass through support columns or other types of load-bearing elements of the rack 902. The embodiments are not limited to this context. Since the inserted sled is connected to the optical switching infrastructure via the MPCM, resources that are typically spent on manually configuring rack cabling to accommodate newly inserted sleds can be saved.

[0039] Figure 10 illustrates an example of a sled 1004 according to some embodiments, which may represent a sled designed to be used in conjunction with Figure 9 the rack 902. The sled 1004 can be characterized by an MPCM 1016 that includes an optical connector 1016A and a power connector 1016B and is designed to couple with a corresponding MPCM in the sled space (in conjunction with inserting the MPCM 1016 into that sled space). Coupling the MPCM 1016 with such a corresponding MPCM can cause the power connector 1016 to couple with the power connector included in the corresponding MPCM. This can generally enable the physical resources 1005 of the sled 1004 to be powered from an external source via the power connector 1016 and a power transmission medium 1024 that conductively couples the power connector 1016 to the physical resources 1005.

[0040] The sled 1004 can also include a multi-mode optical network interface circuit 1026. The multi-mode optical network interface circuit 1026 can generally include circuitry capable of communicating via an optical signaling medium according to each of a plurality of link layer protocols supported by Figure 9 the multi-mode optical switching infrastructure 914. In some embodiments, the multi-mode optical network interface circuit 1026 can be capable of both Ethernet protocol communication and communication according to a second high-performance protocol. In various embodiments, the multi-mode optical network interface circuit 1026 can include one or more optical transceiver modules 1027, each of which can be capable of transmitting and receiving optical signals via each of one or more optical channels. The embodiments are not limited to this context.

[0041] Coupling the MPCM 1016 with its counterpart MPCM in the sled space within a given rack enables coupling of the optical connector 1016A with the optical connector included in the counterpart MPCM. This can generally establish optical connectivity between the multimode optical network interface circuit 1026 and the fiber optic cables of the sled via each of the optical channel set 1025. The multimode optical network interface circuit 1026 can communicate with the physical resources 1005 of the sled 1004 via the telecommunications signaling medium 1028. In addition to the arrangement of the components on the sled and the size of the sled for providing improved cooling and enabling operation at a relatively high thermal envelope (e.g., 250W) as described above with reference to Figure 9 In some embodiments, the sled may include one or more additional features to facilitate air cooling, such as heat pipes and / or heat sinks (arranged to dissipate the heat generated by the physical resources 1005). It is noted that although the exemplary sled 1004 depicted in Figure 10 is not characterized by expansion connectors, any given sled characterized by the design elements of the sled 1004 may also be characterized by expansion connectors in some embodiments. The embodiments are not limited to this context.

[0042] Figure 11 An example of a data center 1100 according to various embodiments is illustrated, which generally may represent a data center in / for which one or more of the techniques described herein may be implemented. As reflected in Figure 11 a physical infrastructure management framework 1150A may be implemented to facilitate management of the physical infrastructure 1100A of the data center 1100. In various embodiments, one function of the physical infrastructure management framework 1150A may be to automate maintenance functions within the data center 1100, such as using robotic maintenance equipment to service computing devices within the physical infrastructure 1100A. In some embodiments, the physical infrastructure 1100A may be characterized by an advanced telemetry system that performs telemetry reporting that is robust enough to support remote automation management of the physical infrastructure 1100A. In various embodiments, the telemetry information provided by such an advanced telemetry system may support features such as fault prediction / prevention capabilities and capacity planning capabilities. In some embodiments, the physical infrastructure management framework 1150A may also be configured to manage the authentication of physical infrastructure components using hardware attestation techniques. For example, a robot may verify the reliability of components by analyzing information collected from radio frequency identification (RFID) or other physical tags associated with each component to be installed before installation. The embodiments are not limited to this context.

[0043] As in Figure 11As shown in, the physical infrastructure 1100A of the data center 1100 may include an optical fabric 1112, which may include a multimode optical switching infrastructure 1114. The optical fabric 1112 and the multimode optical switching infrastructure 1114 may be the same as or similar to the optical fabric 412 of Figure 4 and the multimode optical switching infrastructure 514 of Figure 5 respectively, and may provide high-bandwidth, low-latency, multi-protocol connectivity between the sleds of the data center 1100. As discussed above, with reference to Figure 1 , in various embodiments, the availability of such connectivity may make it feasible to depool and dynamically pool resources such as accelerators, memories, and storage. In some embodiments, for example, one or more pooled accelerator sleds 1130 may be included between the physical infrastructures 1100A of the data center 1100, and each physical infrastructure 1100A may include an accelerator resource pool - such as co-processors, specialized processors, and / or FPGAs - for example - which are globally accessible to other sleds via the optical fabric 1112 and the multimode optical switching infrastructure 1114.

[0044] In another example, in various embodiments, one or more pooled storage sleds 1132 may be included between the physical infrastructures 1100A of the data center 1100, and each physical infrastructure 1100A may include a storage resource pool that is available for global access to other sleds via the optical fabric 1112 and the multimode optical switching infrastructure 1114. In some embodiments, such pooled storage sleds 1132 may include a pool of solid-state storage devices (e.g., solid-state drives (SSDs)). In various embodiments, one or more high-performance processing sleds 1134 may be included between the physical infrastructures 1100A of the data center 1100. In some embodiments, the high-performance processing sled 1134 may include a high-performance processor pool and cooling features (which enhance air cooling to produce a higher thermal envelope of up to 250W or higher). In various embodiments, any given high-performance processing sled 1134 may be characterized by an expansion connector 1117, which may accept a remote memory expansion sled, such that the remote memory available locally to the high-performance processing sled 1134 is depooled from the near memory and processors included on that sled. In some embodiments, such high-performance processing sleds 1134 may be configured with remote memory (using an expansion sled that includes a low-latency SSD storage device). The optical infrastructure allows the computing resources on one sled to utilize remote accelerator / FPGA, memory, and / or SSD resources (which are depooled on sleds on any other rack located in the same rack or data center). The remote resources may be located in the spine-leaf network architecture described above with reference to Figure 5 . Embodiments are not limited to this context.

[0045] In various embodiments, one or more abstraction layers can be applied to the physical resources of the physical infrastructure 1100A to define a virtual infrastructure, such as the software-defined infrastructure 1100B. In some embodiments, the virtual computing resources 1136 of the software-defined infrastructure 1100B can be allocated to support the provisioning of the cloud service 1140. In various embodiments, a specific set of the virtual computing resources 1136 can be grouped for use in provisioning the cloud service 1140 (in the form of the SDI service 1138). Examples of the cloud service 1140 can include - but are not limited to - Software as a Service (SaaS) services 1142, Platform as a Service (PaaS) services 1144, and Infrastructure as a Service (IaaS) services 1146.

[0046] In some embodiments, a virtual infrastructure management framework (SDI) 1150B can be used to manage the software-defined infrastructure 1100B. In various embodiments, the virtual infrastructure management framework 1150B can be designed to implement workload fingerprinting techniques and / or machine learning techniques in conjunction with managing the allocation of the virtual computing resources 1136 and / or the SDI service 1138 to the cloud service 1140. In some embodiments, the virtual infrastructure management framework 1150B can use / consult telemetry data in conjunction with performing such resource allocation. In various embodiments, an application / service management framework 1150C can be implemented to provide QoS management capabilities for the cloud service 1140. The embodiments are not limited to this context.

[0047] Now referring Figure 12 , in an illustrative embodiment, the switch 515 is a multimode optical switch that is connected to the sled (e.g., sled 704, 1004) via an optical connection (e.g., the optical fabric 1112) to provide higher bandwidth and lower latency than a switch that is connected to a computing device using a typical twisted pair cable (e.g., Category 5, Category 5e, Category 6, etc.). The higher bandwidth and lower latency interconnection enables the pooling of resources, such as memory, accelerators (e.g., graphics accelerators, FPGAs, ASICs, etc.), and physically disaggregated data storage devices, for use by computing resources (e.g., processors) to execute workloads as needed. More specifically, the high bandwidth and low latency provided by the optical connection and the corresponding switch 515 enable physical resources 206 located at various locations within the data centers 100, 300, 400 ( Figure 2As shown in [FIGURE], they can provide a similar responsiveness as if they were local to the processors that utilize them. Additionally, in the illustrative embodiment, switch 515 is multi-mode, which means that it can switch (i.e., forward) network traffic formatted according to two or more different link layer protocols (such as the HPC link layer protocol (e.g., Intel Omni-Path), Ethernet, or any other proprietary communication protocol, such as the raw accelerator interconnect protocol, storage protocol, or even an application-specific protocol not embedded / tunneled in the existing Internet Protocol (IP) or Omni-Path protocol).

[0048] Still referring to Figure 12 , switch 515 can be implemented as any type of device capable of performing the functions described herein, which include: receiving network traffic from one or more devices via an optical connection; determining the communication protocol of the network traffic; using the communication protocol to determine the destination device of the network traffic, and forwarding the network traffic to the destination device via another optical connection. For example, switch 515 can be implemented as a computer, a multiprocessor system, or a network facility (e.g., physical or virtual). As Figure 12 shown, illustrative switch 515 includes a central processing unit (CPU) 1202, a main memory 1204, an input / output (I / O) subsystem 1206, communication circuitry 1208, and one or more data storage devices 1212. Of course, in other embodiments, switch 515 may include other or additional components, such as components commonly found in a computer (e.g., a display, peripheral devices, etc.). Additionally, in some embodiments, one or more of the illustrative components may be incorporated in another component or otherwise form a part of another component. For example, in some embodiments, main memory 1204 or a portion thereof may be incorporated in CPU 1202.

[0049] The CPU 1202 can be implemented as any type of processor capable of performing the functions described herein. The CPU 1202 can be implemented as a single-core or multi-core processor, a microcontroller, or other processor or processing / control circuitry. In some embodiments, the CPU 1202 can be implemented as, include, or be coupled to a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), reconfigurable hardware, or hardware circuitry, or other specialized hardware that facilitates the performance of the functions described herein. Similarly, the main memory 1204 can be implemented as any type of volatile (e.g., dynamic random access memory (DRAM), etc.) or non-volatile memory or data storage device capable of performing the functions described herein. In some embodiments, all or part of the main memory 1204 can be integrated into the CPU 1202. In operation, the main memory 1204 can store various software and data used during operation, such as network traffic data, protocol data, address data, operating systems, applications, programs, libraries, and drivers.

[0050] The I / O subsystem 1206 can be implemented as circuitry and / or components that facilitate input / output operations with the CPU 1202, the main memory 1204, and other components of the switch 515. For example, the I / O subsystem 1206 can be implemented as or otherwise include a memory controller hub, an input / output control hub, an integrated sensor hub, a firmware device, a communication link (e.g., a silicon photonics device, a point-to-point link, a bus link, a wire, a cable, an optical waveguide, a printed circuit board trace, etc.), and / or other components and subsystems that facilitate input / output operations. In some embodiments, the I / O subsystem 1206 can form part of a system-on-chip (SoC) and be incorporated on a single integrated circuit chip together with one or more of the CPU 1202, the main memory 1204, and other components of the switch 515.

[0051] The communication circuit 1208 can be implemented as any communication circuit, device, or collection thereof that enables communication between the switch 515 and other devices (e.g., other switches 515 or sled 704) over a network. In an illustrative embodiment, the communication circuit 1208 includes components similar to the multi-mode optical network interface circuit 1026 described above with reference to Figure 10 The communication circuit 1208 can be configured to use a variety of communication technologies (e.g., wired or wireless communication) and associated protocols (e.g., Intel InfiniBand, Ethernet, Bluetooth®, Wi-Fi®, WiMAX, etc.) to effect such communication. In an illustrative embodiment, the communication circuit 1208 is configured to communicate over the optical fabric 1112 described with reference to Figure 11

[0052] ​The illustrative communication circuit 1208 includes one or more port logics 1210. In the illustrative embodiment, each port logic 1210 may be implemented as an optical transceiver module 1027. Each port logic 1210 may be implemented as one or more daughterboards, daughter cards, network interface cards, controller chips, chip sets, or other devices that can be used by the switch 515 to connect to other devices (e.g., other switches 515 and / or sled boards 704) through a network (e.g., multimode optical switching infrastructure 514, 914, 1114). In the illustrative embodiment, one or more port logics 1210 together enable simultaneous communication with multiple other devices (e.g., up to 1024 other devices). Additionally, in the illustrative embodiment, each device is connected to the port logic 1210, where one optical fiber is used for incoming network traffic (e.g., frames), and another optical fiber is used for outgoing network traffic. In some embodiments, the port logic 1210 may be implemented as part of a system-on-chip (SoC) that includes one or more processors, or may be included on a multi-chip package that also contains one or more processors. In some embodiments, each port logic 1210 may include a local processor (not shown) and / or local memory (not shown), which are local to the port logic 1210. In such embodiments, the local processor of the port logic 1210 may be capable of performing one or more of the functions of the CPU 1202 described herein. Additionally or alternatively, in such embodiments, the local memory of the port logic 1210 may be integrated into one or more components of the switch 515 at the board level, socket level, chip level, and / or other levels.

[0053] One or more illustrative data storage devices 1212 may be implemented as any type of device configured for short-term or long-term storage of data, such as, for example, memory devices and circuits, memory cards, hard disk drives, solid state drives, or other data storage devices. Each data storage device 1212 may include a system partition that stores data and firmware code for the data storage device 1212. Each data storage device 1212 may also include an operating system partition that stores executable files of the operating system as well as data files.

[0054] Additionally, switch 515 may include a display 1214. The display 1214 may be implemented as or otherwise use any suitable display technology, such as including liquid crystal displays (LCDs), light emitting diode (LED) displays, cathode ray tube (CRT) displays, plasma displays, and / or other displays that may be used in a computing device. The display 1214 may include a touchscreen sensor that uses any suitable touchscreen input technology to detect a user's tactile selection of information displayed on the display, including but not limited to: resistive touchscreen sensors, capacitive touchscreen sensors, surface acoustic wave (SAW) touchscreen sensors, infrared touchscreen sensors, optical imaging touchscreen sensors, acoustic touchscreen sensors, and / or other types of touchscreen sensors.

[0055] Additionally or alternatively, switch 515 may include one or more peripheral devices 1216. Such peripheral devices 1216 may include any type of peripheral device commonly found in a computing device, such as speakers, mice, keyboards, and / or other input / output devices, interface devices, and / or other peripheral devices.

[0056] Now refer to Figure 13, in an illustrative embodiment, switch 515 may establish environment 1300 during operation. Illustrative environment 1300 includes network communicator 1320, protocol determiner 1330, and network traffic switcher 1340. Each component of environment 1300 may be implemented as hardware, firmware, software, or a combination thereof. As such, in some embodiments, one or more of the components of environment 1300 may be implemented as a collection of electrical devices or circuits (e.g., network communicator circuit 1320, protocol determiner circuit 1330, network traffic switcher circuit 1330, etc.). It should be recognized that in such embodiments, one or more of network communicator circuit 1320, protocol determiner circuit 1330, or network traffic switcher circuit 1340 may form part of one or more of CPU 1202, main memory 1204, I / O subsystem 1206, and / or other components of switch 515. The circuit may be implemented as dedicated hardware or a general-purpose processor that executes code to implement the desired functionality. In an illustrative embodiment, environment 1300 includes network traffic data 1302, which may be implemented as a set of data (e.g., frames) received from a source device (e.g., another switch 515 or physical resource 705 (e.g., a processor) residing on sled 704), destined for another device (e.g., another switch 515 or physical resource 705 (e.g., a processor) residing on sled 704), and formatted according to a corresponding communication protocol. Additionally, illustrative environment 1300 includes protocol data 1304, which indicates the formats (e.g., frame size, positions and sizes of fields within the frame, and / or one or more codes typically embedded within frames of a specific communication protocol) and rules (e.g., the position within the frame identifying the destination address, timing information regarding whether to forward the frame immediately or delay it for a defined period of time, whether and how to acknowledge receipt, etc.) of various communication protocols for exchanging network traffic according to the corresponding protocol. Further, illustrative environment 1300 includes address data 1306, which indicates the unique addresses (e.g., media access control addresses) of the devices (e.g., physical resources 705 (e.g., processors) residing on sled 704 and / or switch 515) connected to the corresponding physical ports of this switch 515 and where the optical fibers associated with each device are connected.

[0057] In the illustrative environment 1300, the network communicator 1320 (which may be implemented as hardware, firmware, software, virtualized hardware, emulation architectures, and / or combinations thereof, as discussed above) is configured to facilitate inbound and outbound network communications (e.g., network traffic, network frames, network packets, network flows, etc.) to and from the switch 515, respectively. To this end, the network communicator 1320 is configured to receive and process network traffic (e.g., frames) from a device (e.g., another switch 515 or sled 704), and forward the network traffic to another device (e.g., another switch 515 or sled 704) using the address data encoded in the network traffic (e.g., in the frame header) according to the corresponding protocol of the network traffic (e.g., HPC communication protocol, Ethernet protocol, etc.). Accordingly, in some embodiments, at least part of the functionality of the network communicator 1320 may be performed by the communication circuitry 1208 and, in the illustrative embodiment, by one or more NICs 1210.

[0058] The protocol determiner 1330 (which may be implemented as hardware, firmware, software, virtualized hardware, emulation architectures, and / or combinations thereof, as discussed above) is configured to analyze the received frame of network traffic data 1302, identify the format of the frame, such as by identifying the size of the frame, fields within the frame (such as headers and payloads), and / or one or more codes within the fields that indicate the specific communication protocol supported by the switch. In so doing, the protocol determiner 1330 may be configured to compare the identified format of the frame with the protocol data 1304 to identify a match and the corresponding rules for exchanging network traffic.

[0059] The network traffic switch 1340 (which may be implemented as hardware, firmware, software, virtualized hardware, emulation architectures, and / or combinations thereof, as discussed above) is configured to identify the address of the device to which each frame of network traffic is to be forwarded by: extracting data from the frame according to the identified network protocol (e.g., by extracting data from a specific location within the frame as specified in the protocol data); reading the address data 1306 to match the identified address with a physical port of the NICs 1210 (e.g., optical channel 1025), wherein the device that matches the identified address is connected to the switch 515; and issuing a request to the network communicator 1320 to transmit a frame of network data to the corresponding device through the physical port.

[0060] Now refer to Figure 14, in use, switch 515 can execute method 1400 for switching network traffic. Method 1400 begins at block 1402, where switch 515 determines whether to switch network traffic. In an illustrative embodiment, if switch 515 is powered on and connected to a network (e.g., multi-mode optical switching infrastructure 514, 914, 1114), then switch 515 determines to switch network traffic. In other embodiments, switch 515 may determine whether to switch network traffic based on other factors. In any case, in response to the determination to switch network traffic, in an illustrative embodiment, method 1400 proceeds to block 1404, where switch 515 receives network traffic to be forwarded (i.e., switched). The network traffic may be formatted according to any of a plurality of supported link layer communication protocols, including high performance computing (HPC) protocols (e.g., Intel Omni-Path, etc.) or another type of link layer communication protocol such as Ethernet.

[0061] When receiving network traffic, switch 515 can receive network traffic through an optical connection (e.g., an optical fiber connected to communication circuit 1208), as indicated in block 1406. Additionally, in an illustrative embodiment, when receiving network traffic to be forwarded, switch 515 receives network traffic through a connection (e.g., a 50 gigabit per second connection of a 200 gigabit per second link, a 100 gigabit per second connection of a 400 gigabit per second link, a 200 gigabit per second connection of an 800 gigabit per second link, etc.) having a portion of the total bandwidth of the link, as indicated in block 1408. As described herein, in an illustrative embodiment, switch 515 is one of a plurality of switches (e.g., four switches) in data center 100 that together provide a total amount of gigabits per second of connection for devices in the data center (e.g., sled 704). When receiving network traffic, switch 515 can receive network traffic from sled 704, as indicated in block 1410. As an example, network traffic received from sled 704 can be the result of workloads assigned to sled 704, or can be instructions from one sled 704 to another sled 704 to perform a specific operation (e.g., retrieve data, compress data, encrypt data, etc.), or the result of such an operation (e.g., retrieved data, compressed data, encrypted data, etc.). As such, switch 515 can receive network traffic from compute sled 704 (such as a sled 704 that includes one or more processors (e.g., physical computing resources 205-4)), as indicated in block 1412. Additionally or alternatively, switch 515 can receive network traffic from storage sled 704 (such as a sled 704 that includes one or more data storage devices (e.g., physical storage resources 205-1)), as indicated in block 1414. Switch 515 can additionally or alternatively receive network traffic from accelerator sled 704 (such as a sled 704 that includes one or more coprocessors, field programmable gate arrays (FPGAs), or other specialized hardware for performing computations (e.g., physical accelerator resources 205-2)), as indicated in block 1416. Additionally or alternatively, switch 515 can receive network traffic from memory sled 704 (such as a sled 704 that includes one or more memory devices (e.g., physical memory resources 205-3)), as indicated in block 1418. Switch 515 can additionally or alternatively receive network traffic from another switch, as indicated in block 1420. In doing so, switch 515 can receive network traffic from a leaf switch in a leaf-spine architecture (e.g., Figure 5 leaf switch 530), as indicated in block 1422. Alternatively, switch 515 can receive network traffic from a spine switch in a leaf-spine architecture (e.g., Figure 5The spine switch 520) receives network traffic as indicated in block 1424.

[0062] Still referring to Figure 14 , after receiving the network traffic, method 1400 proceeds to block 1426 where switch 515 determines the link layer protocol of the network traffic as indicated in block 1426. In doing so, switch 515 can determine whether the network traffic includes Ethernet protocol traffic or HPC network traffic as indicated in block 1428. Additionally or alternatively, switch 515 can determine whether the network traffic includes traffic of a different network protocol (e.g., different from Ethernet or HPC traffic) as indicated in block 1430. As discussed above, switch 515 can determine the type of network traffic by identifying aspects of the frame such as the size of the frame, fields within the frame, and / or one or more codes included in the frame and comparing these aspects to protocol data 1304 to determine if a matching protocol (i.e., a set of reference aspects that match the identified aspects) is included therein. After determining the link layer protocol, method 1400 proceeds to Figure 15 block 1432 where switch 515 forwards the network traffic.

[0063] Now referring to Figure 15 , at block 1432, switch 515 forwards the network traffic according to the link layer protocol (e.g., according to the corresponding rules defined in protocol data 1304). In doing so, in an illustrative embodiment, switch 515 determines the destination address according to the link layer protocol as indicated in block 1434. In doing so, in an illustrative embodiment, switch 515 identifies the address of the device to which each frame of network traffic is to be forwarded, which is done by extracting data from the frame in accordance with the identified communication protocol (e.g., by extracting data from a specific location within the frame as specified in protocol data 1304) and reading address data 1306 to match the identified address to the physical port of NICS 1210 where the device matching the identified address is connected. In block 1436, in an illustrative embodiment, switch 515 forwards the network traffic to the destination address. In an illustrative embodiment, when forwarding the network traffic, switch 515 transmits the frame of network data through the physically matched port. Additionally, in an illustrative embodiment, switch 515 forwards the network traffic through an optical connection (e.g., an optical fiber) as indicated in block 1438. Moreover, in an illustrative embodiment, when forwarding the network traffic, switch 515 forwards the network traffic through a connection (e.g., a 50 gigabit per second connection of a 200 gigabit per second link, a 100 gigabit per second connection of a 400 gigabit per second link, a 200 gigabit per second connection of an 800 gigabit per second link, etc.) having a portion of the total bandwidth of the link (e.g., a quarter or other portion) as indicated in block 1440.

[0064] Still referring to Figure 15 , as indicated in block 1442, when forwarding network traffic, switch 515 can forward the network traffic to sled 704. In doing so, switch 515 can forward the network traffic to compute sled 704, as indicated in block 1444. Alternatively, switch 515 can forward the network traffic to storage sled 704, as indicated in block 1446. As indicated in block 1448, switch 515 can forward the network traffic to accelerator sled 704. Alternatively, as indicated in block 1450, switch 515 can forward the network traffic to memory sled 704. As indicated in block 1452, switch 515 can alternatively forward the network traffic to another switch 515, such as to leaf switch 530, as indicated in block 1454, or to a spine switch, as indicated in block 1456. After forwarding the network traffic, method 1400 loops back to Figure 14 block 1402 to determine whether to continue switching network traffic.

[0065] Example

[0066] Illustrative examples of the technologies disclosed herein are provided below. Embodiments of the technologies may include any one or more of the examples described below and any combination thereof.

[0067] Example 1 includes a network switch that includes: one or more processors; a communication circuit coupled to the one or more processors, wherein the communication circuit assists the one or more processors in switching network traffic of multiple link layer protocols; and one or more memory devices in which multiple instructions are stored, the instructions when executed causing the network switch to: receive, via an optical connection with the communication circuit, network traffic to be forwarded; determine a link layer protocol of the received network traffic, wherein the received network traffic is formatted according to one of the multiple link layer protocols; and forward the network traffic to a destination network device according to the determined link layer protocol.

[0068] Example 2 includes the subject matter of Example 1, and wherein receiving the network traffic includes receiving network traffic via an optical connection that provides one quarter of the total bandwidth of the link.

[0069] Example 3 includes the subject matter of any one of Examples 1 and 2, and wherein receiving the network traffic includes receiving the network traffic from a sled coupled to the optical connection.

[0070] Example 4 includes the subject matter of any one of Examples 1-3, and wherein receiving the network traffic from the skateboard includes receiving network traffic from at least one of a computing skateboard including one or more processors, a storage skateboard including one or more data storage devices, an accelerator skateboard including one or more coprocessors or field programmable gate arrays, or a memory skateboard including one or more memory devices.

[0071] Example 5 includes the subject matter of any one of Examples 1-4, and wherein receiving the network traffic includes receiving the network traffic from another network switch.

[0072] Example 6 includes the subject matter of any one of Examples 1-5, and wherein receiving the network traffic from another network switch includes receiving the network traffic from a leaf switch or a spine switch in a leaf-spine network architecture.

[0073] Example 7 includes the subject matter of any one of Examples 1-6, and wherein the plurality of instructions further cause the network switch to forward traffic of two or more different link layer protocols.

[0074] Example 8 includes the subject matter of any one of Examples 1-7, and wherein determining the link layer protocol of the received network traffic includes: determining whether the link layer protocol is an Ethernet protocol, a high performance computing (HPC) protocol, or another dedicated communication protocol.

[0075] Example 9 includes the subject matter of any one of Examples 1-8, and wherein forwarding the network traffic according to the determined link layer protocol includes: determining the destination address of the network traffic according to the determined link layer protocol.

[0076] Example 10 includes the subject matter of any one of Examples 1-9, and wherein forwarding the network traffic includes forwarding the network traffic through another optical connection to one of the skateboard or another network switch using the communication circuit.

[0077] Example 11 includes the subject matter of any one of Examples 1-10, and wherein forwarding the network traffic includes forwarding the network traffic to one of a leaf switch or a spine switch in a leaf-spine network architecture.

[0078] Example 12 includes the subject matter of any one of Examples 1-11, and wherein forwarding the network traffic includes forwarding the network traffic through an optical connection providing one quarter of the total bandwidth of the link.

[0079] Example 13 includes the subject matter of any of Examples 1 - 12, and wherein forwarding the network traffic includes forwarding network traffic to at least one of a compute sled including one or more processors, a storage sled including one or more data storage devices, an accelerator sled including one or more coprocessors or field programmable gate arrays, or a memory sled including one or more memory devices.

[0080] Example 14 includes a method for switching network traffic, including: receiving, by a network switch via an optical connection, network traffic to be forwarded; determining, by the network switch, a link layer protocol of the received network traffic, wherein the determined link layer protocol is one of a plurality of link layer protocols supported by the switch; and forwarding, by the network switch, the network traffic to a destination network device according to the determined link layer protocol.

[0081] Example 15 includes the subject matter of Example 14, and wherein receiving the network traffic includes receiving network traffic via an optical connection providing one quarter of the total bandwidth of the link.

[0082] Example 16 includes the subject matter of any of Examples 14 and 15, and wherein receiving the network traffic includes receiving the network traffic from a sled coupled to the optical connection.

[0083] Example 17 includes the subject matter of any of Examples 14 - 16, and wherein receiving the network traffic from a sled includes receiving network traffic from at least one of a compute sled including one or more processors, a storage sled including one or more data storage devices, an accelerator sled including one or more coprocessors or field programmable gate arrays, or a memory sled including one or more memory devices.

[0084] Example 18 includes the subject matter of any of Examples 14 - 17, and wherein receiving the network traffic includes receiving the network traffic from another network switch.

[0085] Example 19 includes the subject matter of any of Examples 14 - 18, and wherein receiving the network traffic from another network switch includes receiving the network traffic from a leaf switch or a spine switch in a leaf - spine network architecture.

[0086] Example 20 includes the subject matter of any of Examples 14 - 19, and further includes: forwarding, by a network switch, network traffic of two or more different link layer protocols.

[0087] Example 21 includes the subject matter of any of Examples 14 - 20, and wherein determining the link layer protocol of the received network traffic includes: determining whether the link layer protocol is an Ethernet protocol, a high performance computing (HPC) protocol, or another dedicated communication protocol.

[0088] Example 22 includes the subject matter of any of Examples 14 - 21, and wherein forwarding the network traffic according to the determined link layer protocol includes: determining a destination address of the network traffic according to the determined link layer protocol.

[0089] Example 23 includes the subject matter of any of Examples 14 - 22, and wherein forwarding the network traffic includes forwarding the network traffic through another optical connection to one of a skateboard or another network switch.

[0090] Example 24 includes the subject matter of any of Examples 14 - 23, and wherein forwarding the network traffic includes forwarding the network traffic to one of a leaf switch or a spine switch in a leaf - spine network architecture.

[0091] Example 25 includes the subject matter of any of Examples 14 - 24, and wherein forwarding the network traffic includes forwarding the network traffic through an optical connection providing one - quarter of the total bandwidth of the link.

[0092] Example 26 includes the subject matter of any of Examples 14 - 25, and wherein forwarding the network traffic includes forwarding the network traffic to at least one of a compute skateboard including one or more processors, a storage skateboard including one or more data storage devices, an accelerator skateboard including one or more coprocessors or field - programmable gate arrays, or a memory skateboard including one or more memory devices.

[0093] Example 27 includes one or more machine - readable storage media including a plurality of instructions stored thereon, which when executed cause a network switch to perform the method of any of Examples 14 - 26.

[0094] Example 28 includes a network switch including: one or more processors; a communication circuit coupled to the one or more processors; and one or more memory devices storing a plurality of instructions, which when executed cause the network switch to perform the method of any of Examples 14 - 26.

[0095] Example 29 includes a network switch including components for performing the method of any of Examples 14 - 26.

[0096] Example 30 includes a network switch including: a network communicator circuit for receiving network traffic to be forwarded through an optical connection; a protocol determiner circuit for determining a link layer protocol of the received network traffic, wherein the received network traffic is formatted according to one of a plurality of link layer protocols; and a network traffic switcher circuit for forwarding the network traffic to a destination network device according to the determined link layer protocol.

[0097] Example 31 includes the subject matter of Example 30, and wherein receiving the network traffic includes receiving network traffic via an optical connection that provides one quarter of the total bandwidth of the link.

[0098] Example 32 includes the subject matter of any one of Examples 30 and 31, and wherein receiving the network traffic includes receiving the network traffic from a sled coupled to the optical connection.

[0099] Example 33 includes the subject matter of any one of Examples 30 - 32, and wherein receiving the network traffic from the sled includes receiving network traffic from at least one of a compute sled that includes one or more processors, a storage sled that includes one or more data storage devices, an accelerator sled that includes one or more coprocessors or field programmable gate arrays, or a memory sled that includes one or more memory devices.

[0100] Example 34 includes the subject matter of any one of Examples 30 - 33, and wherein receiving the network traffic includes receiving the network traffic from another network switch.

[0101] Example 35 includes the subject matter of any one of Examples 30 - 34, and wherein receiving the network traffic from another network switch includes receiving the network traffic from a leaf switch or a spine switch in a leaf - spine network architecture.

[0102] Example 36 includes the subject matter of any one of Examples 30 - 35, and wherein the network traffic switch circuit further forwards traffic of two or more different link - layer protocols.

[0103] Example 37 includes the subject matter of any one of Examples 30 - 36, and wherein determining the link - layer protocol of the received network traffic includes: determining whether the link - layer protocol is an Ethernet protocol, a high - performance computing (HPC) protocol, or some other proprietary communication protocol.

[0104] Example 38 includes the subject matter of any one of Examples 30 - 37, and wherein forwarding the network traffic according to the determined link - layer protocol includes: determining the destination address of the network traffic according to the determined link - layer protocol.

[0105] Example 39 includes the subject matter of any one of Examples 30 - 38, and wherein forwarding the network traffic includes forwarding the network traffic via another optical connection to one of a sled or another network switch.

[0106] Example 40 includes the subject matter of any one of Examples 30 - 39, and wherein forwarding the network traffic includes forwarding the network traffic to one of a leaf switch or a spine switch in a leaf - spine network architecture.

[0107] Example 41 includes the subject matter of any of Examples 30-40, and wherein forwarding the network traffic includes forwarding the network traffic over an optical connection that provides one quarter of the total bandwidth of the link.

[0108] Example 42 includes the subject matter of any of Examples 30-41, and wherein forwarding the network traffic includes forwarding network traffic to at least one of a compute sled that includes one or more processors, a storage sled that includes one or more data storage devices, an accelerator sled that includes one or more coprocessors or field programmable gate arrays, or a memory sled that includes one or more memory devices.

[0109] Example 43 includes a network switch that includes: circuitry for receiving network traffic to be forwarded over an optical connection; means for determining the link layer protocol of the received network traffic, wherein the received network traffic is formatted according to one of a plurality of link layer protocols; and means for forwarding the network traffic to a destination network device according to the determined link layer protocol.

[0110] Example 44 includes the subject matter of Example 43, and wherein the circuitry for receiving the network traffic includes circuitry for receiving network traffic over an optical connection that provides one quarter of the total bandwidth of the link.

[0111] Example 45 includes the subject matter of any of Examples 43 and 44, and wherein the circuitry for receiving the network traffic includes circuitry for receiving the network traffic from a sled coupled to the optical connection.

[0112] Example 46 includes the subject matter of any of Examples 43-45, and wherein the circuitry for receiving the network traffic from a sled includes circuitry for receiving network traffic from at least one of a compute sled that includes one or more processors, a storage sled that includes one or more data storage devices, an accelerator sled that includes one or more coprocessors or field programmable gate arrays, or a memory sled that includes one or more memory devices.

[0113] Example 47 includes the subject matter of any of Examples 43-46, and wherein the circuitry for receiving network traffic includes circuitry for receiving network traffic from another network switch.

[0114] Example 48 includes the subject matter of any of Examples 43-47, and wherein the circuitry for receiving the network traffic from another network switch includes circuitry for receiving the network traffic from a leaf switch or a spine switch in a leaf-spine network architecture.

[0115] Example 49 includes the subject matter of any of Examples 43 - 48, and further includes: components for forwarding network traffic of two or more different link layer protocols.

[0116] Example 50 includes the subject matter of any of Examples 43 - 49, and wherein the components for determining the link layer protocol of the received network traffic include: components for determining whether the link layer protocol is an Ethernet protocol, a High Performance Computing (HPC) protocol, or another dedicated communication protocol.

[0117] Example 51 includes the subject matter of any of Examples 43 - 50, and wherein the components for forwarding the network traffic according to the determined link layer protocol include: components for determining the destination address of the network traffic according to the determined link layer protocol.

[0118] Example 52 includes the subject matter of any of Examples 43 - 51, and wherein the components for forwarding the network traffic include components for forwarding network traffic through another optical connection to one of a sled or another network switch.

[0119] Example 53 includes the subject matter of any of Examples 43 - 52, and wherein the components for forwarding the network traffic include components for forwarding the network traffic to one of a leaf switch or a spine switch in a leaf - spine network architecture.

[0120] Example 54 includes the subject matter of any of Examples 43 - 53, and wherein the components for forwarding the network traffic include components for forwarding network traffic through an optical connection that provides one - quarter of the total bandwidth of the link.

[0121] Example 55 includes the subject matter of any of Examples 43 - 54, and wherein the components for forwarding the network traffic include components for forwarding network traffic to at least one of a compute sled including one or more processors, a storage sled including one or more data storage devices, an accelerator sled including one or more coprocessors or field - programmable gate arrays, or a memory sled including one or more memory devices.

[0122] Example 56 includes a data center that includes: a plurality of racks, each rack including a plurality of sleds; one or more multimode optical switches that are coupled to the sleds through optical connections, wherein the racks do not include switches at the top of the racks.

[0123] Example 57 includes the subject matter of Example 56, and wherein one or more switches include a plurality of switches, and each switch is connected to each sled through an upstream optical connection and a downstream optical connection.

[0124] Example 58 includes the subject matter of any of Examples 56 and 57, and each optical connection provides one quarter of the total bandwidth of the switch links.

[0125] Example 59 includes the subject matter of any of Examples 56 - 58, and a first subgroup of the skateboards will communicate using a first link layer protocol; and a second subgroup of the skateboards will communicate using a second link layer protocol different from the first link layer protocol; and one or more switches will exchange network traffic between multiple skateboards using at least the first link layer protocol and the second link layer protocol simultaneously.

[0126] Example 60 includes the subject matter of any of Examples 56 - 59, and the first link layer protocol is a non - Ethernet protocol, while the second link layer protocol is an Ethernet protocol.

[0127] Example 61 includes the subject matter of any of Examples 56 - 60, and one or more switches include a plurality of switches arranged in a leaf - spine architecture.

[0128] Example 62 includes the subject matter of any of Examples 56 - 61, and each skateboard includes one or more physical resources, the one or more switches include four switches, each skateboard is coupled to each of the four switches, and each physical resource of each skateboard is coupled to the four switches.

[0129] Example 63 includes the subject matter of any of Examples 56 - 62, and one or more switches are arranged in a two - tier switch architecture.

[0130] Example 64 includes the subject matter of any of Examples 56 - 63, and at least one switch is a spine switch connected to each skateboard with one quarter of the total switch link bandwidth.

[0131] Example 65 includes the subject matter of any of Examples 56 - 64, and the spine switch is additionally connected to one or more other connections with the total switch link bandwidth.

[0132] Example 66 includes the subject matter of any of Examples 56 - 65, and at least one spine switch is a plurality of spine switches.

[0133] Example 67 includes a data center, which includes: a layer of spine switches; a plurality of racks, where each rack includes a plurality of skateboards, and each skateboard will connect a plurality of other skateboards to the layer of spine switches.

[0134] Example 68 includes a data center that includes a two-tier switch system that includes: a spine switch tier; and a leaf switch tier connected to the spine switch tier; and a plurality of racks, where each rack includes a plurality of sled, and where each sled will connect a plurality of other sleds to the two-tier switch system.

Claims

1. An on-chip system (SoC) network switch integrated circuit (IC) for use in forwarding network traffic according to multiple different link layer communication protocols, the network traffic to be forwarded via a physical port of the SoC network switch IC to an optical communication channel and from there to a destination network device, the SoC network switch IC comprising: Network communication circuitry for use in (1) receiving multi-protocol inbound network traffic to the SoC network switch IC and (2) transmitting multi-protocol outbound network traffic from the SoC network switch IC when the SoC network switch IC is in operation; Network traffic switching circuitry for use in determining a forwarding address for outbound network traffic; Wherein, when the SoC network switch IC is in the operation: Based on frame data extracted by the SoC network switch IC according to one or more of the multiple different link layer communication protocols determined by the SoC network switch IC to correspond to the network traffic, the SoC network switch IC is to determine (1) a forwarding address of the destination network device to which the network traffic is to be forwarded by the SoC network switch IC corresponding to the physical port, (2) communication protocol related code, and (3) rules; The forwarding address, the code, and the rules are to be used by the SoC network switch IC to determine the forwarding of the network traffic by the SoC network switch IC; Format the network traffic according to one or more of the multiple different link layer communication protocols; The network communication circuitry is configurable to transmit the network traffic via the physical port; The optical communication channel is to carry a portion of the total bandwidth of the communication link; The portion is configurable as a quarter of the total bandwidth; and At least part of the telemetry information related to the SoC network switch IC is to be provided by the SoC network switch IC for use in quality of service management and / or software-defined infrastructure management; and Additionally, wherein: The SoC network switch IC is included in a multi-chip package; The SoC network switch IC further includes a processor and a physical storage device; When the SoC network switch IC is in the operation: The physical storage device is to store protocol data to be used by the processor in the operation of the SoC network switch IC; and The processor is to dynamically access other storage devices on an as-needed basis; and The other storage devices are physically decoupled from the processor and the physical storage device and are also coupled to the processor.

2. The SoC network switch IC according to claim 1, wherein: The SoC network switch IC is configurable for use with a circuit board for use in a rack.

3. The SoC network switch IC according to claim 2, wherein: The SoC network switch IC is configurable to be coupled via the physical port to an optical transceiver of the circuit board for coupling to an optical switch infrastructure.

4. The SoC network switch IC according to claim 3, wherein: The optical switch infrastructure includes one or more of the following: Optical fabric; Leaf switch; and Spine switch.

5. The SoC network switch IC according to claim 1, wherein: One or more of the multiple different link layer communication protocols include the Ethernet protocol and another different link layer protocol.

6. The SoC network switch IC according to claim 5, wherein: The another different link layer protocol includes the full path protocol.

7. At least one non-transitory machine-readable storage medium storing instructions for execution by a system-on-chip (SoC) network switch integrated circuit (IC), the SoC network switch IC being used for forwarding network traffic according to multiple different link layer communication protocols when the SoC network switch IC is in operation, the network traffic being to be forwarded via a physical port associated with the SoC network switch IC to an optical communication channel and from there to a destination network device, the instructions, when executed by the SoC network switch IC, causing the SoC network switch to be configured to perform operations including: Using the network communication circuitry of the SoC network switch IC to receive multi-protocol inbound network communication destined for the SoC network switch IC; Using the network communication circuitry to transmit multi-protocol outbound network communication from the SoC network switch IC; Using the network traffic switching circuitry of the SoC network switch IC in outbound network communication forwarding address determination; And Based on frame data accessed by the SoC network switch IC according to one or more of the multiple different link layer communication protocols determined by the SoC network switch IC to correspond to the network traffic, identifying (1) the forwarding address of the destination network device to which the network traffic is to be forwarded by the SoC network switch IC corresponding to the physical port, (2) communication protocol-related code, and (3) rules; and Wherein, when the SoC network switch IC is in the operation: The forwarding address, the code, and the rules are to be used by the SoC network switch IC to determine the forwarding of the network traffic by the SoC network switch IC; Formatting the network traffic according to one or more of the multiple different link layer communication protocols; The network communication circuitry is configurable to transmit the network traffic via the physical port; The optical communication channel is to carry a portion of the total bandwidth of the communication link; The portion is configurable to be one quarter of the total bandwidth; and At least part of the telemetry information related to the SoC network switch IC is to be provided by the SoC network switch IC for use in quality of service management and / or software-defined infrastructure management; and Additionally, wherein: The SoC network switch IC is included in a multi-chip package; The SoC network switch IC further includes a processor and a physical storage device; When the SoC network switch IC is in the operation: The physical storage device is to store protocol data to be used by the processor in the operation of the SoC network switch IC; and The processor is to dynamically access other storage devices on an as-needed basis; and The other storage devices are physically decoupled from the processor and the physical storage device and are also coupled to the processor.

8. The at least one non-transitory machine-readable storage medium of claim 7, wherein: The SoC network switch IC is configurable for use in a circuit board for use in a rack.

9. The at least one non-transitory machine-readable storage medium of claim 8, wherein: The SoC network switch IC is configurable to be coupled via the physical port to an optical transceiver of the circuit board for coupling to an optical switch infrastructure.

10. The at least one non-transitory machine-readable storage medium of claim 9, wherein: The optical switch infrastructure includes one or more of the following: Optical fabric; Leaf switch; and Spine switch.

11. The at least one non-transitory machine-readable storage medium of claim 7, wherein: One or more of the plurality of different link layer communication protocols includes an Ethernet protocol and another different link layer protocol.

12. The at least one non-transitory machine-readable storage medium of claim 11, wherein: The another different link layer protocol includes a full path protocol.

13. A method at least partially implemented by a system-on-chip (SoC) network switch integrated circuit (IC), the SoC network switch IC for use in forwarding network traffic according to a plurality of different link layer communication protocols when the SOC network switch IC is in operation, the network traffic to be forwarded via a physical port associated with the SoC network switch IC to an optical communication channel and from there to a destination network device, the method comprising: Using network communication circuitry of the SoC network switch IC to receive multi-protocol inbound network communication destined for the SoC network switch IC; Using the network communication circuitry to transmit multi-protocol outbound network communication from the SoC network switch IC; Using network traffic switching circuitry of the SoC network switch IC in outbound network communication forwarding address determination; And Based on frame data accessed by the SoC network switch IC according to one or more of the plurality of different link layer communication protocols determined by the SoC network switch IC to correspond to the network traffic, identifying (1) a forwarding address of the destination network device to which the network traffic is to be forwarded by the SoC network switch IC corresponding to the physical port, (2) communication protocol-related code, and (3) rules; and Wherein, when the SoC network switch IC is in the operation: The forwarding address, the code, and the rules are to be used by the SoC network switch IC to determine the forwarding of the network traffic by the SoC network switch IC. Format the network traffic according to one or more of the plurality of different link layer communication protocols; The network communication circuitry is configurable to transmit the network traffic via the physical port; The optical communication channel is to carry a portion of the total bandwidth of the communication link; The portion is configurable as one quarter of the total bandwidth; and Telemetry information that is at least partially associated with the SoC network switch IC is to be provided by the SoC network switch IC for use in quality of service management and / or software defined infrastructure management; and Additionally, wherein: The SoC network switch IC is included in a multi-chip package; The SoC network switch IC further includes a processor and a physical storage device; When the SoC network switch IC is in the operation: The physical storage device is to store protocol data that is to be used by the processor in the operation of the SoC network switch IC; and The processor is to dynamically access other storage devices on an as-needed basis; and The other storage devices are physically disaggregated from the processor and the physical storage device and are also coupled to the processor.

14. The method of claim 13, wherein: The SoC network switch IC is configurable for use with a circuit board for use in a rack.

15. The method of claim 14, wherein: The SoC network switch IC is configurable to be coupled via the physical port to an optical transceiver of the circuit board for coupling to an optical switch infrastructure.

16. The method of claim 15, wherein: The optical switch infrastructure includes one or more of the following: Optical fabric; Leaf switch; and Spine switch.

17. The method of claim 13, wherein: One or more of the plurality of different link layer communication protocols includes an Ethernet protocol and another different link layer protocol.

18. The method of claim 17, wherein: The another different link layer protocol includes a full path protocol.

19. An apparatus at least partially implemented by a system-on-chip SoC network switch integrated circuit IC for use in forwarding network traffic according to a plurality of different link layer communication protocols when the SOC network switch IC is in operation, the network traffic to be forwarded via a physical port associated with the SoC network switch IC to an optical communication channel and from there to a destination network device, the apparatus comprising: Means for receiving multi-protocol inbound network traffic destined for the SoC network switch IC using network communication circuitry of the SoC network switch IC; Means for transmitting multi-protocol outbound network traffic from the SoC network switch IC using the network communication circuitry; Means for using network traffic switching circuitry of the SoC network switch IC in outbound network traffic forwarding address determination; And A component for identifying, based on frame data accessed by the SoC network switch IC according to one or more of the plurality of different link layer communication protocols determined by the SoC network switch IC to correspond to the network traffic, the following: (1) the forwarding address of the destination network device to which the network traffic is to be forwarded by the SoC network switch IC corresponding to the physical port, (2) communication protocol-related code, and (3) rules; and wherein, when the SoC network switch IC is in the operation: the forwarding address, the code, and the rules are to be used by the SoC network switch IC to determine the forwarding of the network traffic by the SoC network switch IC; format the network traffic according to one or more of the plurality of different link layer communication protocols; the network communication circuit is configurable to transmit the network traffic via the physical port; the optical communication channel is to carry a portion of the total bandwidth of the communication link; the portion is configurable as one quarter of the total bandwidth; and at least part of the telemetry information related to the SoC network switch IC is to be provided by the SoC network switch IC for use in quality of service management and / or software-defined infrastructure management; and additionally, wherein: the SoC network switch IC is included in a multi-chip package; the SoC network switch IC further includes a processor and a physical storage device; when the SoC network switch IC is in the operation: the physical storage device is to store protocol data, which is to be used by the processor in the operation of the SoC network switch IC; and the processor is to dynamically access other storage devices on an as-needed basis; and the other storage devices are physically decoupled from the processor and the physical storage device and are also coupled to the processor.

20. The apparatus according to claim 19, wherein: the SoC network switch IC is configurable for use in a circuit board for use in a rack.

21. The apparatus according to claim 20, wherein: the SoC network switch IC is configurable to be coupled via the physical port to an optical transceiver of the circuit board for coupling to an optical switch infrastructure.

22. The apparatus according to claim 21, wherein: the optical switch infrastructure includes one or more of the following: optical fabric; leaf switch; and spine switch.

23. The apparatus according to claim 19, wherein: one or more of the plurality of different link layer communication protocols include an Ethernet protocol and another different link layer protocol.

24. The apparatus according to claim 23, wherein: the another different link layer protocol includes a full path protocol.

Citation Information

Patent Citations

  • Method and system for data traffic handling in a distributed fabric protocol (dfp) switching network architecture

    CN102821038A

  • Method and system for transmission of control data between a network controller and a switch

    CN104052574A