Interconnection Devices and Switching Systems
The interconnection device and switching system address inefficiencies in XPU processing by offloading protocols and using distributed arbitration, enhancing XPU efficiency and reducing latency and power consumption.
Patent Information
- Application Number
- JP2025508714
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-08-18
- Publication Date
- 2025-08-07
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Conventional protocols in high-performance computing systems result in reduced XPU utilization efficiency due to processing overhead, lack of QoS control among incoming packets, and inefficient data transfer, leading to increased latency and power consumption.
An interconnection device that offloads protocol processing from the XPU to a dedicated protocol processor, which directly accesses the XPU memory, and a switching system with an optical switch that performs distributed arbitration to reduce latency and power consumption.
Improves XPU utilization efficiency, reduces processing overhead, and enhances data transfer by minimizing latency and power consumption through distributed arbitration and direct memory access.
Smart Images

Figure 2025526149000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an interconnection device and a switching system for interconnecting processors and switches in a computing system. [Background technology]
[0002] High-performance computing systems are composed of a large number of computing units (XPUs), which require improved resource utilization efficiency and reduced latency. XPU is an abbreviation for different types of processing units (processors), including commonly used central processing units (CPUs) and application-oriented graphical processing units (GPUs).
[0003] This requires improvements in the speed of physical data transfer in transmission links and switching fabrics, and in the speed of data transmission protocol execution. In particular, acceleration of protocol execution is addressed at the hardware and software levels.
[0004] Conventional protocol execution will now be described with reference to Figure 18A, where the protocol is executed on packets arriving at an XPU.
[0005] In an operating system (OS) kernel 400, a device driver 401 provides a network internet card (NIC) 410 with the address of a memory queue in which incoming packets are stored in advance.
[0006] Upon the arrival of a new packet, NIC 410 sends the received packet bits to queue 404 (dotted arrow in the figure) and generates an interrupt signal at XPU 420. XPU memory 421 now connects to XPU 420. At this time, the application running on XPU 420 is suspended and a context switching is performed to save all relevant processing data to designated registers for later resumption of this application. Other processing actions, such as flushing the pipeline, may also be performed.
[0007] The post-interrupt processing of the operating system (OS) kernel 400 performs protocol steps on the received packets, such as verifying header checksums and reassembling fragments of the outgoing data flow at the IP layer 402 and TCP layer 403. When the protocol execution is complete, the application-level payload is delivered to the corresponding endpoint queue 404.
[0008] FIG. 18B shows the processing timeline of a normal XPU (no packet reception) and the timeline of an XPU performing the same processing when a packet is received.
[0009] In the normal XPU processing timeline, a first process 431 is followed by a second process 432 (430 in the figure).
[0010] On the other hand, when a packet is received, the first process 431 is interrupted and a packet reception process 441 is executed. Here, a class selector (CS) 442 is set before and after the packet process 441 (440 in the figure).
[0011] Thus, a large processing overhead is observed when receiving packets. The impact of this overhead is greater for packets with short payloads. Figure 19 shows the relationship between Ethernet bandwidth and transmitted message size on a 1 Gb / s Ethernet link. The relationship is shown for standard-sized Ethernet frames (1500 bytes, 451 in the figure) and large-sized Ethernet frames (9000 bytes, 452 in the figure). In the small message size range (less than about 512 octets), the bandwidth increases as the message size increases, but in the large message size range (512 octets or more), once the upper limit is reached, the bandwidth hardly increases at all.
[0012] In conventional protocol execution, no protocol is executed for a new incoming packet until the running protocol is completed.
[0013] The duration of the short payload of an Ethernet packet is only a small fraction of the protocol execution time. As shown in Figure 19, the effective bandwidth of packet reception is very low and increases with increasing payload size until it reaches an upper limit. [Prior art documents] [Non-patent literature]
[0014] [Non-Patent Document 1] David Riddoch, "Low Latency Distributed Computing," A submitted dissertation for the degree of Doctor of Philosophy, University of Cambridge 2002. Summary of the Invention [Problem to be solved by the invention]
[0015] As mentioned above, the execution of conventional protocols poses the problem of reduced XPU utilization efficiency.
[0016] Furthermore, in the implementation of conventional protocols, the processing of incoming packets is done on a first-come, first-served basis without priority, which causes a problem in computing overhead due to the lack of QoS control among incoming packets.
[0017] For example, when a loss of QoS control occurs, a high priority application will wait longer for execution of a lower priority protocol that arrives first.
[0018] Also, if one or more packets of a data flow are lost or if a queue overflow occurs before the flow reception is complete, a long time will be spent executing the protocol for packets that will never be used by the XPU. [Means for solving the problem]
[0019] In order to solve the above-mentioned problems, the interconnection device of the present invention is an interconnection device that interconnects an XPU connected to an XPU memory and a switching device, and is equipped with a protocol processor that executes protocols offloaded from the XPU and directly accesses the XPU memory. [Effects of the Invention]
[0020] According to the present invention, it is possible to provide an interconnection device and a switching system that can improve the utilization efficiency of XPUs and execute protocols. [Brief explanation of the drawings]
[0021] [Figure 1A] FIG. 1A is a schematic diagram showing the configuration of an interconnection device according to a first embodiment of the present invention. [Figure 1B] FIG. 1B is a diagram for explaining the operation of the interconnection device according to the first embodiment of the present invention. [Figure 2] FIG. 2 is a diagram for explaining the operation of the interconnection device according to the first embodiment of the present invention. [Figure 3] FIG. 3 is a schematic diagram showing the configuration of a switching system according to the second embodiment of the present invention. [Figure 4] FIG. 4 is a schematic diagram showing the basic configuration of a switching device 10 according to the second embodiment of the present invention. [Figure 5A] FIG. 5A is a diagram for explaining the basic operation of the switching device 10 in the second embodiment of the present invention. [Figure 5B] FIG. 5B is a diagram for explaining the basic operation of a conventional switching device. [Figure 6A] FIG. 6A is a diagram for explaining the basic operation of the switching device 10 in the second embodiment of the present invention. [Figure 6B] FIG. 6B is a diagram for explaining the basic operation of the switching device 10 in the second embodiment of the present invention. [Figure 6C] FIG. 6C is a diagram for explaining the basic operation of the switching device 10 in the second embodiment of the present invention. [Figure 7] FIG. 7 is a diagram for explaining the basic operation of a conventional switching device. [Figure 8] FIG. 8 is a schematic diagram showing the basic configuration of a switching device 30 according to the second embodiment of the present invention. [Figure 9] FIG. 9 is a diagram for explaining the basic operation of the switching device 30 in the second embodiment of the present invention. [Figure 10A] FIG. 10A is a diagram for explaining the basic operation of the switching device 30 in the second embodiment of the present invention. [Figure 10B] FIG. 10B is a diagram for explaining the basic operation of the switching device 30 in the second embodiment of the present invention. [Figure 10C] FIG. 10C is a diagram for explaining the basic operation of the switching device 30 in the second embodiment of the present invention. [Figure 10D] FIG. 10D is a diagram for explaining the basic operation of the switching device 30 in the second embodiment of the present invention. [Figure 11] FIG. 11 is a diagram for explaining the basic operation of the switching device 30 in the second embodiment of the present invention. [Figure 12A] FIG. 12A is a diagram for explaining the effect of the switching device 30 according to the second embodiment of the present invention. [Figure 12B] FIG. 12B is a diagram for explaining the effect of the switching device 30 in the second embodiment of the present invention. [Figure 13] FIG. 13 is a schematic diagram showing the basic configuration of a switching system 40 according to the second embodiment of the present invention. [Figure 14] FIG. 14 is a schematic diagram showing the basic configuration of a switching system 50 according to the second embodiment of the present invention. [Figure 15] FIG. 15 is a diagram for explaining the basic operation of the switching system 50 in the second embodiment of the present invention. [Figure 16] FIG. 16 is a schematic diagram showing the configuration of a switching system according to a first embodiment of the present invention. [Figure 17] FIG. 17 is a schematic diagram showing an example of the configuration of a switching system according to the first embodiment of the present invention. [Figure 18A] FIG. 18A is a diagram for explaining a conventional switching system. [Figure 18B] FIG. 18B is a diagram for explaining a conventional switching system. [Figure 19] FIG. 19 is a diagram for explaining a conventional switching system. DETAILED DESCRIPTION OF THE INVENTION
[0022] First Embodiment An interconnection device according to a first embodiment of the present invention will be described with reference to FIGS. 1A to 2. FIG.
[0023] <Configuration of interconnection device> As shown in FIG. 1A, the interconnection device 100 of this embodiment has a processor dedicated to protocol processing (hereinafter referred to as the "protocol processor") 110, and is connected to an XPU 120 and a memory (hereinafter referred to as the "XPU memory") 121 connected to the XPU 120.
[0024] In this way, in the interconnection device 100, protocol processing is offloaded from the XPU 120 to the protocol processor 110.
[0025] The protocol processor 110 is connected to the XPU 120 and the XPU memory 121 .
[0026] The protocol processor 110 also connects to the NIC 130 .
[0027] The payload of the incoming packet is extracted by the protocol processor 110 and made applicable to the XPU 120 .
[0028] When a packet is sent during the execution of a process, the XPU 120 processes the packet's payload in a timely manner without context switching or other conventional overhead processing steps. Figure 1B shows a processing timeline 103 for the XPU 120 and a processing timeline 104 for the protocol processor 110. For example, as shown in Figure 1B, processing 102 of the packet's payload is performed after the completion of processing 101 of a given process without interrupting processing 101.
[0029] The protocol processor 110 has the following features:
[0030] First, the protocol processor 110 must have a sufficiently high processing capability. For example, the protocol processor 110 must be able to operate at high speed. If the protocol processor 110 operates slowly, the time spent on protocol processing increases, which reduces the advantage of using a separate device as the protocol processor 110.
[0031] In a computing system including a variety of target XPUs 120, allocating a high-performance protocol processor 110 to each XPU 120 simultaneously increases power consumption and costs. Therefore, it is necessary to trade off the number of processors to be allocated with power consumption and costs, taking into account the operating conditions of the computing system.
[0032] Second, protocol processor 110 needs to be able to directly access XPU memory 121. This is because if protocol processor 110 cannot directly access XPU memory 121, it will disrupt the operation process of XPU 120 and cause unnecessary overhead.
[0033] In the interconnection device 100, a PCIe (Peripheral Component Interconnect-Express) subsystem is used as an example of a framework that allows the protocol processor 110 to directly access the XPU memory 121.
[0034] As shown in Figure 2, in the PCIe subsystem, a Root Complex (RC) 140 is a main device that connects the XPU 120 and the XPU memory 121 to different endpoints. For example, a NIC is connected to the RC as an endpoint of the RC, and RDMA (remote direct memory access) is configured between the XPU 120 and the NIC, which are located remotely. Here, the PCIe protocol has a simple configuration with only three layers: transaction, link, and physical layers.
[0035] For example, as shown in FIG. 2, a transaction is composed of a memory write, a memory read, and a completion of data (CplD).
[0036] In this transaction, first, the XPU (driver) notifies the NIC that it is preparing to send a message (S11 in the figure).
[0037] Next, memory read and data completion are transmitted between the RC and the NIC, and memory write is performed to indicate completion of the executed process (S12 in the figure). In this way, a series of processes are executed together in a transaction.
[0038] Next, when the NIC receives the payload, it transmits the data to the network. If the transmission is successful, the NIC receives an acknowledgment (ACK) from the destination (S13 in the figure).
[0039] Finally, when the NIC receives the ACK, it directly accesses the XPU memory via the RC and writes a transmission completion message (S14 in the figure).
[0040] Additionally, memory polling is used for XPU 120 and XPU Memory 121. XPU 120 periodically checks designated memory areas to avoid context switching and other overhead processing steps. However, XPU 120 may optionally be interruptible and access input data directly without using memory polling.
[0041] Additionally, implementations that replace round-trip transactions with unidirectional transactions can further reduce latency.
[0042] According to the interconnection device of this embodiment, by offloading protocol processing from the XPU, it is possible to promote the utilization efficiency of the XPU as a diverse computing resource.
[0043] <Second embodiment> A switching system according to a second embodiment of the present invention will be described with reference to FIGS.
[0044] <Switching system configuration> A switching system 200 according to this embodiment includes a switching system 210 and the interconnection device 100 according to the first embodiment.
[0045] As shown in FIG. 3, the switching system 210 includes an optical switch 14 and a plurality of control groups 211_10 to 211_m0.
[0046] In each control group, a transmission group consisting of devices such as an optical transmitter 13 for signal transmission, an electric switch, and a processor, and a reception group consisting of devices such as an optical receiver 15 for signal reception, an electric switch, and a processor are mounted together, for example, on the same mounting board.
[0047] Here, the electrical switch 212 is a single ASIC chip, arranged in groups, that controls all of the switching and data management described above. Here, the ASIC chip 212 functions as a control unit (described below), a transmitting electrical switch, and a receiving electrical switch.
[0048] The basic configuration and operation of switching system 210 will be described below with reference to Figures 4 to 15. For ease of explanation, an example will be described in which a transmission group and a reception group are separately arranged on each of the input and output sides of the optical switch.
[0049] As a basic configuration of the switching system 210, for example, the switching systems 40 and 50 include a switching device 10 and 30, a plurality of sending groups (for example, 3_10 to 3_40), and a plurality of receiving groups (for example, 4_20 and 4_30).
[0050] <Configuration of Switching Device 10> First, the switching device 10 will be described with reference to FIGS.
[0051] As shown in FIG. 4, the switching device 10 includes an optical transmitter 13, an optical switch 14, an optical receiver 15, and a control unit 17.
[0052] Packet 1 is sent from the sending host 3, compressed and divided (packet 2) by the optical transmitter 13, switched by the optical switch 14, and received by the receiving host 4 via the optical receiver 15. Here, a priority is set for packet 1 sent from the sending host 3.
[0053] The optical switch 14 is an all-through switch, and does not require arbitration for switching itself. In the optical switch 14, all packets destined for the same output port are sent simultaneously. Therefore, the destination output port receives different packets at the same time.
[0054] The control unit 17 is an electric chip, and is arranged so as to be connected to the optical receiver 15. The control unit 17 also performs processing such as arbitration on the packets 2 output from the optical switch .
[0055] A control unit 17 is assigned locally to each receiving host 4 and performs self-management of traffic entering the receiving host 4. The control unit 17 is located close to the receiving host 4 and quickly updates data retrieval priorities.
[0056] Furthermore, the control unit 17 manages the processing of communications from a single receiving host 4 and can operate at high speed.
[0057] The information channel, in addition to the data link, connects the receiving host 4 and the control unit 17 .
[0058] <Operation of the Switching Device 10> Figure 5A shows packet processing in switching device 10. For comparison, Figure 5B shows a conventional non-blocking switch 20.
[0059] In a conventional electrical non-blocking switch 20, the bandwidth of all output ports is fixed, and only one packet can be transmitted at a time. For example, if flow B (101_2) and flow C (101_3) are transmitted to the same output port, they are processed by a single band with a fixed bandwidth (Fig. 5B).
[0060] On the other hand, in the optical switch 14 of the switching device 10, the bandwidth of each output port is variable, and all packets input to the switch can be adjusted. For example, as shown in Figure 4A, the bandwidth is changed and processing is performed in two bands. Here, the bandwidth of the output port is half the total throughput (bandwidth) of the switch.
[0061] In this manner, optical switch 14 is capable of supporting a high bandwidth per output port.
[0062] In the switching device 10, as shown in FIG. 5A, the low priority input data (packet C) 102_3 is buffered and held for a long period of time in the RAM of the control unit 17.
[0063] On the other hand, the data (packet B) 102_2 with a higher priority is transmitted first.
[0064] The input data (packet C) 102_3 with a lower priority is transmitted after the transmission of the data (packet B) 102_2 with a higher priority is completed.
[0065] The switching apparatus 10 may also have an optical data link located within the control unit 17, allowing optical data to be transmitted directly between the output ports of the switch and a receiving host 4 capable of processing optical input signals.
[0066] The optical switch 14 can be configured based on a broadcast-and-select system, as shown in FIGS. 6A to 6C.
[0067] Signals entering different switch ports are multiplexed, for example, at different wavelengths. Each input signal is sent to all output ports and selected at each output port based on the desired signal destination.
[0068] For example, as shown in FIG. 6A, an input packet 1_1 from a sending host A (3_1) is branched by a splitter 18, and the branched packets 2_1 to 2_4 are transmitted to receiving hosts 4_1 to 4_4.
[0069] 6B, input packets 1_1 and 1_4 from transmitting hosts A and D (3_1 and 4) are branched by splitter 18, and the branched packets 2_1 to 2_4 are selected by optical selection filter 19 at the output port and transmitted to receiving hosts 4_1 to 4_4. Here, a fast tunable filter or a polarizing filter element can be used as the optical selection filter 19.
[0070] Also, as shown in FIG. 6C, input packets 1_1 and 1_4 from transmitting hosts A and D (3_1, 4) are each split by a splitter 18, and the branched packets 2_1 to 2_4 are received by multiple optical receivers 15 for each packet and transmitted to receiving hosts 4_1 to 4_4.
[0071] <Effects> In the conventional non-blocking packet switch 20, a data packet input to any input port is switched to the desired output port.
[0072] When multiple packets are simultaneously transmitted to the same output port, contention occurs and arbitration is performed for the colliding packets: the higher priority packet is selected for transmission first, while other packets are buffered and transmitted subsequently.
[0073] In this way, in the conventional non-blocking packet switch 20, arbitration is required for scheduling when packets are sent simultaneously to the same destination.
[0074] Typically, the arbitration process follows a centralized control scheme, where information regarding data availability, priority, and selection is collected in a central control unit (not shown) of the system before a decision is made in arbitration. The accuracy of this decision highly depends on the availability of all necessary information that has been updated recently. However, in dynamic computing systems, it is difficult to maintain the accuracy of the recently updated information.
[0075] 7, a receiving host 4 is assigned a computational task to reduce two data flows, flow A (201_1) and flow B (201_2). The host 4 has already processed flow A and is waiting to receive flow B.
[0076] On the other hand, flow C (201_3) is another flow sent to host 4, and arrives at the switch before flow B (201_2) arrives with a time difference. If the time difference is longer than 0 (zero), the switch sends flow C (201_3) to host 4.
[0077] Here, the time difference is represented as ΔT.
[0078] When the optical switch 14 is operated, the transmission of flow B does not start until the transmission of flow C is completely completed.
[0079] When operating the electrical switch, packet arbitration begins between flow C (201_3) and flow B (201_2) when flow B (201_2) arrives. Priority is given to flow B (201_2) only if the arbitrator 22 of the switch has already been notified. If there is a failure or delay in the arbitrator 22 being notified about the host 4's request to give top priority to flow B (201_2), flow B will not be switched fast enough even if the time difference is 0 (zero).
[0080] Since the host 4 is a processing unit, the priorities for retrieving data change rapidly. It is therefore difficult to continuously update the arbitrator 22 with these changing priorities. Therefore, in handling the traffic of all systems, it is difficult for the arbitrator 22 to make the right decision fast enough for the large amount of highly dynamic data.
[0081] As described above, conventional non-blocking packet switches 20 perform arbitration using a centralized control method, and as the number of switch ports and processing power (throughput) increase, the process becomes more complex, resulting in increased latency and power consumption.
[0082] Furthermore, it is difficult to collect information necessary for the arbitration process, such as priority, according to a predetermined rule as the scale of the system increases.
[0083] On the other hand, in the switching device 10, arbitration is performed in a distributed control manner by the output control unit 17. Here, the host 4 connected to each output port decides which packet to process first. In this way, the control unit 17 can perform arbitration locally for each single host.
[0084] All packets are output with a predetermined duration T. In other words, the data rate of the output signal is the same as the data rate of the input signal.
[0085] In this way, packets arriving at the same output port in different slots are converted to the initial (original) data rate, and packets are output in the order requested by the connected output hosts, thus giving the highest priority packets the lowest latency.
[0086] As a result, the switching device 10 can reduce latency and power consumption in switching, and reduce (or eliminate) the burden of collecting information required for the arbitration process.
[0087] <Configuration of Switching Device 30> Next, the switching device 30 will be described with reference to FIGS. 8 to 12B.
[0088] 8, an example of a switching device (packet switch) 30 includes, in order, an input port 11, an input block 12, an optical transmitter 13, an optical switch 14, an optical receiver 15, and an output port 16. It also includes a control unit 17 connected to the optical receiver 15. In the switching device 30 according to this embodiment, the optical switch 14 operates in a time slot manner.
[0089] <Optical switch operation> The operation of the switching device (packet switch) 30 in this embodiment will be described with reference to FIG.
[0090] 9 shows the basic operation of a packet switch 30 that performs non-blocking processing, taking a 4×4 switch as an example. This packet switch 30 is based on a time slot operation, which will be described below.
[0091] First, packets are switched by input group, where arbitration is performed only among input packets of the same input group. This arbitration is performed for a small number of ports and low traffic, so it is fast.
[0092] Furthermore, an electrical packet 1 input to the switch has a bandwidth BW (bit / sec) and a duration T, and a desired output port 16 to which it is to be sent is set.
[0093] In the packet switch 30, a packet switching operation to any of the four output ports 16 is completed in time T. This is because if switching a single packet takes longer than T, the next incoming packet will be blocked, and continuous switching delays will accumulate.
[0094] Also, the input packet 1 is compressed by a factor (here 4) equal to the number of ports (i.e., the number of optical receivers to which the packet is sent) at the optical transmitter 13 to fit into the time slot. That is, the duration of the input packet is divided by a factor equal to the number of ports, which becomes T / 4. Also, in order to preserve the packet data content, the bandwidth is multiplied by the same factor, which becomes 4BW.
[0095] In this way, an input packet of light 2 is generated that satisfies these conditions.
[0096] Each optical input packet 2 is then distributed to a desired output port 16 in a periodic time slot by an optical switch 14. Here, the periodic operation of the switch is divided into four time slots.
[0097] Each time slot has a duration Δt of T / 4.
[0098] The distribution (switching) of the optical input packet 2 is repeatedly performed for each time slot according to a sequence made up of steps S1 to S4 (described later).
[0099] Finally, the packets are converted into electrical packets by the optical receiver 15 and output at a fixed duration from the packet switch 30. In other words, the data rate of the signal output from the packet switch 30 is the same as the data rate of the signal input thereto.
[0100] In this way, packets arriving in different time-reduced time slots will have their data rate changed to the original data rate, in the order of arrival of the packets, or by other arbitration prioritizing the change.
[0101] The switching operation of the optical switch 14 described above will be described with reference to Figures 10A to 10D. Each of Figures 10A to 10D shows an example of a series of switching operations in steps S1 to S4.
[0102] Packets are input to each of the four ports 11_1 to 11_4 in the packet switch 30. The packet input to the port 11_3 (packet C) has the highest priority and its desired output port is the port 16_3.
[0103] The packets (packets A and D) input to the ports 11_1 and 11_4 have the desired output ports 16_2 and 16_1, respectively, and have the second priority.
[0104] The packet (packet B) input to the port 11_2 has the desired output port as the port 16_4 and has the third priority.
[0105] First, packet C has the highest priority and is therefore transmitted to output port 16_3 for the duration of the first time slot (step S1, FIG. 10A).
[0106] Next, since packets A and D have the second priority, they are transmitted simultaneously to output ports 16_2 and 16_1, respectively, for the duration of the second time slot (step S2, FIG. 10B). Here, packets A and D are transmitted to different output ports, so no collision occurs.
[0107] Next, the packet input to port B (packet B) has the third priority and is therefore transmitted to output port 16_4 with the duration of the third time slot (step S3, FIG. 10C).
[0108] Finally, since the transmission (switching) of packets A to D has been completed in the previous step (step 3), no switching is performed during the duration of the fourth time slot (step S4, FIG. 10D).
[0109] Here, the duration of the first time slot, the duration of the second time slot, the duration of the third time slot, and the duration of the fourth time slot are represented by Δt1, Δt2, Δt3, and Δt4, respectively.
[0110] In this way, when the operation cycle (four steps) is completed, all input packets are switched to their desired output ports simultaneously in a non-blocking manner.
[0111] In this switching operation, in every step, each output port 16 is connected to only one input port 11, as shown in Figures 10A to 10D. Also, the input port 11 is connected to the desired output port 16, and in switching a packet input to the input port 11, the packet is placed in the correct (correct) time slot (with a divided duration).
[0112] Moreover, the optical switch 14 operates as shown in FIG.
[0113] Packets (packets A to D) are input to each of four ports 11_1 to 11_4 in the packet switch 30. Packets A to D have the same desired output port (16_2), and packets B, A, D, and C are prioritized in this order.
[0114] First, packet B has the highest priority and is therefore transmitted to output port 16_2 for the duration of the first time slot (step S1).
[0115] Next, since packet A has the second priority, it is transmitted to the output port 16_2 for the duration of the second time slot (step S2).
[0116] Next, since packet D has the third priority, it is transmitted to the output port 16_2 for the duration of the third time slot (step S3).
[0117] Finally, since packet C has the fourth priority, it is transmitted to output port 16_2 for the duration of the fourth time slot (step S4).
[0118] In this way, when the operation cycle (four steps) is completed, all input packets are simultaneously switched to the desired output ports in a non-blocking manner. Here, packets A to D are transmitted in different time slots, so no collisions occur.
[0119] Thus, in packet switch 30, all input packets destined for the same output port are correctly (accurately) switched to that port at time T.
[0120] <Effects> The effects of the switching device 30 in this embodiment will be described below.
[0121] The optical switch in the first embodiment has problems such as a decrease in signal output level with an increase in the number of ports (FIG. 6A), and an increase in the number of constituent units such as optical selection filters 19, such as high-speed wavelength selection filters, and optical receivers 15 (FIGS. 6B and 6C). In particular, controlling a large number of high-speed selection units is technically difficult and increases power consumption.
[0122] By introducing time slots and increasing the bit rate, the optical switch of this embodiment can process large amounts of highly dynamic data sufficiently quickly and reduce power consumption without reducing the signal output level with the number of ports or increasing the number of constituent units.
[0123] Furthermore, in the switching device 10 according to this embodiment, arbitration according to the conventional centralized control method is distributed among the following steps to execute arbitration.
[0124] First, packets from different input groups are sent to the same output group in different time slots in the optical switch 14. The optical switch can perform this step with high data rates and precise time control.
[0125] Arbitration is then performed by the output control unit 17 in a distributed control manner, where the host 4 connected to each output port decides which packet to process first. In this way, the control unit 17 can perform arbitration locally for each single host.
[0126] All packets are output with a predetermined duration T. In other words, the data rate of the output signal is the same as the data rate of the input signal.
[0127] In this way, packets arriving at the same output port in different slots are converted to the initial (original) data rate, and packets are output in the order requested by the connected output hosts, thus giving the highest priority packets the lowest latency.
[0128] As a result, the switching device according to this embodiment can reduce latency and power consumption in switching, and can reduce (or eliminate) the burden of collecting information necessary for the arbitration process.
[0129] Furthermore, the effects of the switching device 30 will be explained in detail in comparison with a conventional non-blocking switch.
[0130] Fig. 12A shows the latency of a flow switched by the switching device 30. For comparison, Fig. 12B shows the latency of a flow switched by the conventional non-blocking switch 20.
[0131] For example, it is assumed that a flow A (1_10) having packets A1 (1_11) to A3 (1_13) and a flow D (1_40) having packets D1 (1_41) to D3 (1_43) are input and transmitted to the same output port.
[0132] In this case, in a conventional non-blocking switch, flows A (2_10) and D (2_40) are processed in a single band, as shown in Figure 12B, resulting in an accumulation of delays. As a result, when the length (time) of one packet is T, the delay is 3T, which is the length (time) of the flow sent immediately before.
[0133] Thus, in conventional non-blocking switches, delay time increases, resulting in increased latency.
[0134] On the other hand, in the switching device 30, flows A (2_10) and D (2_40) are processed in two bands as shown in Fig. 12A, so the delay time is hardly accumulated and is less than T. This delay occurs only in the first packet and is negligible compared to the length of the flow, 3T.
[0135] In this way, the switching device 30 has a short delay time and can reduce latency.
[0136] Also, in a typical electrical switch, an input packet passes through an input port of the switch, where its destination and priority are first inspected, and then a centralized arbitration is performed to determine which packet should be sent first among all packets destined for the same output port.
[0137] The complexity of the centralized arbitration process increases as the number of switch ports and throughput increases, resulting in increased communication latency and power consumption.
[0138] On the other hand, the switching device 30 can switch packets without performing centralized arbitration, which takes a long time, and therefore can reduce communication latency and power consumption.
[0139] Furthermore, because the optical switch handles part of the switching process, it consumes less power than an ASIC using CMOS transistors and can increase the switching capacity.
[0140] Furthermore, because chiplets are used for the input block 12, the area occupied by the input block 12 can be reduced. As a result, even if an optical-electrical interface is implemented, the overall area of the packet switch (chip) does not increase. Therefore, the optical-electrical interface can be implemented without changing the chip area, and the throughput (processing capacity) of the switch can be increased. Furthermore, by using chiplets, power consumption can be reduced.
[0141] It also avoids contention between ports in the same block, allowing non-blocking processing.
[0142] Furthermore, in order to simultaneously transmit multiple packets to the same destination using a conventional packet switch, the same number of parallel optical receivers as the number of packets was required.
[0143] On the other hand, the switching device 30 creates a compact copy of each input packet at a high data rate and transmits the compact packets in short time slots. In this way, packets to the same destination can be transmitted in a time shorter than the actual packet input interval using time interleaving.
[0144] Here, the optical receiver 15 used in the switching device 30 can operate in response to such burst mode transmission.
[0145] The switching device 30 can easily implement a high-speed 4×4 optical switch device by being composed of four 1×4 switching units corresponding to different input ports 11. Here, in the 4×4 optical switch device, the time (transition time) required to transition from one switch mode (e.g., FIG. 10A) to another switch mode (e.g., FIG. 10B) is very short compared to the duration of an input packet.
[0146] For example, assuming that the transition time is negligible, with practically available technology the transition time can be reduced to 10 psec, which is extremely short compared to, for example, a 100 Gb / s Ethernet packet which has a duration of 120 nsec.
[0147] There may also be a short guard time between optical packets to avoid any data loss during switching.
[0148] The bandwidth of packets generated by the host is multiplied by a factor F (the number of switch ports). For example, in an 8x8 switch, 25Gb / s electrical packets need to be converted into 200Gb / s optical packets, which are generated by directly modulated lasers and multilevel modulation formats.
[0149] Here, since the distance between hosts assumed in this embodiment is short, high data rates can be achieved, and the optical dispersion effect is negligible.
[0150] Additionally, high bitrate packets may be generated in other ways. The switch may be used as a core switching unit in a hybrid switching architecture to scale up the number of interconnected hosts without centralized control.
[0151] <Configuration of Switching System 40> Next, the switching system 40 will be described with reference to Fig. 13. The switching system 40 is scaled by grouping hosts.
[0152] As shown in FIG. 13, the switching system 40 includes a plurality of source groups 3_10 to 3_40 on the transmitting side, and each source group includes a plurality of transmitting hosts (for example, 3_11, 3_12) and a switching element 41.
[0153] The receiving side is provided with a plurality of destination groups (e.g., 4_20, etc.), and each destination group is composed of an optical receiver 15, a control unit 17, a receiving-side switch 43, and a plurality of receiving hosts (e.g., 4_21, 4_22, etc.). The other configurations are the same as those of the first embodiment.
[0154] The switching element 41 is a low-radix ASIC switch chip.
[0155] The optical switch 14 has four input / output ports and an operating cycle divided into four time slots, where a group of sending hosts 3_10 to 3_40 connects at the ports of the optical switch 14, rather than individual host units.
[0156] For example, in the transmission group A (3_10), the ASIC switch 41 and the two hosts A1 and A2 (3_11 and 3_12) are arranged close to each other. The ASIC switch 41 and the two hosts 3_11 and 3_12 are electrically linked at short distances.
[0157] Increasing the number of host units per group improves scalability. In order to take advantage of the characteristics of the electrical link, it is desirable to have around 10 host units per group.
[0158] As shown in FIG. 13, destination groups A to D (3_10 to 3_40) are connected to the input ports of the optical switch 14, and destination groups 4_10 to 4_40 are connected to the output ports.
[0159] Hosts in the same group exchange packets using the ASIC switch 41. For example, the ASIC switch 41 in group A (3_10) is used to interconnect hosts A1 and A2 (3_11, 3_12). Packets between hosts in different groups (hereinafter referred to as "inter-group packets") are exchanged via interconnection with the optical switch 14.
[0160] At any time slot, each destination (output) group connects with only one transmit (input) group. Inter-group packet switching of transmit groups is handled by placing these packets (reduced duration optical packets) within the exact time slot at which the transmit group connects with the desired destination group.
[0161] <Operation of Switching System 40> The operation of the switching system 40 will now be described with reference to FIG.
[0162] In the switching system 40, the end-to-end transmission of an inter-group packet from the source host to the destination host consists of the following three steps:
[0163] As a first step, electrical switching is performed on packets at the destination group level rather than the destination host level.
[0164] In particular, electrical switching is performed according to destination groups using the local ASIC switch 41 to split the packets generated by the host among the sending groups 3_10 to 3_40.
[0165] Here, inter-group packets simultaneously transmitted to the same group are collected in the destination virtual queue regardless of differences in destination hosts. Here, queues G1 to G4 (42_1 to 42_4) correspond to destination groups 4_10 to 4_20.
[0166] As an example, consider two packets simultaneously sent from hosts A1 and A2 (3_11, 3_12) to hosts 4_22 and 4_21 in destination group 4_20, respectively. At this time, both packets are switched to queue G2 (42_1).
[0167] For example, packets from hosts A1 and A2 (3_11 and 3_12) are transmitted at twice the bandwidth of 25 Gb / s.
[0168] As a second step, optical switching is performed on the packets (reduced duration optical packets) sent to the desired destination group by placing each packet in a matching time slot in the optical switch 14.
[0169] For example, a packet is divided into four parts, compressed four times, and each part is assigned to the first to fourth time slots within a time period T, and transmitted with a bandwidth of 200 Gb / s.
[0170] In detail, packets are sent to each optical switch 14 only from the corresponding queue in the ASIC switch 41 .
[0171] Here, hosts in the same group generate packets at the same time that are all sent to the same group.
[0172] Also, to avoid contention, all these simultaneous packets are coordinated into the same time slot of the optical switch 14. The bandwidth of the optical transmitter 13 makes this possible.
[0173] Here, it is not necessary to separate packets from different sending hosts using different wavelengths for identification, but a WDM-based optical transmitter can be used to meet the high bandwidth requirements associated with an increasing number of host units per group.
[0174] As a third step, the packets arrive at their desired destination group, e.g., all 25 Gb / s packets are received at time T.
[0175] Subsequently, electrical switching is performed on the packets.
[0176] In particular, multiple packets may be sent to the same end host at the same time, with higher priority packets being processed first. The self-management of incoming data packets described above is performed for the local ASIC switch 41 that is assigned to receive the data.
[0177] On the other hand, in the switching system 40, arbitration according to the conventional centralized control method is divided into the following three steps to execute arbitration.
[0178] In the first step, the input ports of the switching system 40 are divided into groups (e.g., 3_10 to 3_40), and packets in each group are processed independently of other packets. Within an input group, output groups that are input simultaneously and have the same destination are treated as the same group and sent together without arbitration. In this way, grouping the input ports on a small scale allows this step to be processed quickly.
[0179] Thereafter, as in the second embodiment, a processing step by the optical switch is executed as the second step, and an arbitration step is executed as the third step.
[0180] As a result, the switching device according to this embodiment can reduce latency and power consumption in switching, and can reduce (or eliminate) the burden of collecting information necessary for the arbitration process.
[0181] Furthermore, with the switching system 40, the number of interconnected hosts can be increased by grouping the hosts, thereby improving the expandability of the system.
[0182] <Configuration of Switching System 50> Next, the switching system 50 will be described with reference to Figures 14 and 15. The switching system 50 is scaled by optical multicasting (optical multiplexing).
[0183] 14, the switching system 50 includes an optical multiplexing unit 51 between the optical transmitter 13 and the optical switch 14, and a first demultiplexing unit 52 and a second demultiplexing unit 53 between the optical switch 14 and the optical receiver 15. The other configurations are the same as those of the third embodiment.
[0184] A plurality of optical transmitters 13 are connected to the optical multiplexing unit 51 .
[0185] An output port of the optical switch 14 is connected to the first demultiplexing unit 52. An output port of the first demultiplexing unit 52 is connected to the second demultiplexing unit 53.
[0186] Here, as an example of optical multicasting, an example will be shown in which wavelength multiplexing is used to multiplex optical signals.
[0187] An arrayed waveguide grating (AWG) optical coupler is used in the optical multiplexing unit 51 to wavelength-multiplex the optical signals.
[0188] An optical splitter is used in the first demultiplexing section 52 to split the optical signal at a predetermined power ratio.
[0189] Furthermore, an AWG filter is used in the second demultiplexing unit 53 to demultiplex the optical signal for each wavelength.
[0190] <Switching system operation> In the switching system 50, for example, as shown in FIG. 14, an optical transmitter 13_1 connected to a transmission group A (3_10) and an optical transmitter 13_2 connected to a transmission group B (3_20) each output optical packets of different wavelengths.
[0191] Optical packets of different wavelengths are multiplexed by an AWG optical coupler 51 and are simultaneously transmitted to multiple destination groups 4_10 to 4_40 in the same time slot.
[0192] In this way, optical packets can be sent to many groups, for example, many more end hosts, without increasing the number of ports on the switch.
[0193] Higher multicasting ratios are also possible, with the power budget of the optical link determining the maximum achievable ratio.
[0194] For example, if two optical packets with different wavelengths are each transmitted at a bandwidth of 200 Gb / s, they will be transmitted at twice the bandwidth (400 Gb / s).
[0195] The transmitted optical packet is split into destination groups by an optical splitter 52, and then demultiplexed into wavelengths by an AWG filter 53 in each group (e.g., group 4_40) and transmitted to an end host (e.g., destination hosts 4_41, 4_42).
[0196] Thus, when multicasting is used, optical packets from multiple destination groups arrive at the same destination group simultaneously, so a receiver unit with demultiplexing capabilities is used to process packets from different destination groups, increasing the total number of receiver units in the system.
[0197] In this way, the switching system 50 can improve the expandability of the system by using optical multicasting (optical multiplexing).
[0198] 15 shows an example of a timing chart of the switching system 50. In the switching system 50, one switching period is divided into four time slots, of which the first slot is shown on the left and the second slot is shown on the right.
[0199] Here we have 128 25Gb / s hosts, 16 groups (8 hosts per group), and multicast to 4 groups at a time.
[0200] The switching system 50 uses commercially available transceiver units based on the PAM4 multi-level format and processes a total communication volume of 6.4 Tb / s.
[0201] <Effects> As described above, the switching system 210 of this embodiment can process the aggregate bandwidth of data emerging from the electrical switching edges to efficiently reduce the switching energy consumed per bit.
[0202] Furthermore, in the switching system 210, the optical switch does not require arbitration and can improve the scalability of the entire switch system without sacrificing end-to-end latency. That is, the optical switch can improve both system scalability and end-to-end latency without requiring a trade-off between them.
[0203] Additionally, multiple low-radix electrical switches are used in switching system 210. Compared to using a single switch unit with high switching capacity, the use of low-radix switches allows for increased scalability of the computing system at lower power density and allows for larger switching fabrics.
[0204] Each switch handles traffic from a small (low count) group of XPUs.
[0205] The XPUs are placed in close proximity to designated switches, where energy-efficient, high-bandwidth electrical connections are preferred.
[0206] In this embodiment, each XPU can quickly update its designated switch with the priority of data reception, and arbitration can be performed quickly because the switch handles only a small portion of the total data traffic.
[0207] <First Example> A switching system 300 according to a first embodiment of the present invention will be described with reference to FIGS.
[0208] As shown in FIG. 16, a switching system 300 according to this embodiment includes a switching device 301 and an interconnection device 302.
[0209] The switching device 301 includes the optical switch 14, the optical transmitter 13, the optical receiver 15, and the electrical switch (including the control unit) 212, as described above.
[0210] The interconnection device 302 is connected to each of the multiple XPUs 120. In the interconnection device 302, the RC 140 and the protocol processor 110 are connected in turn, and the protocol processor 110 is connected to the electrical switch 212. Here, protocol processing is offloaded from the XPU 120 to the protocol processor 110. In addition, an XPU memory 121 is connected to each XPU 120. In addition, the RC 140 is connected to the XPU memory 121 and can directly access it.
[0211] Thus, to enable an ultra-low latency I / O subsystem, each electrical switch port interfaces with a protocol processor and a device having functionality such as a Root Complex (RC), where functionality such as an RC is not specifically limited to the PCIe standard.
[0212] Also, an example of a conventional electrical switch uses a packet forwarding table, implemented by a Ternary Content Addressable Memory (TCAM), where all simultaneously incoming packets are checked for priority and destination, and the entire packet is sent to the desired output port.
[0213] On the other hand, in this embodiment, the protocol processor may be shared among different ports. In this case, the checksum of the packet bits is checked in a parallel step with the TCAM. Furthermore, only the payload of the packet is sent to the desired output port, not the entire packet.
[0214] Also, allocating a separate high performance processor for each port is difficult and requires a trade-off between the computing power of the protocol processor and other system features.
[0215] Thus, high-power protocol processors may be shared among XPUs of the same group, ie equipped with the same electrical switch.
[0216] Furthermore, as shown in FIG. 17, physical switching and protocol execution may be integrated into the same device 312 for reduced latency and improved power efficiency.
[0217] For example, content addressable memory (CAM) is used to implement lookup tables in high-end ASIC switches, allowing for fast detection of memory addresses that match specific content.
[0218] In a typical switch, this is done by matching the content to the desired output switch port and forwarding the data.
[0219] On the other hand, in the integrated implementation of I / O and switching functions in this embodiment, data with matching content is protocol processed and sent to the correct physical switch port, and only the valid payload part is applied to the XPU.
[0220] An optical transmitter 13 and a receiver 15, each equipped with an FEC encoder and decoder, are placed at the port of the device connected to the optical switch. This enables high-data-rate (100 Gb / s or higher) optical signals to be transmitted with high reliability.
[0221] Here, the optical switch 14 allows all packets sent simultaneously to be sent to the same XPU.
[0222] For example, a low-priority packet and a high-priority packet destined for the same XPU #1 are simultaneously input from the optical switch 14. At this time, the low-priority packet is sent to a shared memory (not shown) of the device. Here, the shared memory is embedded in or attached to the device 312.
[0223] On the other hand, high-priority packets undergo protocol processing and the payload part is sent to XPU#1.
[0224] According to the switching system of this embodiment, signal processing delay and power consumption can be reduced by integrating switching in a low-radix electrical switch with a high-level function device.
[0225] In the embodiments of the present invention, examples of the structure, dimensions, materials, etc. of each component in the configuration of the interconnection device and switching system are shown, but the present invention is not limited to these examples. Anything that can demonstrate the functions and effects of the interconnection device and switching system may be used. [Industrial Applicability]
[0226] The present invention relates to an interconnection device and a switching system that interconnects a processor and a switch, and can be applied to a computing system. [Explanation of symbols]
[0227] 100 Interconnection Device 110 Protocol Processor
Claims
1. An interconnection device interconnecting an XPU connected to an XPU memory and a switching device, comprising: A protocol processor that executes protocols offloaded from the XPU and has direct access to the XPU memory. An interconnection device comprising:
2. RC between the XPU and the protocol processor The interconnection device of claim 1 , comprising:
3. An optical switch, An electric switch, an optical transmitter that converts an electrical signal input from the electrical switch into an optical signal and outputs the optical signal to the optical switch; an optical receiver that converts an optical signal input from the optical switch into an electrical signal and outputs the electrical signal to the electrical switch; a plurality of interconnection devices according to claim 1 connected in parallel to the electrical switch; A switching system comprising:
4. A priority is set for the input electrical signal, a control unit connected to the optical receiver, the control unit holding the converted electrical signals with a lower priority and transmitting the converted electrical signals with a higher priority first; 4. The switching system according to claim 3.
5. a memory for storing the converted electrical signals with low priority; The switching system of claim 4 , comprising:
6. Encoder and decoder The switching system of claim 3 comprising:
7. The optical signals transmitted from the optical transmitter are multiplexed and demultiplexed according to predetermined optical characteristics, and are received by the optical receiver.
4. The switching system according to claim 3.
8. the optical transmitter divides the optical signal by the number of the optical receivers to which the optical signal is transmitted; The optical switch transmits the divided optical packets in the time slots assigned to each of the divided optical packets.
4. The switching system according to claim 3.
Citation Information
Patent Citations
Systems and Methods for Photonic Switching
JP2016519536A
System and method for photonic switching
JP2016527737A
Apparatus and method for monitoring optical gate device, and optical switch system
US20090238574A1
System and method for TCP / IP offload independent of bandwidth delay product
US20100250783A1
Method for transferring data packets in a communication network and switching device
US20110142052A1