MPI set operation
By optimizing the MPI set operation method and optical circuit switching network, the problems of high network overhead and low throughput in the prior art are solved, and efficient distributed computing task execution and high reliability communication are realized.
Patent Information
- Application Number
- CN202380081248.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-24
- Filing Date
- 2023-11-22
- Publication Date
- 2025-07-04
AI Technical Summary
Existing MPI policies and network architectures lead to significant network overhead, high idle time and low operational throughput when performing distributed and parallel computing tasks, and cannot meet the high performance requirements of high-performance computing applications such as distributed deep learning.
By performing MPI set operations in the network, including receiving operation information, determining algorithm steps, initialization and finalization processes, and data transmission between nodes and subsets, communication between nodes is optimized, and optical circuit switching networks are adopted and an optical circuit switching network is used to improve communication efficiency.
Improve network performance, reduce the completion time of collection operations, and maintain high reliability and flexibility in the event of failure, achieving port-level full-to-full communication.
Smart Images

Figure CN120266099A_ABST
Abstract
Description
Technical Field
[0001] The present technology relates to communication between computing nodes that perform parallel and distributed tasks in the field of high-performance computing. More specifically, but not limited to, the present technology relates to the Message Passing Interface MPI, collective operations, and network architectures in which MPI collective operations can be performed. Background Art
[0002] With the growth of parallel and distributed computing, and the increasing complexity of applications running in such systems, it is necessary to improve the performance of computing nodes in a similar manner. In addition, since computing nodes are interconnected to perform distributed tasks or processes, a large part of the system performance depends on the technology used by individual computing nodes to communicate with each other.
[0003] MPI collective operations are used in high-performance computing HPC, such as distributed deep learning applications DDL, to control how messages are exchanged between computing nodes running parallel tasks or processes. However, as discovered by the inventors of the present technology, existing MPI strategies are not optimal for meeting the high-performance requirements of HPC applications (such as DDL that highly depend on network performance). In addition, the inventors have found that existing MPI strategies are not optimal for current network architectures (such as optical circuit switching networks). In fact, existing MPI strategies and network architectures result in significant network overhead, high idle time, and low operation throughput when performing distributed and parallel computing tasks.
[0004] Embodiments of the present disclosure solve the above problems. Example network architectures in which the present technology can be performed will also be described. Summary of the Invention
[0005] The invention is defined in the appended claims.
[0006] From a first aspect, there is provided a method for performing Message Passing Interface MPI collective operations in a network, where the network includes a plurality of interconnected nodes, and the method includes: receiving, at a node among the plurality of interconnected nodes, MPI collective operation information identifying an MPI collective operation to be performed and a graph of the network; determining the number of algorithm steps of the MPI collective operation based on the MPI collective operation and the graph of the network; determining an initialization process for the algorithm steps; determining a final determination process for the algorithm steps; for each algorithm step, determining a subset of nodes in the plurality of interconnected nodes with which the node is to communicate, and one or more data portions that the node is to send to and receive from the nodes within the subset of nodes; and initializing the MPI collective operation based on the determined subset, initialization process, final determination process, and one or more data portions.
[0007] The present inventors have found that various MPI operations (such as reduce scatter, all gather, barrier, all to all, scatter, gather, broadcast, and all reduce) can be characterized by multiple different algorithmic steps (partial collective operations involving subsets of nodes in the network), where each step requires specific nodes to communicate specific information with other nodes in a specific subset of nodes. The present inventors have found that by doing so, MPI operations can be performed more efficiently and the completion time can be reduced compared to comparative examples that do not use the present technology.
[0008] From a second aspect, there is provided a node for performing MPI collective operations on data in a network, where the network includes a plurality of interconnected nodes, and the node includes a processor configured to perform the present technology.
[0009] From a third aspect, there is provided a computer-readable medium including instructions that, when executed by a processor, cause the processor to perform the present technology.
[0010] From a fourth aspect, there is provided an optical circuit switching network including: a plurality of nodes, each node including one or more optical transceivers configured to implement time-division multiplexing such that each node belongs to one of a plurality of transmit groups or one of a plurality of receive groups at a given time; a plurality of one-to-many switches, where each optical transceiver of each node in the transmit group nodes of the plurality of nodes is connected to a one-to-many switch among the plurality of one-to-many switches; a plurality of many-to-one switches, where each optical transceiver of each node in the receive group nodes of the plurality of nodes is connected to a many-to-one switch among the plurality of many-to-one switches; and a plurality of photon subnet units, where each port of each of the one-to-many switches and the many-to-one switches is connected to a different photon subnet unit.
[0011] From a fifth aspect, there is provided an electronic time-division multiplexing circuit switching network including: a plurality of nodes, each node including one or more transceivers and configured to implement time-division multiplexing such that each node belongs to one of a plurality of transmit groups or one of a plurality of receive groups at a given time; a plurality of one-to-many switches, where each transceiver of each node in the transmit group nodes of the plurality of nodes is connected to a one-to-many switch among the plurality of one-to-many switches; a plurality of many-to-one switches, where each transceiver of each node in the receive group nodes of the plurality of nodes is connected to a many-to-one switch among the plurality of many-to-one switches; and a plurality of subnet units, where each port of each of the one-to-many switches and the many-to-one switches is connected to a different subnet unit.
[0012] Thus, the fourth and fifth aspects provide more efficient communication between transmit nodes and receive nodes in the network, thereby improving network performance and reducing the collective operation completion time. In addition, port-level all-to-all communication is achieved and the resilience of the network is improved because there is no single point of failure.
[0013] From a sixth aspect, a method for communication in a network according to the fourth aspect is provided, the method comprising: sending, from an optical transceiver of a transmitter node, via a port of a one-to-many switch connected to the node, light encoding data for transmission to a photonic subnet unit connected to the port; receiving, at a receiver node, light from the photonic subnet unit via a many-to-one switch connected to the receiver node.
[0014] Thus, as a result of using the present technology, the efficiency of data communication in a network can be improved.
[0015] After reading this disclosure, especially after reading the sections on the description of the drawings, the specific implementation manners, and the claims, other aspects will also become apparent. Description of the Drawings
[0016] Examples of the present disclosure will now be described by way of example only with reference to the drawings, in which:
[0017] Figure 1 An example network architecture in which the present technology can be implemented is schematically shown.
[0018] Figure 2 A method according to the present technology is schematically shown.
[0019] Figure 3 An example node that can execute the present technology is schematically shown.
[0020] Figure 4 A node and an algorithm according to the present technology are schematically shown.
[0021] Figure 5 An example network architecture according to the present technology is schematically shown.
[0022] Figure 6 An example method according to the present technology is schematically shown.
[0023] Figure 7 Different subnet units according to the present technology are schematically shown.
[0024] Figure 8 Different subnet units according to the present technology are schematically shown.
[0025] Figure 9 An example connection of an example subnet unit is schematically shown.
[0026] Figure 10 An example network and data plane architecture are schematically shown.
[0027] Figure 11Schematically shows examples of many-to-many communication patterns across different time slots between nodes of a) the same source-destination communication group pair and b) different source-destination communication group pairs, and illustrates the WDM, TDM, and SDM attributes of the architectures of different communication groups by way of example.
[0028] Figure 12 Schematically shows examples of a) one-to-many, b) many-to-one, and c) one-to-one communication patterns between nodes with the same source-destination communication group pair in the same time slot, and illustrates the WDM, TDM, SDM (across multiple transceivers) attributes of the architecture, allowing high-bandwidth (up to full capacity) communication between one or more non-pairs or sets.
[0029] Figure 13 Schematically shows an example node subgroup for each algorithm step of a 54-node network (x = 3, J = 3, Λ = 6).
[0030] Figure 14 Schematically shows an example MPI operation process and node architecture according to the present technology.
[0031] Figure 15 Schematically shows the MPI operation workflow.
[0032] Although the present disclosure is susceptible to various modifications and alternative forms, specific example methods are shown by way of example in the drawings and described in detail herein. However, it should be understood that the appended drawings and specific embodiments herein are not intended to limit the present disclosure to the particular forms disclosed, but rather the present disclosure covers all modifications, equivalents, and alternatives falling within the spirit and scope of the claimed invention.
[0033] It should be recognized that the features of the above examples of the present disclosure can be conveniently and interchangeably used in any suitable combination. Detailed Description
[0034] The present technology described herein relates to initialization (preparing nodes so that they can subsequently perform MPI operations) and performing MPI collective operations. More specifically, the present inventors have discovered a method that allows for more efficient initialization and execution of MPI operations.
[0035] In some examples, a node receives information indicating an MPI operation to be performed and a graph of the network. This information can be received from a job scheduler. Alternatively, this information can be obtained or retrieved by the node.
[0036] In some examples, the node also receives the message size of the messages associated with the MPI collective operation to be performed, and for each step, determines the one or more message sizes of the one or more messages that the node is to send to nodes within a subset of nodes, and initializes further based on the determined one or more message sizes.
[0037] Then, the node determines how many algorithmic steps are needed to be able to perform the MPI operation. This determination is based on which MPI operation is to be performed and the graph of the network (e.g., the number of other nodes). The inventors have found that various MPI operations can be divided into algorithmic steps, where each step requires a particular node to communicate particular information with other nodes. The inventors have further found that doing so can perform MPI operations more efficiently, with a shortened completion time.
[0038] Then, the node determines the initialization process for the algorithmic step. In some examples, the initialization process is performed at the start of each algorithmic step, or on the message before subsequent processing of the received message. For example, the initialization process can be a process to be performed in the algorithmic step before the node sends data to other nodes in the network. Then the final determination process is determined. In some examples, the final determination process can be performed at the end of each algorithmic step, or after the initialization process. In some examples, the final determination process is a process to be performed on the data or message received from other nodes in the network in the algorithmic step.
[0039] Then, for each determined algorithmic step, the node determines the subset of nodes in the network that the node is to communicate with at each algorithmic step, and the one or more data portions that the node is to send to other nodes.
[0040] Initializing the MPI operation can then include storing the determined information (i.e., the determined initialization and final determination processes) in the memory of the node, and for each step storing: the subset of nodes and the one or more data portions. In examples where the node also receives the message size and determines the one or more message sizes for each step, the one or more message sizes can also be stored in the memory. Then, the node can retrieve the stored data when subsequently receiving a message to be processed using the MPI operation.
[0041] This technique can be performed before the MPI operation runs. Thus, when the MPI operation is to be performed, each node involved in the MPI operation has already determined the information required for each node to perform the MPI operation. Thus, subsequent execution of the MPI operation on the received message is more efficient, and is a more efficient process for performing the MPI operation itself.
[0042] It will be understood that the present technology can be implemented in any interconnected network, such as any port-level all-to-all network (without oversubscription). For example, the present technology can be implemented in an existing electrical packet-switching or optical circuit-switching network.
[0043] In some examples, each of the multiple interconnected nodes is configured to perform the method. For example, each node can perform the method simultaneously to determine the information it will need to be able to perform the identified MPI collective operation on subsequently received messages. Thus, the nodes in the network can efficiently process messages using the predetermined information. In some examples, the multiple nodes are fully interconnected.
[0044] In some examples, the MPI collective operation information defines the MPI collective operation to be performed. Thus, in these examples, each node can receive the MPI collective operation to be performed and use the operation to determine the information required to perform the operation.
[0045] In some examples, the graph of the network includes information indicating the levels of the multiple interconnected nodes. For example, the graph can include network-specific coordinates for each node that identify the level of the node. In some examples, the coordinates identify the position of each node relative to other nodes. In this way, each node can efficiently receive information indicating the network topology (in a format optimized for node processing).
[0046] In some examples, each subset of nodes for each algorithm step is unique. Thus, each node communicates with a different set of nodes in each algorithm step, enabling more efficient information sharing and aggregation among the nodes.
[0047] In some examples, determining the number of algorithm steps is based on retrieving stored information associated with the MPI operation. For example, a node may have stored in memory information associated with multiple MPI operations, and this information can identify the number of algorithm steps for each MPI operation. Then, the node can look up the number of algorithm steps for the received MPI operation based on the stored information. In some examples, each node stores a lookup table that includes information indicating, for each of the multiple MPI operations, the number of algorithm steps required to complete the corresponding MPI operation. In this way, the nodes can efficiently and independently determine the number of algorithm steps.
[0048] In some examples, the network is a circuit-switching network. In other examples, the network is an optical circuit-switching network. The present technology can be particularly effective in such examples because the technology has been specifically optimized for such network architectures.
[0049] In some examples, the network includes one or more clusters, each cluster includes one or more groups, and each group includes one or more nodes; each of the multiple interconnected nodes has a node number within the group, a group number within the cluster, and a cluster number; and the graph includes information indicating the node number, group number, and cluster number of each node. Thus, the nodes receive information summarizing the network in an efficient manner. In some examples, the node number, group number, and cluster number are a coordinate system of the network.
[0050] In some examples, a cluster (also referred to as a communication group) is a logical group of nodes or racks equal to the cardinality (the number of transceiver groups of each node in the network). In some examples, a group (also referred to as a rack) is a logical grouping of nodes. In some examples, the grouping of groups can be such that: the total number of nodes in the network = the number of racks x the number of communication groups x the number of nodes per rack, where the number of nodes per rack is a multiple of the number of communication groups, and the number of racks is less than or equal to the number of communication groups.
[0051] As described above, in some examples, the node number, group number, and cluster number are coordinates that identify the following: the position of a given node within the node hierarchy in the above network. For example, each node can have the following coordinates (g, j, λ), where, for the current node, g is the cluster number, j is the group number within the cluster, and λ is the node number within the group. Additionally, as discussed in more detail below, a cluster can be used interchangeably with a communication group, a group can be used interchangeably with a rack, and a node number can be used interchangeably with a device number. Using this coordinate information enables each node to efficiently determine the connectivity of other nodes in the network and is optimized for techniques used to determine the required information before an MPI operation runs, as discussed in more detail herein.
[0052] In some examples, for a first algorithm step, a subset of nodes includes nodes having the same node number, the same group number, and different cluster numbers. Thus, a node can efficiently determine the other nodes with which the node needs to communicate for the first algorithm step. A node can determine the subset of nodes for each step based on a formula stored in the node's memory. For example, a node can store a lookup table that includes a formula for determining the subset of nodes for each algorithm step. This determination can be based on the node coordinate information associated with the graph of the network.
[0053] In some examples, the multiple nodes in each group are divided into node sets including x nodes, where each node has a unique node set number from 1 to x, and x is the number of clusters. The inventors have found that by dividing the nodes in this way, the techniques and formulas discussed later in this document can be executed and used more efficiently. Thus, the nodes can more efficiently determine the information required to perform an MPI operation.
[0054] In some examples, for the second algorithm step, the subset of nodes includes nodes having consecutive node numbers, the same group number, and different cluster numbers within the same node set. Thus, a node can efficiently determine the other nodes with which the node needs to communicate for the second algorithm step.
[0055] In some examples, for the third algorithm step, the subset of nodes includes nodes having the same node number, different group numbers, and different cluster numbers. Thus, a node can efficiently determine the other nodes with which the node needs to communicate in the third algorithm step.
[0056] In some examples, for the fourth algorithm step: the subset of nodes includes nodes having the same node number, different node sets, the same group number, and different clusters within the node set; or, the subset of nodes includes nodes within a consecutive node set having the same node number, the same group number, and different cluster numbers within the node set. Thus, a node can efficiently determine the other nodes with which the node needs to communicate for the fourth algorithm step.
[0057] Thus, for each determined algorithm step, a node can determine the other nodes in the subset with which the node will need to communicate in an efficient manner.
[0058] In some examples, the network includes x clusters; each cluster includes J groups, where J ≤ x; each group includes Λ nodes; each cluster has a cluster number g, defined as 0 ≤ g ≤ x - 1; each node has a node number λ within the group, defined as 0 ≤ λ ≤ Λ - 1; each group has a group number j, defined as 0 ≤ j ≤ J – 1; and the multiple nodes in each group are partitioned into node sets including x nodes, where each node has a unique node number from 1 to x within the node set. In some examples, if Λ > x 2 , then the partitioning of nodes into node sets is performed multiple times such that there are smaller node sets, and in some examples, the fourth step will be repeated the same number of times as the number of times this partitioning is performed so that it works across the partitions. The inventors have found that organizing the nodes in this way is optimized for the present technology and the formulas discussed later in this document, allowing the present technology to be performed more efficiently.
[0059] In some examples, a subset of the nodes in the network can form a network of nodes. In this case, by making x, J, Λ the number of communication groups, racks, and unique node / device IDs used for the subset of nodes in the entire graph (depending on the node arrangement / selection), the algorithm is also applicable to the subset of nodes.
[0060] In some examples, for the first algorithm step, the total number of subsets of nodes = ΛJ, each subset has an identifier i; the total number of nodes in each subset is x; and for the first algorithm step, the identifier of the subset is determined based on the following formula: i = λ + Λ·j. Thus, the nodes can efficiently determine the number of subsets, the identifiers of the subsets, and the number of nodes in each subset for the first algorithm step using the coordinate information and the graph of the network. In examples where each node executes the present technology, each node thus independently has the required information, thereby improving resilience.
[0061] In some examples, for the second algorithm step, the total number of subsets of nodes = ΛJ; the total number of nodes in each subset is x; and for the second algorithm step, the identifier of the subset is determined based on the following formula: Thus, the nodes can efficiently determine the number of subsets, the identifiers of the subsets, and the number of nodes in each subset for the second algorithm step using the coordinate information and the graph of the network.
[0062] In some examples, for the third algorithm step, the total number of subsets of nodes = Λx; the total number of nodes in each subset is J; and for the third algorithm step, the identifier of the subset is determined based on the following formula: i = (λ + Λ(j - g)) mod (Λj). Thus, the nodes can efficiently determine the number of subsets, the identifiers of the subsets, and the number of nodes in each subset for the third algorithm step using the coordinate information and the graph of the network.
[0063] In some examples, for the fourth algorithm step, the total number of subsets of nodes = Jx 2 ; the total number of nodes in each subset is Λ / x; and for the fourth algorithm step, the identifier of the subset is determined based on the following formula: Or Thus, the nodes can efficiently determine the number of subsets, the identifiers of the subsets, and the number of nodes in each subset for the fourth algorithm step using the coordinate information and the graph of the network.
[0064] In some examples, the method further includes: in response to determining that the MPI operation is a reduce-scatter operation, selecting a reshaping process as the initialization process and selecting a reduction process as the final determination process; in response to determining that the MPI operation is an all-gather operation, selecting a copy process as the initialization process and selecting an identity process as the final determination process; in response to determining that the MPI operation is a barrier operation, selecting an identity process as the initialization process and selecting a logical AND process as the final determination process; in response to determining that the MPI operation is an all-to-all operation, selecting a reshaping process as the initialization process and selecting a reshaping process as the final determination process; in response to determining that the MPI operation is a scatter operation, selecting a reshaping process as the initialization process and selecting an identity process as the final determination process; in response to determining that the MPI operation is a gather operation, selecting a copy process as the initialization process and selecting an identity process as the final determination process; in response to determining that the MPI operation is a broadcast operation, selecting a process associated with the scatter operation and the all-gather operation; and in response to determining that the MPI operation is an all-reduce operation, selecting a process associated with the reduce-scatter operation and the all-gather operation.
[0065] Thus, a node can determine the type of MPI operation and use a lookup table to look up the processes to be executed that depend on the MPI operation. As described above, the inventors have found that various MPI operations (e.g., reduce-scatter, all-gather, scatter, gather, etc.) can be characterized by a combination of processes performed on data at each step of the algorithm. This enables more efficient execution of MPI operations, with reduced completion times and higher throughput.
[0066] In some examples, for a first algorithm step, one or more data portions are determined based on the following: For a second algorithm step, the data portion is determined based on the following: For a third algorithm step, the data portion is determined based on the following: data portion = j; and for a fourth algorithm step, the data portion is determined based on the following:
[0067] As described above, at each algorithm step, each node determines the portion of the received message to be sent to another node or other nodes in its subset. In some examples, the matrix is of a defined size. In some examples, the message is a vector, an array, or a matrix, and each element in the vector / array / matrix has an index. Thus, the node can, after performing an initialization process on the received message (whether the original message at the start of the process or the message received from another node in the subset in a previous algorithm step), determine the portion of the message that each other node in the subset should receive. Specifically, since the node has determined the coordinates of all other nodes in the subset, the node can utilize the above formula and, for each coordinate of each node in the subset, determine the portion of the message that the node should receive. For example, portion = 2 can correspond to the third element in [0, 1, 2, n]. Thus, the node knows which portion of the message each other node in the subset needs to receive for each algorithm step.
[0068] At runtime, the node can send the determined one or more portions to the corresponding one or more nodes. Thus, information is shared in an efficient manner, and each node understands how the information is shared within its node subset.
[0069] In some examples, the size of the received message is m, and the method further includes: in response to determining that the MPI operation is a reduce-scatter operation: for the first algorithm step, selecting a message size of m / x; for the second algorithm step, selecting a message size of m / x 2 , for the third algorithm step, selecting a message size of m / (Jx 2 ); for the fourth algorithm step, selecting a message size of m / (JΛx); in response to determining that the MPI operation is an all-gather operation: for the first algorithm step, selecting a message size of m·JΛx; for the second algorithm step, selecting a message size of m·JΛ; for the third algorithm step, selecting a message size of m·JΛ / x; for the fourth algorithm step, selecting a message size of m·Λ / x; in response to determining that the MPI operation is a barrier operation: for the first algorithm step, selecting a message size of 0; for the second algorithm step, selecting a message size of 0; for the third algorithm step, selecting a message size of 0; for the fourth algorithm step, selecting a message size of 0; in response to determining that the MPI operation is an all-to-all operation: for the first algorithm step, selecting a message size of m / x; for the second algorithm step, selecting a message size of m / x; for the third algorithm step, selecting a message size of m / J; for the fourth algorithm step, selecting a message size of m·x / Λ; in response to determining that the MPI operation is a scatter operation: for the first algorithm step, selecting a message size of m / x; for the second algorithm step, selecting a message size of m / x2 For the third algorithm step, the size of the message is selected as m / (Jx 2 ). For the fourth algorithm step, the size of the message is selected as m / (JΛx); in response to determining that the MPI operation is a gather operation: for the first algorithm step, the size of the message is selected as m·JΛx, for the second algorithm step, the size of the message is selected as m·JΛ, for the third algorithm step, the size of the message is selected as m·JΛ / x, and for the fourth algorithm step, the size of the message is selected as m·(Λ / x); and in response to determining that the MPI operation is a broadcast operation: the size of the message for each step is determined based on the message sizes determined from the scatter and allgather operations. In some examples, in response to determining that the MPI operation is a broadcast operation: the message size for each step is determined based on the following: message size = m / k; and the number of steps = s + k - 2s, where k = √((m·(s - 2)) / αβ), where s is the diameter of the tree generated for performing the broadcast, α is the communication setup latency, and β is the reciprocal of the total node capacity. In some examples, the determined message size corresponds to the message that a node will send to other nodes in each communication step. In some examples, the maximum value of s is 3.
[0070] Thus, a node (or each node participating in the collective operation) can determine the message size for each algorithm step and track the total message size. The inventors have found that this series of relationships and formulas allow for the efficient determination of the message size at each step.
[0071] In some examples, the method further includes: in response to determining that the step or MPI operation to be performed is an allgather, reduce - gather, or gather, performing the algorithm steps in the reverse order. In other words, the steps in Tables 1 to 4 presented below are performed in the reverse order, i.e., step 4 is performed first, then step 3, then step 2, and finally step 1.
[0072] In some examples, the method further includes, for each algorithm step, storing in a memory the determined subset of nodes and one or more data portions, and optionally one or more message sizes.
[0073] In some examples, the method further includes:
[0074] after initializing the MPI operation, receiving a message associated with the MPI collective operation;
[0075] performing the first algorithm step in the algorithm steps by:
[0076] processing the message using the determined initialization process;
[0077] Send one or more parts of the determined processed message to the respective one or more nodes of the subset;
[0078] Receive messages from nodes within the subset; and
[0079] Process the received messages using a final determination process, where the processed received messages become messages for subsequent algorithm steps, and where the algorithm steps are repeatedly executed for all determined algorithm steps using the respective information determined for each step.
[0080] Thus, once a node has determined the information for implementing an MPI operation and stored the information in memory, when a subsequent message to be processed using that MPI operation is received, the node can efficiently process the message using the information it has stored.
[0081] In some examples, after initializing the MPI operation, the method further includes performing an MPI operation on the received messages based on the subset of nodes, one or more data parts, the initialization process, and the final determination process determined for each respective step.
[0082] In some examples, performing each algorithm step includes providing the determined subset of nodes and one or more data parts to a network transcoder of the node, and the method further includes: converting, by the network transcoder, the determined subset of nodes and one or more data parts into instructions for configuring one or more transceivers of the node. Thus, the node may include a network transcoder configured to transcode the information determined by the node into instructions for a network interface card or a transceiver of the node.
[0083] In some examples, the network is an optical network including multiple parallel subnets, each subnet connected to a splitter and a combiner and multiple transceivers. Thus, in these examples, the present technology is combined with optical network technology, as discussed herein, which reduces the completion time of MPI operations and reduces contention. Thus, collective operations can be performed more efficiently. Additionally, using an optical network instead of an EPS network further improves the performance of these technologies and reduces the overall energy consumption and infrastructure cost of the network.
[0084] In some examples, the network is an optical network including multiple transceivers having full - to - full connectivity. In this way, the nodes can achieve unrestricted multi - node communication and reliability in terms of network component failures. For example, even if a transceiver or a subnet fails, communication can still occur between any pair of nodes.
[0085] Thus, techniques for efficiently initializing MPI operations and for more efficiently performing MPI operations are described herein. Specific examples will now be described with reference to the accompanying drawings.
[0086] Figure 1 FIG. 1 schematically shows an example network architecture or topology 100 in which the present technology may be implemented. The network architecture 100 includes a plurality of nodes 110, 120, 130, and 140, which are interconnected as shown by solid lines. It will be appreciated that the network may include any number of nodes and may also include non-interconnected nodes. The nodes 110, 120, 130, and 140 are configured to communicate with each other, for example, via a packet-switched or circuit-switched network architecture (not shown), such as an electrical packet-switched or optical circuit-switched (OCS) network. In an example using OCS, the network includes one or more optical devices. The nodes include communication circuitry for communicating with other nodes in the network.
[0087] Each node may implement the present technology, and in some cases, each node may implement the present technology simultaneously. Thus, at each algorithm step, each node may send data to another node or other nodes in its subset and may also receive data from at least one other node in the subset. Subsets 150 and 160 include nodes 110, 130 and 120, 140, respectively. Subsets 150 and 160 (also referred to herein as subgroups) include groups of nodes that will communicate for each algorithm step. Thus, Figure 1 two possible subsets of nodes for an algorithm step are shown.
[0088] Figure 2 FIG. 2 schematically shows a method 200 according to the present technology. The method 200 may be performed by one or each of the nodes 110, 120, 130, 140 in the network 100.
[0089] At S201, a node among a plurality of interconnected nodes receives MPI collective operation information and a graph of the network, and the MPI collective operation information identifies an MPI collective operation to be performed.
[0090] As discussed above, the graph of the network may include the coordinates of each node in the network (indicating the level of the node) and information indicating the following: the total number of clusters, the total number of groups in each cluster, and the total number of nodes in each group.
[0091] At S202, the node determines the number of algorithm steps of the MPI collective operation based on the MPI collective operation and the graph of the network. For example, the node may store a lookup table in memory that is associated with the number of algorithm steps for each of a plurality of MPI operations. Thus, the node may determine the number of algorithm steps based on the lookup table.
[0092] At S203, the node determines an initialization process for the algorithm step. The node can determine the initialization process based on MPI collective operations. The node can store a lookup table in the memory, which is associated with the initialization process for each of the multiple MPI operations. Thus, the node can determine the initialization process based on this lookup table. The initialization process can be a process to be performed on the received data before the received data is split and sent to other nodes in the subset.
[0093] At S204, the node determines a final determination process for the algorithm step. The node can determine the final determination process based on MPI collective operations. The node can store a lookup table in the memory, which is associated with the final determination process for each of the multiple MPI operations. Thus, the node can determine the final determination process based on this lookup table. The final determination process can be a process to be performed on the data received from other node(s) during each algorithm step.
[0094] At S205, the node determines, for each algorithm step, a subset of the nodes in the multiple interconnected nodes with which the node is to communicate. The node can already have stored in the memory a formula for determining the subset of nodes for each algorithm step. Determining the subset can be based on a graph of the network, such as the coordinates of the nodes in the network.
[0095] In some examples, determining the subset of nodes includes:
[0096] a. Determining an identifier of the subset in which the node is located, based on information related to the position of the node in the network;
[0097] b. Determining the number of nodes in the subset; and
[0098] c. Determining the other nodes within the subset, such as based on a graph of the network.
[0099] At S206, the node determines, for each algorithm step, one or more data portions for the node to send to and receive from the nodes within the subset of nodes. For example, each node can already have stored in the memory a formula for determining the data portion that each node in the subset should receive during the algorithm step.
[0100] At S207, the node initializes the MPI collective operation based on the determined subset, initialization process, and final determination process, and one or more data portions.
[0101] In some examples, a node determines, for each algorithm step, one or more message sizes for one or more messages that the node sends to nodes within a subset of nodes. For example, a node can determine one or more message sizes for messages that the node sends to other nodes in the subset based on received message sizes. In some examples, a node can have stored in memory a formula for determining one or more message sizes based on received message sizes.
[0102] As discussed herein, initializing an MPI collective operation can include storing in memory the determined subset, initialization process, and final determination process, one or more data portions, and optionally the determined one or more message sizes (for each relevant step).
[0103] Thus, a node can efficiently determine the information that the node needs to efficiently perform MPI operations.
[0104] Figure 3 An example node 310 that can perform the disclosed techniques (e.g., the techniques of method 200) is schematically illustrated. Node 310 can perform Figure 1 techniques in the network architecture of. Node 310 includes a processor 320 (or processing circuitry) and a memory 330, as well as communication circuitry for communicating with other nodes (not shown).
[0105] As shown, node 310 receives or otherwise obtains an MPI collective operation to perform (or information identifying the MPI collective operation to perform) and a graph of the network. Node 310 includes a processor 320 that is configured to perform the processing required by this technique. Then, node 310 performs method 200. As a result, node 310 determines information 340 and stores information 340 in memory 330. Node 310 has stored in memory 330 a lookup table and formula 335 for determining information 340.
[0106] Information 340 includes a determined initialization process, a determined final determination process, and for each of a number N of algorithm steps: a subset of nodes, and one or more data portions.
[0107] Thus, at a later time, when a message to be processed using an MPI collective operation is received, the node has the required information and can perform the MPI collective operation using this technique.
[0108] The lookup table and formula 335 will now be described, which a node uses to determine the number of algorithm steps, initialization process, final determination process, subset of nodes for each step, one or more data portions for each step, and optionally the message size for each step. These tables will be further described in the "Working Examples" section below.
[0109] As discussed previously, the graph may include the coordinate information of each node participating in the collective operation. The graph may also include information indicating that the network includes x clusters; each cluster includes J groups, where J ≤ x; each group includes Λ nodes; each cluster has a cluster number g, defined as 0 ≤ g ≤ x - 1; each node has a node number λ in the group, defined as 0 ≤ λ ≤ Λ - 1; each group has a group number j, defined as 0 ≤ j ≤ J – 1.
[0110] In fact, as discussed previously, the network may include clusters, groups of nodes within each cluster, and nodes within each cluster. This coordinate information may take the form: (g, j, λ), where, for the current node, g is the cluster number, j is the group number within the cluster, and λ is the node number within the group.
[0111] The following tables may be stored in the memory of the node (or each node) and used to determine information 340. In this document, subgroup and subset of nodes may be used interchangeably.
[0112] Table 1: Shows subgroup ID selection. #SG is the number of subgroups, and #NS is the number of nodes in each subgroup.
[0113]
[0114] Table 2: Shows the message sizes for each step of various MPI collective operations, as well as buffer operations and local operations. In this document, buffer operation and initialization process may be used interchangeably, and local operation and final determination process may be used interchangeably.
[0115]
[0116] Table 3: Describes the formula for which part of the previous message a node should receive at any algorithm step.
[0117]
[0118] Table 4: Formulas for calculating the coordinates (cluster number, group number within the cluster, node number within the group (also known as the communication group), rack number, device number) of other nodes in the current subgroup. The current node has coordinates (g, j, λ) at any algorithm step. The variables show the ranges of the variables used to describe all members of the subgroup.
[0119]
[0120] Table 1 can be used to determine the number of algorithmic steps of the MPI operation. Table 1 can be used to determine the subset of nodes with which a node is to communicate at each algorithmic step. As shown, for each step, the group information (i.e., g, x, λ, J, j, Λ) can be used to determine the number of subgroups, the number of nodes in each subgroup, and the subgroup identifier. The coordinates of the other nodes in the subgroup can be determined using Table 4.
[0121] In some examples, if, for a given algorithmic step, the number of nodes calculated for each subgroup = 1, then that algorithmic step is skipped.
[0122] Table 2 can be used to determine the initialization and finalization processes to be performed. As shown, Buff_op (or buffer operation or initialization process) can be one of a number of processes, depending on the MPI operation to be performed. For example, the initialization process can be a reshape, copy, or identity process. Op (or local operation or finalization process) can be a reduce, identity, reshape, logical AND process. The combination of initialization and finalization (or buffer and local operations) is specific to / depends on the MPI operation to be performed.
[0123] Table 3 can be used to determine one or more data portions for each algorithmic step. The table includes formulas for calculating the portion of the message that each node in the subset should receive at each algorithmic step (e.g., the indices of a vector message). The formulas take coordinates and graph information as input.
[0124] Thus, using Tables 1, 2, 3, and 4 stored in memory, a node can determine the information 340 required to perform an MPI collective operation. As discussed, the information 340 can be stored in memory prior to the runtime of the MPI operation.
[0125] Figure 4 Schematically shown is the algorithm / method executed by node 410 after initializing the node with the above information and upon receiving a message. Node 410 can be node 310, or can be any node 110, 120, 130, 140 in network 100.
[0126] Node 410 receives message 420, which will be used for an MPI operation for which the node has stored all the required information. Message 420 can be an array, vector, or matrix of a defined length. As shown, message 410 is a vector, divided into N = 3 portions, each portion labeled with index 0, 1, or 2. It will be understood that the specific form of the message, its division, and the indices can vary depending on the implementation.
[0127] Steps 1, 2, and 3 illustrate a pseudocode version of the process executed by node 420 after receiving a message. After node 420 receives a message, the node operates as follows: 1. Set the received message = m; and 2. Retrieve information related to the MPI operation to be performed. Specifically, node 410 retrieves from the memory the number of algorithm steps associated with the MPI operation, the initialization process associated with the MPI operation, the finalization process associated with the MPI operation, the subset of nodes to communicate with at each step, and the data portion that each node should receive at each step.
[0128] At step 3, node 410 processes the message. Specifically, for each of the multiple algorithm steps, the node performs the initialization process on m. Then, the node distributes portions of m to the nodes in the subgroup based on the determined one or more portions. For example, the portion of the message with index 0 can be assigned to node x, the portion with index 1 can be assigned to node y, and the portion with index 3 can be assigned to node z (where nodes x, y, and z are members of the subset of nodes for that step). In other words, using the stored one or more data portions, the portion of the message that each node needs to receive (according to the indices of these N portions) is assigned to the corresponding node.
[0129] Then, node 410 sends the respective portions of the message to the corresponding one or more nodes. In the same algorithm step, the node receives messages from the nodes in the subset. Node 410 performs the finalization process on the received messages and sets the processed message = m to be used as the message for the next algorithm step. This process is repeated for all algorithm steps, using information specific to each algorithm step, until the operation is complete.
[0130] Thus, a method for efficiently performing MPI operations has been described.
[0131] Now an example network architecture in which the present technology can be implemented will be described. It will also be understood that the present technology can also be implemented in other network architectures while achieving one or more of the discussed advantages. The example network architecture provides improved performance, for example, improved performance of the MPI collective operations discussed herein. It will be understood that any of nodes 110, 120, 130, 140, 310, 410 can operate in the following network architecture and perform method 200. The network architecture proposed herein can achieve the following advantages:
[0132] a) Large-scale port-level all-to-all connections. For example, each transceiver can be fully connected. In other words, the transceivers in the current architecture are not partially connected.
[0133] b) Full-capacity node-to-node connections. In other words, the current architecture is not restricted by a single transceiver for each source-destination pair.
[0134] c) High-capacity (e.g., >12.8 Tbps, or >10 Tbps) and large-scale (e.g., >4096 nodes) communication can be achieved.
[0135] d) High reliability with no single point of failure.
[0136] e) Suitable for both HPC applications and DCN applications.
[0137] f) Dynamic trade-off of bandwidth and connectivity.
[0138] In fact, an optical circuit switching network is provided, including: multiple nodes, each node including one or more optical transceivers configured to implement time-division multiplexing such that each node belongs to one of multiple transmitting groups or one of multiple receiving groups at a given time; multiple one-to-many switches, where each optical transceiver of each node in the transmitting group nodes of the multiple nodes is connected to an one-to-many switch among the multiple one-to-many switches; multiple many-to-one switches, where each optical transceiver of each node in the receiving group nodes of the multiple nodes is connected to a many-to-one switch among the multiple many-to-one switches; and multiple photon subnet units, where each port of each of the one-to-many switches and the many-to-one switches is connected to a different photon subnet unit.
[0139] In some examples, the optical circuit switching network includes port-level connections. For example, the multiple nodes can have port-level connections.
[0140] In some examples, the transceivers of the nodes in the transmitting group are transmitters, and the transceivers of the nodes in the receiving group are receivers.
[0141] In some examples, each photon subnet unit is configured to connect corresponding different sets or clusters of nodes belonging to the transmitting group and the receiving group.
[0142] In some examples, each node includes multiple optical transceivers. Thus, the number of available paths in the network is increased, the resilience of the network due to the additional paths is improved, and the communication ability of each node is also enhanced. In addition, unrestricted communication between nodes is achieved, and the bandwidth is increased.
[0143] In some examples, each photon subnet unit is an optical coupling subnet unit or a subnet routing unit.
[0144] In some examples, the network includes one or more clusters, each cluster includes one or more groups, each group includes one or more of a plurality of nodes; the number of clusters is x, the number of groups in each cluster is J, and the number of nodes in each group is Λ, where J <= x; and the number of nodes Λ in each group is equal to the number of different wavelength channels available in the network, and the optical transceivers are tunable to transmit and / or receive on the different wavelengths so as to select a given node to communicate with.
[0145] In some examples, each photonic subnet unit is configured to connect the same ports of the respective one-to-many and many-to-one switches of nodes in different cluster pairs. In some examples, for each transceiver, there is a unique subnet unit for each cluster pair. Thus, nodes in different cluster pairs can communicate efficiently.
[0146] In some examples, different photonic subnet units connect different sets of nodes. Thus, full connectivity can be achieved.
[0147] In some examples, each photonic subnet unit includes J ΛxΛ star couplers, each coupler is connected to the ports of J Λ filter arrays, and each filter array is configured to select different wavelengths. Thus, the network can communicate efficiently between nodes.
[0148] In some examples, Λ is equal to the number of different wavelength channels available in each subnet unit. In some examples, Λ is equal to the number of wavelengths that each laser in the system can transmit and / or each tunable filter can receive. In addition, each node may have a fixed transmitter and a tunable receiver. Thus, full utilization of the available wavelengths in the network is provided. Thus, the network and hardware resources are used more efficiently. In addition, the inventors have found that this arrangement provides improved communication performance between nodes.
[0149] In some examples, each node includes bx transceivers, which are grouped into x transceiver groups, each group having b transceivers. In some examples, b = 1, and in other examples, b is greater than 1.
[0150] For example, each transceiver is connected to a different set of subnet units to communicate with the same transceiver of all nodes. The number x of transceiver groups may be equal to the number of communication groups (also referred to as clusters), so that each node can simultaneously send information to all communication groups. This improves the communication efficiency between nodes and is particularly useful for HPC applications.
[0151] In addition, in the example, each transceiver group can operate independently and send and receive from any node at any time step. This means that the same node can use different transceivers to send to multiple nodes simultaneously. Each node can send at full capacity simultaneously to any of the following: nodes in different clusters and groups, nodes in different clusters and the same group, nodes in the same cluster but different groups, different nodes in the same cluster and the same group, the same node in the same cluster and the same group. This makes communication between node pairs unrestricted and increases the bandwidth connectivity between node pairs or sets of nodes by using multiple transceivers in parallel. In the comparative example labeled PULSE, there is a single connection between any node pair, so the bandwidth cannot be increased. Additionally, in PULSE, whenever a connection is established, the node will be unable to communicate with Λx - 1 nodes because only a single transceiver handles connections to / from all nodes with the same group number in all clusters.
[0152] In some examples, b transceivers of a given transceiver group are configured to receive their respective optical inputs from a shared light source circuit. Thus, the resources in the network are used more efficiently. In some examples, the light source circuit is a tunable laser. In some examples, b transceivers of a given transceiver group share the same control. In some examples, all transceivers of a given transceiver group can share a tunable source and control both the switch and the tunable filter as needed.
[0153] In some examples, at a given time: b transceivers of a given transceiver group are configured to send to a given optical transceiver of a given receiving group; and the transceivers of a second given transceiver group are operable to send to at least one of the following: the given optical transceiver of the given receiving group, a second optical transceiver of a second different receiving group.
[0154] Thus, the transceivers in the group can send to the same destination, thereby increasing the total bandwidth. Additionally, since each transceiver group can be independent, the transceiver groups can send to different or the same destinations simultaneously. Thus, the bandwidth and connectivity are further increased.
[0155] In some examples, the total number of photonic subnet units in the network is bx 3 . The inventors have found that this number of subnet units is particularly suitable for the network architecture and increases the connectivity in the network.
[0156] In some examples, the cardinality of each photonic subnet unit is ΛJ x ΛJ. In other words, the number of input / output ports of the subnet unit is ΛJ x ΛJ. The inventors have found that this arrangement provides increased connectivity in the network.
[0157] For example, the architecture can be a subnet-based architecture, where different subnets (which may be interchangeably referred to herein as photon subnet units) connect different sets of nodes. Each subnet connects the same ports of all nodes in different cluster pairs. Each subnet can be a ΛJ×ΛJ network device. It can include the following combinations: J Λ×Λ broadcast (OCS) or routing (wavelength routing OCS) elements, followed by Λ J×J broadcast (OCS) or switching (OCS) elements; or an array of J of Λ fixed filters (single wavelength) or amplifiers (SOA or others) or Λ Jx1 WDM multiplexers, followed by Λ 1x J tunable demultiplexing filters (removing an actively selected wavelength for each port).
[0158] In some examples, the network includes bx paths between the nodes in the sending group and the nodes in the receiving group. Thus, fault tolerance and reliability are improved. If a certain subnet fails, communication can still occur between all nodes, and the only difference is that the transmitters connected to that subnet cannot be used. Additionally, the number of paths and the number of transceivers can create replicated copies of the network.
[0159] In some examples, a one-to-many switch is configured to select a given node in the receiving group to receive the transmitted data; and a many-to-one switch is configured to select a given node in the sending group to transmit the transmitted data. Thus, the switch can efficiently perform source and destination / path selection. In some examples, the ports of the one-to-many switch determine the destination communication group. In some examples, the ports of the many-to-one switch determine the source communication group.
[0160] In some examples, the photon subnet unit is configured to perform one of the following techniques: broadcast and select, route and broadcast, route and switch, broadcast filter amplify and broadcast, broadcast filter and switch, broadcast filter multiplex and demultiplex. The inventors have found that these techniques allow for efficient communication between nodes.
[0161] In some examples, each of the photon subnet units includes one or more of the following: star coupler, filter, space switch, semiconductor optical amplifier, arrayed waveguide grating router, AWGR, multiplexer, and tunable add-drop demultiplexing filter. Thus, the subnet unit can be configured for a specific network configuration, e.g., depending on the fixed / tunable type of transceiver used. Thus, flexibility is improved.
[0162] In some examples, each optical transceiver includes: a tunable transmit element and a fixed-wavelength filtering receive element; a tunable transmit element and a tunable filtering receive element; a fixed-wavelength transmit element and a tunable filtering receive element; or a tunable transmit element and a filterless receive element. Thus, various forms of transceivers can be used according to the usage. Therefore, flexibility is improved. Optionally, the filtering receive (and filterless) elements can be connected to a multi-to-one switch. In other words, filtering can be performed before one or more multi-to-one switches. In other words, the filtering element can be located before each entry port of the multi-to-one switch. In some embodiments, the filtering element is directly connected to each / any / one port of the multi-to-one switch.
[0163] In some examples, each one-to-many switch includes one or more space switches that are configured in use to activate each port of each one-to-many switch to select the corresponding photonic network unit connected to the activated port. Specifically, each port of the subnet unit is connected to a different cluster. In some examples, the space switch can include a semiconductor optical amplifier.
[0164] In some examples, one or more one-to-many switches are semiconductor optical amplifier-based switches, and one or more multi-to-one switches are semiconductor optical amplifier-based switches. In some examples, one or more one-to-many switches are semiconductor optical amplifier gated splitters, and one or more multi-to-one switches are semiconductor optical amplifier gated couplers. Thus, fast switching times can be achieved. In some examples, depending on the type of space switch, splitters and couplers are not necessary.
[0165] The list of network resources accessible to each node can be: transceiver group (2D: b, x), wavelength, space / path, and time slot, (xDM: SDM, WDM, TDM, transceiver).
[0166] In some examples, the one-to-many switch selects the destination cluster to which the node will send, and the multi-to-one switch selects the receiver from which source cluster to receive. Additionally, wavelength selection can be used to select the destination / source node in the group (WDM node selection). Switch port selection can be used to select the destination and source clusters (cluster selection). Broadcasting or switching is performed between nodes having the same node number in all groups within the same cluster in the same subnet (group selection). All of these can be performed at the transceiver level.
[0167] Furthermore, communication can be active in synchronized time slots, where one or more transceivers can communicate with one or more destinations in the same time slot. In this way, one or more transceivers can be used between node pairs and node subsets. This allows communication up to full capacity between node pairs.
[0168] An electronic time-division multiplexing circuit-switched network is also provided, including: a plurality of nodes, each node including one or more transceivers and configured to implement time-division multiplexing such that each node belongs to one of a plurality of transmission groups or one of a plurality of reception groups at a given time; a plurality of one-to-many switches, wherein each transceiver of each node in the transmission group nodes of the plurality of nodes is connected to a one-to-many switch among the plurality of one-to-many switches; a plurality of many-to-one switches, wherein each transceiver of each node in the reception group nodes of the plurality of nodes is connected to a many-to-one switch among the plurality of many-to-one switches; and a plurality of subnet units, wherein each port of each of the one-to-many switches and the many-to-one switches is connected to a different subnet unit.
[0169] In some examples, each subnet unit is configured to connect to corresponding different sets or clusters of nodes belonging to the transmission group and the reception group.
[0170] In some examples, the subnet is J ΛxΛ space switches (electrical), followed by Λ J x J broadcast units (RF couplers or optical couplers) or space switches. Λ in this example is equal to the total number of ports of the subnet switch, which is equal to the number of nodes in each group.
[0171] In some examples, the number of nodes Λ in each group is equal to the number of paths of the space switch in the subnet unit. For example, the space switch in the subnet unit may have ΛxΛ input / output ports.
[0172] It will be understood that the examples can be combined, and the non-optical examples are equally applicable to the network electronic time-division multiplexing circuit-switched network.
[0173] A further discussion of the network architectures of the fourth and fifth aspects will now be presented.
[0174] In some examples, the network is a port-level fully connected network. This is different from the comparative example at the node level because in the comparative example, a single transceiver is used for communication between any pair of nodes.
[0175] In some examples, the network includes a plurality of nodes (x, J, Λ), which are organized into clusters, groups, and nodes in each group. In these examples, a cluster contains one or more groups, each group contains one or more nodes. The number of clusters in the system is x, the number of groups in each cluster is J, and the number of nodes in each group is Λ. The number of nodes Λ in each group is equal to the number of wavelengths available in the OCS (optical circuit switching) system or the number of paths in the e-TDM (electronic time-division multiplexing) system switch. In these examples, J ≤ x and Λ mod x = 0.
[0176] In some examples, the architecture is a subnet-based architecture where different subnets connect different sets of nodes. Each subnet connects the same ports of all nodes in different cluster pairs (whereas in the comparative example, only group pairs are connected). Each subnet is a ΛJ×ΛJ network device. It can include the following combination: Λ×Λ broadcast (OCS) or routing (wavelength-routed OCS / space-switched e-TDM) elements, followed by ΛJ×J broadcast (OCS / e-TDM) or switching (OCS / e-TDM) elements. The comparative example (e.g., PULSE) has Λ×Λ broadcast or wavelength-routed elements. In some examples, there are a total of bx 3 subnets, while in the comparative example (e.g., PULSE), it can be x 4 .
[0177] In some examples, through the use of specific transceivers, port-level all-to-all communication can be achieved. Each transmitter and receiver can be connected to 1×x and x×1 space switches respectively, and each port of each switch is connected to a different subnet. The switch port at the transmitting end selects the destination cluster to which to send. The switch port at the receiving end selects the source cluster from which to receive. In the comparative example, only groups within a specific cluster are selected.
[0178] In some examples, the transceivers allow wavelength tunability (across Λ wavelengths) at either / both the transmitting and receiving ends of the OCS. The wavelength selection at either end enforces the source-destination nodes for each group pair for each communication. For an e-TDM system, this can be performed by selecting a path in the subnet space switch. The wavelength selection for each source-destination pair can be the same or independent, depending on the transceiver group used. The selection can depend on the selection of the x:1 switch. If the switch is formed by an SOA-gated multiplexer, each transceiver group can use a different mapping.
[0179] In some examples, each node in the system is equipped with bx transceivers. These transceivers are grouped into x transceiver groups, each group having b transceivers. Each transceiver can be connected to a different set of subnets to communicate with the same transceiver of all nodes. The number x of transceiver groups can be equal to the number of communication groups, so that each node can simultaneously send information to all communication groups (which is useful for HPC applications). Each transceiver in a transceiver group can share the same tunable laser (if the OCS has a tunable tx) and the same controller (for OCS and e-TDM). All transceivers in the same transceiver group can send to the same node.
[0180] In some examples, each transceiver group can operate independently and send and receive from any node at any time step. This means that the same node can use different transceivers to send to multiple nodes simultaneously. Each node can send at full capacity simultaneously to any of the following: nodes in different clusters and groups, nodes in different clusters and the same group, nodes in the same cluster but different groups, different nodes in the same cluster and the same group, the same node in the same cluster and the same group. This means that communication between node pairs is unrestricted, and bandwidth connectivity between node pairs or sets of nodes can be increased by using multiple transceivers in parallel.
[0181] Therefore, since there is no single point of failure, the present method improves network resilience. When a network component (trx / subnet) fails, additional paths can be available, and the network is reconfigured so as to avoid using the failed resources.
[0182] The comparison with the comparative example PULSE is now shown in Comparative Tables 1 and 2 below:
[0183]
[0184]
[0185]
[0186] Figure 5 An example network 500 according to the present technology is schematically shown. The network 500 can be an optical circuit switching network, or alternatively an electronic time division multiplexing circuit switching network. The network 500 includes a plurality of nodes 501, 504, 507, 510. These nodes can implement the present technology described herein, such as method 200. Each node is configured to implement time division multiplexing such that each node belongs to a transmitting group or a receiving group at a given time. It will be understood that the transmitting group and the receiving group can vary over time, and the nodes can generally transmit and receive. In Figure 5 this case, nodes 501 and 504 can be regarded as the transmitting group, and nodes 507 and 510 can be regarded as the receiving group.
[0187] Each node can include one or more transceivers 502, 505, 508, 511. In some examples, each node includes a plurality of transceivers. These transceivers can be optical transceivers in some examples. The network also includes a plurality of one-to-many switches 503, 506 and a plurality of many-to-one switches 509, 512. Again, it will be understood that this is defined by the connection direction between the nodes of the transmitting group and the receiving group.
[0188] The transceiver can be an integrated or non-integrated part of the node, or can be a connected circuit. Each transceiver (i.e., 502 and 505) of each node in the sending group of nodes is connected to a one-to-many switch (i.e., 503 and 506). Each transceiver (i.e., 508 and 511) of each node in the receiving group of nodes is connected to a many-to-one switch (i.e., 509 and 512).
[0189] Network 500 also includes a plurality of subnet units 513, 514 (also referred to as subnets). The subnet units can be optical subnet units, and / or coupling units or routing units. As Figure 5 shown, each port of each of the one-to-many switches (503, 506) and the many-to-one switches (509, 512) is connected to a different subnet unit 513, 514. It will be understood that Figure 5 an example configuration is shown, and the number of nodes, transceivers, switches, and subnet units can vary.
[0190] Figure 6 A method 600 for communicating in a network (e.g., network 500) is shown. In this example, the network is an optical circuit switching network.
[0191] At S601, light encoding data for transmission is sent from the optical transceiver of the transmitter node, via a port of the one-to-many switch connected to the node, to the optical subnet unit connected to the port. Referring to Figure 5 , the light can be sent from the optical transceiver 502 of node 501, via a port of the one-to-many switch 503 connected to node 501, to the optical subnet unit 513 connected to the port.
[0192] At S602, light from the optical subnet unit is received at the receiver node via the many-to-one switch connected to the receiver node. Referring again to Figure 5 , the light from the optical subnet unit 513 can be received at the receiver node 510 via the many-to-one switch 512 connected to the receiver node 510. Thus, in this way, light can be transmitted from the transmitter node to the receiver node through the network.
[0193] It will be understood that a version of method 600 can be executed in the electronic time division multiplexing architecture described herein.
[0194] In some examples, a method for communicating in an electronic time division multiplexing architecture network is provided, the method including: sending data from a transceiver of a transmitter node to a subnet unit connected to a port via a port of a one-to-many switch connected to the transmitter node; and receiving data at a receiver node from the subnet unit via a many-to-one switch connected to the receiver node.
[0195] Figure 7and Figure 8 illustrates an example subnet unit type.
[0196] Figure 7 a depicts the broadcast and select (B&S) type. Figure 7 The passive subnet shown in a includes an N×N star coupler that connects the i-th transmitter and receiver of two different communication groups. The number of ports N in this subnet can be different from the number of wavelengths Λ (scaling independent of the wavelength channel map). The number of ports N can be up to Λx. The larger the number of ports of the star coupler, the higher the possible loss. The system can implement using wavelength tunability at the transmitter and / or receiver end. This example can use either or both of a tunable receiver / tunable transmitter. In some examples, the transmitter is tunable and the receiver is fixed, which may be preferred in certain use cases. The star coupler connects all nodes between cluster pairs.
[0197] Figure 7 b depicts the route and broadcast (R&B) type. The route and broadcast architecture is an N×N port subnet that includes two main components: an arrayed waveguide grating (AWGR) and a star coupler. At the input stage, there are J Λ×Λ AWGRs, each of which routes the information from the i-th port of each individual rack. All the l-th output ports of each AWGR are then connected to Λ J×J star couplers. The j-th output port of the k-th star coupler is connected to the i-th port of the k-th device in the j-th rack. Since the maximum subnet size is N = Λx, this network requires x Λ×Λ AWGRs and Λ x×x star couplers. In the case of J = 1, the subnet can be a single AWGR. The wavelength routing followed by the broadcast requires wavelength tunability at both the transmitter and receiver ends. This example can use tunability at both the transmitter and receiver. The J Λ×Λ AWGRs, followed by Λ J×J star couplers, are connected between the same port numbers of each AWGR. All ports 1 of the J AWGRs are connected in star coupler 1.
[0198] Figure 7c depicts the routing and switching (R&S) type. In the R&S example, each output port of any AWGR is followed by a gated 1×J splitter, for a total of N splitters. Then, the output ports of the SOAs are connected to an array of N J×1 combiners. The output of the SOA is connected to the k-th port of the splitter of the λ-th output port of the AWGR, which routes the information of the j-th transmitting rack, and is connected to the j-th port of the combiner connected to the λ-th node of the k-th receiving rack. This effectively creates an array of Λ J×J space switches. In this type of system, a single path at the two combiners will actually receive the information, so no tunability is required at the receiver. Routing through the AWGR requires transmitter tunability. This example can use tunability at both the transmitter and the receiver. J Λ×Λ AWGRs, followed by Λ J×J space switches, are connected between the same port numbers of each AWGR. All port 1s of the J AWGRs are connected in space switch 1.
[0199] Figure 8 d depicts a broadcast, filtering, amplification, and broadcast subnet with fixed receivers. In this example, the following configuration is used: star coupler + filter + amplifier + star coupler, with a tunable transmitter and fixed receivers. J Λ×Λ star couplers, each followed by J Λ×Λ filters, such that all ports with the same number have the same wavelength. These can each be followed by an amplification stage, and then Λ J×J star couplers connected between all filters with the same wavelength. All port 1s of the J coupler + filter + amplifier are connected in star coupler 1. This figure shows a configuration with a tunable transmitter and fixed receivers. In this configuration, the port that has filtered the i-th wavelength will be coupled by the i-th star coupler.
[0200] Figure 8 e depicts a broadcast, filtering, amplification, and broadcast subnet with tunable receivers. In this example, the following configuration is used: star coupler + filter + amplifier + star coupler, with a tunable transmitter and tunable receivers. J Λ×Λ star couplers, each followed by J Λ×Λ filters, such that all ports with the same number have the same wavelength. These can each be followed by an amplification stage, and then Λ J×J star couplers connected between all filters with the same wavelength, in a manner similar to Figure 8 the example in d. Figure 8 e shows a configuration with a tunable transmitter and tunable receivers. In this configuration, the subsequent ports of different filters (port 1 of filter 0 and port 1 of filter 1, etc., port 1 of filter 0 and port 2 of filter 1...) are connected in a cyclic manner by Λ J×J star couplers. Figure 9 a and Figure 9An example of such a connection is shown in b.
[0201] A further implementation of this could be A) using tunability at both ends. B) For the second - stage coupler (after the filter and possibly the amplifier), the wavelengths are not the same (for all port 1 of all J - couplers), but are connected to different wavelengths in a cyclic manner (for the first coupler of the second stage, port / w 1 of coupler 1 is connected to port / w 2 of coupler 2, etc.; for the second coupler of the second stage, port / w 2 of coupler 1 is connected to port / w 3 of coupler 2, etc.). This can reduce contention. Figure 9 a and Figure 9 An example of such a connection is shown in b.
[0202] Figure 8 Figure f depicts a broadcast, filtering, and switching subnet. In this example, the configuration is as follows: star coupler + filter + switch, with a tunable transmitter and a fixed receiver. J Λ×Λ star couplers, each followed by J Λ×Λ filters such that all ports with the same number have the same wavelength. These are followed by Λ J×J space switches, connected between all filters with the same wavelength. All port 1 of the J couplers + filters are connected in the space switch.
[0203] Figure 8 Figure g depicts a broadcast, filtering, multiplexing, and demultiplexing subnet. In this example, the following configuration is used: star coupler + filter + multiplexer + demultiplexer, with a tunable transmitter and a high - bandwidth receiver. J Λ×Λ star couplers, each followed by J Λ×Λ filters such that all ports with the same number have the same wavelength. The subsequent ports of different filters (port 0 of filter 0 is connected to port 1 of filter 1, etc., port 1 of filter 0 is connected to port 2 of filter 1...) are connected to Λ J×1 multiplexers. Each of these multiplexers is followed by Λ tunable add - drop 1×J filters. Each of these filters extracts the wavelength for each port. These components can be located in the subnet or at the edge. Each port connects to devices with the same node ID in different racks of the same communication group. Figure 8 Figure g shows a filter located at the edge. Figure 9 An example connection for this subnet is shown in b.
[0204] Figure 8h depicts a broadcast, filtering, multiplexing, and demultiplexing subnet. In this example, the following configuration is used: star coupler + filter + multiplexer + demultiplexer, with a tunable transmitter and a high-bandwidth receiver. There are J Λ×Λ star couplers, each followed by J Λ×Λ filters such that all ports with the same number have the same wavelength. The subsequent ports of different filters (port 0 of filter 0 and port 1 of filter 1, etc., port 1 of filter 0 and port 2 of filter 1...) are connected to J Λ×1 multiplexers. Each of these multiplexers is followed by J 1×Λ filters for tunable add-drop. Each of these filters extracts the wavelength for each port. These components can be located in the subnet or at the edge. Each port is connected to devices with different node IDs in the same rack of the same communication group. Figure 8 h shows the filters located at the edge. Figure 9 An example connection for this subnet is shown in c.
[0205] For Figure 8 g and Figure 8 In the example of h, the second stage can be formed by J Λ:1 multiplexers. Each multiplexer can select the subsequent node of each device in a cyclic manner. In this way, for the same switch port of all node IDs in the same rack and the same communication group, the add-drop cascading of the filters can be performed.
[0206] Therefore, the example subnet connection can be as follows:
[0207] Broadcast and Selection (As Figure 7 shown in a): Each subnet can be formed by a ΛJ×ΛJ star coupler. Thus, in some examples, one or more photonic subnet units are configured to perform broadcast and selection, and one or more of the photonic subnet units include a ΛJ×ΛJ star coupler.
[0208] Routing and broadcast (as Figure 7 shown in b): Each subnet can be formed by an array of J Λ×Λ AWGRs connected to Λ J×J star couplers such that the i-th port of the j-th AWGR is connected to the j-th port of the i-th star coupler. Thus, in some examples, one or more photonic subnet units are configured to perform routing and broadcast, and one or more of the photonic subnet units include an array of J Λ×Λ AWGRs connected to Λ J×J star couplers such that the i-th port of the j-th AWGR is connected to the j-th port of the i-th star coupler.
[0209] Routing and Switching (As Figure 7As shown in FIG. c: Each subnet includes an array of J Λ×Λ AWGRs and an array of Λ J×J space switches, which are connected in such a way that the i-th port of the j-th AWGR is connected to the j-th port of the i-th space switch. The connection order can be the AWGR array followed by the switch array, or it can be the switch array followed by the AWGR array. Thus, in some examples, one or more photonic subnet units are configured to perform routing and switching, and wherein, one or more photonic subnet units include an array of J Λ×Λ AWGRs and an array of Λ J×J space switches, which are connected in such a way that the i-th port of the j-th AWGR is connected to the j-th port of the i-th space switch.
[0210] Broadcast Filtering Amplification and Broadcasting (as shown in Figure 8 FIG. d): Each subnet includes J Λ×Λ star couplers followed by an array of J Λ×Λ optical filter arrays, which are configured such that the i-th port of the j-th star coupler is connected to the i-th port of the j-th filter, which retrieves the i-th channel for all J filter arrays. After the array of filter arrays, an amplification stage can optionally be provided at each port by using an array of J Λ×Λ semiconductor optical amplifiers. After the array of filter arrays and the optional amplification, there is an array of Λ J×J star couplers. This can be orthogonally connected or can be loop-connected. Orthogonal connection: The i-th port of the j-th array (filter or SOA) is connected to the j-th port of the i-th star coupler. Loop connection: Each port of each star coupler is connected to a different port of a different array. Figure 9 FIG. a and Figure 9 FIG. b show examples of such connections. Thus, in some examples, one or more photonic subnet units are configured to perform broadcasting, filtering, amplification, and broadcasting, and wherein one or more photonic subnet units include J Λ×Λ star couplers followed by an array of J Λ×Λ optical filters, which are configured such that the i-th port of the j-th star coupler is connected to the i-th port of the j-th filter, followed by an array of Λ J×J star couplers.
[0211] Broadcast and Switching (as shown in Figure 8 FIG. f): Each subnet includes two stages:
[0212] 1) An array of J Λ×Λ star couplers followed by an array of J Λ×Λ optical filters, which are configured such that the i-th port of the j-th star coupler is connected to the i-th port of the j-th filter, which retrieves the i-th channel for all J filter arrays. After the array of filters, an amplification stage can optionally be provided at each port by using an array of J Λ×Λ semiconductor optical amplifiers.
[0213] 2) An array of Λ J×J space switches.
[0214] The i-th port of the j-th array (filter or SOA) is connected to the j-th port of the i-th space switch. The connection order can be a star coupler + filter (+ optional SOA) array followed by a switch array, or it can be a switch array followed by a star coupler + filter (+ optional SOA) array. Thus, in some examples, one or more photonic network units are configured to perform broadcast and switching, and one or more of the photonic network units include J Λ×Λ star couplers, followed by an array of J Λ×Λ optical filters configured such that the i-th port of the j-th star coupler is connected to the i-th port of the j-th filter, followed by Λ J×J space switches.
[0215] Broadcast Filtering Multiplexing and Demultiplexing (as Figure 8 g Figure 8 and
[0216] h show): Each subnet includes an array of J Λ×Λ star couplers, followed by an array of J Λ×Λ optical filters configured such that the i-th port of the j-th star coupler is connected to the i-th port of the j-th filter, which retrieves the i-th channel for all J filter arrays. After the array of filter arrays, an amplification stage can optionally be provided at each port by using an array of J Λ×Λ semiconductor optical amplifiers. Then it can be followed by: Figure 9 1) An array of Λ J×1 multiplexers, connected to the previous stage in a cyclic manner (an example connection is shown in
[0217] b). Each of the Λ filters is connected to an array of Λ 1×J tunable demultiplexers formed by a series of cascaded J add-drop filters such that the j-th multiplexer is connected to the j-th demultiplexer. The demultiplexer stage can be located within the subnet or at the edge of the network. Figure 9 or 2) An array of J Λ×1 multiplexers, connected to the previous stage in a cyclic manner (an example connection is shown in
[0218] Thus, in some examples, one or more photonic network units are configured to perform broadcasting, filtering, multiplexing, and demultiplexing, and one or more of the photonic network units include J Λ×Λ star couplers followed by an array of J Λ×Λ optical filters configured such that the i-th port of the j-th star coupler is connected to the i-th port of the j-th filter, followed by: an array of Λ J×1 multiplexers, each multiplexer connected to an array of Λ 1×J tunable demultiplexers formed by a series of cascaded J add-drop filters such that the j-th multiplexer is connected to the j-th demultiplexer; or an array of J Λ×1 multiplexers, each multiplexer connected to an array of J 1×Λ tunable demultiplexers formed by a series of cascaded J add-drop filters such that the j-th multiplexer is connected to the j-th demultiplexer.
[0219] It will be understood that one or more photonic network units present in the network can be configured to perform the above techniques, and subnet units performing different techniques can be combined in the same network.
[0220] A further example subnet unit is implemented as follows:
[0221] i) Either or both of a tunable receiver / a tunable transmitter. In some examples, the transmitter is tunable and the receiver is fixed, which may be preferred in certain use cases. The star coupler connects all nodes between the cluster pairs.
[0222] ii) Tunability at both the transmitter and the receiver. J Λ×Λ AWGRs are followed by Λ J×J star couplers connected between the same port numbers of each AWGR. All port 1s of the J AWGRs are connected in star coupler 1.
[0223] iii) Tunability at both the transmitter and the receiver. J Λ×Λ AWGRs are followed by Λ J×J space switches connected between the same port numbers of each AWGR. All port 1s of the J AWGRs are connected in space switch 1.
[0224] For the current subnet example, an amplification stage using an SOA may be present after the filtering stage.
[0225] The system before any port of a multi-to-one switch can have any of the following: no filter, a fixed filter, or a tunable filter. A no-filter system can be used for: A) coherent detection where low amplification is required at the switching stage; B) cases where filtering has been performed in the subnet, such as in a fixed receiver B+F+A+B system. Fixed filter: Can be used for both coherent detection communication and direct detection communication. It can be used in cases where the system requires fixed reception. The filter can be selected such that it will be able to retrieve a single wavelength from multiple wavelengths. The filter can be selected such that: no two filters retrieve the same wavelength for the same switch port of all nodes in a rack. The wavelength selection for each switch port in a single node (and transceiver) can be the same or can be different. This selection may be important for the selection of multi-to-one switch technology (different wavelengths can allow the use of SOA-gated AWG as a switch). Tunable filter: Can be used for both coherent detection communication and direct detection communication. It can be used in cases where: the system requires a tunable receiver (direct detection) and the signal requires edge amplification in a coherent system.
[0226] In some examples, the xgs is arranged before the multi-to-one switch element.
[0227] The tunable filter can be an add-drop filter that connects devices, such as in the case of a star coupler + filter + multiplexer + demultiplexer subnet (f above).
[0228] Working examples
[0229] The present technology will now be described with reference to working examples. The inventors have found that various aspects of the working examples can improve the execution efficiency of MPI collective operations and reduce the completion time of collective operations. As part of this working example, a network architecture is described. It will be understood that the present technology can (and in some cases) be implemented in this network architecture. In some examples, implementing the present technology in the following architecture can further improve the execution efficiency of MPI operations. In fact, the inventors have found that compared with comparative examples where the present technology is not used, the present technology has at least achieved the following advantages:
[0230] 1. High-capacity communication between nodes (e.g., >12.8 Tbps), making the network architecture suitable for HPC and DDL application requirements.
[0231] 2. High scalability (e.g., >4096 nodes). Capable of handling increasingly complex workloads.
[0232] 3. Achieve nanosecond-level circuit reconfiguration through wavelength switching and B&S (Broadcast and Select). This allows each node to communicate with any other node with little limitation on the communication degree; allows set operations using logic graphs with significantly lower diameters without sacrificing bandwidth; allows the proposed architecture to handle the rapidly changing circuits required for DCN traffic.
[0233] 4. Port-level all-to-all connectivity and rearrangeable or strictly non-blocking communication. Any transceiver can send to / receive information from any node. The communication blocking probability only depends on the selection of the subnet.
[0234] 5. A completely passive interconnection system. Remove the complexity from the network core and transfer it to the edge.
[0235] 6. Unrestricted multi-node communication and reliability with no single point of failure. Each node can communicate with other nodes using multiple possible paths, and full-to-full communication is still allowed for any failure of the transceiver / network component, although the capacity will decrease slightly.
[0236] Example Network Architecture
[0237] In some examples, this technology is implemented in an electrical packet-switching network. However, in other examples, this technology is implemented in an optical circuit-switching OCS network. Now, specific examples of this network will be referred to Figure 10 for discussion. It will be understood that any combination of features from the example architectures can be combined, and features can be omitted.
[0238] In this example, the network architecture is a switchless OCS architecture that supports full bipartite bandwidth and high-capacity communication between node pairs, thus providing a fast reconfiguration time (nanosecond-level) and high scalability. This network architecture achieves all-to-all connectivity at the port level, allows unrestricted multi-node communication, and reliability in terms of network component failures. Therefore, this example architecture is optimal for HPC and DDL operations that require high-bandwidth communication between node pairs. In addition, the nanosecond-level circuit reconfiguration time and all-to-all connectivity allow each node to communicate with little limitation on the communication degree.
[0239] As Figure 10As shown, the network architecture includes parallel subnets, which are arranged by communication groups (also known as clusters) and transceivers (or transmitters and receivers). In this example, there are x communication groups CG (or clusters), where each communication group contains J racks (also referred to as groups in this article), and the maximum number of racks in each communication group is J = x. Each rack contains Λ devices or nodes, where Λ is also the total number of available wavelength channels. Therefore, the maximum number of nodes in a communication group is N = Λx. Each node is equipped with x transceiver groups, and each transceiver group contains b transceivers that share the same light source. In Figure 10 the example, b = 1. Each transceiver is connected to a 1:x splitter, creating x possible paths for each transceiver. Each path is selected by activating the SOA (Semiconductor Optical Amplifier) attached to each port of the 1:x splitter and connected to a different subnet (and thus a different communication group). In this way, each transceiver can communicate with each communication group. Each receiver (or transceiver) is connected to an x:1 combiner, enabling each receiver to receive information from each communication group. Under the proposed network configuration, the i-th transmitter of any node can send information to the i-th receiver of each node, thus achieving all-to-all per-transceiver communication. In this example, the topology requires a total of bx 3 subnets, that is, one subnet is required for each communication group pair for each transceiver.
[0240] As Figure 10 shown in the right table, the example architecture can be scaled up to Λx 2 nodes, providing a total capacity of bBΛx 2 , where B is the effective line rate of each transceiver. The bipartite bandwidth is ΛJx 3 / 2, and the total number of physical links required is 2Jx 2 , because the paths can be grouped by rack and source-destination communication group. Source-destination selection and circuit reconfiguration are performed through path / transceiver, wavelength, and time slot mapping.
[0241] There are many example selections for subnets: (i) a star coupler with N ports (Broadcast and Select, B&S); (ii) J parallel Λ×Λ Arrayed Waveguide Gratings AWGRs, followed by Λ parallel J×J star couplers, mixing information between the same ports of each AWGR (Route and Broadcast, R&B); or (iii) the same AWGRs, followed by a SOA-based J×J crossbar switch (Route and Switch, R&S). Other example selections include J arrays of Λ fixed filters (single wavelength) or amplifiers (SOA or others), or Λ Jx1 WDM multiplexers followed by Λ 1xJ tunable demultiplexing filters (removing one actively selected wavelength per port).
[0242] Each node in the example architecture can have coordinates defined by a communication group, a rack number within the communication group, and a node number within the rack (or cluster, group number, and node number), as discussed in more detail herein. For example, each node can be identified based on (communication group, rack number, node number).
[0243] Figure 11 and Figure 12 illustrate how the example architecture handles different communication patterns. The example architecture shown in these figures is a fixed receiver broadcast and select (B&S) architecture, but it will be understood that other techniques can be implemented alternatively.
[0244] In Figure 11 , a many-to-many communication pattern in multiple time slots across multiple sources and destinations is shown within: a) a single source-destination communication group pair, and b) multiple communication groups. For Figure 11 .a), communication occurs between multiple source nodes 1, λ, Λ in rack j and communication group c and destination nodes 1, γ, Λ in rack k and communication group d by using the t-th transceiver. At the sending end, each node has a tunable transmitter followed by a 1:x space switch (implemented by an SOA gated splitter), while at the receiving end, each receiver has a filtering (single wavelength) x:1 switch (SOA gated coupler) in front of it, making it a fixed receiver.
[0245] Each node in the rack receives at a different wavelength, and the different wavelengths are represented by the receiving node, receiver, and filter colors in Figure 11 and Figure 12 . There is a single subnet: c;d;t between the communication group pair (c-d) of the t-th transceiver, which allows all transmitters t of all source nodes in communication group c to communicate with all destination nodes in communication group d. To perform communication and transmission through the correct subnet, the correct switch ports need to be selected at both the sending and receiving ends. At transmission, the switch port corresponds to the destination communication group (port d is used to communicate with the d-th communication group), and at reception, the switch port corresponds to the source-destination group. For Figure 11 and Figure 12 , the color of the transmit switch port and subnet matches the color of the destination communication group, and similarly, the color of the receive switch port matches the color of the source communication group from which the port receives.
[0246] At each time slot, each node sets its destination by selecting its receiving wavelength, as shown in Figure 11.As shown at the transmitting end of.a), where the transmitting node (c,j,λ) sends information to nodes (d,k,γ) and (d,k,1) by selecting wavelengths γ and 1 for time slots 1 and 2 respectively. In each subnet, due to the broadcast principle, each active wavelength is available at each output port (represented by the rainbow color in Figure 11 and Figure 12 ), and the correct wavelength for each destination is restored by the filter before each port of the 1:x switch. For these two time slots, since the communication group pairs of the source node and the destination node are constant, ports d of the transmitting - end switch and port c of the receiving - end switch are selected respectively. In a similar way, node (d,k,γ) receives from nodes (c,j,λ) and (c,j,1) in different time slots, and their transmitters are tuned to the γ - th wavelength. In other words, the source - destination communication group pairs remain the same across different time slots, but the communication uses different node pairs. Since the source - destination communication group pairs are constant, the port switches at the transmitting and receiving ends are also constant.
[0247] Figure 11 .b) shows a similar many - to - many pattern between different nodes (for tx: (1,λ,Λ) and for rx: (1,γ,Λ)) in different racks (for tx: (i,j,k) and for rx: (l,m,n)) for different communication groups (for tx: (1,c,x)) and for rx: (1,d,x)). Each communication group pair is connected by a subnet and accessed through specific source and destination switch - port selections. As Figure 11 .a) shows, the node selection in the rack is performed by wavelength selection for each time slot, while different communication groups are accessed by gating different ports of the transmitting and receiving switches. In Figure 11 .b), node (c,j,λ) communicates with nodes (d,m,γ) and (1,l,1) in different time slots by selecting wavelengths 1, γ and gating ports d, 1 and c, c of the transmitting and receiving switches respectively in each time slot. Selecting different switch - port pairs at each time slot results in different communication - group communications, thus allowing efficient port - level all - to - all communication with fast reconfiguration.
[0248] In addition, Figure 11can be considered as showing an example of a many-to-many communication mode for a network that has a star coupler-based network with tunable transmitters and fixed receivers. A source node (c,j,λ) transmits to a node (d,k,γ) using a transceiver group t by: selecting a wavelength γ for transmission (selecting the destination node number in the receive cluster), and using port d of a 1×x switch such that the information is routed to a subnet (c,d,t) that handles communication between the t-th transmitter of all nodes in cluster c and the t-th receiver of all nodes in cluster d. A destination node (d,k,γ) receives from a source node (c,j,λ) by: selecting the switch port c of its x×1 switch (which allows reception from the t-th transmitter of all nodes in cluster c), and recovering its receive wavelength by filtering. In this figure, multiple paths in different time slots are shown within the same cluster and across multiple clusters and groups.
[0249] Figure 12 shows different communication modes for each same time slot: 12.a) one-to-many, 12.b) many-to-one, and 12.c) one-to-one. For all communication modes, Figure 12 depicts communication between multiple source nodes (1,λ,Λ) in rack j and communication group c and destination nodes (1,γ,Λ) in rack k and communication group d by using multiple transceivers.
[0250] Figure 12 .a) shows a one-to-many communication mode from a source node (c,j,λ) to all nodes in rack k of communication group d. Each transceiver of the source node transmits to different destinations in the same time slot by selecting different wavelengths. If the destinations are in different communication groups, different transmit and destination switch ports are selected for each time slot, similar to Figure 12 .b).
[0251] Figure 12 .b) shows a many-to-one communication mode where a destination node (d,k,γ) receives simultaneously from multiple sources by using different transceivers.
[0252] Figure 12 .c) shows multiple one-to-one communication modes between different source-destination pairs. In this figure, all transmitters of each source node are used to communicate with all receivers of the same destination node, thus ensuring full-capacity communication between the node pairs at any time slot. It should be noted that in some examples, only a subset of the transceivers can be used between node pairs according to application requirements.
[0253] In addition, Figure 12Shows how a network can simultaneously use multiple transceivers to send data to multiple nodes and receive data from multiple nodes, and how multiple transceivers can be used simultaneously between pairs or sets of devices to increase bandwidth. This figure uses the same wavelength selection and switch port selection principles as Figure 11 the same. Although this figure only shows the communication between two rack pairs, Figure 11 and Figure 12 the principles shown in can be generalized to any node.
[0254] The described principles can be used simultaneously to adjust network requests, and they are widely combined for set operations. It should be noted that rack selection is not performed in Figure 11 and Figure 12 . This is because the signals between nodes with the same node number in different racks are coupled together and the same information is broadcast to all racks. This may actually cause contention in each subnet. However, multiple paths between each source-destination pair allow communication to be rearrangeably non-blocking and achieve full bandwidth under correct scheduling. In Figure 10 , Figure 11 and Figure 12 , an example architecture with b = 1 is shown, where the transceiver group is equivalent to one transceiver, although it can be understood that b can be greater than 1.
[0255] Switching in the example architecture can be achieved by configuring wavelengths / time slots / paths at the end-node transceivers. For wavelength switching, a wavelength-tunable source (WTS) can be used at the transmitting end. For example, a time-interleaved tunable laser with a gated SOA (e.g., over a wide range of 122 wavelength channels) can be used. These have been shown to achieve an effective switching time of less than 1 nanosecond. At the destination end, the receiver can be tunable or fixed, depending on the subnet implementation. If B&S is implemented, the receiver can operate at a fixed wavelength by using a passive filter. However, wavelength tunability is required when considering subnets with wavelength routing capabilities. Tunability can be achieved by a wavelength filter gated by an SOA or by using an additional tunable laser for coherent detection.
[0256] For space switching, B&S filters based on SOA-gated couplers and combiners can be used. Using SOA-based gating as a space switching mechanism allows sub-nanosecond path selection. In addition, the SOA is also used for amplification.
[0257] Time division multiplexing can be achieved by using predefined time slots. Synchronization and clock data recovery (CDR) use the same principles as those known in the art, in particular PULSE and Sirius. The duration of the time slot can be selected such that the maximum reconfiguration overhead is 5%, resulting in a minimum data transmission time slot of 20 nanoseconds.
[0258] Using a low-power silicon-organic hybrid (SOH) modulator, a transceiver node capacity of (B =) 400 Gbps can be achieved. With this example data rate, the minimum message size that each transceiver can send in one time slot is 950 B. Such small messages are common in large-scale DCN traffic and HPC MPI collective operations.
[0259] In fact, a fast circuit reconfiguration time is desirable for HPC applications, especially a circuit reconfiguration time in the nanosecond range, as it allows for the efficient transmission of small message sizes and allows the use of dynamic collective strategies for MPI operations. Specifically, when the circuit reconfiguration time is less than the node I / O time (transceiver and compute latency), it will not incur any overhead in the transmission time. Since the transceiver latency (and I / O latency) can be as low as tens of nanoseconds, the switching reconfiguration time should be as well. The example architecture and technology achieve this desired switching reconfiguration time. Compared to existing known architectures and technologies, the example architecture and technology also offer better scalability, reduced cost, and reduced power consumption.
[0260] Star couplers can be used as broadcast technologies both at the network edge and in the network core. At the edge, they can be used in the form of SOA-gated splitters and combiners (to create 1:N and N:1 switches). In the core, an N:N star coupler can be used, which has been shown to scale to 1024 ports as a single component and to even more ports when using a cascading method. This approach makes the network passive and cost-effective.
[0261] The wavelength routing component in the network core can be an arrayed waveguide grating router, which has been shown to scale to hundreds of ports with low loss.
[0262] The combination of these above-mentioned technologies allows the example network to achieve a nanosecond-level circuit reconfiguration time while achieving a high node capacity. Thus, in some examples, the present method provides a higher-performance network and a more efficient execution of MPI operations. Compared to existing network architectures, these technologies also improve scalability, reduce component costs, and reduce power consumption. It will be understood that the present technology (e.g., method 200) can be combined with this network architecture to further enhance their respective advantages in terms of network and operational performance.
[0263] Set Operations
[0264] The network can be controlled by a scheduler. In some examples, the scheduler is configured to handle dynamic traffic. For deterministic operations, the scheduler can interface with distributed hardware to transform information for network interface cards.
[0265] The collective operations and MPI operations discussed herein are designed to avoid contention and minimize the collective operation completion time. Each MPI collective operation follows a set of unscheduled reconfiguration steps based on the following: a) parallel subgroup mapping (a subset of nodes that execute the collective operation in parallel); b) information and message mapping at each node at each algorithm step; c) wavelength and subnet selection; and d) time slot mapping.
[0266] The operations discussed can be implemented on any all-to-all network (e.g., any port-level all-to-all large-scale network) without oversubscription. While various advantages can be achieved by using this operation in known EPS or OCS networks, performance will be maximized when this operation is combined with aspects of the example network architecture. For example, the collective operation completion time, cost, and power consumption are reduced.
[0267] Hereinafter, and as previously discussed, 0 ≤ g ≤ x - 1, 0 ≤ j ≤ J - 1, and 0 ≤ λ ≤ Λ - 1 correspond to local communication groups, racks, and device numbers (represented by colors in Figure 13 (or cluster numbers, group numbers within a cluster, and node numbers within a group). Example MPI operations and strategies can be executed in three or four algorithm steps, but it will be understood that the number of algorithm steps will vary depending on the implementation.
[0268] An example process will now be described. Figure 13 A reduce-scatter strategy is shown, where Λ = 6 and J = x = 3. In this figure, four columns represent steps 1 - 4 of the algorithm. At each algorithm step, parallel logical graphs (called subgroups) are created between a unique subset of devices, which are represented as lines in Figure 13 . Figure 13 The left side of Figure 13 shows the chord diagram of the example network for each step, where nodes are grouped by communication group, rack, and device ID. The right side of the figure shows the connection matrix for each node at each step. The numbering of each node for the connection matrix represents the number within each vertex shown as the chord diagram. Although the graph is sparse, the network resources are maximally utilized because each node uses x - 1 transceivers for the first 3 steps and x transceivers for the last 1 step. In
[0269] In the first step (step 1) of the reduce-scatter operation, for each node, the overall message is divided into three parts and sent to different destinations in the subgroup. Then, the received information is summed (reduced) in each node. The part of the information that each node needs to send / receive (see Table 3) is determined by the information graph, and the transformation operation (such as summing) is determined by the MPI operation. Each node now contains the sum of the unique 1 / 3 of the messages in each subgroup.
[0270] Track the location of the parts of the information in each node after each communication step (see Table 3). For subsequent steps, select subgroups such that they only include nodes with the same combination of information parts. In the second step (step 2), the message is further divided into 3 parts (1 / 9 of the original message), sent to the correct nodes in each subgroup and processed. In the same way, the third step (step 3) is executed so that each device contains the sum of the unique 1 / 27 of the original information (global reduce-scatter). In the fourth step (step 4), information is exchanged between node pairs to complete the information update across all 54 devices. This last step can vary depending on the formula chosen for subgroup selection.
[0271] A similar process executed in reverse (from step 4 to step 1) applies to allgather, where the unique parts of the information are shared and gathered (concatenated) at each algorithm step in each subgroup. In this way, starting from 1 / 54 of the overall message, each node will contain the complete 1 / 27, 1 / 9, 1 / 3, and all the information after steps 4, 3, 2, and 1 respectively.
[0272] MPI Process
[0273] In some examples, this example network architecture or aspects thereof can be combined with this technology regarding MPI operations and will thus provide a particularly high-performance solution for performing HPC operations. However, it will also be understood that the MPI operations discussed further below can be applied to any electrical packet-switching network or OCS network, or any port-level all-to-all network (without oversubscription), and the technical advantages discussed herein can still be achieved.
[0274] Reference Figure 14 , now an example node architecture and process for performing MPI collective operations will be discussed. It will be understood that each node in the network can execute this process simultaneously.
[0275] Distributed tasks / jobs are arranged by a network job scheduler, and thereafter, information related to the rank / coordinates of the nodes and the MPI operations to be performed is shared with all relevant nodes. Specifically, a node may receive MPI collective operation information that identifies the MPI collective operation to be performed on data. The job scheduler may also provide time profile information to the nodes. In some examples, the rank of the node is included in a graph of the network, which is received by the node. The received information may be processed by an engine (e.g., labeled as the RAMP engine in Figure 14 ). The engine (or RAMP engine) includes two components: an MPI engine 1 and a network transcoder 2 (discussed hereinbelow).
[0276] The MPI engine 1 uses the physical topology of the network (i.e., the graph of the network) and the MPI operations to generate instructions required for the application 3 (the processor of the node) and the network transcoder 2 to complete the collective operation. The MPI engine 1 and the network transcoder 2 handle scheduling and communication, while the processing is handled by the application 3.
[0277] As Figure 14 shown, the MPI engine 1 uses the physical graph G and the MPI operation information to calculate the number of algorithmic steps required to perform the MPI operation. This can be performed based on a lookup. In some examples, the MPI engine 1 compares the MPI operation identified by the MPI operation information with multiple MPI operations stored in a memory and the number of associated algorithmic steps thereof. Based on this comparison, the number of algorithmic steps required for the MPI operation can be determined.
[0278] Then, the MPI engine 1 may generate information 1.a and information 1.b for each of the determined number of algorithmic steps. Information 1.a includes the information required for the application 3 to correctly process and retrieve data / messages for each step. The application 3 is the processing module of the computing node. In some examples, information 1.a includes only the information required by the application 3. As Figure 14 shown, information 1.a includes an information graph, local operations, buffer operations, and the number of nodes for each algorithmic step. This will be discussed in more detail hereinafter.
[0279] Information 1.b includes the algorithmic information required for the network transcoder 2 to convert the information into information suitable for the network interface card (NIC) 4. Specifically, information 1.b includes the data size and subgroup 1.c for each algorithmic step. Subgroup 1.c represents the logical graph of the nodes (derived from the graph G of the network) that perform a partial MPI operation at each algorithmic step. In other words, for each algorithmic step, the MPI engine 1 determines a subset or subgroup of network nodes with which the node running the MPI engine 1 should communicate to complete the MPI operation. As Figure 14As shown, the physical graph G is a graph of node connections in the network, and 1.c indicates a subgroup or subset of the nodes of the physical graph G. As can be seen below 1.c, the current node executing the MPI engine 1 is the node with a lighter color, and its subgroup is the nodes immediately below it and connected to it.
[0280] The network transcoder 2 receives the information 1.b from the MPI engine 1 and the physical graph G, and converts (transcodes) it into instructions for the network interface card 4. For each algorithm step, the network transcoder 2 generates instructions 2.b for each individual transceiver (such as Figure 1 or Figure 10 the transceivers in) to select the time slot size and quantity, transmit / receive wavelengths, and paths. After processing these instructions, the network transcoder 2 sends a'ready' flag / signal 2.a to the application 3, indicating that the NIC 4 is ready for transmission. The application 3 retrieves and converts the data using 1.a so that the data can be correctly processed and transmitted by the NIC 4 to perform MPI operations. The application 3 shares the processed data with the NIC 4, and the NIC 4 converts it into a signal 4.a on the physical system using the information 2.b. The NIC 4 tunes the transceiver according to the indicated wavelength and selects the correct SOA path (to conduct) for a given time slot size.
[0281] The quantities of the information 1.a and 1.b are based on Table 1-4 above. As described above, the quantities cited in the table are defined as follows: The network includes x communication groups; each communication group includes J racks, where J ≤ x; each rack includes Λ nodes; each node has a device number λ in the rack, defined as 0 ≤ λ ≤ Λ - 1; each rack has a rack number j, defined as 0 ≤ j ≤ J – 1; and the multiple nodes in each rack are divided into device groups including x nodes, where each node has a unique device group number ranging from 1 to x. It will be understood that the exact numbering scheme can vary, for example, the numbering can start from 1 instead of 0.
[0282] Information Map
[0283] The information graph includes a set of formulas that describe the parts of the information that each node should send, receive, and process at each algorithm step. The formulas are described in Table 3, which describes the information graph for the policies related to data transmission at each algorithm step. The combination of the values generated by this table at each algorithm step represents the node rank. This also represents a part of the original message, or represents the collected information available at the node after the last operation according to the selected operation. The decimal representation of the information values at all algorithm steps represents the rank of each node in the set.
[0284] In some examples, the message is a vector or matrix with a defined length. Each node does not need all the information in the message at each algorithm step. For example, if the number of nodes communicating in a certain algorithm step is N, the message will be divided into N smaller parts. The formula listed in Table 3 is used to find the part of the information in the message that the node needs to receive (according to the index of these N smaller parts). Thus, in an example where the information part is 2, the node will need to receive the third smaller part among the N smaller parts of the message (e.g., starting from 0).
[0285] Local Operations
[0286] The local operation (Loc_op(DATA)) is a transformation performed on the received data after the communication step. The local operation is specific to the MPI operation being performed, as shown in Table 2. The information graph (info) of the current step is used to arrange the information from the NIC in the correct order. There are four operations:
[0287] Reduce: An associative operation between vectors received from different sources, usually a sum.
[0288] Reshape: Only used for all-to-all operations. Transposes the information in the source (viewed as a 3D array), permutes the dimensions, and flattens it into a one-dimensional vector. This operation places the information to be sent in the correct rank order in a contiguous part of the memory.
[0289] Logical AND between boolean values, indicating the presence of a correct message. Only used for barrier operations.
[0290] Identity: Does not perform any transformation.
[0291] Buffer Operations
[0292] The buffer operation (Buff_op) corresponds to a transformation performed on the message generated by the MPI engine and defined by the MPI operation before transmission. The buffer operation is specific to the MPI operation being performed, as shown in Table 2. It takes three parameters: the message to be processed (DATA), the number of nodes in the current subgroup (nodes), and the information graph (info) of the current step. Info is used to sort the message so that the correct part of the information is given to the correct transceiver.
[0293] As shown in Table 2, there are three types of operations:
[0294] Reshape: The information vector is reshaped so that it can be divided into contiguous segments addressable by the same-sized nodes.
[0295] Duplication: The buffer size is increased in multiples of nodes and reshaped as described above. The original information will be located in the array segment corresponding to the local rank of the nodes in the subgroup.
[0296] Identity: No transformation is performed.
[0297] Number of Nodes
[0298] For each algorithm step, the number of nodes in each subgroup can be determined based on Table 1. In other words, the number of nodes in each subgroup refers to the number of other nodes to which the current node will send data and receive data from at each algorithm step.
[0299] Communication Subgroup Map
[0300] A subgroup (or subset) describes the set of nodes (logical graph) with which each node needs to share information (communicate) at any algorithm step. An overview and formula describing how each node is mapped to any subgroup at any communication step are shown in Table 1. For this mapping, the nodes in the rack are further divided into groups of x devices, called device groups, where each node has a unique device group number ranging from 1 to x. In fact, as discussed above, each of the multiple interconnected nodes has a unique node number, device number within the rack, rack number within the communication group, and communication group number.
[0301] The communication subgroup at each algorithm step corresponds to the communication performed between unique device groups in different system dimensions. These steps include:
[0302] Step 1: Nodes with the same node number, rack, and different communication groups;
[0303] Step 2: Nodes with consecutive node numbers in the same device group, rack, and different communication groups;
[0304] Step 3: Nodes with the same node number, different racks, and communication groups;
[0305] Step 4: Nodes with the same device group number, different device groups, racks, and communication groups, or nodes in consecutive device groups with the same device group number, rack, and different communication groups.
[0306] Depending on the selection of the formula for the subgroup in Step 4, two different operations can be used. When the first formula is selected, the algorithms considered for the last step use a strategy with one-to-one communication (such as ring, recursive halving / doubling, and Bruck algorithms), which may incur additional steps if the number of nodes is greater than 2 (the value at the maximum scale).
[0307] Subgroup selection defines the logical circuits to which each node belongs for each algorithm step, i.e., the group of nodes that will communicate. The number of nodes in each subgroup (as shown in Table 1) selects which of the four steps is active (#NS>1). Based on the subgroup information, each node can know all the sources and destinations that are active at any algorithm step, as described in Table 1.
[0308] Using the information provided in Table 1, each node can find the members of each subgroup from each algorithm step. The formulas for finding the coordinates (cluster number, group number in the cluster, node number in the group) of the other members of the same subgroup of the current node for each algorithm step are shown in Table 4.
[0309] MPI Operation Algorithm
[0310] The combination of Buff_op and Loc_op is defined by the MPI operations (Table 2), and the application will perform the operation on the message. The pseudocode for a single MPI operation running on a separate node is shown in Algorithm 1 below. In Algorithm 1, starting from the local message, each node requests information from the MPI engine based on the rank of the current node and the active nodes and the MPI operation (line 2). For each step indicated by the MPI engine, the data is first transformed by Buff_op (line 6), and after receiving the confirmation from the transcoder that the NIC is ready (line 7), data is pushed / received to / from the NIC, and this data will be transformed by the local operation (Loc_op, line 9) and will be used as the data for the next step.
[0311] The selection of Buff_op and Loc_op for each MPI operation is shown in Table 2. The message size for each step and operation in Table 2 is derived from the combination of Buff_op and Loc_op following Algorithm 1 below.
[0312]
[0313] The Reduce and All-Reduce operations are not included in Table 2. These operations are implemented by following a method similar to the known Rabenseifner algorithm, where the Reduce and All-Reduce operations are considered as reduce scatter operations, followed by gather operations and all-gather operations respectively.
[0314] For the broadcast operation in Table 2, the optical properties of the system are used. Using SOA gating, a device can broadcast to x at the full node capacity 2 or x 3A number of nodes (depending on the selected system configuration) multicast data. Given this feature, a pipelined tree broadcast is created where the root node can communicate with up to x 2 nodes, where λ - 1 nodes will transmit to another x 2 devices using different wavelengths. This creates a logical tree with a diameter of 3. The number k of stages of the pipeline under consideration is:
[0315]
[0316] where s is the diameter of the tree generated for performing the broadcast, α is the communication setup delay (propagation delay and node / software related delay), and β is the reciprocal of the total node capacity. The total number of steps required to perform the operation is k + s - 2, and the total number of messages sent in each stage is message / k.
[0317] In this way, each node performing a collective operation can receive the graph of the network and the MPI collective operation information identifying the MPI collective operation to be performed on the data, and determine the number of algorithmic steps required to perform the MPI collective operation. For each step, the node determines the subset of nodes with which the node is to communicate, the portion of the data that the node is to send, the process to be performed on the data portion before sending the message including the data portion to the other nodes in the subset, and the size of the message, which includes the data portion that the node is to send to the nodes in the subset of nodes. Then, based on the determined data portion, the determined process, the determined size of the message, and the determined subset of nodes, an MPI collective operation is initiated.
[0318] Network Transcoder
[0319] Network transcoder 2 uses the information from MPI engine 1 and the collective operation, and converts this information into instructions for NIC 4 to establish an optical circuit by only configuring the transceiver (wavelength) and 1:x switch (path) of the node (see Figure 14 ).
[0320] 1) Wavelength mapping: In an OCS network, wavelength selection is crucial for correctly routing information and avoiding contention. In addition to subgroup selection, a color / wavelength is assigned to each node in order to convey appropriate information at each algorithmic step. Wavelength mapping varies for various subnets and uses a lookup table. For subnets with only star couplers, the mapping is determined by the node receiving the wavelength, while for subnets with AWGRs, the mapping is enforced by the source / destination pair.
[0321] 2) Subnet / Path / Transceiver Selection: For any source-destination pair, there are bx possible paths and subnets that allow communication. Among the parallel subgroups in the first three algorithm steps, there can be up to bx communications sharing the same set of subnets using the same wavelength. To avoid contention, a wavelength can only be used once within the same subnet.
[0322] To minimize the control complexity, the transceiver used by any node to perform the set operation is pre-determined. The transceiver group selected between any source-destination pair is as follows:
[0323]
[0324] where g src -g dst , j src -j dst and λ src -λ dst are the source and destination communication group numbers, rack numbers, and node numbers respectively. Transceiver selection enforces subnet selection because each subnet is defined by the combination of g src , g dst and T rx .
[0325] In fact, whenever the number of devices in each subgroup is less than the number of the communication group, multiple transceiver groups can be used to communicate between the same source-destination pair. The number of additional transceiver groups that can be used for each communication in the set operation is:
[0326]
[0327] where d is the number of devices in the active subgroup. If #TRX additional is not 0, additional transceiver groups are used for communication. The transceiver group for any communication pair is as follows:
[0328]
[0329] where Trx(d src , d dst ) is the original transceiver group described in Equation 2.
[0330] According to Equation 4, the effective I / O unidirectional bandwidth of a node can be defined as:
[0331] B IO Eff = B·b·(1 + #TRX additional )(d - 1). (5)
[0332] For the fourth step, transceiver selection can vary depending on the selected subgroup formula (Table 2). For the first formula, the number of transceiver groups used for each communication is x, as there will be no contention for a single job. When the second formula is selected, the transceiver mapping follows Formula 4, where the maximum number of transceiver groups that can be used for each communication is
[0333] 3) Time slot mapping: The time slot diagram is given by the data transmitted in each step (Table 4) and the effective bandwidth of each transceiver (Formula 5), and a deterministic communication delay is given. Different subnets can be selected to further increase the number of parallel jobs (e.g., an AWGR-based subnet allows different device number settings to be supported for the same reason as the communication group settings).
[0334] Example MPI Collective Operations and Strategies
[0335] In some examples, to perform a full MPI operation, each node performs the following operations as described in Figure 15 Each node first receives from the job allocator / scheduler the collective operation, message size, active nodes of the collective, and network coordinates represented as communication group (x), rack (J), and node number (Λ) (or cluster number, group number in the cluster, and node number in the group). Using this information, each node calculates its subgroup ID and the number of nodes in each subgroup for each algorithm step based on Table 1 (stored in the memory of each node) and as described in the 'Communication subgroup graph' section above. While calculating these, the active steps (the steps that must run for the collective operation) are also selected, as the number of their nodes will be greater than 1. In other words, the combination of local operations and buffer operations is determined based on Table 2 and as described in the 'Buffer operations', 'Local operations', and 'MPI operation algorithm' sections. Then, for each active step, based on Table 1 and Table 4 and as discussed in the 'Communication subgroup graph' section, the logical circuit or subgroup (nodes with the same subgroup ID) is found.
[0336] Once the logical circuit / subgroup of the nodes (i.e., the nodes to communicate at that algorithm step) is determined, the information part to be sent to each node is calculated based on Table 3 and as discussed in the 'Information graph' section, and is stored in a lookup table. Based on the information part and buffer operations, the message size for each source-destination pair is calculated. Using the graph of the network and the logical circuit information (subgroup), the transceiver for each source-destination pair is selected, which determines the effective bandwidth of the node-to-node communication. Based on the message size and effective bandwidth, the number of time slots for each communication is determined, and the wavelength and path of each active transceiver are selected. The received data is processed by local operations and is considered a message for the next active step.
[0337] All information is deterministic and pre-computed at application setup such that it can be used as a lookup table at runtime following the principles described in the 'MPI operation algorithm' section.
[0338] Accordingly, various techniques for improving network performance in HPC applications have been described. Specifically, techniques for initializing and executing MPI operations have been described, and the network architectures in which these techniques can be executed have been described. It will be understood that the techniques related to MPI initialization and execution and the network architectures can be implemented separately.
[0339] The methods discussed above can be executed under the control of a computer program executable on a computing node / device (such as any node described in the accompanying figures herein). The computing node can include one or more processors, a memory, and communication circuitry. Accordingly, the computer program can include instructions for controlling the computing device / node to execute any of the methods discussed above. The program can be included in a computer-readable medium. The computer-readable medium can include a non-transitory type of medium, such as a physical storage medium, such as a storage disk and a solid-state device. The computer-readable medium can additionally or alternatively include a transient medium, such as a carrier signal and a transmission medium, which can be used, for example, to carry instructions between multiple separate computer systems and / or between components within a single computer system.
[0340] The various embodiments described herein are presented only to aid in understanding and teaching the claimed features. These embodiments are provided only as representative examples of embodiments and are not exhaustive and / or exclusive. It is to be understood that the advantages, embodiments, examples, functions, features, structures, and / or other aspects described herein should not be considered as limitations on the scope of the disclosure defined by the claims or on equivalents of the claims, and that other embodiments can be used and modifications can be made without departing from the scope of the invention defined by the claims.
Claims
1. A method for performing Message Passing Interface (MPI) collective operations in a network, wherein, The network includes a plurality of interconnected nodes, and the method includes: Receiving, at one of the plurality of interconnected nodes, MPI collective operation information identifying a MPI collective operation to be performed and a graph of the network; Determining, based on the MPI collective operation and the graph of the network, the number of algorithmic steps of the MPI collective operation; Determining an initialization process for the algorithmic steps; Determining a finalization process for the algorithmic steps; For each of the algorithmic steps, determining: A subset of the nodes in the plurality of interconnected nodes with which the node is to communicate; and One or more data portions that the node is to send to and receive from the nodes within the subset of the nodes; and Initializing the MPI collective operation based on the determined subset, initialization process, finalization process, and one or more data portions.
2. The method according to claim 1, wherein Each of the plurality of interconnected nodes is configured to perform the method.
3. The method according to any one of the preceding claims, wherein, The subset of nodes for each of the algorithmic steps is unique.
4. The method according to any one of the preceding claims, wherein Determining the number of algorithmic steps is based on retrieving stored information associated with the MPI operation.
5. The method according to any one of the preceding claims, wherein, The network is a circuit-switched network, where optionally the network is an optical circuit-switched network.
6. The method according to any one of the preceding claims, wherein: The network includes one or more clusters, each cluster including one or more groups, each group including one or more nodes; Each of the plurality of interconnected nodes has a node number within a group, a group number within a cluster, and a cluster number; and The graph includes information indicating the node number, group number, and cluster number of each node.
7. The method according to claim 6, wherein For a first algorithmic step among the algorithmic steps, the subset of nodes includes nodes having the same node number, the same group number, and different cluster numbers.
8. The method according to claim 6 or 7, wherein The plurality of nodes in each group are divided into node sets including x nodes, where each node has a unique node set number from 1 to x, and where x is the number of clusters.
9. The method according to claim 8, wherein, For a second algorithmic step among the algorithmic steps, the subset of nodes includes nodes having consecutive node numbers within the same node set, the same group number, and different cluster numbers.
10. The method according to claim 8 or 9, wherein, For a third algorithmic step among the algorithmic steps, the subset of nodes includes nodes having the same node number, different group numbers, and different cluster numbers.
11. The method according to any one of claims 8 to 10, wherein, For a fourth algorithmic step among the algorithmic steps: The subset of nodes includes nodes having the same node number within a node set, different node sets, the same group number, and different clusters; or The subset of nodes includes nodes within consecutive node sets having the same node number within a node set, the same group number, and different cluster numbers.
12. The method according to any one of the preceding claims, wherein, For each of the algorithmic steps, determining the subset of nodes includes: Determining an identifier of the subset of nodes; Determining the number of nodes within the subset of nodes; and Determining other nodes within the subset of nodes with which the node is to communicate.
13. The method according to claim 12, wherein: The network includes x clusters; Each cluster includes J groups, where J ≤ x; Each group includes Λ nodes; Each cluster has a cluster number g, defined as 0 ≤ g ≤ x - 1; Each node has a node number λ in the group, defined as 0 ≤ λ ≤ Λ - 1; Each group has a group number j, defined as 0 ≤ j ≤ J – 1; and The multiple nodes in each group are divided into node sets including x nodes, where each node has a unique node number from 1 to x in the node set.
14. The method according to claim 13, wherein: For the first algorithm step in the algorithm steps, the total number of subsets of nodes = ΛJ, and each subset has an identifier i; The total number of nodes in each subset is x; And For the first algorithm step in the algorithm steps, the identifier of the subset is determined based on the following formula: i = λ + Λ·j.
15. The method according to any one of claims 13 to 14, wherein: For the second algorithm step in the algorithm steps, the total number of subsets of nodes = ΛJ; The total number of nodes in each subset is x; and For the second algorithm step in the algorithm steps, the identifier of the subset is determined based on the following formula:
16. The method according to any one of claims 13 to 15, wherein: For the third algorithm step in the algorithm steps, the total number of subsets of nodes = Λx; The total number of nodes in each subset is J; and For the third algorithm step in the algorithm steps, the identifier of the subset is determined based on the following formula: i = (λ + Λ(j - g)) mod (Λj).
17. The method according to any one of claims 13 to 16, wherein: For the fourth algorithm step in the said algorithm steps, the total number of subsets of nodes = Jx 2 ; The total number of nodes in each subset is Λ / x; and For the fourth algorithm step in the algorithm steps, the identifier of the subset is determined based on the following formula: a) Or b) 18. The method according to any one of claims 13 to 17, wherein, Determining other nodes within the subset of the nodes further includes determining the node number, group number, and cluster of the other nodes based on the following: For the first algorithm step in the algorithm steps: Node number = λ; Group number = j; and Cluster = (g + γ) mod x, where 0 ≤ γ ≤ x – 1; For the second algorithm step in the algorithm steps: Group number = j; and Cluster = (g + λ) mod x, where 0 ≤ γ ≤ x – 1; For the third algorithm step in the algorithm steps: Node number = λ; Group number = [(j + γ) mod J – j] mod J; and Cluster = [(g - j) mod x + γ] mod x, where 0 ≤ γ ≤ J – 1; and For the fourth algorithm step in the algorithm steps: Among them 19. The method according to any one of the foregoing claims, the method further comprising: In response to determining that the MPI operation is a reduce-scatter operation, selecting a reshaping process as the initialization process and a reduction process as the final determination process; In response to determining that the MPI operation is an all-gather operation, selecting a replication process as the initialization process and an identity process as the final determination process; In response to determining that the MPI operation is a barrier operation, selecting an identity process as the initialization process and a logical AND process as the final determination process; In response to determining that the MPI operation is an all-to-all operation, select the reshaping process as the initialization process and select the reshaping process as the final determination process; In response to determining that the MPI operation is a scatter operation, select the reshaping process as the initialization process and select the identity process as the final determination process; In response to determining that the MPI operation is a gather operation, select the copy process as the initialization process and select the identity process as the final determination process; In response to determining that the MPI operation is a broadcast operation, select the processes associated with the scatter operation and the all-gather operation; And In response to determining that the MPI operation is a reduce-scatter operation, select the processes associated with the reduce-scatter operation and the all-gather operation.
20. The method according to any one of claims 13 to 19, wherein: For a first algorithm step among the algorithm steps, determine the one or more data portions based on the following; For a second algorithm step among the algorithm steps, determine the one or more data portions based on the following; For a third algorithm step among the algorithm steps, determine the one or more data portions based on the following: data portion = j; and For a fourth algorithm step among the algorithm steps, determine the one or more data portions based on the following; 21. The method according to any one of claims 13 to 20, wherein, The received message size is m, and the method further includes: In response to determining that the MPI operation is a reduce-scatter operation: For the first algorithm step in the algorithm steps, select the size of the message as m / x. For the second algorithm step in the algorithm steps, select the size of the message as m / x 2 , for the third algorithm step in the algorithm steps, select the size of the message as m / (Jx 2 ), for the fourth algorithm step in the algorithm steps, select the size of the message as m / (JΛx); In response to determining that the MPI operation is an all-gather operation: For the first algorithm step among the algorithm steps, select the message size to be m·JΛx, for the second algorithm step among the algorithm steps, select the message size to be m·JΛ, for the third algorithm step among the algorithm steps, select the message size to be m·JΛ / x, and for the fourth algorithm step among the algorithm steps, select the message size to be m·Λ / x; In response to determining that the MPI operation is a barrier operation: For the first algorithm step among the algorithm steps, select the message size to be 0, for the second algorithm step among the algorithm steps, select the message size to be 0, for the third algorithm step among the algorithm steps, select the message size to be 0, and for the fourth algorithm step among the algorithm steps, select the message size to be 0; In response to determining that the MPI operation is an all-to-all operation: For the first algorithm step among the algorithm steps, select the message size to be m / x, for the second algorithm step among the algorithm steps, select the message size to be m / x, for the third algorithm step among the algorithm steps, select the message size to be m / J, and for the fourth algorithm step among the algorithm steps, select the message size to be m·x / Λ; In response to determining that the MPI operation is a scatter operation: For the first algorithm step in the algorithm steps, the size of the message is selected as m / x, and for the second algorithm step in the algorithm steps, the size of the message is selected as m / x 2 , for the third algorithm step in the algorithm steps, the size of the message is selected as m / (Jx 2 ), and for the fourth algorithm step in the algorithm steps, the size of the message is selected as m / (JΛx); In response to determining that the MPI operation is a gather operation: For the first algorithm step in the algorithm steps, select the size of the message to be m·JΛx; for the second algorithm step in the algorithm steps, select the size of the message to be m·JΛ; for the third algorithm step in the algorithm steps, select the size of the message to be m·JΛ / x; for the fourth algorithm step in the algorithm steps, select the size of the message to be m·(Λ / x); and In response to determining that the MPI operation is a broadcast operation: Determine the size of the message for each step based on the message sizes determined from scatter and allgather operations.
22. The method according to any one of the preceding claims further comprises: For each of the algorithm steps, store in memory the determined subset of nodes and one or more data portions, wherein the method further comprises: After initializing the MPI operation, receive messages associated with the MPI collective operation; Execute the first algorithm step in the algorithm steps by: Processing the message using the determined initialization process; Sending one or more portions of the determined processed message to the corresponding one or more nodes of the subset; Receiving messages from nodes within the subset; and Processing the received messages using the final determination process, wherein the processed received messages become the messages for subsequent algorithm steps, and wherein the algorithm steps are repeated for all determined algorithm steps using the corresponding information determined for each step.
23. The method according to claim 22, wherein, Executing each of the algorithm steps includes providing the determined subset of nodes and one or more data portions to a network transcoder of the nodes, the method further comprising: converting, by the network transcoder, the determined subset of nodes and one or more data portions into instructions for configuring one or more transceivers of the nodes.
24. A node for performing MPI collective operations on data in a network, wherein, The network includes a plurality of interconnected nodes, the nodes including a processor configured to execute the method according to any one of claims 1 to 23.
25. A computer-readable medium comprising instructions that, when executed by a processor, cause the processor to execute the method according to any one of claims 1 to 23.