Node, processing method, and processing system
The described processing system addresses the inefficiency of port-limited nodes by using optical connections and selectors to facilitate direct communication among nodes, reducing port requirements and enhancing the efficiency of distributed computing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2026-03-26
AI Technical Summary
Existing methods for exchanging distributed computing results require many steps or transfer to other nodes or switches due to limited ports per node, leading to inefficiencies as the number of computing nodes increases.
A processing system with nodes equipped with L-1 transmit and receive ports, along with L-1 transmit-side and receive-side selectors, each having k branch destinations, allows efficient exchange of results through recursive doubling without forwarding to other nodes or switches by using optical connections and selectors to switch communication paths.
The system reduces the number of ports required per node and improves efficiency by enabling direct optical communication between nodes, allowing for faster and more efficient exchange of results across multiple nodes using a k-branch selector.
Smart Images

Figure JP2024033330_26032026_PF_FP_ABST
Abstract
Description
Node, processing method, and processing system
[0001] The present disclosure relates to a node, a processing method, and a processing system.
[0002] In order to improve the efficiency of distributed computing, it is required to exchange the results of distributed computing among more nodes. Since the number of ports provided by each node is limited, a device for distributing the results to each node is necessary.
[0003] As a method for exchanging the results of distributed computing, there is a method that requires many steps (Non-Patent Document 1). There is a method of transferring the results calculated by the computing node to other nodes or switches (Non-Patent Document 2).
[0004] All-reduce is an algorithm that calculates the sum of data held by all nodes and distributes it to each node. Recursive doubling is an algorithm that efficiently executes the all-reduce algorithm (Non-Patent Document 3). The recursive doubling algorithm can also speed up various operations such as all-gather (Non-Patent Document 4), reduce-scatter (Non-Patent Document 4), and prefix-sum (Non-Patent Document 5) in addition to the all-reduce operation.
[0005] There is a method of connecting a plurality of nodes via a plurality of electrical switches (Non-Patent Document 6).
[0006] W. Wang, M. Khazraee, Z. Zhong, M. Ghobadi, Z. Jia, D. Mudigere, Y. Zhang, and A. Kewitsch, “TopoOpt: Co-optimizing network topology and parallelization strategy for distributed training jobs,” in 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), 2023, pp. 739-767.H. Ballani, P. Costa, R. Behrendt, D. Cletheroe, I. Haller, K. Jozwik, F. Karinou, S. Lange, K. Shi, B. Thomsen et al., “Sirius: A flat datacenter network with nanosecond optical switching,” in Proceedings of the Annual conference of the ACM Special Interest Group on Data Communication on the applications, technologies, architectures, and protocols for computer communication, 2020, pp. 782-797.W. Li, G. Yuan, C. Wu, P. Huang, and Z. Wang, “A reconfigurable optical network for distributed deep learning,” in 2023 Opto-Electronics and Communications Conference (OECC). IEEE, 2023, pp. 1-3.E. Chan, M. Heimlich, A. Purkayastha, and R.Van De Geijn, “Collective communication: theory, practice, and experience,” Concurrency and Computation: Practice and Experience, vol. 19, no. 13, pp. 1749-1783, 2007.T. HUANG, “Generic implementation of parallel prefi sums and their applications,” Ph.D. dissertation, Texas A&M University, 2007.W. Wang, M. Ghobadi, K. Shakeri, Y. Zhang, and N. Hasani, “How to Build Low-cost Networks for Large Language Models (without Sacrificing Performance)?,” arXiv preprint arXiv:2307.12169v3, 2023.
[0007] However, the methods described in Non-Patent Documents 1-2 require many steps or transfer to other nodes or switches. The efficiency of the distributed computing results is limited.
[0008] This disclosure is made in view of the above circumstances, and the purpose of this disclosure is to provide a technology that enables the efficient exchange of distributed computing results.
[0009] A node according to one aspect of the present disclosure is used in a processing system having a number of nodes equal to L raised to the power of k, and comprises L-1 number of transmit ports, L-1 number of receive ports, L-1 number of transmit-side selectors connected to each of the transmit ports, and L-1 number of receive-side selectors connected to each of the receive ports, wherein one transmit-side selector has k branch destinations and one receive-side selector has k branch destinations.
[0010] A processing method according to one aspect of the present disclosure is used in a processing system having a number of nodes equal to L raised to the power of k, wherein each node comprises L-1 transmit ports, L-1 receive ports, L-1 transmit-side selectors connected to each of the transmit ports, and L-1 receive-side selectors connected to each of the receive ports, one transmit-side selector having k branch destinations, one receive-side selector having k branch destinations, and the identifier of the plurality of nodes in the processing system is identified by a vector having k elements of any natural number from 1 to L, wherein all but the i-th element of the L nodes are the same as the node, and the i-th step is defined as the set of nodes. Each of the L-1 transmitting selectors is connected to one of the L-1 receiving selectors that are connected to the receiving ports of the other L-1 nodes belonging to the i-step set, and each of the L-1 receiving selectors is connected to one of the L-1 transmitting selectors that are connected to the transmitting ports of the other L-1 nodes belonging to the i-step set, and the partial results held by the node are transmitted to the other connected node via each of the L-1 transmitting ports, and the partial results held by the other node are received from the other connected node via each of the L-1 receiving ports, and the partial results held by the node are processed using the partial results held by the other node and the partial results held by the node.
[0011] A processing system in one aspect of the present disclosure is a processing system comprising a number of nodes equal to L (where L is a natural number) raised to the power of k, wherein the identifiers of the plurality of nodes are identified by a vector having k elements that are any natural number from 1 to L, and when L nodes that are the same except for the i-th element are considered a set of the i-th step, each of the plurality of nodes comprises L-1 transmit ports, L-1 receive ports, L-1 transmit-side selectors connected to each of the transmit ports, and L-1 receive-side selectors connected to each of the receive ports, and one transmit-side selector comprises k Each receiving selector has k branch destinations, and in the i-th step, each of the L-1 transmitting selectors connects to one of the L-1 receiving selectors that connects to the receiving port of another L-1 node belonging to the i-th step set, and each of the L-1 receiving selectors connects to one of the L-1 transmitting selectors that connects to the transmitting port of another L-1 node belonging to the i-th step set, and in each step from the first to the k-th step, each of the plurality of nodes simultaneously switches its branch destination.
[0012] A processing system according to one aspect of the present disclosure comprises a switch having a number of nodes equal to L to the power of k (where L is a natural number) and a number of ports equal to L to the power of k, wherein each node comprises one transmit port, one receive port, one transmit-side selector connected to the transmit port, and one receive-side selector connected to the receive port, the transmit-side selector having k branch destinations, the receive-side selector having k branch destinations, the identifiers of the plurality of nodes are identified by a vector having k elements that are any natural number from 1 to L, and a set of L nodes that are the same except for the i-th element is defined as the i-th step set, and among the ports of the switch, in the i-th step, the ports that accommodate each node in the i-th step set are defined as the i-th step port set, in the i-th step, the transmit-side selector and the receive-side selector of the node are respectively connected to ports of the i-th step port set that accommodate L-1 other nodes in the i-th step set to which the node belongs, and in each step from the first to the k-th step, each of the plurality of nodes simultaneously switches its branch destination.
[0013] This is used in a processing system that has a number of nodes equal to L to the power of k (where L is a natural number) and a switch having a number of ports equal to L to the power of k. Each node has one transmit port, one receive port, one transmit-side selector connected to the transmit port, and one receive-side selector connected to the receive port. The transmit-side selector has k branch destinations, and the receive-side selector has k branch destinations. The identifiers of the multiple nodes are identified by a vector having k elements that are any natural number from 1 to L. A set of L nodes that are the same except for the i-th element is defined as the i-th step, and the switch In the i-th step, if the ports of the switch that accommodate each node in the i-th step set are designated as the i-th step port set, then in the i-th step, the transmitting selector and receiving selector of the node are each connected to the ports of the i-th step port set that accommodate L-1 other nodes in the i-th step set to which the node belongs. The node transmits a partial result to the ports of the i-th step port set, the ports of the i-th step port set transmit the partial result received from the node to the other ports of the i-th step port set, the other ports of the i-th step port set transmit the partial result received from the L-1 other nodes in the i-th step set to the ports of the i-th step port set, the ports of the i-th step port set transmit the partial result received from the other ports of the i-th step port set to the node, and the node processes the partial result held by the L-1 other nodes in the i-th step set and the partial result held by the node.
[0014] According to this disclosure, it is possible to provide a technology that enables the efficient exchange of results from distributed computing.
[0015] Figure 1 is a diagram illustrating a system configuration comprising three nodes and the functional blocks of the nodes in a processing system according to the present disclosure. Figure 2 is a diagram illustrating the case in which each node exchanges results with many other nodes. Figure 3 is a diagram illustrating the case in which each node has few ports and can exchange results with a small number of nodes at once. Figure 4 is a diagram illustrating all-reduce. Figure 5 is a diagram illustrating an example of an identifier assigned to the nodes of the processing system. Figure 6 is a diagram illustrating the algorithm of recursive doubling. Figure 7 is a diagram illustrating the system configuration of a processing system comprising 16 nodes. Figure 8 is a diagram illustrating the system configuration of a processing system according to a second embodiment. Figure 9 is a diagram illustrating the system configuration of a processing system according to a third embodiment. Figure 10 is a diagram illustrating the connection between a node and an Ethernet switch in a processing system according to a third embodiment. Figure 11 is a diagram illustrating the hardware configuration of a computer used in a node.
[0016] Embodiments of this disclosure will be described below with reference to the drawings. In the drawings, the same parts are denoted by the same reference numerals and their descriptions are omitted.
[0017] (First Embodiment) The processing system 1 according to the first embodiment comprises a plurality of nodes NO1, NO2, and NO3. When the plurality of nodes NO1, NO2, and NO3 are not particularly distinguished, they may be referred to as node NO.
[0018] The processing system 1 in Figure 1 is illustrated with three node numbers, but is not limited to this. The processing system 1 can have any number of node numbers, such as several thousand.
[0019] Processing system 1 is suitable for computationally intensive systems such as HPC (High Performance Computing) or large-scale AI (Artificial Intelligence) training. Multiple nodes share the execution of large-scale calculations that cannot be handled by a single node. Nodes exchange intermediate calculation results as the calculation progresses.
[0020] In the processing system 1 according to the first embodiment, each node NO is connected by optical connection. Compared to electrical connection, optical connection achieves high bandwidth and low latency inter-node NO connection. Optical connection is also superior to electrical connection in terms of power consumption. The processing system 1 can reduce the cost required for processing.
[0021] Generally, as shown in Figure 2, the efficiency of distributed computing can be increased if results can be exchanged between more nodes. However, due to limitations such as power capacity or spatial capacity, there is a limit to the number of communication ports that each node can have.
[0022] As shown in Figure 3, when there are few ports and the number of nodes that can be exchanged at once is small, many steps are required for the results to reach all nodes. It is possible to seemingly exchange results with many nodes at once by "transferring" the results to other computing nodes or electrical switches (Ethernet switches). However, because the links between nodes are shared by multiple communications, and the optical connection is interrupted by inefficient electrical processing, the time required for each exchange (step) becomes long.
[0023] Thus, from the perspective of distributed computing, we want to exchange results between many nodes at once, but the limited number of ports per node has been a conventional challenge. For this reason, conventional technologies have compromised by using methods that require many steps to exchange results or by forwarding them to other nodes or switches.
[0024] Referring to Figure 4, all-reduce is explained to illustrate the recursive doubling used in this disclosure. In all-reduce, each node pre-stores its own calculation results (partial results). These partial results are generally large vectors, and passing them to other nodes requires considerable communication time. For simplicity, this disclosure explains the case where each partial result is a scalar.
[0025] As shown in Figure 4(a), each node holds its partial results. As shown in Figure 4(b), each node collects its respective partial results into a representative node. As shown in Figure 4(c), the representative node calculates the sum of the partial results collected from each node. As shown in Figure 4(d), the representative node distributes the sum calculated in Figure 4(c) to each node.
[0026] Figure 4 illustrates the case where a representative node calculates a sum, but any calculation is possible. The representative node can perform various calculations, such as product, arithmetic mean, and geometric mean, using the partial results of each node. For example, to calculate a product, the representative node may take the logarithm of the partial results, calculate the sum, and then convert back to the exponential. All-reduce can be considered a general-purpose operation. All-reduce is used, for example, to distribute the weights of a neural network.
[0027] In simple terms, all-reduce can achieve its goal by collecting all partial results on a single representative node, calculating the sum, and distributing it to all nodes. However, collecting the partial results on the representative node and then distributing the sum requires an enormous amount of communication time. For example, let N be the number of computing nodes and P be the number of ports on each node. Assume that it takes one step to exchange partial results using one port. For the representative node to collect partial results from the other N-1 nodes, it takes at least (N-1) / P steps. The same applies to distribution.
[0028] As distributed computing becomes more large-scale, with the number of computing nodes N reaching several thousand, all-reduce does not scale.
[0029] Therefore, the processing system 1 according to this disclosure uses recursive doubling. Recursive doubling will now be explained.
[0030] First, an identifier (ID) is determined for each node NO in the processing system 1, as shown in Figure 5. Assume that multiple grids with side lengths of 1 are formed within a k-dimensional cube with side lengths L, and node NO is placed at each point. The number of nodes in the processing system 1 is N = L. kThis is the result. The ID of each node number is the coordinates of each node number. Note that this arrangement is used only to determine the ID. The arrangement of each node number used to determine the ID is unrelated to the physical location of each node number. Also, the adjacency relationship on the coordinates is unrelated to the connection relationship between node numbers.
[0031] The recursive doubling algorithm will be explained with reference to Figure 6. In recursive doubling, the number of nodes N in processing system 1 is L k In this case, the all-reduce operation is completed in just k steps. In the i-th step, a group of nodes whose IDs match except for the i-th element are called a "pair". In the first step, each column in Figure 6(a) becomes a pair. In the second step, each row becomes a pair. Each pair consists of L-th nodes. In the i-th step, each node in the pair from the i-th step shares the partial result at that point. That is, each node sends its partial result to L-1 nodes and receives those partial results. Then, all nodes in the pair can calculate and hold the "sum" of that pair.
[0032] As shown in Figure 6(a), in the first step, for example, the leftmost node shares the partial results of each node in its set: 2, 2, 1, and 2. As shown in Figure 6(b), each node in the leftmost column calculates the sum of the shared values, which is 7. Similarly, the sum of the second column from the left is 10. The sum of the third column from the left is 6. The sum of the rightmost column is 9.
[0033] Repeating these calculation steps for all dimensions, as shown in Figure 6(c), all nodes can hold the "sum" of the partial results of all nodes. In the second step of Figure 6(c), each node calculates the sum of the partial results of each column. Each node calculates the overall sum of 32.
[0034] In the recursive doubling algorithm, each step involves communicating with L-1 nodes. Over k steps, each node communicates with k(L-1) nodes in total. In Figure 6, each node communicates with L-1 = 3 nodes in each step. In total, each node communicates with k(L-1) = 6 nodes throughout all steps.
[0035] To perform this calculation without forwarding to other nodes or switches, each node must have k(L-1) ports. For example, in a system with 4096 nodes where L=8 and k=4, each node has 28 ports. Given that the current number of ports on each node is 4-8, it is generally difficult to implement the recursive doubling algorithm without forwarding to other nodes or switches.
[0036] Strictly speaking, the term "recursive doubling" refers only to the case where L=2, but this disclosure describes a more generalized algorithm. The recursive doubling algorithm is known to be able to speed up various operations such as all-gather, reduce-scatter, and prefix-sum, in addition to the all-reduce operation. The processing system 1 according to this disclosure can also be applied to any of these operations.
[0037] Each node NO of the processing system 1 according to this disclosure reduces the number of ports required to implement the recursive doubling algorithm without forwarding at other nodes or switches. Each node NO solves the port shortage using a k-branch selector, as shown in Figure 1. The k-branch selector S is a 1 × kOCS (Optical Circuit Switch). The number of branches of the selector S connected to the port of node NO is at least the number of steps k in recursive doubling. The selector S is a pair of a transmitting selector STx and a receiving selector SRx. In this disclosure, a port may include a transmit / receive transceiver (TRx) function.
[0038] The processing system 1 includes a plurality of nodes NO of the number of L to the k-th power (L is a natural number). One node NO includes L - 1 transmission ports Tx, L - 1 reception ports Rx, L - 1 transmission-side selectors STx connected to each of the transmission ports Tx, and L - 1 reception-side selectors SRx connected to each of the reception ports Rx. One transmission-side selector STx has k destinations. One reception-side selector SRx has k destinations.
[0039] In the first embodiment, the transmission ports Tx and the reception ports Rx of the node NO transmit and receive optical signals. Each of the L - 1 transmission ports Tx transmits an optical signal. Each of the L - 1 reception ports Rx receives an optical signal.
[0040] The transmission-side selector STx is mounted on the transmission port Tx. The transmission-side selector STx is connected to k destinations. Different from a port (transceiver), the transmission-side selector STx does not require a large amount of power and can be added.
[0041] When the node NO transmits a result, it selects one from k destinations and performs communication by optical connection. The first node NO1 in FIG. 1 is in a state where the uppermost transmission-side selector STx selects the second node NO2 connected by a solid line as the destination. The first node NO1 can communicate with the second node NO2, but cannot communicate with the third node NO3 connected by a broken line.
[0042] One node NO can communicate with only one of the k destinations at a time, but by shifting the time zone and switching the connection destination of the transmission-side selector STx, communication by optical connection with all k destinations becomes possible. The node NO1 in FIG. 1 may later switch the transmission-side selector STx and communicate with the third node NO3.
[0043] Similarly on the reception side, the reception-side selector SRx is mounted on the reception port Rx. The reception-side selector SRx is connected to k transmission sources. The reception-side selector SRx can communicate with any one of the k by optical connection.
[0044] In the present disclosure, for the sake of simplicity, transmission / reception ports may be handled in pairs without distinction. That is, when a port or a selector is connected in the figure, it is considered that the reception side opposite to one transmission side is connected and also connected in the reverse direction. A pair of a transmission-side selector STx and a reception-side selector SRx may be denoted as a selector S.
[0045] In the present disclosure, the ID of each node NO is determined as described with reference to FIG. 5. The identifiers of the plurality of nodes NO included in the processing system 1 are specified by a vector having k elements with values of natural numbers from 1 to L. In the present disclosure, L nodes with the same elements except the i-th element are defined as the i-th step set.
[0046] Each node NO is assumed to have ports with a connection number of L - 1 or more within the set (P ≥ L - 1). The i-th destination of each selector S of each node NO is connected to any node NO in the i-th set, and the wiring is performed such that the i-th destination of each selector S is connected to each node in the i-th set. Since there is no distinction in ports, from the perspective of distributed computing, any port may be connected to any node. By connecting the i-th respective connection destinations of each selector S of one node NO to the respective nodes NO in the i-th set, a full mesh (complete graph) is connected.
[0047] FIG. 7 shows an example of the processing system 1a according to the present disclosure. In FIG. 7, the optical connections between two node NOs of the block nodes are shown only in the leftmost column and the uppermost row, and the others are omitted. The upper left node (1, 1) includes k = 2 selectors S for each of the P = 3 ports. The first destination of each selector S is a plurality of nodes NO in the vertical column as shown by the solid line. The second destination of each selector S is a plurality of nodes NO in the horizontal row as shown by the broken line. The same applies to other nodes NO.
[0048] As shown in FIG. 1, the node NO includes a switching unit 10, a transmission / reception unit 20, and a processing unit 30. In FIG. 1, only the first node NO1 is depicted with the switching unit 10, the transmission / reception unit 20, and the processing unit 30, but each node NO of the processing system 1 includes the switching unit 10, the transmission / reception unit 20, and the processing unit 30.
[0049] The switching unit 10 switches the destination of the transmitting selector STx and the receiving selector SRx for each step. In the i-th step, the switching unit 10 switches the destination of the transmitting selector STx so that each of the L-1 transmitting selector STx is connected to one of the L-1 receiving selector SRx that is connected to a receiving port Rx of another node NO belonging to the set of L-1 in the i-th step. The switching unit 10 switches the destination of the receiving selector SRx so that each of the L-1 receiving selector SRx is connected to one of the L-1 transmitting selector STx that is connected to a transmitting port Tx of another node NO belonging to L-1.
[0050] In each step from the first to the kth step, the switching unit 10 of each of the multiple node NOs simultaneously switches the branch destination. In a certain step, a certain node NO and the node NO to which the selector S of that node NO is connected are connected in a way that allows them to communicate with each other. Here, "simultaneously" means that each node NO sends and receives partial results after the switching of the selector S of each node NO is completed.
[0051] In processing system 1, each node NO simultaneously switches the connection destination of the transmitting selector STx and the receiving selector SRx. Each node NO, for example, shares a clock and switches at a predetermined time. Alternatively, each node NO notifies the control server (not shown) of the completion of the transition to each step or the transmission and reception of data, and the control server checks the processing status of each node NO and instructs each node NO to proceed to the next step.
[0052] In this disclosure, the switching of the connection destinations of the transmitting selector STx and the receiving selector SRx is described in software, but this is not the only method. The switching of the transmitting selector STx and the receiving selector SRx can be performed by any mechanism.
[0053] The transmitting / receiving unit 20 transmits and receives partial results after the switching of each node NO in the processing system 1 by the switching unit 10 is completed. The transmitting / receiving unit 20 transmits the partial results held by node NO to the other connected node NO via each of the L-1 number of transmission ports Tx. The transmitting / receiving unit 20 receives the partial results held by the other connected node NO from each of the L-1 number of receiving ports Rx.
[0054] The processing unit 30 processes the partial results obtained by the transmitting / receiving unit 20. The processing unit 30 processes the partial results held by this node NO using the partial results held by other node NOs.
[0055] When performing recursive doubling, the corresponding pair of nodes are directly connected by optical connection by setting the selector S to the i-th destination in the i-th step. In the example in Figure 7, partial results are shared column by column in the first step and row by row in the second step. The processing system 1 according to this disclosure provides optical connections along the data flow shown in Figure 6. Therefore, since the processing system 1 according to this disclosure switches the communication destination for each calculation step, it is an optimal system configuration for realizing recursive doubling.
[0056] In this disclosure, each node NO only needs to have L-1 ports. As described above, the conventional recursive doubling algorithm required each node to have k(L-1) ports. Thus, the processing system 1 according to this disclosure can reduce the number of ports for each node NO. Furthermore, since the processing system 1 optically connects the nodes NO at the timing when each node NO exchanges partial results, the efficiency of distributed computing is improved.
[0057] (Second Embodiment) Referring to Figure 8, the processing system 1b according to the second embodiment will be described. The processing system 1b is equipped with 256 nodes NO, but the connection from node NO 1 to node NO 16 will be shown in focus.
[0058] In processing system 1b, L=4, k=4, and P=3. Processing system 1b has 256 node numbers. Four node numbers belong to each set. Each node number has three ports on both the transmitting and receiving sides. Each port has four switching destinations.
[0059] In the processing system 1b shown in Figure 8, the selector S for each node NO is housed in OCS boxes B1, B2, and B3, respectively. Each selector S is connected within each OCS box B.
[0060] The ID for each node number is set as follows: 1: (1,1,1,1) 2: (2,1,1,1) 3: (3,1,1,1) 4: (4,1,1,1) 5: (1,2,1,1) 6: (2,2,1,1) 7: (3,2,1,1) 8: (4,2,1,1) 9: (1,3,1,1) ... 13: (1,4,1,1) ... 256: (4,4,4,4)
[0061] In the first step, the group to which the first node NO1 belongs contains nodes that are the same except for the first element, specifically, from the first node NO1 to the fourth node NO4. In the first step, the first node NO1 connects to the second node NO2 in the first OCS box B1, to the fourth node NO4 in the second OCS box B2, and to the third node NO3 in the third OCS box B3.
[0062] In the second step, the group to which the first node NO1 belongs contains nodes that are the same except for the second element, specifically the first node NO1, the fifth node NO5, the ninth node NO9, and the thirteenth node NO13. In the second step, the first node NO1 connects to the fifth node NO5 in the first OCS box B1, to the thirteenth node NO13 in the second OCS box B2, and to the ninth node NO9 in the third OCS box B3.
[0063] In the third step, each selector S on the lower side of each OCS box B is connected to each selector S on the upper left side. The first OCS box B1 connects each of nodes NO. 1-16 to one of nodes NO. 17-32. The second OCS box B2 connects each of nodes NO. 1-16 to one of nodes NO. 49-64. The third OCS box B3 connects each of nodes NO. 1-16 to one of nodes NO. 33-48.
[0064] In the fourth step, each selector S on the lower side of each OCS box B is connected to each selector S on the upper right side. The first OCS box B1 connects each of nodes NO. 1-16 to one of nodes NO. 65-128. The second OCS box B2 connects each of nodes NO. 1-16 to one of nodes NO. 193-256. The third OCS box B3 connects each of nodes NO. 1-16 to one of nodes NO. 129-192.
[0065] In the processing system 1b shown in Figure 8, wiring work can be reduced by bundling fibers traveling in the same direction into a multi-core optical fiber. Figure 8 shows bundles of optical fibers in units of 16 or 64. The number of optical fibers to bundle should be determined according to the system size and the number of multi-core fibers.
[0066] (Third Embodiment) In the first embodiment, each node NO transmits and receives optical signals via optical connection, and the case in which the destination of the optical fiber to which the optical signal is transmitted is switched to the destination node using only the selector S was described. In the processing system 1c according to the third embodiment, each node NO is housed in switches E1, E2, etc., and transmits and receives signals via switches E1, E2, etc.
[0067] In the third embodiment, optical signals are transmitted and received between node NO and electrical switch E via an optical connection. The transmit port Tx of node NO transmits an optical signal to electrical switch E. The receive port Rx of node NO receives an optical signal.
[0068] In the first embodiment, each node NO has a number of ports equal to P = L-1. However, in reality, it is conceivable that the number of ports for each node NO may be less than L-1. In that case, the processing system 1b can handle the case where the number of ports for each node NO is less than L-1 by using switch E.
[0069] Each node NO in the third embodiment comprises one transmit port Tx, one receive port Rx, one transmit-side selector STx connected to the transmit port Tx, and one receive-side selector SRx connected to the receive port Rx. The transmit-side selector STx has k branch destinations, and the receive-side selector SRx has k branch destinations.
[0070] In the third embodiment, the identifiers of the multiple nodes are identified by a vector having k elements that are natural numbers from 1 to L, similar to the first embodiment. A set of L nodes that are the same except for the i-th element is defined as the i-th step set.
[0071] In the processing system 1c shown in Figure 9, switch E is described as an Ethernet switch (electrical switch), but any switch that can forward the signals transmitted and received by a port to a specific port is acceptable. Switch E may also be an electrical switch that converts optical signals into electrical signals for processing. Switch E may also be a switch formed by an optical splitter or the like that processes the signals as they are.
[0072] Switch E has L to the power of k ports. L to the power of k ports may be housed in one switch E or in multiple switch E units. The number of switch E units is not limited as long as it can accommodate each of the L to the power of k node numbers of the processing system 1c.
[0073] In the third embodiment, among the ports of switch E, the port that accommodates each node NO of the i-th step group in the i-th step is defined as the i-th step port group. The ports of switch E are divided into L to the power of k-1 switch groups. In the i-th step, each port belonging to the i-th step port group enables the exchange of partial results between each node of the i-th step group.
[0074] In step i, the transmitting selector STx and receiving selector SRx of node NO are each connected to ports of the i-th step port set that accommodate L-1 other nodes in the i-th step set to which node NO belongs. Multiple i-th step port sets are formed, and each node in the i-th step set is connected to a port in the same port set.
[0075] Each of the multiple node NOs of the processing system 1c simultaneously switches its branch destination in each step from the first step to the kth step, as described in the first embodiment.
[0076] In the third embodiment, the processing related to a predetermined node number will be described.
[0077] Here, "ports in the i-th step port group" are the ports to which the transmitting selector STx and receiving selector SRx of a predetermined node NO are connected. "Other ports in the i-th step port group" are the other ports in the port group to which the port of switch E to which the predetermined node NO is connected belongs. To each of "other ports in the i-th step port group," L-1 other nodes from the i-th step group, other than the predetermined node NO, are connected.
[0078] Node NO sends a partial result to a port in the i-th step port set. The port in the i-th step port set sends the partial result received from Node NO to the other ports in the i-th step port set. The other ports in the i-th step port set send partial results received from L-1 other nodes in the i-th step set to the ports in the i-th step port set.
[0079] A port in the i-th step port set sends the partial result received from the other ports in the i-th step port set to node NO. Node NO retrieves the partial results held by L-1 other nodes in the i-th step set. Node NO processes the partial results held by the L-1 other nodes in the i-th step set and the partial result held by node NO.
[0080] Here, we will explain the case where switch E is an electrical switch E. Each node NO sends an optical signal to electrical switch E, and the electrical switch E converts the optical signal into an electrical signal and switches the destination.
[0081] In Figure 7, groups of nodes in the same column or row were directly connected by optical connections. In contrast, each node NO in the processing system 1c according to the third embodiment shown in Figure 9 is connected via an electrical switch (Ethernet switch) E.
[0082] Referring to Figure 10, the connection between node NO and electrical switches E1-E4 in the processing system 1c according to the third embodiment will be explained. Each node NO is connected to electrical switches E1 and E2 via two selectors, one on the node NO side and one on the electrical switch E side. In Figure 10, the transmit / receive transceivers (TRx) are depicted together, but in reality, as in the first embodiment, a transceiver, selector, and optical fiber are provided for both transmission and reception. In Figure 10, the selector S is shown as a white triangular symbol. In this embodiment, a port may include transmit / receive transceiver (TRx) functionality.
[0083] Figure 10 shows an example where L=4 and k=2, attempting to perform the same calculation with the same number of nodes as in the first embodiment. However, since each node NO has only one port (less than L-1), the connection configuration shown in the first embodiment cannot be realized. However, by housing each node NO in an electrical switch as shown in Figure 10, in the first step, the selector is switched to enable the solid line connection, and data can be exchanged with other nodes (L-1 units) in the same group via the electrical switch. Similarly, in the second step, the dashed line connection is enabled to exchange data. In this way, by using the electrical switch E, it is possible to handle cases where the number of ports on each node is less than L-1.
[0084] In the third embodiment, each node NO is connected via an electrical switch E. In the processing system 1c shown in Figure 10, each node NO has only one port. Each node NO can exchange partial results with other nodes NO of the same set via the electrical switch E.
[0085] For example, with a port bandwidth of 400 Gbps, each node NO can communicate with each of the other three node NOs at 133 Gbps.
[0086] In the example shown in Figure 10, the total number of ports on the electrical switch E is 16. The 16 ports are housed in four electrical switches E, each having four ports.
[0087] Electrical switch E is a port that receives an optical signal from node NO, converts the optical signal into an electrical signal, and processes the electrical signal. Electrical switch E is a port that connects to a predetermined destination node, converts the electrical signal into an optical signal, and transmits it to the destination node.
[0088] Node NO transmits the optical signal, including the partial result, to the electrical switch E.
[0089] Electrical switch E is a receiving port connected to node NO, and receives an optical signal from node NO. Electrical switch E converts the received optical signal into an electrical signal. From the converted electrical signal, electrical switch E identifies the identifier of the destination node NO. The destination node NO is the node in the set to which the node NO of the optical signal source belongs.
[0090] There are several possible ways for the electrical switch E to identify the identifier of the destination node number. First, the source node number may include the identifier of the destination node number in the optical signal, and the electrical switch E may convert the optical signal into an electrical signal and then identify the identifier of the destination node number from the electrical signal.
[0091] Alternatively, the electrical switch E may share the transitions of each step from the first step to the kth step along with multiple node numbers, and may also hold data that identifies the identifier of the node number that sends and receives the partial result at each step. The electrical switch E may transfer the partial result to a node other than the source node number among the node numbers that send and receive the partial result at the current step, based on the current step and the identifier of the node number that transmits the optical signal. The control server may notify the electrical switch E of the timing of the transition and the identifier of the node to share, etc., at the same time that it notifies each node of the transition of each step, etc.
[0092] Electrical switch E transmits an electrical signal to a transmission port connected to a specified destination node. The transmission port converts the transmitted electrical signal into an optical signal, and then transmits the converted optical signal to the specified destination node.
[0093] The processing system 1c can set large sets such as L = 32 or 64 by using an electrical switch E. This allows the processing system 1c to accommodate many node numbers even when k is small. For example, when L = 64 and k = 2, the processing system 1c can accommodate node numbers Lk = 4096. This is effective when using a selector S that takes time to switch, as k-1 connection destinations need to be switched during calculation.
[0094] In Figure 10, separate electrical switches E are used for each group, but if there are many ports on the electrical switch E, multiple groups can be accommodated with a single electrical switch E.
[0095] The processing system 1c according to the third embodiment has lower communication efficiency compared to the processing system 1 according to the first embodiment, which realizes optical connection, when using an electrical switch E. However, the method disclosed in Non-Patent Literature 6 can pass through multiple electrical switches in communication between nodes. In contrast, the processing system 1c according to the third embodiment passes through only one electrical switch E, so an improvement effect can be expected.
[0096] In the third embodiment, the case where each node has 1 port was described, but it is not limited to this. In the first embodiment, it is assumed that each node has L-1 or more ports, whereas in the third embodiment, it is shown that partial calculations can be swapped even if the number of ports on each node is less than L-1. In the third embodiment, the number of ports on each node only needs to be 1 or more.
[0097] Furthermore, each node in the third embodiment may be combined with a technique that uses multiple ports to treat multiple links as a single entity, such as link aggregation.
[0098] The node NO described above in this disclosure uses, for example, a general-purpose computer system comprising a CPU (Central Processing Unit, processor) 901, memory 902, storage 903 (HDD: Hard Disk Drive, SSD: Solid State Drive), communication device 904, input device 905, and output device 906. In this computer system, each function of node NO is realized by the CPU 901 executing a program loaded onto memory 902.
[0099] Note that Node NO may be implemented on one computer or on multiple computers. Furthermore, Node NO may be a virtual machine implemented on a computer.
[0100] Node NO's programs can be stored on computer-readable storage media such as HDDs, SSDs, USB (Universal Serial Bus) memory, CDs (Compact Discs), and DVDs (Digital Versatile Discs), or distributed over a network. Computer-readable storage media are, for example, non-transitory storage media.
[0101] This disclosure is not limited to the embodiments described above, and numerous modifications are possible within the scope of its essence.
[0102] 1 Processing System 10 Switching Unit 20 Transmit / Receive Unit 30 Processing Unit 901 CPU 902 Memory 903 Storage 904 Communication Device 905 Input Device 906 Output Device NO Node Rx Receive Port S Selector SRx Receiver Selector STx Transmitter Selector Tx Transmitter Port
Claims
1. Used in a processing system having a number of nodes equal to L (where L is a natural number) raised to the power of k, comprising L-1 number of transmit ports, L-1 number of receive ports, L-1 number of transmit-side selectors connected to each of the transmit ports, and L-1 number of receive-side selectors connected to each of the receive ports, wherein one transmit-side selector has k branch destinations, and one receive-side selector has k branch destinations.
2. The node according to claim 1, wherein each of the L-1 transmitting ports transmits an optical signal, and each of the L-1 receiving ports receives an optical signal.
3. The node according to claim 1, wherein the identifiers of the plurality of nodes provided in the processing system are identified by a vector having k elements of any natural number from 1 to L, and all but the i-th element are the same as the node, forming a set of L nodes in the i-th step, and in the i-th step, each of the L-1 transmitting selectors is connected to one of the L-1 receiving selectors that are connected to the receiving ports of the other L-1 nodes belonging to the set of the i-th step, a switching unit that connects each of the L-1 receiving selectors to one of the L-1 transmitting selectors that are connected to the transmitting ports of the other L-1 nodes belonging to the set of the i-th step, a transmitting and receiving unit that transmits the partial results held by the node to the other connected nodes via each of the L-1 transmitting ports, and receives the partial results held by the other nodes from the other connected nodes via each of the L-1 receiving ports, and a processing unit that processes using the partial results held by the other nodes and the partial results held by the node.
4. Used in a processing system having a number of nodes equal to L (where L is a natural number) raised to the power of k, the node comprises L-1 transmit ports, L-1 receive ports, L-1 transmit side selectors connected to each of the transmit ports, and L-1 receive side selectors connected to each of the receive ports, one transmit side selector having k branch destinations, one receive side selector having k branch destinations, the identifier of the plurality of nodes in the processing system is identified by a vector having k elements of any natural number from 1 to L, the i-th step set consists of L nodes, all but the i-th element being the same as the node, and the node is, A processing method comprising: step i, each of the L-1 transmitting selectors is connected to one of the L-1 receiving selectors that are connected to the receiving ports of the other L-1 nodes belonging to the set of step i; each of the L-1 receiving selectors is connected to one of the L-1 transmitting selectors that are connected to the transmitting ports of the other L-1 nodes belonging to the set of step i; the partial results held by the node are transmitted to the other connected node via each of the L-1 transmitting ports; the partial results held by the other node are received from the other connected node via each of the L-1 receiving ports; and processing is performed using the partial results held by the other node and the partial results held by the node.
5. A processing system comprising a number of nodes equal to L (where L is a natural number) raised to the power of k, wherein the identifiers of the plurality of nodes are identified by a vector having k elements that are any natural number from 1 to L, and when L nodes that are the same except for the i-th element are considered a set of the i-th step, each of the plurality of nodes comprises: L-1 transmit ports, L-1 receive ports, L-1 transmit side selectors connected to each of the transmit ports, and L-1 receive side selectors connected to each of the receive ports, one transmit side selector having k branch destinations, one receive side selector having k branch destinations, in the i-th step, each of the L-1 transmit side selectors is connected to one of the L-1 receive side selectors connected to the receive ports of other L-1 nodes belonging to the set of the i-th step, each of the L-1 receive side selectors is connected to one of the L-1 transmit side selectors connected to the transmit ports of other L-1 nodes belonging to the set of the i-th step, A processing system in which each of the multiple nodes simultaneously switches the branch destination in each step from the first to the kth step.
6. A switch comprising a number of nodes equal to L (where L is a natural number) raised to the power of k, and a number of ports equal to L raised to the power of k, wherein each node comprises one transmit port, one receive port, one transmit-side selector connected to the transmit port, and one receive-side selector connected to the receive port, the transmit-side selector having k branch destinations, the receive-side selector having k branch destinations, the identifiers of the multiple nodes are identified by a vector having k elements that are any natural number from 1 to L, a set of L nodes that are the same except for the i-th element is defined as the i-th step, and among the ports of the switch, in the i-th step, the ports that accommodate each node of the i-th step set are defined as the i-th step port set, in the i-th step, the transmit-side selector and the receive-side selector of the node are each connected to the ports of the i-th step port set that accommodate L-1 other nodes of the i-th step set to which the node belongs. A processing system in which each of the multiple nodes simultaneously switches the branch destination in each step from the first to the kth step.
7. The processing system according to claim 6, wherein the node transmits an optical signal including a partial result to the switch, the switch receives the optical signal from the node at a receiving port connected to the node, identifies an identifier of a destination node from an electrical signal converted from the optical signal, transmits the electrical signal to a transmitting port connected to the identified destination node, and transmits the optical signal converted from the electrical signal to the identified destination node.
8. Used in a processing system comprising a number of nodes equal to L (where L is a natural number) raised to the power of k, and a switch having a number of ports equal to L raised to the power of k, wherein each node comprises one transmit port, one receive port, one transmit-side selector connected to the transmit port, and one receive-side selector connected to the receive port, the transmit-side selector has k branch destinations, the receive-side selector has k branch destinations, the identifiers of the multiple nodes are identified by a vector having k elements that are any natural number from 1 to L, a set of L nodes that are the same except for the i-th element is defined as the i-th step set, and among the ports of the switch, the ports that accommodate each node of the i-th step set in the i-th step are defined as the i-th step port set, in the i-th step, the transmit-side selector and the receive-side selector of the node are each connected to the ports of the i-th step port set that accommodate L-1 other nodes of the i-th step set to which the node belongs, and the node transmits a partial result to the ports of the i-th step port set. A processing method in which a port in the i-step port set transmits the partial result received from the node to another port in the i-step port set; another port in the i-step port set transmits the partial result received from L-1 other nodes in the i-step set to a port in the i-step port set; a port in the i-step port set transmits the partial result received from another port in the i-step port set to the node; and the node processes the partial result held by the L-1 other nodes in the i-step set.
Citation Information
Patent Citations
Filtering repeat function
JP2007067922A
Relay device and relay method
JP2011035728A
Method for enlarging communication network, node and communication band, and communication band enlarging program
WO2006107087A1