Network computer with two embedded rings
By adopting an improved topology in computer networks, processing nodes are arranged into multi-layer ring connections to form embedded one-dimensional paths, which solves the problem of low efficiency of all-reduce aggregates in existing technologies and achieves more efficient data transmission and resource utilization, which is particularly suitable for machine learning and neural network training.
Patent Information
- Application Number
- CN202180004037.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-30
- Filing Date
- 2021-03-24
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2041-03-24
AI Technical Summary
Existing computer networks suffer from inefficiency and resource waste when implementing all-reduce collectives, especially in machine learning applications. In particular, in ring interconnect architectures, existing technologies fail to efficiently utilize bandwidth and reduce latency.
An improved topology is adopted to arrange the processing nodes in a multi-layer configuration. Each layer includes at least four processing nodes, which are connected into a ring through intra-layer and inter-layer links, forming two embedded one-dimensional paths. This allows data to be operated simultaneously in two logical rings without sharing links, utilizing the bandwidth of each link and optimizing data transmission.
It improves the efficiency of the full-reduction aggregate, reduces resource waste, and optimizes the data transmission process. It is particularly suitable for data exchange and training neural networks in machine learning.
Smart Images

Figure CN114026551B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to data exchange between connected processing nodes in a computer, particularly but not exclusively for optimizing data exchange and machine learning / artificial intelligence applications. Background Art
[0002] Collectives are routines commonly used when processing data in a computer. They are routines that enable data to be shared and processed across multiple different processes, which may be running on the same processing node or different processing nodes. For example, if a process reads data from a data store, it can use a "broadcast" process to share that data with other processes. Another example is when the result of a particular function is needed by multiple processes. A "reduction" is a result that requires applying a computational function to data values from each of multiple processes. "Gather" and "Scatter" collectives process more than one data item. Certain collectives have become increasingly important in processing machine learning applications.
[0003] MPI (Message Passing Interface) is a message passing standard that can be applied to a variety of parallel computing architectures. MPI defines a number of collectives suitable for machine learning. Two such collectives are called "Reduce" and "Allreduce". The reduction operation enables the results of a computational function acting on multiple data values from different source processes to be provided at a single receiving process. Note that the receiving process can be one of the source processes. The allreduce collective reduces data values from multiple source processes and distributes the results to all source processes (which serve as receiving processes for the reduction results). For the "Reduce" or "Allreduce" operation, the reduction function can be any desired combinatorial function, such as sum, maximum, or minimum. According to the MPI standard, the allreduce collective can be implemented in the following way: the data values from all source processes are reduced in the reduce collective (for example, at one of the multiple processes) and the result is then broadcast to each source process.
[0004] Figure 1 1 is a schematic block diagram of a distributed architecture for training a neural network. A training data source 100 is provided. This can be a database or any other type of data store capable of retaining training data suitable for the neural network model being trained. The processing of the neural network model itself is distributed across multiple processing units 110a, 110b, 110c, etc. Figure 1Only three units are shown, but it will be readily appreciated that any number of processing units may be used. Each processing unit 110a, 110b, 110c receives batches of training data from a training data source 100. Each processing unit 110a, 110b, 110c retains a set of parameters 112a, 112b, 112c that define the model. The incoming batch of training data is processed in a calculation function 114 with the current set of parameters, and the result of the calculation function is used to generate a so-called delta, which represents the difference between the original parameters and the new parameters as a result of applying the calculation function to the batch of training data and the current set of parameters. In many neural networks, these parameters are called "weights", and so the incremental values are called "delta weights". In Figure 1 In FIG, weights are labeled 112a, 112b, 112c, and incremental weights are labeled 116a, 116b, 116c. It should be understood that in practice, weights and incremental weights are stored in a suitable storage device accessible to the processing unit. If weights and incremental weights can be kept in local memory, this will make the training process more efficient.
[0005] Figure 1 The architecture is not designed to train three separate models, but rather to train a single model in a distributed manner. Therefore, the goal is to have the model parameters (or weights) converge to a single common set in each processing unit. Obviously, starting from any particular set of weights, and assuming that the batch of training data received at each processing unit is different, the incremental weights produced by each computation function in each processing unit will be different. Therefore, a method is needed to combine and distribute the incremental weights across the processing units after each iteration of the batch training data. This is done in Figure 1 , where a combination function 118 receives incremental weights from each processing unit and performs a mathematical function (such as an average function) that reduces the incremental weights. The output of the combination function 118 is then fed back to the combination circuits 120a, 120b, and 120c within each processing unit. Thus, a new set of weights is generated, which is a combination of the original weights and the combined output from the combination function 118, and the new weights 118a, 118b, and 118c are stored back in the local memory. Then, the next batch of training data is provided to each processing unit, and the process is repeated multiple times. Obviously, if the starting weights of the processing units are the same, then after each iteration, they will be reset to the same new value again. It is easy to see that the above is an example of a particularly useful full reduction function. The incremental weights are provided to the combination function 118a, where they are reduced, and then these incremental weights are provided back to each processing unit in their reduced form, where they can be combined with the original weights.
[0006] Figure 1A is a schematic diagram illustrating how to implement an all-reduce aggregate in a line connection topology of six processing nodes N0...N5. These processing nodes may correspond to Figure 1 processing units, where the combination functions are distributed among these processing nodes, so there is no longer Figure 1 The processing nodes are shown as connected in a row configuration, where each processing node is connected via a "forward" link L F and the "backward" link L B connected to its neighbors. As shown in the figure, and as the directional word implies, the forward link will process the node in Figure 1A From left to right in the connection, the backward link will process the node in Figure 1A 200 and memory capacity. The processing nodes are connected from right to left. Each processing node has processing capacity, designated as 200, and memory capacity, designated as 202. The processing capacity and memory capacity can be implemented in any of a wide variety of ways. In one particular representation, a processing node may include multiple tiles, each individual tile having its own processing capacity and associated memory capacity. Each processing node also has one or more link interfaces to enable it to communicate via link L F / L B Connect to its neighbors.
[0007] To understand the implementation of the full-reduce collective, assume that the first node N0 has generated a partial vector labeled Δ0. A "partial" can be a data structure that includes an array of incremental weights (such as a vector or tensor). A partial vector is an array of multiple parts, each "part" corresponding to a calculation on a processing node. Each "part" can be a set of incremental weights. This is stored in storage capacity 202 ready to be exchanged in the full-reduce collective. In a simple "streaming" full-reduce algorithm, the forward link is used for "reduction" and the backward link is used for "broadcasting". The algorithm starts with a processing node N0 ( Figure 1A The left node in the figure sends its portion Δ0 to its neighboring node N1. At this node, the incoming portion (in this case Δ0) is reduced with the corresponding portion (Δ1) that has been generated by the computing power 200 at the processing node N1. This reduction is then Figure 1A The result after the addition of the function (shown as the add function in the figure) is sent from processing node N1 to the next connected node N2. As mentioned further in this article, the add function can be replaced by any combination function that can be used to reduce the parts. This process occurs at each processing node until the final processing node (at Figure 1A At N5), the reduction of the part is completed. At this time, the backward link L B, the reduction (the sum Δ) is sent back to each processing node. It is received at each node and stored in the memory capacity at that node, and then also transmitted to the next node. In this way, each processing node eventually ends up with the result of the reduction.
[0008] Figure 1B shows the timing diagram of the reduction and broadcast phases. Note that a processing node does not send the reduced result to the next node until it has received the incoming data from the previous node. Therefore, for each outgoing transmission on the forward link, there is a R inherent delay.
[0009] Furthermore, the backward link is not broadcast until the fully reduced result has been obtained at the end node. However, if the partial vectors are large, due to pipeline effects, the leading data item of the result, i.e., the first part of the reduction of the partial vector at each node, will be returned to the starting node before the starting node has finished sending its partial data item, so there may be substantial overlap in activity on all forward and backward links.
[0010] In a modification of this algorithm (which represents a minor improvement), the processing nodes at each end of the row can begin transmitting their portions to the central node, where the reduction is performed. In this case, the results are broadcast back to the end nodes. Note that in this scenario, there will be a reversal of direction of movement, for example at node N2 on both the forward and backward links. If a row is closed into a ring (by connecting the last node N5 to the first node N0 on both the backward and forward links), the pipeline algorithm can serialize the reduction and broadcast in the same direction, so that the two logical rings formed by the bidirectional links can each operate independently on half of the data. That is, each partial vector is split in half, and the first half Δ A is reduced on the forward link (e.g. Figure 1A As shown), and is broadcasted on the connection leg between N5 and N0. The other half vector Δ B It is reduced on the backward link and then broadcast on the connection ring of the backward link so that each node receives a copy of the full reduction.
[0011] Figure 1D The corresponding timing diagrams for the forward and reverse links are shown.
[0012] Figure 1C and Figure 1D The principles of the one-dimensional ring shown in can be extended to two-dimensional rings, such as in a torus or ring-connected computer.
[0013] Using a two-dimensional torus, an alternative approach is to implement all-reduce using a reduce-scatter collective followed by an allgather collective. Nikhil Jain and Yogish Sabharwal's paper, "Optimal Bucket Algorithms for Large MPI Collectives on Torus Interconnects" (ICS', June 2-4, Tsukuba), proposes bucket-based algorithms for allgather, reduce-scatter, and all-reduce collectives, assuming bidirectional links between processing nodes in a torus interconnect. The approach operates on the assumption that multiple data values (fragments) are processed in each step. As mentioned earlier, these fragments may be parts of partial vectors. In a reduce-scatter collective, each process starts with an initial partial vector. References to processes here are assumed to refer to processes executed on processing nodes. The partial vectors can be partitioned into multiple elements, or fragments. The corresponding elements of all processes are reduced, and then the reduced elements are distributed across the processes. In an all-gather collective, each process receives all elements from all other processes. Reduce-scatter collectives perform reductions on all parts and store each reduction on the corresponding node - see Figure 2 An all-reduce collective operation can be implemented by performing a reduce-scatter collective followed by an all-gather collective operation.
[0014] As discussed in Jain's paper, the torus interconnect is an attractive interconnect architecture for distributed memory supercomputers. In the discussion above, collectives have been explained in the context of communication between processes. In a distributed supercomputer, processing nodes are interconnected, and each processing node may be responsible for one or more processes (in the context of a collective). The torus interconnect is a type of mesh interconnect in which processing nodes are arranged in an n-dimensional array, with each node connected to its nearest neighbor and corresponding nodes on opposite edges of the array also connected. Bidirectional communication links exist between the interconnected processing nodes.
[0015] The algorithm for implementing the collective discussed in the aforementioned paper by Jain and Sabharwal is applied to a torus-connected architecture. This allows the collective to process different segments of a vector simultaneously in rings of different dimensions, making the processing bandwidth efficient. Despite the accepted view in the art that this is the case, the inventors have determined that the techniques proposed by Jain and Sabharwal are not optimal for symmetric or asymmetric toroids. A symmetric toroid is understood to be one in which the number of nodes in the non-axial rings of the torus matches the number of nodes in the axial rings. An asymmetric toroid is understood to be one in which the number of nodes in the non-axial rings does not match the number of nodes in the axial rings. Note that in both cases, the number of axial rings equals the number of nodes in the non-axial rings.
[0016] An object of the present disclosure is to propose an improved topology and method for implementing collectives, such as all-reduce functions, particularly but not exclusively for processing functions in machine learning. Summary of the Invention
[0017] Although embodiments of the present invention are described in the context of aggregates (such as all-reduce functions), it should be understood that the improved topologies and methods described herein have broader applicability.
[0018] According to one aspect of the present invention, a computer is provided, comprising a plurality of interconnected processing nodes arranged in a configuration having a plurality of layers arranged axially, each layer comprising at least four processing nodes, wherein the at least four processing nodes are connected into a ring via corresponding intra-layer links between each pair of adjacent processing nodes, wherein the processing nodes in each layer are connected to corresponding nodes in one or more adjacent layers via corresponding inter-layer links, and the computer is programmed to transmit data around two embedded one-dimensional paths, each logical path using all the processing nodes of the computer so that the embedded one-dimensional paths operate simultaneously without sharing links.
[0019] According to another aspect of the present invention, a computer is provided, comprising a plurality of interconnected processing nodes arranged in a configuration in which a plurality of layers of interconnected nodes are arranged axially, each layer comprising at least four processing nodes connected into a non-axial ring by at least corresponding intra-layer links between each pair of adjacent processing nodes, wherein each of the at least four processing nodes in each layer is connected to corresponding nodes in one or more adjacent layers by corresponding inter-layer links, the computer being programmed to provide two embedded one-dimensional paths in the configuration and to transmit data around each of the two embedded one-dimensional paths, each embedded one-dimensional path using all processing nodes of the computer so that the two embedded one-dimensional paths operate simultaneously without sharing links.
[0020] Embodiments of the present invention may provide one or more of the following, alone or in combination:
[0021] A computer, wherein the plurality of tiers includes a first-most tier and a second-most tier and at least one intermediate tier between the first-most tier and the second-most tier, wherein each processing node in the first-most tier is connected to a corresponding one of the processing nodes in the second-most tier;
[0022] The computer of claim 1 , wherein the configuration is a toroid configuration, wherein corresponding nodes of the respective connections of the plurality of layers form at least four axial rings;
[0023] A computer wherein the plurality of tiers includes a first-most tier and a second-most tier and at least one intermediate tier between the first-most tier and the second-most tier, wherein each processing node in the first-most tier is connected to non-adjacent nodes in the first-most tier in addition to its adjacent nodes, and each processing node in the second-most tier is connected to non-adjacent nodes in the second-most tier in addition to its adjacent nodes;
[0024] A computer wherein each processing node is configured to output data on its respective intra-layer link and inter-layer link, wherein there is equal bandwidth utilization on each of the intra-layer link and inter-layer link of the processing node;
[0025] A computer in which each of a plurality of layers has exactly four nodes;
[0026] A computer comprising a number of layers arranged along the axis, the number of layers being greater than the number of processing nodes in each layer;
[0027] A computer in which the number of layers arranged along the axis is the same as the number of nodes in each layer;
[0028] A computer wherein the intra-layer links and the inter-layer links comprise fixed connections between processing nodes;
[0029] A computer wherein at least one of the inter-layer link and the intra-layer link includes switching circuitry operable to selectively connect one of the processing nodes to one of a plurality of other processing nodes;
[0030] A computer wherein at least one of an inter-layer link and an intra-layer link of a processing node in a first end-most layer includes a switching circuit operable to disconnect the processing node from its corresponding node in a second end-most layer and connect it to a non-adjacent node in the first end-most layer;
[0031] A computer wherein at least one interlayer link of a processing node in a first end-most layer includes switching circuitry operable to disconnect the processing node from its neighboring node in the first end-most layer and connect it to a corresponding node in a second end-most layer;
[0032] A computer wherein each embedded one-dimensional path comprises an alternating sequence of one of the inter-layer links and one of the intra-layer links;
[0033] A computer programmed to transmit data in a transmission direction in each layer, the transmission direction being the same in all layers within each one-dimensional path;
[0034] A computer according to any preceding claim, wherein each one-dimensional embedded path comprises a sequence of processing nodes visited in each layer that is the same in all layers within each one-dimensional path.
[0035] A computer programmed to transmit data in a transmission direction in each layer, the transmission direction being different for transmission in successive layers around each one-dimensional path;
[0036] A computer according to any one of claims 1 to 12, wherein each one-dimensional embedded path comprises a sequence of processing nodes that are visited in a direction in each layer, the direction being different in successive layers of each one-dimensional path.
[0037] A computer comprising six layers, each layer having four processing nodes connected in a ring;
[0038] A computer comprising eight layers, each layer having eight processing nodes connected in a ring;
[0039] A computer comprising eight layers, each layer having four processing nodes connected in a ring;
[0040] A computer comprising four layers, each layer having four processing nodes connected in a ring;
[0041] A computer in which the rings of each layer into which the processing nodes are connected are non-axial;
[0042] A computer wherein each processing node is programmed to divide a corresponding portion vector of the processing node into segments and to transmit data in successive segments around each one-dimensional path;
[0043] A computer programmed to operate each path as a set of logical rings, wherein successive segments are transmitted around each logical ring in simultaneous transmission steps;
[0044] A computer programmed to transmit data in data transmission steps, wherein in each data transmission step, each link of a processing node is utilized with the same bandwidth as other links of the processing node, i.e., there is symmetric bandwidth utilization;
[0045] A computer wherein each processing node is configured to simultaneously output a corresponding fragment on each of two links, wherein the fragments output on each of the links have the same size or approximately the same size;
[0046] A computer wherein each processing node is configured to reduce a plurality of incoming segments with a plurality of corresponding locally stored segments; and / or
[0047] A computer wherein each processing node is configured to transmit fully reduced segments simultaneously on each of its intra-layer links and inter-layer links in an all-gather phase of an all-reduce collective.
[0048] Another aspect of the present invention provides a method for generating a set of programs to be executed in parallel on a computer, the computer comprising a plurality of processing nodes, the plurality of processing nodes being connected in a configuration of a plurality of axially arranged layers, each layer comprising at least four processing nodes, the at least four processing nodes being connected in a non-axial ring via corresponding intra-layer links between each pair of adjacent processing nodes, wherein the processing nodes in each layer are connected to corresponding corresponding nodes in each adjacent layer via inter-layer links, the method comprising:
[0049] generating at least one data transfer instruction for each program to define a data transfer phase for transferring data from a processing node executing the program, wherein the data transfer instruction includes a link identifier that defines an output link from which data is to be transferred from the processing node in the data transfer phase; and
[0050] Link identifiers are determined to transmit data around each of two embedded one-dimensional paths provided by the configuration, each path utilizing all processing nodes of the computer such that the embedded one-dimensional logic paths operate simultaneously without sharing links.
[0051] In some embodiments of the method, each program includes one or more instructions to deactivate any of its inter-layer links and intra-layer links that are not used in the data transmission step.
[0052] In some embodiments of the method, each program includes one or more instructions for dividing a corresponding portion of a vector of a processing node on which the program is executed into segments and transmitting data in the form of consecutive segments over respectively defined links.
[0053] In some embodiments, in each data transfer step, each link of a processing node is utilized with the same bandwidth as the other links of the processing node, that is, the configuration operates with symmetric bandwidth utilization.
[0054] Another aspect of the present invention provides a method for executing a set of programs in parallel on a computer, the computer comprising a plurality of processing nodes, the processing nodes being connected in a configuration having a plurality of layers arranged along an axial direction, each layer comprising at least four processing nodes, the at least four processing nodes being connected in a ring via corresponding intra-layer links between each pair of adjacent processing nodes, wherein the processing nodes in each layer are connected to corresponding corresponding nodes in each adjacent layer via inter-layer links, the method comprising:
[0055] executing at least one data transfer instruction in each program to define a data transfer phase in which data is transferred from a processing node executing the program, wherein the data transfer instruction includes a link identifier defining an output link over which data is to be transferred in the data transfer phase;
[0056] Link identifiers have been determined to transmit data around each of the two embedded one-dimensional paths, each logical ring using all processing nodes of the computer so that the embedded one-dimensional paths operate simultaneously without sharing links. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] For a better understanding of the invention and to show how it may be put into practice, reference will now be made, by way of example, to the accompanying drawings.
[0058] Figure 1 is a diagram illustrating distributed training in neural networks.
[0059] Figure 1A is a diagram showing a row of processing nodes for implementing a simple "streaming" row-line reduction algorithm.
[0060] Figure 1B This is the timing diagram of the "streaming" row-reduce algorithm.
[0061] Figure 1C is a diagram of a row with end nodes connected in a ring.
[0062] Figure 1D It is the timing diagram of the ring-reduce algorithm.
[0063] Figure 2 is a diagram showing the implementation of an all-reduce function via a reduce-scatter step followed by an all-gather step.
[0064] Figure 3A and Figure 3B A bucket-based all-reduce algorithm is shown.
[0065] Figure 4A A computer network in the form of a 4×6 ring is shown, in which two isomorphic rings are embedded.
[0066] Figure 4B and Figure 4C Each isomorphic embedded ring is shown according to one embodiment.
[0067] Figure 4D It shows Figure 4A A three-dimensional graph of one of two embedded rings in a computer network.
[0068] Figure 4E It shows Figure 4A 3D illustration of two embedded rings within another embedded ring in a computer network.
[0069] Figure 5A and Figure 5B Two isomorphic embedded rings are shown, which can be embedded in a 4×4 computer network connected in a ring.
[0070] Figure 6A and Figure 6B Represents each of two isomorphic embedded rings on a 4×8 computer network connected in a ring.
[0071] Figure 7A and Figure 7B Represents each embedded ring that can be embedded into two isomorphic embedded rings on an 8×8 computer network connected in a ring.
[0072] Figure 8A A computer network in the form of a 4x6 diagonal closed prism is shown.
[0073] Figure 8B and Figure 8C Shown embedded in Figure 8A Two isomorphic rings on the network.
[0074] Figure 8D It shows Figure 8A A 3D graph of one of two embedded rings in a computer network. DETAILED DESCRIPTION
[0075] Aspects of the present invention have been developed in the context of a multi-tile processor designed to act as an accelerator for machine learning workloads. The accelerator comprises a plurality of interconnected processing nodes. Each processing node may be a single multi-tile chip, a package of multiple chips, or a rack of multiple packages. The goal of this paper is to design a machine that is efficient in deterministic (repeatable) computation. The processing nodes are interconnected in such a way that aggregates, in particular but not limited to broadcasts and all-reduces, can be implemented efficiently. However, it should be noted that the embodiments of the invention described herein may have other applications.
[0076] One particular application is updating models when training neural networks using distributed processing. In this context, distributed processing utilizes multiple processing nodes located in different physical entities (such as chips, packages, or racks). That is, data transmission between processing nodes requires exchanging messages over physical links.
[0077] The challenges in developing topologies specifically for machine learning differ from those in the general field of high-performance computing (HPC) networks. HPC networks typically emphasize on-demand asynchronous, all-to-all, personalized communication, so dynamic routing and bandwidth overprovisioning are normal. Excess bandwidth can be overprovisioned in HPC networks to reduce latency rather than to provide bandwidth. Overprovisioning active communication links wastes power that could contribute to computing performance. The most common link types used in computing today consume power when they are active, regardless of whether they are being used to transmit data.
[0078] The inventors have developed a machine topology that is particularly well-suited for ML workloads and addresses the following properties of ML workloads. This embodiment provides a different structure in which two rings are embedded in an m×n computer network, where m is the number of nodes in each of multiple layers of the network, n is the number of layers, and each ring accesses all nodes in the network.
[0079] In ML workloads, inter-chip communication is currently dominated by broadcast and all-reduce collectives. Broadcast collectives can be implemented as scatter collectives followed by all-gather collectives, and all-reduce collectives can be implemented as reduce-scatter collectives followed by all-gather collectives. In this context, the term "inter-chip" refers to any communication between processing nodes connected via external communication links. As mentioned earlier, these processing nodes can be chips, packages, or chassis.
[0080] Note that the communication link can be between chips on one printed circuit board or between chips on different printed circuit boards.
[0081] Workloads can be compiled so that within a single Intelligent Processing Unit (IPU) machine, all-to-all communication is primarily inter-chip.
[0082] The all-reduction aggregate has been described above and is shown in Figure 2 middle. Figure 2 A set of partial values or "partial" vectors P0, P1, P2, P3 at each of the four nodes in the starting state S1 is shown. In this context, a node is a processing node in a network of processing nodes. Note that each node N0, N1, N2, N3 has four "corresponding" parts, which are labeled accordingly (large diamond grid, wide stripes diagonally downward, large square grid, wide stripes diagonally upward). That is, each part has a position in its part vector, so that P0(n) at node n has the same position in its vector as P0(n+1) at node n+1. The suffix (n) is used to indicate the node where the part is located - so, P0(0) is part P0 at node N0. In a reduce-scatter pass, the corresponding parts are reduced, and the reduction is given to one of the nodes. For example, parts P0(0), P0(1), P0(2), P0(3) are reduced (reduced to r0) and placed at node N0. Similarly, parts P1(0), P1(1), P1(2), and P1(3) are reduced (to r1) and placed on node N1. And so on, so that: in the intermediate state S2, each node has one of the reductions r0, r1, r2, and r3. As explained, the reduction can be performed by any combinatorial function f(Pi0), which can include independent operators (such as max) or associative operators = P1(0)*P1(1)*P1(2)*P1(3).
[0083] Then, in an all-gather pass, each reduction is given to all nodes to activate state S3, where each node now holds all four reductions. Note that in S1, the "corresponding" parts (e.g., P0(0), P0(1), P0(2), and P0(3)) might all be different, whereas in state S3, each reduction (e.g., r0) is the same at all nodes, where r i =f{(P i (0),P i (1),P i (2),P i(3))}. In machine learning, a set of parts P0, P1, P2, P3 is a vector. Each pass through the model during training produces a vector of parts (e.g., updated weights). In state S3, the reduction r0, r1, r2, r3 shown by the diamond grid, diagonal downward stripes, square grid, and diagonal upward stripes at each node is the fully reduced vector, that is, the "result" or vector of the fully reduced parts. In the context of machine learning, each part can be an updated increment of a parameter in the model. Alternatively (in an arrangement not further described herein), it can be an updated parameter.
[0084] Figure 3A and Figure 3B The diagram illustrates a bucket-based algorithm for reduce-scatter / all-gather, which assumes six "virtual" rings. These rings are also referred to herein as "logical" rings. Figure 3A is a diagram illustrating the reduction of parts in multiple virtual rings. Each part is divided into six fragments. Figure 3A In , the capital letters R, Y, G, B, P, and L each represent a different fragment of the part stored at each node, represented by a shaded diamond grid, diagonal upward stripes, a square grid, horizontal stripes, diagonal downward stripes, and vertical stripes. The letters represent corresponding fragments to be reduced together, and define a "virtual" ring or "logical" ring for these fragments. See Figure 3A , the “R” segments in each part P0, P1, P2, P3 and P4 are reduced to a single segment (R∑) in the result vector. The same is true for the Y, G, B, P and L segments.
[0085] Figure 3B A timing diagram is shown, where the time on the horizontal axis represents the data exchange and computation in each step of the all-reduce process. Figure 3A and Figure 3B In the process, the full reduction process is completed through the reduce-scatter phase followed by the full gather phase.
[0086] exist Figure 3B In the image, each segment is represented by the following different shading: R-diamond grid, Y-slanting upward stripes, G-square grid, B-horizontal stripes, P-slanting downward stripes, L-vertical stripes.
[0087] Figure 3A and Figure 3BThe representation in is as follows. Each part is denoted as P0, P1, P2, P3, P4, and P5. At the beginning of the process, each part is stored on the corresponding node N0, N1, N2, N3, N4, and N5. Each fragment is labeled according to its fragment number and its position in the virtual ring in which the fragment is considered to be reduced. For example, RA0 represents the R fragment in part P0 because it is the first fragment in the virtual ring formed by nodes N0-N1-N2-N3-N4-N0.
[0088] RA1 represents the R segment at node N1, which is located in the second position in its virtual ring. YA0 represents the Y segment at node N1. The "0" suffix indicates that it is the first segment in its virtual ring, and the Y ring is N1-N2-N3-N4-N0-N1. Note that the A suffix reflects the virtual ring and does not correspond to a physical node (or part). Note that Figure 3A Only the virtual ring on the forward link is shown. Figure 3B An equivalent process is shown taking place on the reverse link, where the segment is designated B.
[0089] In step one, the first segment (A0) in each virtual ring is transmitted from its node to the next adjacent node, where the first segment is reduced to the corresponding segment at the adjacent node. That is, RA0 moves from N0 to N1, where RA0 is reduced to R(A0+A1) at N1. Again, the "+" symbol is used here as an abbreviation for any combination function. Note that in the same step, the 10 segments of A0 for each virtual ring will be transmitted simultaneously. That is, the link between N1 and N2 is used to transmit YA0, the link between N2 and N3 is used to transmit GA0, and so on. In the next step, the corresponding reduced segments are transmitted to their next adjacent node via the forward link. For example, R(A0+A1) is transmitted from N1 to N2, and Y(A0+A1) is transmitted from N2 to N3. Note that for clarity, Figure 3A Not all 15 segments are numbered, nor are all transmissions numbered. The complete set of segments and numbers is shown in Figure 3B This process is performed in five steps. After all five steps, all segments at each node have been reduced. At the end of the fifth step, the reduction is performed on the last node of each corresponding ring for that segment. For example, R is reduced to node N5.
[0090] The starting point of the full gathering phase is the transfer from the last node in each virtual ring to the first node. Therefore, the final reduction of the R segment ends at node N5, thus preparing for the first step of the full gathering phase. The final reduction of the Y segment ends at node N0 accordingly. In the next step of the full gathering phase, the reduced segments are again transferred to their next neighboring node. Therefore, the fully reduced R segment is now also at N2, the fully reduced Y segment is now also at N3, and so on. In this way, at the end of the full gathering phase, each node ends up with all the fully reduced segments R, Y, G, B, P, L of the partial vector.
[0091] If the computation required for the reduction can be hidden behind pipeline delays, then the implementation of the algorithm is efficient. The inventors have noticed that in forming a suitable ring in a computer to implement an all-reduction, efficiency is highest if the tour of the ring visits each node in the ring only once.
[0092] Therefore, the rows with bidirectional links ( Figure 1A ) is not the most efficient ring.
[0093] An improved topology for the interconnection network of processing nodes will now be described that allows efficient exchange of partial sum results between processing nodes to implement all-reduce collectives.
[0094] Figure 4A is a diagram showing the connection topology of multiple processing nodes. Figure 4AIn the example, there are twenty-four processing nodes connected in a ring formation, but it should be understood that these principles can be extended to different numbers of nodes, some of which are illustrated in the following description. Furthermore, the principles described here can be extended to different topologies of diagonally closed square prisms, as described later. Other configurations employing these principles are contemplated. For ease of reference, each processing node is numbered. In the following description, the prefix N will be inserted when referring to a node. For example, N0 represents the processing node on the left. The processing nodes are connected via links in the manner to be described. Each link can be bidirectional, meaning it can transmit data in both directions. Links can operate so that this bidirectional functionality can occur simultaneously (that is, the link can be used in both directions simultaneously). Note that there is physical interconnectivity and logical connectivity. Logical connectivity is used to form two embedded continuous rings. Note that the embedded rings are also referred to here as "paths." These terms are interchangeable, but it is important to recognize that the term "virtual ring" is reserved for the scenario outlined above, where multiple segments may operate within the virtual ring on each embedded ring or path. In some embodiments, each embedded ring (or path) can operate in both directions. First, the physical connectivity will be described. The processing nodes are connected in a ring configuration. Each processing node along the y-axis is connected to its neighboring nodes via a single bidirectional link. For clarity, Figure 4A Not all links are labeled in the diagram. However, links from node N0 are shown. Link L04 connects processing node N0 to processing node N4, located below it on the y-axis. Note that the reference to "below" implies a specific orientation of the computer network. In practice, no orientation of the computer network is implied; any directional description is purely for explanation purposes with reference to the accompanying figures. The network consists of multiple layers organized along the y-axis. In each layer, there are four processing nodes connected to form a ring via corresponding bidirectional links. Each layer ring is considered non-axial because it does not extend along the y-axis. For example, processing node N0 has link L01 connecting it to its adjacent node in its layer. Node N0 also has link L03 connecting it to its other adjacent node N3 in the layer. The ring structure is completed by corresponding processing nodes in the "end-most" layer connected by bidirectional links. Note that the term "end-most" is a convenient reference to the accompanying figures. In reality, in a ring, corresponding connected nodes in adjacent layers form a continuous axial ring. For example, node N0 in the first end-most layer is connected to node N20 in the second end-most layer via link L020. Note that in the figure, the end-most layers are distinct from the middle layers (those formed by nodes N4 to N7, N8 to N11, N12 to N15, and N16 to N19, as they are connected together at their corresponding processing nodes). In reality, they are part of a continuous ring.
[0095] Figure 4AThe links shown in FIGURE 1 may be embodied in various representations. Some specific examples will be discussed later. However, it is important to note that each link can be a single physical link structure providing a bidirectional communication path over that physical link structure. Alternatively, each direction of the link structure can have a separate physical representation. Note that a link can be a fixed link. That is, in the case where the link connects two processing nodes together, it is fixed in place after the network has been established and configured. Alternatively, the link can be connected to or include switching circuitry that enables the connectivity of the network to be changed after it has been established.
[0096] According to the novel principles described here, Figure 4A The physical connectivity shown enables two logical embedded rings (each optionally bidirectional) to be embedded in the network. Figure 4B The first such ring R1 is shown. Figure 4B Not all nodes are shown with reference numerals, but it should be understood that they are Figure 4A The nodes are the same as shown in . Figure 4B Ring R1 in Figure 1 extends through the following nodes in a continuous path along which data can be transmitted. Ring R1 extends through a series of nodes, from node N0 to N1 to N5 to N6 to N10 to N11 to N15 to N12 to N16 to N17 to N21 to N22 to N2 to N3 to N7 to N4 to N8 to N9 to N13 to N14 to N18 to N19 to N23 to N20, and back to N0. Ring R2 extends from N0 to N3 to N23 to N22 to N18 to N17 to N13 to N12 to N8 to N11 to N7 to N6 to N2 to N1 to N21 to N20 to N16 to N19 to N15 to N14 to N10 to N9 to N5 to N4, and back to N0, visiting each node in turn.
[0097] Each ring includes all twenty-four processing nodes. Note also that two rings can be used simultaneously because no links in the same ring are in use. Therefore, there are no conflicts on any single path between processing nodes. This means there are no shared links between the rings. These rings are called homogeneous rings because they each have the same length and traverse the same number of processing nodes.
[0098] Figure 4D A three-dimensional diagram showing ring R1 is shown. Note that the other is the same but rotated 90 degrees around the y-axis. When the program implements the previously described full reduction ring algorithm, consider using Figure 4D The structure shown. Each node outputs size, where n is the number of nodes and v is the size of the data structure being reduced - scatter or all-gather - at a particular stage. Initially, v is the size of the partial vector. Before each step around the ring, the number of segments is equal to the number of nodes in the ring. In most embodiments, each segment has the same size. However, there may be cases where, for example, the number of elements in the vector is not divisible evenly, and the segments may have slightly different sizes. In this case, they are roughly the same size - they may differ by one or two elements depending on the division factor. Note that in contrast to the structure described in the aforementioned Jain paper, each ring passes through all nodes and all links are used at all times. Each processing node can output its data on four links simultaneously and can be configured to operate with full bandwidth utilization. That is, if the node bandwidth is specified as B, then the bandwidth utilization of each link is B / 4. This is symmetric bandwidth utilization for each processing node. Consider Figure 4D At the first, last layer of the network shown, data is transmitted from N0 to N1 along link L01. The arrow indicates this direction of data transmission. As mentioned earlier, the ring may also transmit data in the opposite direction. However, considering the forward direction indicated by the arrow, the next step in the path is from node N1 to node N5. Therefore, the path uses an intra-layer link from N0 to N1 and an inter-layer link from N1 to N5. The next step in the path is the intra-layer link (N5 to N6), followed by an inter-layer link from N6 to N10. Therefore, the path includes a continuous sequence of intra-layer links and inter-layer links. In each layer, the nodes can be accessed in one of two directions: clockwise and counter-clockwise. In Figure 4D In , the arrows indicate that in the first end layer, the nodes are visited in a clockwise direction. Similarly, the nodes are visited in the next middle layer and all subsequent layers in a clockwise direction.
[0099] However, please note that this is not necessarily the case. That is, the direction in which nodes are visited around a particular layer can be the same in each layer, or different in each layer. In some embodiments, it is the same in each layer, while in other embodiments, it is different in different layers, such as in successive layers. Note that if the link is also bidirectional, data can be transmitted in either direction around each path. The following references are to one direction of data transmission to explain the order in which nodes are visited in each embedded path. For example, in Figure 4E In the embodiment of FIG, data is transmitted in the ring from node N0 to node N1 and then to node N5, and then transmitted in a counterclockwise direction along the intra-layer links in the intermediate layer. Then, it passes through the inter-layer links to the next intermediate layer, and then transmitted in a clockwise direction on the intra-layer links of the next intermediate layer.
[0100] It is clear that both symmetric and asymmetric architectures can achieve symmetric bandwidth utilization - where the symmetry of the configuration is determined by the relative number of processing nodes in a tier relative to the number of tiers in the configuration.
[0101] Figure 5A and Figure 5B Two embedded paths in a 4×4 network configuration are shown. Figure 5A and Figure 5B The node numbers in are taken from Figure 4A The network configuration (first four lines). This is just an example. By disconnecting and reconnecting Figure 4A Nodes in a 4×6 configuration can be provided in a 4×4 configuration, in which case the nodes will correspond. However, it is also possible to build a 4×4 configuration with your own nodes. Figure 5A and Figure 5B The interconnections between the nodes show the two embedded paths in the configuration.
[0102] Figure 6A and Figure 6B Two embedded paths in a 4×8 network configuration are shown. The node numbers are Figure 4A The same as in , with the additional nodes N24 to N31 in the bottom two rows. As mentioned before, it is possible to expand Figure 4A The 4×6 configuration forms a Figure 6A and Figure 6B The 4×8 configuration shown is possible, but it is also possible to build a 4×8 configuration from its own network nodes.
[0103] Figure 6A and Figure 6B The interconnection between each node in represents the corresponding two embedded paths in the configuration.
[0104] refer to Figure 7A and Figure 7B , Figure 7A and Figure 7B Two embedded paths in an 8×8 network configuration are shown. Figure 7A and Figure 7B The nodes in Figure 6A The nodes in the configuration are marked with additional nodes N32 to N63 in the four additional columns. Figure 6A configuration, thus forming Figure 7A and Figure 7B Configuration shown. Alternatively, Figure 7A and Figure 7B The configuration in can be built from its own source node.
[0105] Figure 7A and Figure 7B The interconnection between each node in the two embedded rings in the network configuration is shown separately.
[0106] Figure 8A Another embodiment of a computer network consisting of twenty-four processing nodes arranged in a 4×6 diagonal square prism is shown. Figure 4A The ring configuration shown is similar. However, there are some differences. The nodes are again arranged in successive layers arranged along the axis, with each layer consisting of four nodes connected to form a ring through corresponding links between processing nodes. The structure and behavior of the links can be as described above with reference to Figure 4A The corresponding processing nodes are each connected to their neighboring nodes in the next layer through the corresponding layer links. Figure 8A In the , nodes are called "N'1", "N'2", etc. to distinguish them from Figure 4A However, in practice, processing nodes can be Figure 4A The same type of processing nodes in .
[0107] Figure 8A The network configuration in Figure 4A The difference in the configuration of the network in is the way the nodes in the last layer are connected. Figure 4A In , each node in the last layer is connected to its corresponding node in the other last layer. This forms a ring. In contrast, in Figure 8A In the first terminal layer, diagonally opposite processing nodes are connected to each other. That is, node N'0 is connected to node N'2, and node N'1 is connected to node N'3.
[0108] Accordingly, at another endmost layer, node N'20 is connected to node N'22, and node N'21 is connected to node N'23.
[0109] Figure 8A The network can be configured to embed two isomorphic rings R'1 and R'2, respectively. Figure 8B and Figure 8C As shown, the ring R'1 passes through nodes N'0 to N'1 to N'5 to N'4 to N'8 to N'9 to N'13 to N'12 to N'16 to N'17 to N'21 to N'20 to N'23 to N'19 to N'18 to N'14 to N'15 to N'11 to N'10 to N'6 to N'7 to N'3 to N'2, and then back to N'0.
[0110] Ring R'2 extends from node N'0 to N'3 to N'1 to N'2 to N'6 to N'5 to N'9 to N'10 to N'14 to N'13 to N'17 to N'18 to N'22 to N'21 to N'23 to N'20 to N'16 to N'19 to N'15 to N'12 to N'8 to N'11 to N'7 to N'4 to N'0.
[0111] Again, for clarity, note that Figure 8B and Figure 8C Not all nodes are marked.
[0112] like Figure 4A In the network shown, the bandwidth utilization of each processing node is symmetrical. For example, consider processing node N'3. It has four links, and the bandwidth utilization of each link is B / 4, where B is the total node bandwidth.
[0113] Figure 8D is a three-dimensional diagram showing ring R'1. Another identical ring is rotated 90 degrees around the y-axis. Again, the arrows on the links indicate the direction of data transmission in one direction along the ring. Data can also be transmitted in the opposite direction. In this case, data is transmitted from node N'0 to node N'1, through the diagonal connecting link to node N'3, and then clockwise in the layer to node N'4. The data is then transmitted to the next layer through the inter-layer links and is transmitted in a counter-clockwise direction along the layer in the intra-layer links before extending into the inter-layer links to connect to the next layer. Therefore, the path again includes consecutive intra-layer and inter-layer connections. At the next layer, the data is shown to be transmitted in a clockwise direction. However, note that unlike Figure 4A Like a ring, the direction of visiting nodes around a layer may change. For example, it can be the same in all layers or different in different layers.
[0114] The capacity of the computer can be expanded by adding additional processing nodes. These can be added in the form of additional layers in the y-axis direction, or in the form of additional nodes in each layer in the x-axis direction. Note that the term x-axis is used here - although this refers to the "non-axial" ring mentioned earlier. To do this, the interconnectivity of the processing nodes can be changed. For example, consider adding an additional layer at the end of the bottom layer, such as Figure 4A As shown. The links from nodes N20, N21, N22, N23 will be disconnected and each connected to the corresponding processing node in the extra layer. These nodes are not in Figure 4A , but the principle is obvious. These additional nodes will then have links connecting them back to the end-most layers N0, N1, N2, N3. The intra-layer links between the additional processing nodes connect to the additional processing nodes in the ring. Note that the connectivity of the rest of the configuration remains unchanged.
[0115] The ring configuration can be reconnected to form a diagonally closed square prism. To achieve this, the links connecting the end layers together are broken. Figure 4A, link L020 is disconnected and connected between nodes N0 and N2. The link extending between nodes N23 and N3 is disconnected, and node N3 is connected to node N1. Similarly, at the bottom layer, node N23 is connected to node N21, and node N22 is connected to node N20.
[0116] Thus, by reconnecting these links, a diagonally closed square prism can be created from the ring.
[0117] In some embodiments described herein, the computer network has a 4×n structure, where 4 represents the number of processing nodes in each layer and n represents the number of layers. In each case, two isomorphic data transmission rings are embedded, each passing through all processing nodes of the network.
[0118] Each processing node has symmetric bandwidth utilization. That is, each link from a processing node has the same bandwidth utilization as other links from that processing node.
[0119] The two embedded homogeneous rings use all bandwidth, and there are no shared links between the two rings. In other words, due to the lack of link sharing, each ring is able to have full link bandwidth.
[0120] As described above, in one embodiment, the computer network is hardwired into a fixed configuration. In other embodiments, the links are switchable. That is, each link can be connected to a switch, or a switch can be part of the link. In particular, if the top and bottom links are switchable, they can be used to extend the network or switch between rings or diagonal prisms. Note that switching between fixed, hardwired configurations can be done by manually disconnecting wires. If a switch is used, there may be automatic changes between configurations.
[0121] The advantage of the diagonally closed square prism configuration is that the maximum cable length required between processing nodes can be shorter than that of a ring. It can be easily seen that the number of cables required to connect to the same layer ( Figure 8A The cable length closed between processing nodes in the top and bottom end layers (in a toroidal configuration) is less than the "wraparound" link required to connect nodes in the top end layer to nodes in the bottom end layer in a toroidal configuration. As already mentioned, the cable length in a toroidal coil can be reduced by adopting a folded structure.
[0122] However, the advantage of the ring configuration is that the worst-case path for exchanging data between any two processing nodes is shorter than that in the prism case where the diagonals are closed.
[0123] Note that the network can be fault-tolerant in different ways. For example, two physical links can be provided on each link path between processing nodes.
[0124] In another example, each physical link may have multiple lanes (such as in the case of PCI Express) so that the link automatically adapts to a failure on one lane of the link. The link may run slower, but will still function.
[0125] Note that by embedding two rings in the fabric, each passing through all of the fabric's processing nodes, in the event of a complete failure of one ring (e.g., due to a broken link), the other processing ring may still be able to operate. In the case of implementing a machine learning algorithm (such as an all-reduce), the operations of one ring may still be able to make a certain amount of data amenable to the all-reduce operation. In some training contexts, this may be sufficient to support continued operation of the algorithm until the failed ring can be repaired.
[0126] Each node is capable of performing processing or computing functions. Each node can be implemented as a single processor. However, it is more likely that each node will be implemented as a single chip or chip package, where each chip includes multiple processors. There are many possible different manifestations of each individual node. In one example, the node can be composed of intelligent processing units of the type described in British Patent Applications (Publication Nos. GB1816891.4; 1816892.2; 1717299.0; the contents of which are incorporated herein by reference). However, the techniques described herein can be used on any type of processor that constitutes a node. This article outlines a method for exchanging data in an efficient manner to implement specific exchange patterns that are useful in machine learning models. In addition, the links can be represented in any suitable manner. Advantageously, these links are bidirectional, and preferably, these links can operate in both directions simultaneously, although this is not a requirement. A specific class of communication links is the SERDES link, whose power requirements are independent of the amount of data transmitted over the link or the time it takes to transmit the data. SERDES is an acronym for Serializer / Deserializer, and this type of link is known. In order to transmit a signal on the wires of such a link, power needs to be applied to the wires to change the voltage to generate the signal. A SERDES link has the following characteristics: power is continuously applied to the wires to keep them at a certain voltage level so that the signal can be conveyed by changes in that voltage level (rather than by changes between 0 and the applied voltage level). Therefore, the bandwidth capacity on a SERDES link has a fixed power regardless of whether it is used. A SERDES link is implemented at each end by a circuit that connects the link layer devices to the physical link (such as a copper wire). This circuit is sometimes called a PHY (physical layer). PCIe (Peripheral Component Interconnect Express) is an interface standard for connecting high-speed computers.
[0127] It is possible to dynamically deactivate links to effectively not consume power when not in use. However, the activation times and non-deterministic nature of machine learning applications often make dynamic activation during program execution problematic. Therefore, the inventors have determined that it may be better to exploit the fact that for any particular configuration, the power consumption of chip-to-chip links is essentially constant, so the best optimization is to maximize the use of physical links by keeping chip-to-chip traffic as concurrent as possible with IPU activity.
[0128] The SERDES PHY is full-duplex (i.e., 16 Gbit / s, the PHY supports 16 Gbit / s in each direction simultaneously), so full link bandwidth utilization means balanced bidirectional traffic. Also note that using direct chip-to-chip communication has significant advantages over indirect communication such as via a switch. Direct chip-to-chip communication is more power efficient than switched communication.
[0129] Another factor to consider is the bandwidth requirements between nodes. One goal is to have enough bandwidth to hide inter-node communication behind the computations performed on each node for distributed machine learning.
[0130] When optimizing machine architectures for machine learning, all-reduce ensembles can be used as a metric for required bandwidth. An example of an all-reduce ensemble was given above in the treatment of parameter updates for model averaging. Other examples include gradient averaging and computing norms.
[0131] As an example, consider the all-reduce requirement of residual learning networks. Residual learning networks are a type of deep convolutional neural network. In deep convolutional neural networks, multiple layers are used to learn corresponding features within each layer. In residual learning, residuals can be learned instead of learned features. A specific residual learning network called ResNet implements direct connections between different layers of the network. It has been shown that in some cases, training such residual networks can be easier than traditional deep convolutional neural networks.
[0132] ResNet 50 is a 50-layer residual network. ResNet 50 has 25M weights, so a full reduction of all weight gradients in the single-position floating-point format F16 involves a 50-megabyte portion. To illustrate bandwidth requirements, assume that each complete batch requires a full reduction. This is likely (but not necessarily) a full reduction of the gradients. To achieve this, each node must output 100 megabits per full reduction. ResNet 50 requires 250 gigaflops (billion floating-point operations per second) per image for training. If the subbatch size per processing node is 16 images, each processor performs 400 gigaflops per full-reduction ensemble. If the processor achieves 100 teraflops (trillion floating-point operations per second), approximately 25 gigabits per second are required between all links to maintain concurrency between computation and full-reduction communication. For subbatches of 8 images per processor, the required bandwidth nominally doubles, which is partially alleviated by reducing the achievable Teraflops per second to process smaller batches.
[0133] Implementing an all-reduce collective between p processors, where each processor starts with a part of size m megabytes (equal to the reduction size), requires that at least 2m.(p-1) megabytes be sent over the link. So if each processor has l links on which it can send simultaneously, the asymptotic minimum reduction time is 2m.(p-1).(p-1) divided by (p1).
[0134] The above-described concepts and techniques can be used in a number of different examples.
[0135] In one example, a fixed configuration for use as a computer is provided. In this example, processing nodes are interconnected as described and illustrated in the various embodiments discussed above. In such an arrangement, only necessary intra-layer links and inter-layer links are placed between processing nodes.
[0136] A fixed configuration can be constructed from a precise number of processing nodes in that configuration. Alternatively, it can be provided by partitioning it from a larger structure. That is, a set of processing nodes can be provided arranged as stacked layers. The processing nodes in each stacked layer can have: inter-layer links to corresponding processing nodes in adjacent stacked layers; and intra-layer links between adjacent processing nodes in that layer.
[0137] A fixed configuration of a desired number of stacked layers can be provided by disconnecting each inter-layer link in a given stacked layer in the original set of stacked layers and connecting it to an adjacent processing node in the given stacked layer to provide an intra-layer link. In this manner, a given stacked layer in the original set of stacked layers can be made to form one of the first and second most end layers of the structure. Note that the layers of the original set of layers can be partitioned into more than one fixed configuration structure in this manner.
[0138] Inter-layer links and intra-layer links are physical links provided by appropriate buses or lines as described above. In one embodiment, each processing node has a set of lines extending from it for connecting it to another processing node. This can be accomplished, for example, through one or more interfaces in each processing node, each of which has one or more ports connected to the physical lines.
[0139] In another manifestation, the links can be formed by on-board wiring. For example, a single board can support a group of chips, such as four chips. Each chip has an interface with ports that can connect to other chips. Connections between chips can be formed by soldering wiring to the board according to a predetermined method. Note that the concepts and techniques described here are particularly useful in this context because they maximize the use of links between chips that are already pre-soldered on the printed circuit board.
[0140] The concepts and techniques described with reference to some embodiments may be particularly useful because they enable optimized use of non-switchable links. Configurations can be constructed by connecting the processing nodes described herein using fixed, non-switchable links between the nodes. In some manifestations, there is no need to provide additional links between processing nodes if these additional links are not to be used.
[0141] In order to use the configuration, a set of parallel programs is generated. The set of parallel programs contains node-level programs, that is, programs that are specified to work on a specific processing node in the configuration. A set of parallel programs to operate on a specific configuration can be generated by a compiler. It is the compiler's responsibility to generate node-level programs that correctly define the links to be used for each data transfer step for certain data. These programs include one or more instructions for implementing data transfer effects in a data transfer phase, which use link identifiers to identify the links to be used for that transfer phase. For example, a processing node may have four active links at any time (doubled if the links are also bidirectional). The link identifier enables the correct link to be selected for the data item in that transfer phase. Note that each processing node may not be aware of the actions of its neighboring nodes - the exchange activities are precompiled for each exchange phase.
[0142] Note also that the links do not necessarily have to be switched - the data item does not need to be actively routed while it is being transmitted, nor does the connectivity of the link need to be changed. However, as described, a switch may be provided in some embodiments.
[0143] As described above, the configuration of the computer network described herein is intended to enhance parallelism in computing. In this context, parallelism is achieved by loading node-level programs into the processing nodes of the configuration intended to be executed in parallel, for example, in a distributed manner as described above to train artificial intelligence models. However, it will be readily understood that this is only one application of the parallelism achieved by the configuration described herein. One approach to achieving parallelism is called "bulk synchronous parallel" (BSP) computing. According to the BSP protocol, each processing node performs a computation phase and an exchange phase following the computation phase. During the computation phase, each processing node performs its computational tasks locally but does not exchange its computational results with other processing nodes. During the exchange phase, each processing node is allowed to exchange the computational results of the previous computation phase with other processing nodes in the configuration. A new computation phase will not begin until the exchange phase of the configuration is completed. In this form of the BSP protocol, a barrier is placed for synchronization at the juncture of transitioning from the computation phase to the exchange phase, or from the exchange phase to the computation phase, or both.
[0144] In this embodiment, when the exchange phase is initiated, each processing node executes instructions for exchanging data with its neighboring nodes using the link identifier determined by the compiler for that exchange phase. The nature of the exchange phase can be determined using the aforementioned MPI message passing standard. For example, a collective, such as an all-reduce collective, can be called from a library. In this way, the compiler has a precompiled node-level program that controls the links over which partial vectors (or corresponding fragments of partial vectors) are transferred.
[0145] Obviously other synchronization protocols can be used.
[0146] While specific embodiments have been described, other applications and variations of the disclosed technology may become apparent to those skilled in the art once given the disclosure herein.The scope of the present disclosure is not limited by the described embodiments, but only by the claims that follow.
Claims
1. A computer comprising a plurality of interconnected processing nodes arranged in a configuration in which a plurality of layers of interconnected processing nodes are arranged axially, Each layer includes at least four processing nodes connected in a non-axial ring via at least one corresponding intra-layer link between each pair of adjacent processing nodes. in, Each of the at least four processing nodes in each layer is connected to a corresponding node in one or more adjacent layers via a corresponding inter-layer link. The computer is programmed to provide two embedded one-dimensional paths in a configuration and to transmit data around each of the two embedded one-dimensional paths, each embedded one-dimensional path utilizing all processing nodes of the computer in a manner such that the two embedded one-dimensional paths operate simultaneously without sharing links, Each processing node is programmed to divide the corresponding partial vector of the processing node into segments and transmit data around each embedded one-dimensional path in the form of consecutive segments.
2. The computer according to claim 1, wherein The configuration is a ring configuration, wherein corresponding nodes of the respective connections of the plurality of layers form at least four axial rings.
3. The computer according to claim 1, wherein The multiple layers include a first end-most layer and a second end-most layer and at least one intermediate layer between the first end-most layer and the second end-most layer, wherein each processing node in the first end-most layer is connected to non-adjacent nodes in the first end-most layer in addition to its adjacent nodes, and each processing node in the second end-most layer is connected to non-adjacent nodes in the second end-most layer in addition to its adjacent nodes.
4. The computer according to claim 1, wherein Each processing node is configured to output data on its respective intra-layer links and inter-layer links with equal bandwidth utilization on each of the intra-layer links and inter-layer links of the processing node.
5. The computer according to claim 1, wherein Each layer of the plurality of layers has exactly four nodes.
6. The computer according to claim 1, comprising several layers arranged in an axial direction, the number of the layers being greater than the number of processing nodes in each layer. 7 . The computer according to claim 1 , comprising several layers arranged along an axial direction, the number of the layers being the same as the number of nodes in each layer.
8. The computer according to claim 1, wherein The intra-layer links and inter-layer links include fixed connections between processing nodes.
9. The computer according to claim 1, wherein: At least one of the inter-layer link and the intra-layer link includes switching circuitry operable to selectively connect one of the processing nodes to one of a plurality of other processing nodes.
10. The computer according to claim 3, wherein At least one of the inter-layer links and intra-layer links of the processing node in the first-most layer includes a switching circuit that is operable to disconnect the processing node from its corresponding node in the second-most layer and connect it to a non-adjacent node in the first-most layer.
11. The computer according to claim 3, wherein: At least one of the inter-layer links of the processing node in the first-most layer includes switching circuitry operable to disconnect the processing node from its neighboring node in the first-most layer and connect it to a corresponding node in the second-most layer.
12. The computer according to claim 1, wherein Each embedded one-dimensional path includes an alternating sequence of one of the inter-layer links and one of the intra-layer links.
13. The computer according to claim 1, wherein Each embedded one-dimensional path includes a sequence of processing nodes that are visited in a direction in each layer that is the same in all layers of each one-dimensional path.
14. The computer according to claim 1, wherein Each embedded one-dimensional path includes a series of processing nodes, the sequence of which is visited in a direction in each layer, the direction being different in successive layers within each one-dimensional path.
15. The computer of claim 1 comprising six layers, each layer having four processing nodes connected in a non-axial ring.
16. The computer of claim 1 comprising eight layers, each layer having eight processing nodes connected in a non-axial ring.
17. The computer of claim 1 comprising eight layers, each layer having four processing nodes connected in a ring.
18. The computer of claim 1 comprising four layers, each layer having four processing nodes connected in a ring.
19. The computer of claim 1 programmed to operate each path as a set of logical rings, wherein: The consecutive segments are transmitted around each logical ring in a simultaneous transmission step.
20. The computer according to claim 1, wherein Each processing node is configured to simultaneously output a corresponding fragment on each of the two links, wherein the fragments output on each link have approximately the same size.
21. The computer according to claim 1, wherein Each processing node is configured to reduce a plurality of incoming segments using a plurality of respective corresponding locally stored segments.
22. The computer according to claim 21, wherein Each processing node is configured to transmit fully reduced segments simultaneously on each of its intra-layer links and inter-layer links in the all-gather phase of the all-reduce collective.
23. A computer according to any one of claims 1 to 19, programmed to transmit data in the data transmission steps such that in each data transmission step, each link of a processing node is utilized with the same bandwidth as other links of the processing node.
24. A method of generating a set of programs to be executed in parallel on a computer, the computer comprising a plurality of processing nodes connected in a configuration having a plurality of layers arranged in an axial direction, Each layer includes at least four processing nodes connected into a non-axial ring via corresponding intra-layer links between each pair of adjacent processing nodes. in, The processing nodes in each layer are connected to corresponding nodes in each adjacent layer via inter-layer links, the method comprising: generating at least one data transfer instruction for each program to define a data transfer phase for transferring data from a processing node executing the program, wherein the data transfer instruction includes a link identifier that defines an output link over which data is to be transferred from the processing node in the data transfer phase; and determining link identifiers for transmitting data around each of two embedded one-dimensional paths provided by the configuration, each embedded one-dimensional path utilizing all processing nodes of the computer in a manner such that the embedded one-dimensional logic paths operate simultaneously without sharing links, Each program includes one or more instructions to divide a corresponding partial vector of a processing node on which the program is executed into segments and transmit data in the form of continuous segments on respectively defined links.
25. The method according to claim 24, wherein Each program includes one or more instructions to deactivate any one of its inter-layer links and intra-layer links that are not used in the data transmission step.
26. The method of claim 24, comprising transmitting data in the data transmission steps, wherein in each data transmission step, each link of a processing node is utilized with the same bandwidth as other links of the processing node.
27. A method of executing a set of programs in parallel on a computer, the computer comprising a plurality of processing nodes connected in a configuration having a plurality of layers arranged in an axial direction, Each layer includes at least four processing nodes connected into a non-axial ring via corresponding intra-layer links between each pair of adjacent processing nodes. in, The processing nodes in each layer are connected to corresponding nodes in each adjacent layer via inter-layer links, the method comprising: executing at least one data transfer instruction in each program to define a data transfer phase for transferring data from a processing node executing the program, wherein the data transfer instruction includes a link identifier that defines an output link over which the data is to be transferred in the data transfer phase; Link identifiers have been determined for transmitting data around each of two embedded one-dimensional paths, each embedded one-dimensional path operating simultaneously with the embedded one-dimensional paths without sharing the links using all processing nodes of the computer, Each program includes one or more instructions to divide a corresponding partial vector of a processing node on which the program is executed into segments and transmit data in the form of continuous segments on respectively defined links.
Citation Information
Patent Citations
Instruction set
GB2569275A
Synchronization in a multi-tile processing array
GB2569430A
Scheduling tasks in a multi-threaded processor
GB2569843A
System-On-a-Chip and Multi-Chip Systems Supporting Advanced Telecommunication Functions
US20100158005A1