Collective offload cost management for networks
The collective engine optimizes network topology and communication connections to address scalability issues in collective networks, enhancing efficiency and reducing costs through strategic configuration and cost model-based management.
Patent Information
- Application Number
- JP2025531742
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2023-12-14
- Publication Date
- 2025-12-15
AI Technical Summary
Collective communication in networks faces scalability issues as the number of nodes increases, leading to inefficiencies in communication processing.
A collective engine manages communication costs by optimizing network topology and configuring communication connections based on cost models, using parameters like number of ports, tree depth, and stride value to minimize latency and improve efficiency.
The solution enhances network communication efficiency by reducing costs and preventing deadlocks, ensuring optimal data transmission and processing across interconnected nodes.
Smart Images

Figure 2025540529000001_ABST
Abstract
Description
[Background technology]
[0001] Various improvements to computing performance, such as increasing the number of processing cores, result in increased performance but may reach scalability limits. Collective communication enables global communication operations between all processes / nodes in a system (e.g., a collective network) that includes networked nodes. As the number of nodes increases, collective communication may suffer from scalability issues. To ensure better scalability, certain communication processing is offloaded from the nodes of the collective network (e.g., its processors) to other nodes (e.g., network adapters and switches, etc.), which may be managed by a collective engine that may reside on a server or other connected computing device.
[0002] The accompanying drawings illustrate several exemplary embodiments and are a part of this specification and, together with the following description, demonstrate and explain various principles of the present disclosure. [Brief explanation of the drawings]
[0003] [Figure 1] FIG. 1 is a block diagram of an exemplary system for a collective engine. [Figure 2] FIG. 1 is a block diagram of an exemplary collective network. [Figure 3] FIG. 1 is a simplified diagram of a broadcast and reduce operation. [Figure 4] 1 is a simplified diagram of scattering and collecting operations. [Figure 5] 1 is a simplified diagram of a full collection operation and a reduced scatter operation. [Figure 6] FIG. 1 is a schematic diagram of a total reduction operation. [Figure 7A] FIG. 1 is a network topology diagram of the initial phase of a full reduction operation. [Figure 7B] FIG. 1 is a network topology diagram of an intermediate phase of a full reduction operation. [Figure 7C]FIG. 10 is a network topology diagram of the final phase of the full reduction operation. [Figure 8] 1 is a table of network topology cost models for all reduce operations. [Figure 9A] FIG. 1 illustrates a network topology when chunks are routed. [Figure 9B] FIG. 1 illustrates a network topology when chunks are routed. [Figure 9C] FIG. 1 illustrates a network topology when chunks are routed. [Figure 9D] FIG. 1 illustrates a network topology when chunks are routed. [Figure 9E] FIG. 1 illustrates a network topology when chunks are routed. [Figure 9F] FIG. 1 illustrates a network topology when chunks are routed. [Figure 9G] FIG. 1 illustrates a network topology when chunks are routed. [Figure 10] FIG. 1 is a flow diagram of an example method for collective offload cost management in a network. [Figure 11] FIG. 1 is a flow diagram of an example method for collective offload routing management in a network. [Figure 12] FIG. 1 is a flow diagram of an example method for collective offload message fragmentation management in a network. DETAILED DESCRIPTION OF THE INVENTION
[0004] Throughout the drawings, like reference numerals and descriptions indicate similar, but not necessarily identical, elements. While the exemplary embodiments described herein are susceptible to various modifications and alternative forms, specific embodiments have been shown by way of example in the drawings and are described in detail herein. However, the exemplary embodiments described herein are not intended to be limited to the particular forms disclosed. Rather, the present disclosure covers all modifications, equivalents, and alternatives falling within the scope of the appended claims.
[0005] The present disclosure is generally directed to collective offload management for a network. As described in more detail below, embodiments of the present disclosure provide a collective engine that can manage various aspects of collective offload to networked nodes that communicate with each other to share data and / or processing of data.
[0006] The present disclosure is generally directed to collective offload cost management for networks. As described in more detail below, embodiments of the present disclosure configure a topology of a collective network to perform collective operations based on a communication cost model, which may evaluate and / or otherwise optimize for communication costs (e.g., latency associated with transmitting data), such as to reduce and / or minimize the cost for communicating between nodes of the collective network when performing the collective operations. The systems and methods provided herein may improve the efficiency of the collective network, for example, by establishing a more efficient topology that reduces network communication costs.
[0007] In one embodiment, a device for collective offload cost management of a network includes control circuitry configured to select a communication cost model for a collective operation and configure a topology of the collective network for performing the collective operation based on the selected communication cost model.
[0008] In some examples, the control circuitry is configured to select the communication cost model by optimizing a communication cost of the collective operation by evaluating a plurality of communication cost models for the collective operation and selecting a cost model from the plurality of communication cost models corresponding to the optimized communication cost. In some examples, the communication cost model includes a plurality of parameters, and optimizing the communication cost includes determining an optimized parameter with respect to the plurality of parameters. In some examples, configuring the topology is based on using the optimized parameter as a topology parameter.
[0009] In some examples, the plurality of parameters includes at least one of a number of upstream ports, a number of downstream ports, a number of processors, a number of ports per processor, a tree depth, and a stride value. In some examples, configuring the topology includes configuring communication connections between nodes at a level of the collective network and nodes at an adjacent level of the collective network based at least on an optimized stride value corresponding to a number of connected nodes at the next level.
[0010] In some examples, optimizing the communication costs includes flattening a tree associated with the cost model. In some examples, the control circuitry is configured to configure a portion of the topology based on the selected communication cost model. In some examples, the control circuitry is configured to configure a second portion of the topology based on a second communication cost model.
[0011] In one embodiment, a system for collective offload cost management of a network includes a memory, a processor, and control circuitry configured to evaluate cost parameters of a communication cost model for a collective operation and configure a topology of a collective network to perform the collective operation based on the cost parameters.
[0012] In some examples, the control circuitry is configured to evaluate cost parameters of the communication cost model by optimizing a communication cost of the collective operations. In some examples, optimizing the communication cost includes determining optimized parameter values for the cost parameters, and configuring the topology is based on using the optimized parameter values as topology parameters. In some examples, the cost parameters include at least one of a number of upstream ports, a number of downstream ports, a number of processors, a number of ports per processor, a tree depth, and a stride value.
[0013] In some examples, configuring the topology includes configuring communication connections between nodes at a level of the collective network and nodes at an adjacent level of the collective network based at least on an optimized stride value corresponding to a number of connected nodes at the next level. In some examples, optimizing communication costs includes flattening a tree associated with the cost model.
[0014] In some examples, the control circuitry is further configured to configure a portion of the topology based on the evaluated communication cost model and to configure a second portion of the topology based on a second communication cost model.
[0015] In one embodiment, a method for collective offload cost management of a network includes: (i) evaluating a plurality of communication cost models for a collective operation; (ii) determining values for a plurality of cost model parameters based on the evaluation; and (iii) configuring a topology of the collective network to perform the collective operation using a plurality of topology parameters corresponding to the values for the plurality of cost model parameters.
[0016] In some examples, evaluating the plurality of communication cost models includes optimizing a communication cost of the collective operation, and values for the plurality of cost model parameters are determined based on the optimized communication costs. In some examples, configuring the topology includes configuring communication connections between nodes at a level of the collective network and nodes at an adjacent level of the collective network based on the plurality of topology parameters. In some examples, the method further includes configuring a portion of the topology based on the plurality of topology parameters and configuring a second portion of the topology based on a second communication cost model.
[0017] Any features of the embodiments described herein may be used in combination with each other in accordance with the general principles described herein. These and other embodiments, features, and advantages will be more fully understood from the following detailed description, taken in conjunction with the accompanying drawings and claims.
[0018] A detailed description of collective offload management for a network is provided below with reference to Figures 1-10. A detailed description of an exemplary system and network for collective operations is provided with reference to Figures 1 and 2. Exemplary diagrams of collective operations are provided with reference to Figures 3, 4, 5, and 6. An exemplary diagram of collective network routing for an allreduce operation is provided with reference to Figures 7A-7C. Exemplary costs associated with collective network node routing for an allreduce operation are provided with reference to Figure 8. Exemplary data routing incorporating message fragmentation and deadlock prevention for an allreduce operation is provided with reference to Figures 9A-9G. A detailed description of a corresponding computer-implemented method is also provided with reference to Figures 10, 11, and 12.
[0019] 1 illustrates an exemplary collective network 100 (e.g., an interconnected network) for implementing aspects of the present disclosure. Collective network 100 includes computing device 102A, network 104, and computing device 102B. Each of computing device 102A and / or computing device 102B may be, and / or part of, network devices such as network switches and routers, and / or client or user devices such as desktop computers, laptop computers, tablet devices, smartphones, or other computing devices such as servers. Computing device 102A includes physical processor 110A, and computing device 102B includes physical processor 110B. Physical processor 110A and / or physical processor 110B may correspond to one or more instances of any type or form of hardware-implemented processing unit capable of interpreting and / or executing computer-readable instructions, including, but not limited to, a chiplet (e.g., smaller and, in some examples, more specialized processing units that may cooperate as a single chip), a microprocessor, a microcontroller, a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA) implementing a soft-core processor, an application-specific integrated circuit (ASIC), a system-on-chip (SoC), a coprocessor such as a digital signal processor (DSP), a neural network engine (NNE), an accelerator, a graphics processing unit (GPU), one or more portions thereof, one or more variations or combinations thereof, and / or any other suitable physical processor. In some embodiments, a physical processor referred to herein (e.g., physical processor 110A and / or physical processor 110B) may correspond to a host processor together with a coprocessor, which, in some examples, may be a separate processor.
[0020] In some examples, processor 110A and / or processor 110B access and / or modify data and / or instructions stored in memory 120A and / or memory 120B. Memory 120A and / or memory 120B each correspond to an instance of any type or form of volatile or non-volatile storage device or medium capable of storing data and / or computer-readable instructions, including, but not limited to, random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid-state drive (SSD), optical disk drive, cache, any variation or combination of one or more of these, and / or any other suitable storage memory. In some embodiments, memory 120A and / or memory 120B each may correspond to internal memory of processor 110A and / or processor 110B.
[0021] In particular embodiments, memory 120A may store one or more processes, such as process 122A, and memory 120B may store one or more processes, such as process 122B. Process 122A and / or process 122B may represent one or more software applications, programs, and / or processes that, when executed by a computing device, cause the computing device to perform one or more tasks, such as portions of the collective operations described herein. For example, as described in more detail below, one or more of process 122A and / or process 122B may represent instructions (e.g., corresponding to collective operations and / or portions thereof) stored and configured to be executed on one or more computing devices, such as the devices shown in FIG. 1 (e.g., computing device 102A and / or computing device 102B). In some examples, each process (e.g., an instance of process 122A and / or process 122B) may correspond to a localized version of the collective operation (e.g., instructions for a node to perform that portion of the collective operation).
[0022] Computing device 102A is communicatively connected to computing device 102B through network 104. Network 104 represents any type or form of communication network, such as the Internet, and may include one or more physical connections, such as a LAN, and / or wireless connections, such as a WAN.
[0023] 1, computing device 102A includes control circuitry 112. In some embodiments, control circuitry 112 corresponds to and / or is an embodiment of a collective engine for offloading collective communications (e.g., collective operations) from nodes in a collective system corresponding to collective network 100 (e.g., locally from one or more physical processors 110A and / or processes 122A and / or remotely from one or more computing devices 102B and one or more physical processors 110B and / or processes 122B). While FIG. 1 shows computing device 102A and computing device 102B, in other examples, collective network 100 may include additional computing devices that computing device 102A may manage in collective operations. Additionally, in some embodiments, one or more instances of computing device 102A and / or computing device 102B may correspond to devices (e.g., a CPU, a GPU, a network interface card (NIC), etc.) that are part of a larger computing device.
[0024] Control circuitry 112 may correspond to circuitry and / or instructions for managing collective communications / operations of nodes, such as components of computing device 102A and / or computing device 102B and other instances thereof. In some examples, a node may correspond to an application (e.g., software that may use / process data involved in a collective operation, such as may be represented by process 122A and / or process 122B) and / or a device (e.g., a computing or networking device, such as a processor, network interface, switch, etc., that may perform portions of the collective operation locally by receiving data, processing it as needed, and forwarding it to an appropriate node and / or a shared memory device or, in some examples, a device including addressable memory that may not process the data itself), and / or a portion of a device (e.g., different processing components / circuitry within a device). In some embodiments, control circuitry 112 may coordinate or otherwise manage which nodes perform which collective operations and / or portions thereof (e.g., coordinate which nodes execute which instructions required for the collective operation). In some embodiments, control circuitry 112 may further manage network topology and / or other routing aspects of the collective network.
[0025] The collective engine (e.g., control circuitry 112) can establish a network topology for the collective network, which in some examples refers to the connection pattern of nodes in an interconnected network and is often represented by a graph. A node or vertex in some examples refers to a machine (e.g., a server, switch, and / or router, which in some embodiments may correspond to computing device 102A and / or computing device 102B) or device (e.g., a central processing unit (CPU) and / or its cores, a graphics processing unit (GPU) and / or its cores, a network interface or network interface card (NIC), a switch, etc., which in some embodiments may correspond to computing device 102A and / or computing device 102B and / or portions thereof) in the network topology, which may have various physical and / or virtual ports (e.g., corresponding to interfaces for wired and / or wireless communication) for communicating with other nodes in the collective network. A server in some examples refers to a machine including devices such as a processor and a network interface card. A switch, in some examples, refers to a machine or device that can receive packets from any number of input ports and send packets to any number of output ports, which, in some examples, may further be capable of processing the data. A router, in some examples, refers to a machine or device that can forward packets between networks (e.g., via a physical port). A network interface, in some examples, refers to a device used by a machine to connect to a network and that can further forward packets. An edge, in some examples, refers to a connection between one or more vertices / nodes. In interconnected networks, a single edge may include multiple links, and each link may have multiple channels. At the root of a topology, multiple processes may be performed (e.g., by a machine / device) simultaneously.
[0026] 2 illustrates a simplified exemplary collective network system 200 corresponding to collective network 100. System 200 includes server 202 (corresponding to an instance of computing device 102A and / or computing device 102B), GPU 214A (corresponding to an instance of computing device 102A and / or computing device 102B), GPU 214B (corresponding to an instance of computing device 102A and / or computing device 102B), NIC 216A (corresponding to an instance of computing device 102A and / or computing device 102B), NIC 216B (corresponding to an instance of computing device 102A and / or computing device 102B), and switch 218 (corresponding to an instance of computing device 102A and / or computing device 102B). The collective engine (e.g., control circuitry 112) may establish a topology (e.g., the ring of FIG. 2) for the nodes (e.g., server 202, GPU 214A, GPU 214B, NIC 216A, NIC 216B, and switch 218) by establishing directional links (which may be bidirectional in some examples) between the nodes. In other words, the collective engine may establish which nodes communicate with which other nodes.
[0027] The collective engine may further help coordinate processes executing on various nodes in performing collective communication operations that represent (e.g., as instructions) distributed operations to be performed on data. Examples of collective operations include broadcast, reduce (to-one), scatter, gather, allgather, reduce-scatter, and allreduce, as described further below.
[0028] 3 illustrates collective operations 300 corresponding to broadcast and / or reduce operations for process 302A, process 302B, process 302C, process 302D, process 302E, process 302F, process 302G, process 302H, and process 302I for phases 304 and 306. Each of processes 302A-302I may correspond to a process (e.g., one or more iterations of process 122A and / or process 122B) executed on one or more nodes (e.g., one or more iterations of computing device 102A and / or computing device 102B, respectively) in an interconnected network (e.g., collective network 100).
[0029] In a broadcast operation, data belonging to one process (e.g., process 302B) is sent to all processes (e.g., process 302A, process 302C, process 302D, process 302E, process 302F, process 302G, process 302H, and process 302I). More specifically, the broadcast operation can begin with phase 304, during which process 302B has a set of data that can be a complete vector or a message, but in other examples, can correspond to one or more data elements. As a result of the broadcast operation, in phase 306, the set of data (e.g., the same vector) is broadcast to all other processes such that all of processes 302A-302I have the same set of data. FIG. 3 illustrates each index of the data vector with a different pattern.
[0030] In a (-to-one) reduce operation, data (e.g., a complete vector / message that may be common to all processes) belonging to all processes (e.g., processes 302A-302I) may be reduced to a single set / vector and sent to a single process (e.g., process 302B). More specifically, as shown in FIG. 3, the reduce operation may begin with phase 306, during which data held by processes 302A-302I may be reduced to a single vector in phase 304 and sent to a single target process. In some examples, the reduce may be the inverse of the broadcast, as shown in FIG. 3.
[0031] 4 illustrates collective operations 400 corresponding to scatter and collect operations for process 402A, process 402B, process 402C, process 402D, process 402E, process 402F, process 402G, process 402H, and process 402I for phases 404 and 406. Each of processes 402A-402I may correspond to a process (e.g., one or more iterations of process 122A and / or process 122B) performed on one or more nodes (e.g., one or more iterations of computing device 102A and / or computing device 102B, respectively) in an interconnected network (e.g., collective network 100).
[0032] In a scattering operation, data (e.g., a complete vector) belonging to one process (e.g., process 402B in phase 404) is parsed out (e.g., by index) to all processes (e.g., processes 402A-402I in phase 406). More specifically, as shown in Figure 4, each process is indexed such that in phase 406, each of processes 402A-402I receives data at a corresponding index, with the first process (e.g., process 402A) receiving the first data element of the vector, the second process (e.g., process 402B) receiving the second data element of the vector, and so on. Figure 4 illustrates each index of the data vector with a different pattern.
[0033] In a collect operation, data belonging to all processes (e.g., processes 402A-402I in phase 406) is collected into a single vector and sent to a single process (e.g., process 402B in phase 404). More specifically, as shown in FIG. 4, in phase 406, each process has data elements that can be collected and arranged into a vector of data elements in process order, with a data element from a first process (e.g., process 402A) providing the first data element of the vector, a data element from a second process (e.g., process 402B) providing the second data element of the vector, and so on. The collected vector can be sent to a target process (e.g., process 402B in phase 404). In some examples, as shown in FIG. 4, collecting can be the inverse of scattering.
[0034] 5 illustrates collective operations 500 corresponding to all-gather and reduce-scatter operations for process 502A, process 502B, process 502C, process 502D, process 502E, process 502F, process 502G, process 502H, and process 502I for phases 504 and 506. Each of processes 502A-502I may correspond to a process (e.g., one or more iterations of process 122A and / or process 122B) executed on one or more nodes (e.g., one or more iterations of computing device 102A and / or computing device 102B, respectively) in an interconnected network (e.g., collective network 100).
[0035] In a collect all operation, data belonging to all processes (e.g., processes 502A-502I in phase 504) is collected (e.g., based on index) and sent to all processes (e.g., processes 502A-502I in phase 506). More specifically, as shown in FIG. 5, each of processes 502A-502I in phase 504 has data elements that may be collected and arranged into a vector of data elements in process order, with a data element from a first process (e.g., process 502A) providing the first data element of the vector, a data element from a second process (e.g., process 502B) providing the second data element of the vector, and so on. The collected vector of data may be sent to each of processes 502A-502I in phase 506. FIG. 5 illustrates each index of the data vector with a different pattern.
[0036] In a reduce-scatter operation, data belonging to (which may be common to) all processes (e.g., processes 502A-502I in phase 506) is parsed out (e.g., based on vector index) to all processes (e.g., processes 502A-502I in phase 504). More specifically, as shown in FIG. 5, the first data element of the common vector may be sent to the first process (e.g., process 502A in phase 504), the second data element may be sent to the second process (e.g., process 502B), and so on. In some examples, as shown in FIG. 5, a reduce-scatter operation may be the inverse of a gather-all operation.
[0037] 6 illustrates collective operations 600 corresponding to the overall reduce operations for process 602A, process 602B, process 602C, process 602D, process 602E, process 602F, process 602G, process 602H, and process 602I for phase 604 and phase 606. Each of processes 602A-602I may correspond to a process (e.g., one or more iterations of process 122A and / or process 122B) executed on one or more nodes (e.g., one or more iterations of computing device 102A and / or computing device 102B, respectively) in an interconnected network (e.g., collective network 100).
[0038] In a full reduce operation, data belonging to all processes (e.g., processes 602A-602B in phase 604) may be collected and operated on (e.g., reduced), and the results may be sent to all processes (e.g., processes 602A-602I in phase 606). Data may be reduced by index, such that reduction occurs for the corresponding index value for each process (e.g., the first data element across all processes is reduced, the second data element across all processes is reduced, etc.). Figure 6 shows each index of the data vector with a different pattern.
[0039] The resulting vector includes a reduction for each index. Examples of reduction operations include maximum (e.g., selecting the largest value), minimum (e.g., selecting the smallest value), sum (e.g., adding all values), intersection (e.g., multiplying all values), logical AND (e.g., Boolean), bitwise AND (e.g., performing a bitwise AND operation), logical OR, bitwise OR, exclusive logical OR, bitwise exclusive OR, the location of the maximum and maximum, and the location of the minimum and minimum. Thus, in phase 606, processes 602A-602I may all have the same vector with a reduction for each index.
[0040] While Figures 3-6 depict source and destination processes to illustrate the start and end phases of an operation, data may travel different routes based on the operation, such that various processes may be interconnected in different ways for different operations. In some embodiments, a full reduce operation requires an interleaved node, such as a switch, to perform the operation. In contrast, a broadcast operation, for example, requires an interleaved node to forward the data but does not require any operation on the data. Thus, the network topology for performing the full reduce operation may affect the outcome of the full reduce operation compared to a broadcast operation.
[0041] 7A-7C illustrate an example network topology of collective operation 700 that may correspond to the overall reduce operation for process 702A, process 702B, process 702C, process 702D, process 702E, process 702F, process 702G, and process 702H. Each of processes 702A-702H may correspond to a process (e.g., one or more iterations of process 122A and / or process 122B) running on one or more nodes (e.g., one or more iterations of computing device 102A and / or computing device 102B, respectively) in an interconnected network (e.g., collective network 100). Figures 7A-7C further illustrate node 716A, node 716B, and node 718, which each correspond to a node (e.g., one or more iterations of computing device 102A and / or computing device 102B), as described herein, that may be used to transmit intermediate data to / from processes 702A-702H.
[0042] For a full reduce operation, various processes (e.g., processes 702A-702H) may hold the data to be reduced (e.g., each a complete vector, as shown in FIG. 7A). Processes 702A-702H may be organized such that half of the processes may be linked to corresponding nodes. More specifically, as shown in FIGS. 7A-7C, processes 702A-702D may be linked to node 716A, and processes 702E-702H may be linked to node 716B. Furthermore, nodes 716A-716B may be linked to node 718 to form a tree topology as shown in FIGS. 7A-7C, although other topologies may be used in other examples, which may further depend on the collective operation being performed.
[0043] 7A-7C show downstream links (e.g., links away from processes 702A-702H) and upstream links (e.g., towards processes 702A-702H, as opposed to the downstream links). Although FIGS. 7A-7C show downstream links that mirror the upstream links (e.g., each downstream link has an inverse of the corresponding upstream link), in other examples, such as for other collective operations, the downstream links and upstream links may be different.
[0044] Figure 7A corresponds to the start of a full reduction operation (e.g., corresponding to phase 604 of Figure 6), with each of processes 702A-702H having its own set / vector of data. Figures 7A-7C show each index of the data vector in a different pattern. As further shown in Figure 7A, each link is idle (shown by a solid line).
[0045] The overall reduction operation involves each process contributing its data, reducing the data across the process's various data sets, and communicating the results to all processes, as described herein. The topologies shown in Figures 7A-7C illustrate examples of how various processes may coordinate data transmission and processing. In Figure 7B, each process transmits its data downstream (using dotted and dashed transmission links) to its corresponding node for processing. For example, processes 702A-702D each transmit their data to node 716A, and processes 702E-702H transmit their data to node 716B.
[0046] Intermediate nodes may perform reductions (represented by Σ in FIG. 7B) on received data sets / vectors that are subsets of the total data. Specifically, in FIG. 7B, node 716A reduces half of the total data received from processes 702A-702D, and node 716B reduces the other half of the total data received from processes 702E-702H.
[0047] After reducing the received data vectors, nodes 716A and 716B each provide their intermediate results for further processing to downstream node 718. Specifically, node 718 reduces two halves of the total data received from nodes 716A and 716B to produce a final result (represented by Σ) corresponding to the reduction made across all data vectors from processes 702A-702H.
[0048] Returning to FIG. 7C, node 718 communicates the results to each of processes 702A-702H via nodes 716A-716B. More specifically, based on the topology shown in FIG. 7C, node 718 transmits the results to upstream nodes 716A and 716B (via outgoing links shown by dotted and dashed lines). Each of nodes 716A and 716B can transmit the results to upstream processes 702A-702H, with node 716A transmitting the results to processes 702A-702D, and node 716B transmitting the results to processes 702E-702H. Thus, each of processes 702A-702H has the final result of the reduction (e.g., corresponding to phase 606 in FIG. 6).
[0049] For the various collective operations described herein, data is passed and operated on interleaved nodes and links between processes. The collective engine (e.g., control circuitry 112), in some embodiments, may establish a network or routing topology, which may be different for each collective operation, and further improve its performance using the systems and methods described herein. As noted above, for all reduce operations, network topology can affect performance. The collective engine provided herein may establish the appropriate network topology and further establish routes between nodes, for example, when initializing a collective network.
[0050] To improve performance, the collective engine may optimize aspects of the network topology. For example, the collective engine may calculate communication costs (e.g., latency associated with transmitting data between nodes) for different network topologies and determine an optimal topology using cost model parameters. FIG. 8 shows a table 800 of exemplary communication cost models for calculating communication cost time T for various topologies, such as binary trees (e.g., FIGS. 7A-7C), recursive duplexing, rings, and in-networks (see FIGS. 9A-9G, which illustrate in-networks with multiple links between layers / levels of nodes, as described further below). The cost models herein may be represented by cost functions with parameters, which may be further implemented as instructions performed by any device / system described herein, such as the collective engine. For the cost model parameters in Figure 8, α is the latency (e.g., latency per GPU network in time, such as seconds), p is the number of GPUs in the collective communication operation, β is the inverse bandwidth (e.g., bandwidth per interconnection network in 1 / bit / second), ε is the up / down ratio (e.g., ratio of downstream ports to upstream ports per network switch, such as downstream / upstream), and n is the vector size (e.g., size in bytes of the vector to be reduced).
[0051] In some embodiments, the collective engine may optimize the topology using cost models for different topology types, such as those in FIG. 8 . For example, optimizing the topology may correspond to optimizing (e.g., minimizing or otherwise reducing) communication costs as calculated from the cost models. In some examples, optimizing costs may include using different cost models for different levels of the topology (e.g., corresponding to different topology types). In some examples, optimizing costs may include optimizing topology structure, such as flattening the tree (e.g., simplifying the tree structure by rearranging leaf nodes and / or child nodes to similar levels as their corresponding parent nodes), rearranging nodes / links, etc. For example, the collective engine may use an optimizer function (e.g., statistical and / or other mathematical analysis of one or more parameters, which may be implemented with instructions, machine learning models, etc.). Thus, the collective engine may optimize communication costs for a collective operation by evaluating multiple communication models (e.g., as shown in FIG. 8 ), selecting a cost model that corresponds to the optimized communication costs (e.g., corresponding to the lowest feasible cost value calculated from the cost models), and configuring the topology for performing the collective operation using the selected cost model.
[0052] In some embodiments, the cost model may be used to construct the topology itself, such as by using values for cost model parameters (e.g., used to optimize communication costs) as topology parameters. For example, (log- εThe cost model parameter of p) may be related to the number of layers (e.g., tree depth) of the in-network topology. The collective engine may then build the topology (e.g., when initializing the collective network, which may include configuring the nodes of the collective network) by establishing appropriate links (e.g., communication connections) between layers (e.g., nodes in adjacent layers) using heuristics according to topology parameters (which may use values of the cost model parameters corresponding to optimized costs), such as the number of downstream and upstream ports, the stride value (e.g., the number of devices connected in a round-robin manner to nodes in the next layer), the number of processors (e.g., processors and / or other devices that may process data), the number of ports per processor (e.g., input / upstream ports and / or output / downstream ports), and the tree depth (e.g., the number of layers), as further described below with respect to the example in-network topologies of FIGS. 9A-9G.
[0053] Although this disclosure describes optimizing a network topology for an overall reduction operation, in other embodiments, the systems and methods described herein may be applied to other collective operations. Additionally, in some embodiments, a network topology may include portions that are optimized using different cost models, such that the overall network topology may include portions of different topologies (e.g., combinations of different topology types).
[0054] The collective engine (e.g., control circuitry 112) may also establish communication routes for the network topology. For example, FIGS. 9A-9G illustrate an example network topology 900 for in-network all-reduction operations, which, in some examples, are based on the network topology optimizations described herein. The collective engine may include a network resource scheduler or aggregation manager that schedules resources in the network capable of performing collective offload operations and / or portions thereof. In some embodiments, the network resource scheduler may reside within a node (e.g., a GPU) and / or in a separate device connected to the network. The network resource scheduler may set up flows that organize and prioritize specific links / switches. In some embodiments, the network resource scheduler may request state (e.g., metadata and / or other information about the node's current state) from each node in the collective network to initialize the collective network for the collective operation. For example, when a resource (e.g., a node) indicates that it is ready (e.g., when initializing a collective network), it sends its status (e.g., as a message) which may include a readiness message accompanied by a job profile (e.g., metadata which may indicate processing power, memory capacity, type of processing / operation preferred for the node, location in the topology, other preferences / limitations, etc.) to the network resource scheduler to indicate its readiness and establish a flow / path according to the job profile. More specifically, the network resource scheduler establishes a topology of data routes for the collective operation, which may include identifying the processing required at each level of the collective network to perform the collective operation, identifying destination nodes at each level that can process the data required for the collection operation (e.g., as further determined from the job profile and / or readiness), and establishing links between them.Additionally, if a route is blocked (e.g., a node is down or otherwise unavailable), the network resource scheduler may establish a different route accordingly, for example, by prioritizing links between levels of available nodes.
[0055] Deadlock, in some instances, refers to a scenario in which a link to a node sends two packets to the node, but the node has only one buffer available, preventing either packet from being received. In the context of collective operations, where various nodes rely on results provided by other nodes and operate in parallel, a deadlock scenario can significantly impact operations and outcomes. A network resource scheduler can establish routes to prevent deadlock.
[0056] In some examples, the network resource scheduler may implement the route. For example, the network resource scheduler may use segment routing, in which route / path information travels with the packet so that nodes can forward packets to their destinations based on the route information. In some embodiments, the packet's header may include the route information. Thus, a given node may be prevented from arbitrarily selecting a free link to transmit a packet or otherwise transmitting a packet on a link that does not match an established route.
[0057] By establishing routes, particular packets of data may be routed to particular nodes as needed for collective operations. Additionally, in some examples, the network resource scheduler may dynamically update or otherwise re-establish the topology, for example, in response to changes in node availability (e.g., in response to changes in node status). Furthermore, routes allow data to be fragmented (e.g., broken down into smaller segments) and routed to appropriate nodes for reduction with other corresponding nodes.
[0058] A network resource scheduler (e.g., which may be part of a collective engine such as control circuitry 112) provides routes that can guarantee data reaches a particular node. Data can be fragmented (e.g., divided) into smaller segments (e.g., chunks corresponding to portions of the original message size, such as subvectors of a vector) to improve bandwidth and performance. In some embodiments, the fragmentation scheme can be based on a network topology, which can be optimized as described herein. The topology can ensure that corresponding chunks (e.g., chunks to be processed together) reach the same destination as needed to perform the collective operation, as explained further below. In one embodiment, the fragmentation scheme can be based on factors similar to those used to optimize the network topology, such as the number of ports. For example, each node that performs a gather operation or portion thereof on data (which in some examples is received via an input port from a node in a previous level of the collective network) can fragment the resulting data into a number of chunks equal to the number of downstream (e.g., output) ports and transmit the chunks through respective ports of the downstream ports (e.g., to a node in a subsequent level of the collective network). In other examples, other fragmentation schemes can be used.
[0059] 9A-9G illustrate example routing (e.g., preventing deadlocks) and subdivision of an example in-network topology 900 for a total reduce operation. FIGS. 9A-9G illustrate various servers (e.g., corresponding to instances of server 202), GPUs (e.g., corresponding to instances of GPU 214A and / or GPU 214B, which may further correspond to processes), switches (e.g., corresponding to instances of switch 218), NICs (e.g., instances of NIC 216A and / or NIC 216B), and other nodes (e.g., corresponding to instances of any node described herein). In-network topology 900 illustrates particular ways in which a topology may be constructed using topology parameters such as the number of downstream ports (e.g., ports for connecting to downstream nodes), the number of upstream ports (e.g., ports for connecting to upstream nodes), and / or a stride value (which, in some examples, may use different values at different levels). The number of downstream ports may correspond to the number of links heading downstream from a node, and the number of upstream ports may correspond to the number of links heading upstream from a node. The stride value may correspond to the number of links that a link in a node at a given level can connect to different nodes at the next level before repeating the node, or in other words, the stride value may correspond to the number of nodes at the next level that a set of nodes at a given level can connect to.
[0060] In-network topology 900 illustrates, for example, a GPU level with two downstream ports per node (the GPU level may correspond to the root level, and therefore no upstream ports). By having two downstream ports, each GPU may be connected to two different switches at the next level (e.g., the switch level) to further match the stride value. At the switch level, each switch has two upstream ports (mirroring the GPU's downstream links) and two downstream ports (connecting to two different NICs). At the NIC level, each NIC has one upstream port (mirroring the switch's downstream links) and one downstream port (e.g., corresponding to forwarding data). In some examples, the interconnected GPU to NIC level may correspond to a single machine (e.g., server_0 and server_1). The NIC downstream ports may be connected to nodes at the next level based on a stride value of 4. As shown in FIGS. 9A-9G, each NIC may be connected to the nodes sequentially in order, and the order may repeat after the stride value is reached. More specifically, nic_0_0_0, nic_0_1_0, nic_0_0_1, and nic_0_1_1 may each be linked to node_0_0_0, node_0_0_1, node_0_0_2, and node_0_0_3 in sequence, with subsequent NICs repeating links to nodes in a round-robin fashion in sequence. Network topology 900 further includes additional nodes (e.g., node_0_1_0, node_0_1_1, node_0_1_2, and node_0_1_3) that represent how the topology may be extended, which are shown with only a single upstream / downstream link for simplicity of illustration.
[0061] As described herein, the collective engine (e.g., control circuitry 112) may establish different topologies for different collective operations. In other words, nodes in in-network topology 900 may have links not shown in FIGS. 9A-9G , and for all reduce operations, the collective engine may implement a route according to in-network topology 900, for example, using segment routing, destination lists (e.g., associated with ports / links), and / or other routing implementations. Additionally, in some embodiments, the topology (e.g., one or more levels) may be further limited by available links and / or component availability (e.g., in-network topology 900 shows each server with two GPUs), which may further relate to which components reside in which devices. Furthermore, in some embodiments, the topology may be based on node availability and / or type (e.g., certain nodes available to process data and certain other nodes available to forward data).
[0062] To better illustrate how data / chunks may be routed through in-network topology 900 and how it may relate to each node, each level is shown with a specific indexing notation. For example, the GPU level may be indexed based on server_gpu, such that the server index and GPU index for that server identify the GPU (e.g., Gpu_0_0 and Gpu_0_1 branching off of server_0, and Gpu_1_0 and Gpu_1_1 branching off of server_1). The switch level may be indexed based on server_switch, such that the server index and switch index for that server identify the switch. The NIC level may be indexed based on server_GPU_NIC, such that the server index, corresponding GPU index, and NIC index for that server / GPU identify the NIC. Nodes at outer levels may be indexed based on pod_layer_node, allowing the pod of servers (e.g., “0” in FIGS. 9A-9G ), the layer within the node (e.g., “0” for the inner layer and “1” for the outer layer), and the node index to identify the node (which may correspond to a switch or any other node, as described herein). Further, exemplary data (e.g., chunks) may be indexed by server_GPU_chunk_subchunk, allowing the server index, GPU index, and chunk / subchunk index to identify the chunk. In some embodiments, the indexing described herein may be used to identify nodes and / or chunks when data is routed through the in-network topology 900, although other indexing and / or identification may be used. Furthermore, for each level, nodes at that level may transmit data through particular ports that have links to particular destinations, as described herein, which may be further implemented via segment routing (e.g., having the destination node and / or port as part of the transmitted data).
[0063] In FIG. 9A (e.g., time 0 for sending / subdividing data), each GPU holds a full message size (e.g., size=1) of data (e.g., message). Thus, each message is indexed similarly to its corresponding originating GPU. In FIG. 9B (e.g., time 0 for sending / subdividing data), each GPU can subdivide its data according to a subdivision scheme, which in some examples corresponds to splitting the data vector into equal-sized chunks (e.g., subvectors from the original vector) for each downstream port. In some examples, the subvectors can be defined based on vector indices, such as the lower half of the vector index and the upper half of the vector index, and odd and even vector indices, to generate appropriately sized subvectors. As described further below, by applying consistent subdivision to the vectors and performing routing, data elements of the original vector can be reduced with corresponding data elements of the same vector index from other data vectors (e.g., data vectors across nodes).
[0064] Thus, each GPU with two downstream ports may fragment its data into two similarly sized chunks (e.g., chunk size = 1 / 2 of message size) to send one chunk per port. In FIG. 9B, each half of the message may be indexed as "0" and "1" as shown (e.g., message 0_0 is fragmented into chunk 0_0_0 and chunk 0_0_1, and message 0_1 is fragmented into chunk 0_1_0 and chunk 0_1_1, etc.). By applying the fragmentation scheme across GPUs, each corresponding chunk half / index may contain the same vector index (e.g., all chunks ending in "0" may correspond to the same vector index as from the original message). At time 0, all links are idle, as shown in FIGS. 9A and 9B.
[0065] In FIG. 9C (e.g., time 1), chunks from the GPU are routed to switches such that the first switch has the first half of the original message, the second switch has the second half of the original message, and so on, with corresponding halves arriving at the same switch (e.g., switch_0_0 has chunks 0_0_0 and 0_1_0, while switch_0_1 has chunks 0_0_1 and 0_1_1, etc.). As shown in FIG. 9C, links corresponding to downstream ports are used once. In some embodiments, the switches may read data from GPU memory, while in other embodiments, data may be passed on as needed. As described further below, the described fragmentation schemes may ensure that corresponding chunks arrive at the same destination to enable reduction (e.g., values for a given vector index in the original message are reduced with other values for the same vector index from other messages).
[0066] 9D (e.g., time 2), the switch reduces the data (e.g., performs a reduce operation using two received chunks) and generates a single chunk-sized result. The result is again subdivided according to a subdivision scheme, e.g., into two similarly sized sub-chunks (e.g., sub-chunk size = 1 / 2 chunk size) corresponding to the two downstream ports. For example, the result for switch_0_0 is split into sub-chunk0_Σ_0_0 and sub-chunk0_Σ_0_1. For purposes of illustrating how chunks are reduced and routed, the subchunk index indicates that Gpu_0_0 and Gpu_0_1 have been reduced for a given chunk index (e.g., the trailing "0" corresponds to chunks 0_0_0 and 0_1_0) (the "0_Σ" in other words corresponds to the same vector index from reduced message 0_0 and message 0_1, and more specifically indicates that data from two GPUs of the server has been reduced), which is now split (e.g., at the end of the subchunk index) into two subchunks "0" and "1." As described herein, the subchunks may be split based on the vector index values. By applying the subdivision scheme across the switch, each corresponding subchunk half / index may contain the same vector index (e.g., all subchunks ending in "0_0" may correspond to the same vector index as from the original message; more specifically, "0_x" corresponds to the same half of the vector index from the original message, and "x_0" corresponds to the same half of the vector index from the first subdivision).
[0067] In Figure 9E (e.g., time 3), the subchunk is sent from the switch to a downstream NIC based on the route, as shown. Because the NIC does not operate on the data, the subchunk traverses the NIC and is forwarded to a node (e.g., a layer 0 node) as shown in Figure 9F (e.g., time 4). Routing ensures that the appropriate corresponding chunk is forwarded to the same node, thereby enabling reduction. For example, by using stride values as described above, nic_0_0_0 may send subchunk 0_Σ_0_0 to node_0_0_0, and nic_1_0_0 may send subchunk 1_Σ_0_0 to node_0_0_0. Thus, node_0_0_0 has the same vector index of the subchunk (as indicated by the common suffix "0_0" as described above), and similarly for other nodes, so that node_0_0_0 can reduce its data. Furthermore, as mentioned above, a leading "0" or "1" in the subchunk index indicates the originating server, allowing node_0_0_0 to reduce the same vector index of subchunks of data from two servers, and similarly for the other nodes. As a result, each of the four nodes has a reduced subchunk of data (e.g., 1 / 4 of the message size) from each of the original GPUs.
[0068] After the Layer 0 nodes reduce the data, in Figure 9G (e.g., time 6), the resulting sub-chunks may be sent deeper into the network, represented by Layer 1 nodes, if desired. Although not shown in Figure 9G, the sub-chunks may be subdivided into smaller chunks and sent deeper into the network, if desired (e.g., with a topology appropriate for further subdivision).
[0069] As shown in FIG. 9G, the subchunks are indexed with a leading "Σ_Σ" to indicate that the subchunks on both servers and both GPUs have been reduced, and trailing "0_0," "0_1," "1_0," and "1_1," corresponding to the vector indexes split into the four subchunks. Thus, these four subchunks, when reassembled into a single message, represent the reduction of the original message (e.g., the result of the reduction). To distribute the results back to the GPUs (e.g., see FIGS. 6 and 7C), the subchunks can be transferred back through the in-network topology 900 via the upload link (e.g., reversing the route and subdivision shown in FIGS. 9A-9G without further processing / reduction).
[0070] Thus, by establishing a topology with a given level, subdividing the data, and distributing the chunks in a round-robin fashion to the next level with a number of nodes corresponding to the number of chunks (e.g., the GPU level subdivides the data into two chunks for transmission to the switch level with two switches, and from the switch level, further subdivides the data into four subchunks for transmission via the NIC level to four nodes at the node level), the topology enables more efficient processing of smaller data sets (e.g., reducing chunks or subchunks at the switch level and node layer 0, respectively). This topology may further improve bandwidth efficiency. For example, the topology itself may be optimized based on a cost model, as described above. Furthermore, the subdivision allows smaller data sets to be communicated, further avoiding having nodes hold results and distributing complete results back to the process.
[0071] While FIGS. 9A-9G illustrate how fragmented data may be sent via routing that may prevent deadlocks, in other examples, data may be sent through network topology 900 via similar routes without fragmenting the data. Additionally, while the examples described herein generally follow data traveling downstream until it reaches a root node and then reversing upstream (e.g., one direction at a time), in other examples, routing may include downstream and upstream travel (e.g., in both directions) as needed. For example, certain nodes may be unavailable (e.g., lack storage and / or operational capabilities) to receive and / or transmit data, such that routing may include returning to layers, skipping layers, and the like. Additionally, FIGS. 9A-9G illustrate a simplified example of a two-server pod. In other examples, the topology building, routing, and / or fragmentation described herein may be for more and / or fewer pods of servers and / or other devices.
[0072] Figure 10 is a flow diagram of an example computer-implemented method 1000 for collective offload routing management of a network. The steps shown in Figure 10 may be performed by any suitable computer-executable code and / or computing system, including the systems shown in Figures 1, 2, 7A-7C, and / or 9A-9G. In one example, each of the steps shown in Figure 10 represents an algorithm whose structure includes and / or is represented by multiple sub-steps, examples of which are provided in more detail below.
[0073] 10, one or more of the systems described herein evaluate multiple communication cost models for a collective operation at step 1002. For example, control circuitry 112 may evaluate multiple communication cost models for a collective operation (e.g., FIG. 8 for an all-reduce operation).
[0074] Systems described herein may perform step 1002 in various ways. In one example, evaluating multiple communication cost models can include optimizing a communication cost of the collective operation, and values for multiple cost model parameters are determined based on the optimized communication costs. As described above, control circuitry 112 may use an optimizer function (e.g., circuitry and / or instructions for evaluating and optimizing cost functions) for the cost models (e.g., having cost functions) to determine values for the cost model parameters for the optimized communication costs. As described further herein, in some examples, the optimization may include using some of the different cost models in a piecewise manner (e.g., using a particular cost model for a particular level / node of the collective network).
[0075] At step 1004, one or more of the systems described herein determine values for a plurality of cost model parameters based on the evaluation. For example, control circuitry 112 may determine values for the cost model parameters based on the evaluation.
[0076] At step 1006, one or more of the systems described herein configure a topology of the collective network to perform the collective operation using a plurality of topology parameters corresponding to values for a plurality of cost model parameters. For example, control circuitry 112 may configure nodes according to the topology using the topology parameters determined from the cost model parameters.
[0077] The systems described herein may perform step 1006 in various ways. In one example, values of cost model parameters may be used as values of corresponding topology parameters to construct a topology. For example, values for cost model parameters such as the number of upstream ports, the number of downstream ports, the number of processors (e.g., nodes with job profiles matching processing needs), the number of ports per processor, the tree depth, and / or a stride value that produces optimized communication costs may be used as values for topology parameters, as appropriate (e.g., constructing a topology using similar values). Figures 9A-9G show an example topology using topology parameters as described above.
[0078] In some examples, configuring the topology includes configuring communication connections between nodes at a level of the collective network and nodes at an adjacent level of the collective network based on a plurality of topology parameters. For example, control circuitry 112 may establish links (e.g., upstream and downstream links on respective upstream and downstream ports) between nodes at a GPU level with a switch level, a switch level with a NIC level, and similar levels in FIGS. 9A-9G.
[0079] In some examples, control circuitry 112 may configure one portion of the topology based on multiple topology parameters and configure a second portion of the topology based on a second communication cost model. For example, Figures 9A-9G illustrate how certain levels may be configured differently (e.g., the NIC level, which may be further based on port availability, and / or the node level, which may be configured as needed based on node availability and / or collective operation processing needs). Additionally, control circuitry 112 may further configure and / or reconfigure the topology (e.g., dynamically) based on any other factors described herein for establishing the topology.
[0080] Figure 11 is a flow diagram of an example computer-implemented method 1100 for collective offload message granularity management in a network. The operations illustrated in Figure 11 may be performed by any suitable computer-executable code and / or computing system, including the systems illustrated in Figures 1, 2, 7A-7C, and / or 9A-9G. In one example, each of the operations illustrated in Figure 11 represents an algorithm whose structure includes and / or is represented by multiple sub-operations, examples of which are provided in more detail below.
[0081] 11, one or more of the systems described herein receive multiple data sets at step 1102. For example, a node such as computing device 102A or computing device 102B (e.g., corresponding switch_0_0 in FIG. 9C or any other node / process described herein) may receive multiple data sets (e.g., message chunks).
[0082] The systems described herein may perform step 1102 in various ways. In one example, the number of the plurality of datasets corresponds to the number of input ports according to the subdivision scheme. For example, in FIG. 9C , switch_0_0 may receive a first chunk (0_0_0) from a first input port (connected to Gpu_0_0) and a second chunk (0_1_0) from a second input port (connected to Gpu_0_1).
[0083] At step 1104, one or more of the systems described herein perform a collective network collective operation on the multiple data sets to generate a result data set. For example, switch_0_0 may perform a collective operation on the received chunks (e.g., see FIG. 9D).
[0084] At step 1106, one or more of the systems described herein divide the result data set into chunks. For example, in Figure 9D, switch_0_0 may divide the result into sub-chunks.
[0085] The systems described herein may perform step 1106 in various ways. In one example, the number of datasets corresponds to the number of input ports, and the number of chunks corresponds to the number of output ports, according to a subdivision scheme. For example, in FIG. 9D, switch_0_0, which has two downstream ports, may divide the results into two sub-chunks, as further described above.
[0086] At step 1108, one or more of the systems described herein transmit multiple chunks to multiple destinations based on a topology corresponding to a collective operation that defines paths for the chunks across the collective network so that corresponding chunks reach the same destination. For example, in FIG. 9E , switch_0_0 may transmit subchunk 0_Σ_0_0 to Nic_0_0_0 and subchunk 0_Σ_0_1 to Nic_0_1_0 according to the topology.
[0087] Systems described herein may perform step 1108 in various ways. In one example, the topology defines connections between nodes of the collective network based on collective operations, such that multiple data sets are received from a previous level of the collective network and multiple chunks are transmitted to a subsequent level of the collective network, which in some embodiments may be established by a collective engine (e.g., control circuitry 112) as described herein. Additionally, control circuitry 112 may further configure and / or reconfigure the topology (e.g., dynamically) based on any other factors described herein for establishing the topology. In some examples, the path for chunks traversing the collective network ensures that chunks of matching vector indices reach the same destination for the collective operation (e.g., see Figures 9F-9G, as further described above).
[0088] Figure 12 is a flow diagram of an example computer-implemented method 1200 for collective offload routing and deadlock prevention management of a network. The operations shown in Figure 12 may be performed by any suitable computer-executable code and / or computing system, including the system(s) shown in Figures 1, 2, 7A-7C, and / or 9A-9G. In one example, each of the operations shown in Figure 12 represents an algorithm whose structure includes and / or is represented by multiple sub-operations, examples of which are provided in more detail below.
[0089] 12, one or more of the systems described herein receive a job profile and a readiness status for initializing the collective network for collective operation from each of a plurality of nodes in the collective network at step 1202. For example, control circuitry 112 may receive the job profile and / or the readiness status from a node (e.g., one or more iterations of computing devices 102A-102B) in the collective network.
[0090] Systems described herein may perform step 1202 in various ways. In one example, control circuitry 112 may request a state (e.g., including a job profile and / or a readiness status) from each node in the collective network as part of initializing the collective network, and more specifically, initializing the collective network for a collective operation (e.g., the reduce all operation of FIGS. 9A-9G). In some examples, the job profile corresponds to the processing power of the node or the memory capacity of the node. In some examples, the readiness status indicates whether the node is ready to receive, process, and / or transmit data.
[0091] At step 1204, one or more of the systems described herein select, for each level of the collective network, nodes that can process the collective operation as needed for that level based on the job profile and readiness. For example, control circuitry 112 may select appropriate nodes for each level of the collective network that can process the collective operation (and / or portions thereof) as needed.
[0092] The systems described herein may perform step 1204 in a variety of ways. In one example, control circuitry 112 may determine what processing is required at each approximation level to determine processing needs based on collective operations and select nodes (e.g., based on job profile and / or readiness) that can perform processing, such as reducing chunks at the switch level and reducing subchunks at the node level in Figures 9A-9G.
[0093] At step 1206, one or more of the systems described herein determine a topology of data routes for the collective operation based on the nodes selected for each level. For example, control circuitry 112 may determine a topology including the data paths (e.g., based on the links between the nodes).
[0094] Systems described herein may perform step 1206 in various ways. In one example, determining the topology includes determining links to selected nodes at a level to nodes at a previous level and to nodes at a subsequent level. For example, control circuitry 112 may determine that several levels of reduction are necessary for the overall reduction operation (e.g., at the switch level and node level in FIGS. 9A-9G) and further determine intermediate levels (e.g., the NIC level for forwarding data). In some embodiments, control circuitry 112 may further use topology parameters to determine links between ports / nodes (e.g., using a stride value to connect the NIC level to the node level, as described above). Additionally, control circuitry 112 may further configure and / or reconfigure the topology (e.g., dynamically) based on any other factors described herein for establishing the topology.
[0095] At step 1208, one or more of the systems described herein configure the plurality of nodes based on the determined topology. For example, the control circuitry 112 configures the iterations of the computing devices 102A-102B based on the topology (see, e.g., FIGS. 9A-9G).
[0096] Systems described herein may perform step 1208 in various ways. In one example, configuring the plurality of nodes includes establishing links between selected nodes at each level based on the determined topology. For example, control circuitry 112 may instruct the nodes to establish links to appropriate nodes at adjacent levels according to the topology. In some examples, links may be dynamically established when performing respective collective operations. For example, using segment routing, each node may forward data to the appropriate node via the appropriate port, allowing each node to perform different collective operations by processing / forwarding the data according to the topology. In some examples, control circuitry 112 may re-establish and / or dynamically change the topology (e.g., links) in response to changes in the collective network, such as changes in readiness and / or job profile (e.g., processing power and / or memory capacity).
[0097] Additionally, while Figures 10, 11, and 12 depict various steps, in some examples, one or more of the steps depicted in Figures 10, 11, and / or 12 may be combined, intermingled, and / or repeated. For example, certain steps and / or aspects of method 1000 may be incorporated with certain steps and / or aspects of method 1200 for configuring / establishing a topology, which may further be incorporated with certain steps and / or aspects of method 1100 for segmenting and transmitting data.
[0098] As noted above, the computing devices and systems described and / or illustrated herein broadly represent any type or form of computing device or system capable of executing computer-readable instructions, such as those contained within the modules described herein. In their most basic configurations, these computing devices each include at least one memory device and at least one physical processor.
[0099] In some examples, the term "memory device" generally refers to any type or form of volatile or non-volatile storage device or medium capable of storing data and / or computer-readable instructions. In one example, a memory device stores, loads, and / or maintains one or more of the modules and / or circuits described herein. Examples of memory devices include, but are not limited to, random access memory (RAM), read-only memory (ROM), flash memory, hard disk drives (HDDs), solid-state drives (SSDs), optical disk drives, caches, variations or combinations of one or more of these, or any other suitable storage memory.
[0100] In some examples, the term "physical processor" generally refers to any type or form of hardware-implemented processing unit capable of interpreting and / or executing computer-readable instructions. In one example, a physical processor accesses and / or modifies one or more modules stored in the memory devices described above. Examples of physical processors include, but are not limited to, a microprocessor, a microcontroller, a central processing unit (CPU), a field programmable gate array (FPGA) implementing a soft-core processor, an application-specific integrated circuit (ASIC), a system-on-chip (SoC), a digital signal processor (DSP), a neural network engine (NNE), an accelerator, a graphics processing unit (GPU), one or more portions thereof, one or more variations or combinations thereof, or any other suitable physical processor.
[0101] Although shown as separate elements, the modules described and / or illustrated herein may represent portions of a single module or application. Additionally, in certain embodiments, one or more of these modules may represent one or more software applications or programs that, when executed by a computing device, may cause the computing device to perform one or more tasks. For example, one or more of the modules described and / or illustrated herein represent modules stored and configured to operate on one or more of the computing devices or systems described and / or illustrated herein. In some embodiments, a module may be implemented as a circuit or circuitry. Also, one or more of these modules may represent all or part of one or more special-purpose computers configured to perform one or more tasks.
[0102] Additionally, one or more of the modules described herein may transform data, physical devices, and / or representations of physical devices from one form to another. For example, one or more of the modules described herein may receive one or more data vectors to be transformed, transform the data, output the result of the transformation and send it to other modules, perform a collective operation using the result of the transformation, and store the result of the transformation to perform a collective operation. Additionally or alternatively, one or more of the modules listed herein may execute on a computing device, store data on a computing device, and / or otherwise interact with a computing device to transform a processor, volatile memory, non-volatile memory, and / or any other portion of a physical computing device from one form to another.
[0103] In some embodiments, the term "computer-readable medium" generally refers to any form of device, carrier, or medium capable of storing or carrying computer-readable instructions. Examples of computer-readable media include, but are not limited to, transmission-type media such as carrier waves, and non-transitory-type media such as magnetic storage media (e.g., hard disk drives, tape drives, and floppy disks), optical storage media (e.g., compact disks (CDs), digital video disks (DVDs), BLU-RAY disks), electronic storage media (e.g., solid-state drives and flash media), and other distribution systems.
[0104] The process parameters and order of steps described and / or illustrated herein are given by way of example only and can be changed as desired. For example, although the steps illustrated and / or described herein are illustrated or described in a particular order, these steps do not necessarily have to be performed in the order illustrated or described. The various exemplary methods described and / or illustrated herein may omit one or more of the steps described or illustrated herein or may include additional steps in addition to those disclosed.
[0105] The foregoing description is provided to enable those skilled in the art to best utilize various aspects of the exemplary embodiments disclosed herein. This exemplary description is not intended to be exhaustive or to be limited to any precise form disclosed. Many changes and modifications are possible without departing from the spirit and scope of the present disclosure. The embodiments disclosed herein are to be considered in all respects as illustrative and not restrictive. In determining the scope of the present disclosure, reference should be made to the appended claims and their equivalents.
[0106] Unless otherwise specified, the terms "connected to" and "coupled to" (and their derivatives) as used in this specification and claims should be interpreted as allowing both direct and indirect connections (i.e., via other elements or components). Additionally, the terms "a" or "an" as used in this specification and claims should be interpreted as meaning "at least one of." Finally, for ease of use, the terms "including" and "having" (and their derivatives) as used in this specification and claims are interchangeable with the term "comprising," and have the same meaning.
Claims
1. A device, A control circuit is provided, The control circuit selecting a communication cost model for the collective operation; configuring a topology of a collective network for performing the collective operation based on the selected communication cost model; configured to: device.
2. The control circuit optimizing the communication cost of the collective operation by evaluating multiple communication cost models of the collective operation; selecting a cost model corresponding to an optimized communication cost from the plurality of communication cost models; and selecting the communication cost model by The device of claim 1.
3. the communication cost model includes a plurality of parameters; optimizing the communication cost includes determining an optimized parameter for the plurality of parameters. The device of claim 2.
4. and configuring the topology based on using the optimized parameters as topology parameters. The device of claim 3.
5. The plurality of parameters are: The number of upstream ports, The number of downstream ports, The number of processors, The number of ports per processor and The tree depth and The stride value, at least one of The device of claim 3.
6. configuring the topology includes configuring communication connections between nodes at a level of the collective network and nodes at an adjacent level of the collective network based at least on an optimized stride value corresponding to a number of nodes to be connected at the next level; The device of claim 5.
7. optimizing the communication costs includes flattening a tree associated with the cost model; The device of claim 2.
8. the control circuitry is configured to configure a portion of the topology based on the selected communication cost model. The device of claim 1.
9. the control circuitry is configured to configure a second portion of the topology based on a second communication cost model. The device of claim 8.
10. 1. A system comprising: Memory and a processor; a control circuit; The control circuit Evaluating cost parameters of a communication cost model of the collective operation; configuring a topology of a collective network for performing the collective operation based on the cost parameters; configured to: system.
11. the control circuitry is configured to evaluate the cost parameters of the communication cost model by optimizing a communication cost of the collective operation. The system of claim 10.
12. optimizing the communication cost includes determining optimized parameter values for the cost parameters; constructing the topology is based on using the optimized parameter values as topology parameters. The system of claim 11.
13. The cost parameters are: The number of upstream ports, The number of downstream ports, The number of processors, The number of ports per processor and The tree depth and The stride value, at least one of The system of claim 12.
14. configuring the topology includes configuring communication connections between nodes at a level of the collective network and nodes at an adjacent level of the collective network based at least on an optimized stride value corresponding to a number of nodes to be connected at the next level; The system of claim 13.
15. optimizing the communication costs includes flattening a tree associated with the cost model; The system of claim 11.
16. the control circuitry is configured to configure a portion of the topology based on the evaluated communication cost model and to configure a second portion of the topology based on a second communication cost model. The system of claim 10.
17. 1. A method comprising: Evaluating multiple communication cost models of collective behavior; determining values for a plurality of cost model parameters based on the evaluation; and configuring a topology of a collective network for performing the collective operation using a plurality of topology parameters corresponding to values for the plurality of cost model parameters. method.
18. evaluating the plurality of communication cost models includes optimizing a communication cost of the collective action; values for the plurality of cost model parameters are determined based on an optimized communication cost; 18. The method of claim 17.
19. configuring the topology includes configuring communication connections between nodes at a level of the collective network and nodes at an adjacent level of the collective network based on the plurality of topology parameters.
18. The method of claim 17.
20. constructing a portion of the topology based on the plurality of topology parameters; configuring a second portion of the topology based on a second communication cost model.
18. The method of claim 17.