Network collective offload cost management

Through collective engine optimization of communication cost model and configuration of network topology, and adopting block and segmented routing strategies, the scalability problem in collective communication is solved, achieving more efficient communication and lower costs.

CN120303916APending Publication Date: 2025-07-11ADVANCED MICRO DEVICES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380083288.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-14
Filing Date
2023-12-14
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Collective communications face scalability problems as the number of nodes increases in the system, resulting in an increase in communication costs and affecting system performance.

Method used

Optimize the communication cost model through collective engines, configure network topology to reduce communication costs, adopt block and segment routing strategies, prevent deadlocks, and dynamically adjust network topology to improve efficiency.

Benefits of technology

Optimize the communication efficiency of collective networks, reduce communication costs, and improve the scalability and performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303916A_ABST
    Figure CN120303916A_ABST
Patent Text Reader

Abstract

The disclosed apparatus includes a collective engine that may select a communication cost model for a collective operation from a plurality of communication cost models and configure a topology of a collective network for performing the collective operation using the selected communication cost model. Various other methods, systems, and computer readable media are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Various improvements to computing performance, such as increasing the number of processing cores, have improved performance but may reach scalability limits. Collective communication allows global communication operations among all processes / nodes (including networked nodes) in a system (e.g., a collective network). As the number of nodes increases, collective communication may encounter scalability issues. To ensure better scalability, certain communication processing can be offloaded from nodes (e.g., their processors) to other nodes of the collective network (e.g., network adapters, switches, etc.), which can be managed by a collective engine that may reside in a server or other connected computing device. Brief Description of the Drawings

[0002] The drawings illustrate multiple exemplary embodiments and are part of the specification. Together with the following description, these drawings show and explain the various principles of the present disclosure.

[0003] Figure 1 is a block diagram of an exemplary system of a collective engine.

[0004] Figure 2 is a block diagram of an exemplary collective network.

[0005] Figure 3 is a simplified diagram of a broadcast operation and a reduction operation.

[0006] Figure 4 is a simplified diagram of a scatter operation and a gather operation.

[0007] Figure 5 is a simplified diagram of an all-gather operation and a reduce-scatter operation.

[0008] Figure 6 is a simplified diagram of an all-reduce operation.

[0009] Figure 7A is a network topology diagram of an initial stage of an all-reduce operation.

[0010] Figure 7B is a network topology diagram of an intermediate stage of an all-reduce operation.

[0011] Figure 7C is a network topology diagram of a final stage of an all-reduce operation.

[0012] Figure 8 is a table of a network topology cost model for an all-reduce operation.

[0013] Figures 9A to 9G illustrates the network topology when a block is routed through.

[0014] Figure 10 is a flowchart of an exemplary method for network collective offload cost management.

[0015] Figure 11 is a flowchart of an exemplary method for network collective offloading routing management.

[0016] Figure 12 is a flowchart of an exemplary method for network collective offloading message chunk management.

[0017] In all the figures, the same reference numerals and descriptions indicate similar but not necessarily identical elements. While the exemplary embodiments described herein are susceptible to various modifications and alternative forms, specific embodiments have been shown by way of example in the figures and will be described in detail herein. However, the exemplary embodiments described herein are not intended to be limited to the particular forms disclosed. Rather, this disclosure covers all modifications, equivalents, and alternatives falling within the scope of the appended claims. Detailed Description

[0018] This disclosure generally relates to network collective offloading management. As will be explained in more detail below, embodiments of this disclosure provide a collective engine that can manage various aspects of collective offloading onto networked nodes that communicate with each other to share data and / or data processing.

[0019] This disclosure generally relates to network collective offloading cost management. As will be explained in more detail below, embodiments of this disclosure configure the topology of a collective network for performing collective operations based on a communication cost model that can be evaluated for communication costs (e.g., with respect to the latency of transmitting data) and / or otherwise optimized, such as reducing and / or minimizing the communication costs between the nodes of the collective network when performing collective operations. The systems and methods provided herein can improve the efficiency of the collective network, for example, by establishing a more efficient topology that reduces network communication costs.

[0020] In one embodiment, a device for network collective offloading cost management includes control circuitry configured to: select a communication cost model for a collective operation and configure the topology of a collective network for performing the collective operation based on the selected communication cost model.

[0021] In some examples, the control circuitry is configured to select the communication cost model by: optimizing the communication cost of the collective operation by evaluating multiple communication cost models for the collective operation and selecting the cost model corresponding to the optimized communication cost from the multiple communication cost models. In some examples, the communication cost model includes multiple parameters, and optimizing the communication cost includes determining optimized parameters for the multiple parameters. In some examples, configuring the topology is based on using the optimized parameters as topology parameters.

[0022] In some examples, the plurality of parameters includes at least one of the number of upstream ports, the number of downstream ports, the number of processors, the number of ports per processor, the tree depth, and the step value. In some examples, configuring the topology includes configuring a communication connection between nodes of a level of the collective network and nodes of an adjacent level of the collective network based at least on an optimized step value corresponding to the number of nodes connected in a next level.

[0023] In some examples, optimizing the communication cost includes flattening a tree associated with the cost model. In some examples, the control circuit is configured to configure a portion of the topology based on a selected communication cost model. In some examples, the control circuit is configured to configure a second portion of the topology based on a second communication cost model.

[0024] In one particular implementation, a system for network collective offload cost management includes a memory, a processor, and a control circuit configured to evaluate cost parameters of a communication cost model for a collective operation and configure a topology of a collective network for performing the collective operation based on the cost parameters.

[0025] In some examples, the control circuit is configured to evaluate the cost parameters of the communication cost model by optimizing the communication cost of the collective operation. In some examples, optimizing the communication cost includes determining an optimized parameter value for the cost parameter and configuring the topology based on using the optimized parameter value as a topology parameter. In some examples, the cost parameter includes at least one of the number of upstream ports, the number of downstream ports, the number of processors, the number of ports per processor, the tree depth, and the step value.

[0026] In some examples, configuring the topology includes configuring a communication connection between nodes of a level of the collective network and nodes of an adjacent level of the collective network based at least on an optimized step value corresponding to the number of nodes connected in a next level. In some examples, optimizing the communication cost includes flattening a tree associated with the cost model.

[0027] In some examples, the control circuit is further configured to configure a portion of the topology based on the evaluated communication cost model and configure a second portion of the topology based on a second communication cost model.

[0028] In a specific implementation, a method for network collective offloading cost management, the method includes: (i) evaluating a plurality of communication cost models for collective operations, (ii) determining values of a plurality of cost model parameters based on the evaluation, and (iii) configuring a topology of a collective network for performing the collective operation using a plurality of topology parameters corresponding to the values of the plurality of cost model parameters.

[0029] In some examples, evaluating the plurality of communication cost models includes optimizing the communication cost of the collective operation, and the values of the plurality of cost model parameters are determined based on the optimized communication cost. In some examples, configuring the topology includes: configuring communication connections between nodes of a level of the collective network and nodes of an adjacent level of the collective network based on the plurality of topology parameters. In some examples, the method includes: configuring a part of the topology based on the plurality of topology parameters and configuring a second part of the topology based on a second communication cost model.

[0030] According to the general principles described herein, the features of any specific implementation described herein can be used in combination with each other. After reading the following detailed description in conjunction with the drawings and the claims, these and other specific implementations, features, and advantages will be more fully understood.

[0031] Reference Figures 1 to 10 , a detailed description of network collective offloading management will be provided below. It will be described in conjunction with Figure 1 and Figure 2 A detailed description of example systems and networks for collective operations will be provided. It will be described in conjunction with Figure 3 , Figure 4 , Figure 5 and Figure 6 Example diagrams of collective operations will be provided. It will be described in conjunction with Figures 7A to 7C An example diagram of collective network routing for full reduction operations will be provided. It will be described in conjunction with Figure 8 Example costs associated with collective network node routing for full reduction operations will be provided. It will be described in conjunction with Figures 9A to 9G Example data routing for merge message chunking and deadlock prevention for full reduction operations will be provided. It will also be described in conjunction with Figure 10 , Figure 11 and Figure 12 A detailed description of the corresponding computer-implemented method will be provided.

[0032] Figure 1Exemplary collective network 100 (e.g., an interconnection network) implementing aspects of the present disclosure is illustrated. Collective network 100 includes computing device 102A, network 104, and computing device 102B. Computing device 102A and / or computing device 102B may each be a network device (such as a network switch, router, etc.) and / or a client device or user device (such as a desktop computer, laptop computer, tablet device, smart phone) or other computing device (such as a server) and / or a part thereof. Computing device 102A includes physical processor 110A, and computing device 102B includes physical processor 110B. Physical processor 110A and / or physical processor 110B may correspond to one or more instances of any type or form of hardware-implemented processing unit capable of interpreting and / or executing computer-readable instructions, including but not limited to: die (e.g., smaller and in some examples more specialized processing units that can be coordinated as a single chip), microprocessors, microcontrollers, central processing units (CPUs), graphics processing units (GPUs), field-programmable gate arrays (FPGAs) implementing soft-core processors, application-specific integrated circuits (ASICs), systems-on-a-chip (SoCs), coprocessors such as digital signal processors (DSPs), neural network engines (NNEs), accelerators, graphics processing units (GPUs), parts of one or more of the above, variations or combinations of one or more of the above, and / or any other suitable physical processor. In some particular implementations, the physical processors referred to herein (e.g., physical processor 110A and / or physical processor 110B) may correspond to a host processor and a coprocessor, which in some examples may be separate processors.

[0033] In some examples, processor 110A and / or processor 110B access and / or modify data and / or instructions stored in memory 120A and / or memory 120B. Memory 120A and / or memory 120B each correspond to an instance of any type or form of volatile or non-volatile memory device or medium capable of storing data and / or computer-readable instructions, including but not limited to: random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid state drive (SSD), optical disc drive, cache, variations or combinations of one or more of the above, and / or any other suitable memory. In some particular implementations, memory 120A and / or memory 120B may correspond to the internal memory of processor 110A and / or processor 110B, respectively.

[0034] In some specific implementations, the memory 120A may store one or more processes such as process 122A, and the memory 120B may store one or more processes such as process 122B. Process 122A and / or process 122B may represent one or more software applications, programs, and / or processes that, when executed by a computing device, cause the computing device to perform one or more tasks, such as portions of the collective operations described herein. For example, and as will be described in more detail below, one or more of processes 122A and / or 122B may represent instructions that are stored and configured to run on one or more computing devices (such as Figure 1 the devices illustrated in (e.g., computing device 102A and / or computing device 102B)) and that correspond to (e.g., are associated with the collective operation and / or portions thereof). In some examples, each process (e.g., an instance of process 122A and / or process 122B) may correspond to a localized version of the collective operation (e.g., instructions for a node to perform its portion of the collective operation).

[0035] Computing device 102A is communicatively coupled to computing device 102B via network 104. Network 104 represents any type or form of communication network such as the Internet and may include one or more physical connections (such as a LAN) and / or wireless connections (such as a WAN).

[0036] As Figure 1 further shown, computing device 102A includes control circuit 112. In some specific implementations, control circuit 112 corresponds to and / or is an implementation of a collective engine for offloading collective communication (e.g., collective operations) from nodes in a collective system corresponding to collective network 100 (e.g., locally from one or more physical processors 110A and / or process 122A, and / or remotely from one or more computing devices 102B and one or more physical processors 110B and / or process 122B). Although Figure 1 computing device 102A and computing device 102B are illustrated, in other examples, collective network 100 may include additional computing devices that computing device 102A may manage in a collective operation. Additionally, in some specific implementations, one or more instances of computing device 102A and / or computing device 102B may correspond to devices that are part of a larger computing device (e.g., a CPU, a GPU, a network interface card (NIC), etc.).

[0037] The control circuit 112 may correspond to circuitry and / or instructions for managing the collective communication / operation of nodes such as components of computing device 102A and / or computing device 102B and other instances thereof. In some examples, a node may correspond to an application (e.g., software that may use / process data involved in collective operations such as represented by process 122A and / or process 122B) and / or a device (e.g., a computing device or networking device such as a processor, network interface, switch, etc., that may locally perform portions of the collective operation by receiving, processing data as needed, and forwarding the data to appropriate nodes and / or shared memory devices or addressable memory devices that in some examples cannot process data on their own) and / or a portion of a device (e.g., different processing components / circuitry within a device). In some implementations, the control circuit 112 may coordinate or otherwise manage which nodes perform which collective operations and / or portions thereof (e.g., coordinate which nodes execute which instructions required for a collective operation). In some implementations, the control circuit 112 may further manage the network topology and / or other routing aspects of the collective network.

[0038] A collective engine (e.g., control circuit 112) can establish a network topology for a collective network, which, in some examples, refers to a connection pattern of nodes in an interconnected network and is typically represented by a graph. In some examples, a node or vertex refers to a machine (e.g., a server, a switch, and / or a router, which may correspond to computing device 102A and / or computing device 102B in some specific implementations) or a device (e.g., a central processing unit (CPU) and / or its core, a graphics processing unit (GPU) and / or its core, a network interface or network interface card (NIC), a switch, etc., which may correspond to computing device 102A and / or computing device 102B and / or parts thereof) in the network topology, and the machine or device may have various physical ports and / or virtual ports for communicating with other nodes in the collective network (e.g., corresponding to interfaces for wired communication and / or wireless communication). In some examples, a server refers to a machine that includes devices such as a processor and a network interface card. In some examples, a switch refers to a machine or device that can receive packets from any number of input ports and send the packets to any number of output ports, and in some examples, the machine or device may further be capable of processing data. In some examples, a router refers to a machine or device that is capable of forwarding packets between networks (e.g., via physical ports). In some examples, a network interface refers to a device used by a machine to connect to a network and further capable of forwarding packets. In some examples, an edge refers to a connection between one or more vertices / nodes. In an interconnected network, a single edge may include multiple links, and each link may have multiple channels. At the root of the topology, multiple processes may run concurrently (e.g., via a machine / device).

[0039] Figure 2System 200 corresponding to collective network 100 is illustrated with a simplified example collective network. System 200 includes a server 202 (corresponding to an instance of computing device 102A and / or computing device 102B), a GPU 214A (corresponding to an instance of computing device 102A and / or computing device 102B), a GPU 214B (corresponding to an instance of computing device 102A and / or computing device 102B), a NIC 216A (corresponding to an instance of computing device 102A and / or computing device 102B), a NIC 216B (corresponding to an instance of computing device 102A and / or computing device 102B), and a switch 218 (corresponding to an instance of computing device 102A and / or computing device 102B). A collective engine (e.g., control circuit 112) can establish a topology (e.g., Figure 2 the ring in) for the nodes (e.g., server 202, GPU 214A, GPU 214B, NIC 216A, NIC 216B, and switch 218) by establishing directed links (which can be bidirectional in some examples) between the nodes. In other words, the collective engine can establish which nodes communicate with which other nodes.

[0040] The collective engine can further assist in coordinating the processes running on the respective nodes when performing collective communication operations, which represent distributed operations (e.g., as instructions) performed on data. Examples of collective operations include broadcast, reduce (to one), scatter, gather, allgather, reduce scatter, and allreduce, which will be further described below.

[0041] Figure 3 Collective operation 300 corresponding to broadcast operations and / or reduce operations for processes 302A, 302B, 302C, 302D, 302E, 302F, 302G, 302H, and 302I for phase 304 and phase 306 is illustrated. Each of processes 302A through 302I can correspond to a process (e.g., one or more iterations of process 122A and / or process 122B) running on one or more nodes (e.g., one or more iterations of computing device 102A and / or computing device 102B, respectively) in an interconnection network (e.g., collective network 100).

[0042] In a broadcast operation, data belonging to one process (e.g., process 302B) is transmitted to all processes (e.g., processes 302A, 302C, 302D, 302E, 302F, 302G, 302H, 302I). More specifically, the broadcast operation may start at stage 304, during which process 302B has a data set, which may be a complete vector or a message, although in other examples the data set may correspond to one or more data elements. As a result of the broadcast operation, at stage 306, the data set (e.g., the same vector) is broadcast to all other processes such that all processes 302A through 302I have the same data set. Figure 3 Each index of the data vector with different patterns is shown.

[0043] In a reduction (to one) operation, data belonging to all processes (e.g., processes 302A through 302I) (e.g., a complete vector / message that may be common to all processes) may be reduced to a single set / vector and transmitted to a single process (e.g., process 302B). More specifically, as Figure 3 illustrated, the reduction operation may start at stage 306, during which the data held by processes 302A through 302I may be reduced to a single vector and transmitted to a single target process at stage 304. In some examples, reduction may be the inverse operation of broadcast, as Figure 3 illustrated.

[0044] Figure 4 Illustrated is a collective operation 400 corresponding to scatter and gather operations for processes 402A, 402B, 402C, 402D, 402E, 402F, 402G, 402H, and 402I for stage 404 and stage 406. Each of processes 402A through 402I may correspond to a process (e.g., one or more iterations of process 122A and / or process 122B) running on one or more nodes (e.g., one or more iterations of computing device 102A and / or computing device 102B) in an interconnection network (e.g., collective network 100).

[0045] In a scatter operation, data belonging to one process (e.g., process 402B at stage 404) (e.g., a complete vector) is parsed (e.g., by index) to all processes (e.g., processes 402A through 402I at stage 406). More specifically, as Figure 4As illustrated, each process can be indexed such that at stage 406, each of processes 402A through 402I receives data corresponding to the index, such as the first process (e.g., process 402A) receiving the first data element of the vector, the second process (e.g., process 402B) receiving the second data element of the vector, and so on. Figure 4 Shows each index of a data vector having different patterns.

[0046] In a gather operation, data belonging to all processes (e.g., processes 402A through 402I at stage 406) is collected into a single vector and transmitted to a single process (e.g., process 402B at stage 404). More specifically, as Figure 4 illustrated, at stage 406, each process has data elements that can be collected and arranged in a vector of data elements in process order such that the data elements from the first process (e.g., process 402A) provide the first data element of the vector, the data elements from the second process (e.g., process 402B) provide the second data element of the vector, and so on. The collected vector can be transmitted to a destination process (e.g., process 402B at stage 404). In some examples, gather can be the inverse operation of scatter, as Figure 4 illustrated.

[0047] Figure 5 Illustrates a collective operation 500 corresponding to all-gather operations and reduce-scatter operations for processes 502A, 502B, 502C, 502D, 502E, 502F, 502G, 502H, and 502I for stages 504 and 506. Each of processes 502A through 502I can correspond to a process (e.g., one or more iterations of process 122A and / or process 122B) running on one or more nodes (e.g., one or more iterations of computing device 102A and / or computing device 102B, respectively) in an interconnect network (e.g., collective network 100).

[0048] In an all-gather operation, data belonging to all processes (e.g., processes 502A through 502I at stage 504) is collected (e.g., based on an index) and transmitted to all processes (e.g., processes 502A through 502I at stage 506). More specifically, as Figure 5As illustrated, at stage 504, each of processes 502A through 502I has data elements that can be collected in process order and arranged in a vector of data elements such that the data elements from the first process (e.g., process 502A) provide the first data element of the vector, the data elements from the second process (e.g., process 502B) provide the second data element of the vector, and so on. At stage 506, the collected data vector can be transmitted to each of processes 502A through 502I. Figure 5 Shows each index of a data vector with different patterns.

[0049] In a reduction scatter operation, data that belongs to all processes (e.g., processes 502A through 502I at stage 506), which can be data common to all processes, is parsed (e.g., based on vector indices) to all processes (e.g., processes 502A through 502I at stage 504). More specifically, as Figure 5 illustrated, the first data element of a common vector can be transmitted to the first process (e.g., process 502A at stage 504), the second data element can be transmitted to the second process (e.g., process 502B), and so on. In some examples, the reduction scatter operation can be the inverse of a all-gather operation, as Figure 5 illustrated.

[0050] Figure 6 Illustrates a collective operation 600 corresponding to an all-reduce operation for processes 602A, 602B, 602C, 602D, 602E, 602F, 602G, 602H, and 602I for stages 604 and 606. Each of processes 602A through 602I can correspond to a process (e.g., one or more iterations of process 122A and / or process 122B) running on one or more nodes (e.g., one or more iterations of computing device 102A and / or computing device 102B, respectively) in an interconnect network (e.g., collective network 100).

[0051] In an all-reduce operation, data that belongs to all processes (e.g., processes 602A through 602B at stage 604) is collected and operated on (e.g., reduced), and the result can be transmitted to all processes (e.g., processes 602A through 602I at stage 606). The data can be reduced by index such that the reduction is performed on the values of the corresponding index for each process (e.g., the first data element across all processes is reduced, the second data element across all processes is reduced, etc.). Figure 6 Shows each index of a data vector with different patterns.

[0052] The resulting vector includes a reduction for each index. Examples of reduction operations include maximum (e.g., selecting the maximum value), minimum (e.g., selecting the minimum value), sum (e.g., adding all values), product (e.g., multiplying all values), logical sum (e.g., Boolean operation), bitwise sum (e.g., performing a sum operation bitwise), logical, or bitwise, or logical exclusive, or bitwise exclusive, or the maximum value and the position of the maximum value, and the minimum value and the position of the minimum value. Thus, at stage 606, all processes 602A through 602I may have the same vector, which has a reduction for each index.

[0053] Although Figures 3 to 6 illustrates the source and destination processes at the start and end stages of the illustrated operation, the data may travel in different routes based on the operation, such that the various processes may be interconnected in different ways for different operations. In some specific implementations, a full reduction operation requires an interleaving node such as a switch to perform the operation. In contrast, a broadcast operation, for example, requires an interleaving node to forward the data but does not require an operation on the data. Thus, compared to a broadcast operation, the network topology used to perform a full reduction operation may affect the performance of the full reduction operation.

[0054] Figures 7A to 7C illustrates an example network topology for a collective operation 700, which may correspond to a full reduction operation for processes 702A, 702B, 702C, 702D, 702E, 702F, 702G, and 702H. Each of processes 702A through 702H may correspond to a process (e.g., one or more iterations of process 122A and / or process 122B) running on one or more nodes (e.g., one or more iterations of computing device 102A and / or computing device 102B, respectively) in an interconnection network (e.g., collective network 100). Figures 7A to 7C Further illustrates nodes 716A, 716B, and 718, each node corresponding to a node (e.g., one or more iterations of computing device 102A and / or computing device 102B) that may be used to transfer intermediate data to / from processes 702A through 702H as described herein.

[0055] For a full reduction operation, various processes (e.g., processes 702A through 702H) may hold the data to be reduced (e.g., each complete vector, as Figure 7A illustrated). Processes 702A through 702H may be organized such that half of the processes may be linked to the corresponding nodes. More specifically, as Figures 7A to 7CAs illustrated, processes 702A to 702D can be linked to node 716A, and processes 702E to 702H can be linked to node 716B. Additionally, nodes 716A to 716B can be linked to node 718, forming a tree topology as illustrated in Figures 7A to 7C Although in other examples, other topologies can be used, which can further depend on the collective operations to be performed.

[0056] Additionally, Figures 7A to 7C downstream links (e.g., links away from processes 702A to 702H) and upstream links (e.g., opposite to the downstream links and towards processes 702A to 702H) are illustrated. Figures 7A to 7C Downstream links that mirror the upstream links are illustrated (e.g., each downstream link has an inversion of the corresponding upstream link), although in other examples, such as for other collective operations, the downstream and upstream links can vary.

[0057] Figure 7A corresponding to the start of the all-reduce operation (e.g., corresponding to phase 604 in Figure 6 ), where each of processes 702A to 702H has its own data set / vector. Figures 7A to 7C Each index of the data vectors with different patterns is illustrated. As further illustrated in Figure 7A , each link in the link is idle (indicated by a solid line).

[0058] As described herein, the all-reduce operation involves each process providing its data, reducing the data across various data sets of the processes, and propagating the result to each process. Figures 7A to 7C The topology illustrated in Figure 7B illustrates an example of how various processes can coordinate the transfer and processing of data. In

[0059] , each process transfers its data (using the transfer link indicated by a dashed line) downstream to its corresponding node for processing. For example, processes 702A to 702D each transfer their data to node 716A, and processes 702E to 702H transfer their data to node 716B. Figure 7B The intermediate node can perform a reduction (represented by ∑1 in Figure 7B ) on the received data sets / vectors, which are subsets of the total data. Specifically, in

[0060] After reducing the received data vectors, each of nodes 716A and 716B provides its intermediate result downstream to node 718 for further processing. Specifically, node 718 reduces two halves of the total data received from nodes 716A and 716B to produce a final result (represented by ∑2) corresponding to the reduction performed across all data vectors from processes 702A through 702H.

[0061] Go to Figure 7C , node 718 propagates the result to each of processes 702A through 702H via nodes 716A through 716B. More specifically, as Figure 7C illustrated, based on the topology, node 718 propagates the result (via the transmission links indicated by the dashed lines) upstream to nodes 716A and 716B. Each of nodes 716A and 716B may propagate the result upstream to processes 702A through 702H, with node 716A propagating the result to processes 702A through 702D and node 716B propagating the result to processes 702E through 702H. Thus, each of processes 702A through 702H has the final result of the reduction (e.g., corresponding to Figure 6 phase 606 in ).

[0062] For the various collective operations described herein, data is passed and operated on at the interleaving nodes and links between processes. The collective engine (e.g., control circuit 112) may establish a network or routing topology, which in some examples may be different for each collective operation, and further improve its performance using the systems and methods described herein. As described above, for all-reduce operations, the network topology can affect performance. The collective engine provided herein may establish an appropriate network topology, for example, when initializing the collective network, thereby further establishing the routing between nodes.

[0063] To improve performance, the collective engine may optimize aspects of the network topology. For example, the collective engine may compute the communication costs (e.g., the latency associated with transmitting data between nodes) of different network topologies to determine the optimal topology using cost model parameters. Figure 8 Table 800 illustrates an example communication cost model for computing the communication cost time T for various topologies such as binary trees (e.g., Figures 7A to 7C ), recursive doubling, ring, and in-network (see, for example, Figures 9A to 9G illustrates an in-network with multiple links between layers / levels of nodes, as will be further described below)). The cost model herein may be represented by a cost function with parameters, which may be further implemented as instructions executed by any device / system described herein (such as the collective engine). For Figure 8In the cost model parameters, α is the latency (e.g., per-GPU network latency in units of time such as seconds), p is the number of GPUs in a collective communication operation, β is the inverse bandwidth (e.g., 1 / per-interconnect network bandwidth in bits per second), and ε is the downstream-to-upstream ratio (e.g., per network switch, the ratio of downstream ports to upstream ports, such as downstream / upstream), and n is the vector size (e.g., size in bytes of the vector to be reduced).

[0064] In some embodiments, the collective engine may use, for example Figure 8 a cost model of different topology types in to optimize the topology. For example, optimizing the topology may correspond to optimizing (e.g., minimizing or otherwise reducing) the communication cost as calculated according to the cost model. In some examples, optimizing the cost may include using different (e.g., corresponding to different topology types) cost models for different levels of the topology. In some examples, optimizing the cost may include optimizing the topology structure, such as flattening a tree (e.g., simplifying the tree structure by, for example, rearranging leaf nodes and / or child nodes to a level similar to the corresponding parent node), rearranging nodes / links, etc. For example, the collective engine may use an optimizer function (e.g., statistical and / or other mathematical analysis of one or more parameters, which can be implemented with instructions, machine learning models, etc.). Thus, the collective engine may optimize the communication cost of a collective operation (e.g., as exemplified in Figure 8 by evaluating multiple communication models, selecting a cost model corresponding to the optimized communication cost (e.g., corresponding to the feasible minimum cost value calculated according to the cost model), and using the selected cost model to configure the topology for performing the collective operation.

[0065] In some embodiments, the cost model may be used to construct the topology itself, such as by using the values of the cost model parameters (e.g., as used to optimize the communication cost) as topology parameters. For example, the cost model parameter (log-ε(p)) may be related to the number of layers (e.g., tree depth) of the in-network topology. Then, the collective engine may construct the topology (e.g., when initializing the collective network, the topology may include nodes configuring the collective network) by establishing appropriate links (e.g., communication connections) between layers (e.g., nodes of adjacent layers), for example using a heuristic following the topology parameters (which may use the values of the cost model parameters corresponding to optimizing the cost), such as the number of downstream and upstream ports, step values (e.g., the number of devices connecting to nodes of the next layer in a polling manner), the number of processors (e.g., processors and / or other devices capable of processing data), the number of ports per processor (e.g., input / upstream ports and / or output / downstream ports), tree depth (e.g., number of layers), etc., as will be further discussed below with respect to the Figures 9A to 9G example in-network topology.

[0066] Although this disclosure discusses optimizing network topologies for allreduce operations, in other specific implementations, the systems and methods described herein can be applied to other collective operations. Additionally, in some specific implementations, the network topology can include portions optimized using different cost models such that the overall network topology can include portions of different topologies (e.g., a combination of different topology types).

[0067] The collective engine (e.g., control circuit 112) can also establish communication routes for the network topology. For example, Figures 9A to 9G An example in-network topology 900 for allreduce operations within a network is illustrated. In some examples, the example in-network topology is optimized based on the network topologies described herein. The collective engine can include a network resource scheduler or aggregation manager that schedules network resources capable of performing collective offload operations and / or portions thereof. In some specific implementations, the network resource scheduler can reside in a node (e.g., a GPU) and / or in a separate device connected to the network. The network resource scheduler can establish flows that orchestrate and prioritize a particular link / switch. In some specific implementations, the network resource scheduler can request the status (e.g., metadata and / or other information about the current condition of the node) from each node in the collective network to initialize the collective network for collective operations. For example, if a resource (e.g., a node) indicates readiness (e.g., when initializing the collective network), the resource transmits its status (e.g., as a message) (the status can include a readiness status message with a job profile (e.g., metadata that can indicate, for example, processing power, memory capacity, node-preferred processing / operation type, location in the topology, other preferences / constraints, etc.)) to the network resource scheduler to establish a flow / path based on the readiness and job profile. More specifically, the network resource scheduler can establish a topology for data routing for collective operations, which can include identifying each level of the collective network required to perform the collective operation, identifying the destination nodes in each level capable of processing the data required according to the collective operation (e.g., as further determined according to the job profile and / or readiness status), and establishing links between these nodes. Additionally, if a route is blocked (e.g., a node is offline or otherwise unavailable), the network resource scheduler can thus establish a different route, for example, by prioritizing links between available nodes in the level.

[0068] In some examples, deadlock refers to the following scenario: The links to a node deliver two packets to that node, but the node has only one available buffer, such that one of the packets will not be received. In the context of collective operations, since each node depends on the results provided by other nodes and operates in parallel, deadlock scenarios can significantly affect the operation and its performance. The network resource scheduler can establish routes to prevent deadlocks.

[0069] In some examples, the network resource scheduler can enforce routes. For example, the network resource scheduler can use segment routing, where the routing / path information travels with the packet, such that nodes can forward the packet to the destination based on the routing information. In some embodiments, the header of the packet can include the routing information. Thus, a given node can be prevented from arbitrarily selecting an idle link to deliver the packet or otherwise delivering the packet to a link inconsistent with the established route.

[0070] By establishing routes, certain data packets can be routed to specific nodes as required by the collective operation. Additionally, in some examples, the network resource scheduler can dynamically update or otherwise re - establish the topology in response to changes in node availability (e.g., in response to changes in node state). Further, routing allows data to be chunked (e.g., split into smaller segments) and routed to appropriate nodes for reduction with other corresponding nodes.

[0071] The network resource scheduler (e.g., the network resource scheduler can be part of a collective engine such as control circuit 112) provides routes that can ensure data reaches specific nodes. The data can be chunked (e.g., split) into smaller segments (e.g., chunks corresponding to portions of the original message size, such as sub - vectors of a vector) to improve bandwidth and performance. In some embodiments, the chunking scheme can be based on the network topology that can be optimized as described herein. The topology can ensure that corresponding chunks (e.g., chunks that should be processed together) reach the same destination as needed to perform the collective operation, as will be further described below. In one embodiment, the chunking scheme can be based on similar factors for network topology optimization, such as the number of ports. For example, each node that performs a gather operation or a part thereof on data (in some examples, the data is received via an input port from a node in a previous tier of the collective network) can chunk its resulting data into a number of chunks equal to the number of downstream (e.g., output) ports, and transmit the chunks through the corresponding ports of the downstream ports (e.g., to nodes in a subsequent tier of the collective network), although in other examples, other chunking schemes can be used.

[0072] Figures 9A to 9G Illustrates an example route (e.g., for deadlock prevention) and chunking of an example in - network topology 900 for a full reduction operation. Figures 9A to 9GIllustrates various servers (e.g., corresponding to instances of server 202), GPUs (e.g., corresponding to instances of GPU 214A and / or GPU 214B which may further correspond to processes), switches (e.g., corresponding to instances of switch 218), NICs (e.g., instances of NIC 216A and / or NIC 216B), and other nodes (e.g., corresponding to instances of any of the nodes described herein). The in-network topology 900 illustrates how topology parameters (such as the number of downstream ports (e.g., ports for connecting to downstream nodes), the number of upstream ports (e.g., ports for connecting to upstream nodes), and / or a step value) can be used to construct a topology (which may use different values at different levels in some examples). The number of downstream ports can correspond to the number of links traveling downstream from the node, and the number of upstream ports can correspond to the number of links traveling upstream from the node. For the links of a node at a given level, the step value can correspond to how many links can connect to different nodes at the next level before repeating nodes. In other words, the step value can correspond to how many nodes a set of nodes at a given level can connect to at the next level.

[0073] The in-network topology 900 illustrates, for example, a GPU level where each node has two downstream ports (and no upstream ports as the GPU level may correspond to a root level). Each GPU with two downstream ports can connect to two different switches at the next level (e.g., the switch level), thus further matching the step value. At the switch level, each switch has two upstream ports (mirroring the downstream links of the GPUs) and two downstream ports (connecting to two different NICs). At the NIC level, each NIC has one upstream port (mirroring the downstream link of the switch) and one downstream port (e.g., corresponding to forwarding data). In some examples, the interconnect level of the GPU and the NIC can be relative to a single machine (e.g., server_0 and server_1). The NIC downstream ports can connect to nodes at the next level based on the step value of four. As Figures 9A to 9G illustrated, each NIC can connect to nodes sequentially in order and repeat the order after reaching the step value. More specifically, nic_0_0_0, nic_0_1_0, nic_0_0_1, and nic_0_1_1 can be linked to node_0_0_0, node_0_0_1, node_0_0_2, and node_0_0_3 respectively in order, and subsequent NICs can repeat the linking to nodes in order in a round-robin manner. The in-network topology 900 further includes additional nodes (e.g., node_0_1_0, node_0_1_1, node_0_1_2, and node_0_1_3) representing how the topology can be extended, and is illustrated with only a single upstream / downstream link to simplify the illustration.

[0074] As described herein, a collective engine (e.g., control circuit 112) can establish different topologies for different collective operations. In other words, the nodes of the in-network topology 900 can also have Figures 9A to 9G links not illustrated in Figures 9A to 9G , and for full reduction operations, the collective engine can implement routing according to the in-network topology 900 using, for example, segment routing, a destination list (e.g., associated with a port / link), and / or other routing implementations. Additionally, in some embodiments, the topology (e.g., one or more levels) can be further limited by the availability of available links and / or components (e.g., the in-network topology 900 illustrates each server having two GPUs), and the topology can be further related to which components reside in which devices. Additionally, in some embodiments, the topology can be based on the availability and / or type of nodes (e.g., certain nodes can be used to process data, and certain other nodes can be used to forward data).

[0075] To better illustrate how data / blocks are routed through the in-network topology 900 and how that data / block relates to each node, each level is illustrated with a specific indexing symbol. For example, the GPU level can be indexed based on server_gpu such that the server index and the GPU index of that server can identify the GPU (e.g., Gpu_0_0 and Gpu_0_1 branching from server_0 and Gpu_1_0 and Gpu_1_1 branching from server_1). The switch level can be indexed based on server_switch such that the server index and the switch index of that server can identify the switch. The NIC level can be indexed based on server_GPU_NIC such that the server index, the corresponding GPU index, and the NIC index of that server / GPU can identify the NIC. Outer-level nodes can be indexed based on pod_layer_node such that the pod of the server (e.g., Figures 9A to 9G the "0" in Figures 9A to 9G ), the layer within the node (e.g., "0" for the inner layer, "1" for the outer layer), and the node index can identify the node (which can correspond to a switch or any other node described herein). Additionally, example data (e.g., a block) can be indexed by server_GPU_block_subblock such that the server index, the GPU index, and the block / subblock index can identify the block. In some embodiments, although other indexing and / or identification can be used, the indexing described herein can be used to identify nodes and / or blocks as data is routed through the in-network topology 900. Additionally, for each level, the nodes of that level can transmit data through a specific port of a link having a specific destination as described herein, which can be further implemented via segment routing (e.g., having the destination node and / or port as part of the transmitted data).

[0076] In Figure 9AAt time 0 (e.g., regarding the time for transmitting / chunking data), each GPU holds data (e.g., a message) of the full message size (e.g., size = 1). Thus, each message is indexed similar to its corresponding originating GPU. At Figure 9B At time 0 (e.g., also regarding the time for transmitting / chunking data), each GPU can chunk its data according to a chunking scheme which, in some examples, corresponds to dividing the data vector into equally sized chunks (e.g., sub-vectors from the original vector) for each downstream port. In some examples, the sub-vectors can be defined based on vector indices (such as the lower half and upper half of the vector index, odd and even vector indices, etc.) to produce sub-vectors of appropriate size. As will be further described below, by applying a consistent chunking to the vectors and implementing routing, the data elements of the original vector can be reduced with the corresponding data elements of the same vector index from other data vectors (e.g., data vectors across nodes).

[0077] Thus, each GPU with two downstream ports can chunk its data into two similarly sized chunks (e.g., chunk size = 1 / 2 of the message size) for transmitting one chunk per port. At Figure 9B At time 0, each half of the message can be indexed as illustrated by "0" and "1" (e.g., message 0_0 is chunked into chunks 0_0_0 and 0_0_1, message 0_1 is chunked into chunks 0_1_0 and 0_1_1, etc.). By applying the chunking scheme across GPUs, each corresponding chunk half / index can contain the same vector index (e.g., all chunks ending with "0" can correspond to the same vector index from the original message). At time 0, all links in the link are idle, as Figure 9A and Figure 9B illustrated.

[0078] At Figure 9C At time 1 (e.g.), the chunks from the GPUs are routed to the switches such that the corresponding halves reach the same switch, such that the first switch has the first half of the original message and the second switch has the second half of the original message (e.g., switch_0_0 has chunks 0_0_0 and 0_1_0, while switch_0_1 has chunks 0_0_1 and 0_1_1, etc.). As Figure 9C illustrated, the links corresponding to the downstream ports have been used once. In some embodiments, the switch can read data from the GPU memory, although in other embodiments the data can be propagated as needed. As will be further described below, the described chunking scheme can ensure that the corresponding chunks reach the same destination to allow reduction (e.g., the value of a given vector index of the original message is reduced with other values of the same vector index from other messages).

[0079] At Figure 9DIn the example of FIG. 1 , the switch reduces the data (e.g., performs a reduce operation using two received blocks) to produce a result of a single block size. The result is again blocked according to the blocking scheme, for example, into two similarly sized sub-blocks corresponding to the two downstream ports (e.g., sub-block size = 1 / 2 block size). For example, the result of switch_0_0 is divided into sub-block 0_∑_0_0 and sub-block 0_∑_0_1. For the purpose of illustrating how blocks are reduced and routed, the sub-block index indicates that for a given block index (e.g., "0" below corresponds to block 0_0_0 and block 0_1_0), Gpu_0_0 and Gpu_0_1 have been reduced ("0_∑", in other words, this "0_∑" corresponds to the same vector index from message 0_0 and message 0_1 having been reduced, and more specifically indicates that data from two GPUs of the server has been reduced), and that block 0_0_0 and block 0_1_0 are now split into two sub-blocks "0" and "1" (e.g., at the end of the sub-block index). As described herein, sub-blocks may be divided based on vector index values. By applying the blocking scheme across switches, each corresponding sub-block half / index may contain the same vector index (e.g., all sub-blocks ending with "0_0" may correspond to the same vector index as from the original message, and more specifically "0_x" corresponds to the same half of the vector index from the original message, and "x_0" corresponds to the same half of the vector index from the first block).

[0080] exist Figure 9E In (e.g., time 3), the sub-block is transmitted downstream from the switch to the NIC based on the illustrated routing. Because the NIC does not operate on the data, the sub-block passes through the NIC and is forwarded to the node (e.g., layer 0 node), such as Figure 9F As illustrated in (e.g., time 4). Due to routing, appropriate corresponding blocks are forwarded to the same node, making reduction possible. For example, by using the step value as described above, nic_0_0_0 can transmit sub-block 0_∑_0_0 to node_0_0_0, and nic_1_0_0 can transmit sub-block 1_∑_0_0 to node_0_0_0. Therefore, node_0_0_0 has the same vector index of the sub-block (indicated by the common ending "0_0" as described above), so that node_0_0_0 can reduce its data, and so on for other nodes. In addition, as described above, the leading "0" or "1" in the sub-block index indicates the originating server, so that node_0_0_0 can reduce the same vector index of data sub-blocks from two servers, and so on for other nodes. Therefore, each of the four nodes has a data sub-block (e.g., 1 / 4 of the message size) that is reduced according to each of the original GPUs.

[0081] After reducing the data at layer 0 nodes, Figure 9GAt a certain point (e.g., time 6), as needed, the resulting sub - blocks can be transmitted deeper into the network represented by the layer 1 nodes. Although Figure 9G is not illustrated in, but if needed, the sub - blocks can be chunked into smaller blocks and transmitted deeper into the network (e.g., having a topology suitable for additional chunking).

[0082] As Figure 9G is illustrated in, the sub - blocks are indexed with a leading "∑_∑" and suffixes "0_0", "0_1", "1_0", and "1_1". The leading indicates that the sub - blocks of two servers and two GPUs have been reduced, and the suffixes correspond to the vector indices that are split into four sub - blocks. Thus, if these four sub - blocks are reconstructed into a single message, the four sub - blocks represent the reduction of the original message (e.g., the result of the reduction). To distribute the result back to the GPUs (see, for example Figure 6 and Figure 7C ), the sub - blocks can be forwarded back via the upload link through the in - network topology 900 (e.g., reversing the routing and chunking illustrated in Figures 9A to 9G without further processing / reduction).

[0083] Thus, by establishing a topology with a given level, the topology chunks its data and distributes the chunks in a round - robin manner to the next level with a number of nodes corresponding to the number of chunks (e.g., the GPU level chunks the data into two chunks to be transmitted to the switch level with two switches, and the switch level further chunks the data into four sub - blocks to be transmitted to four nodes at the node level via the NIC level). This topology allows for more efficient processing of smaller data sets (e.g., reducing chunks or sub - blocks at the switch level and node level 0 respectively). The topology can further improve bandwidth efficiency. For example, the topology itself can be optimized based on the cost model described above. Additionally, chunking allows for the conveyance of smaller data sets and further avoids keeping results at nodes and distributing the full results back to the process.

[0084] Although Figures 9A to 9G illustrates how chunked data is transmitted via deadlock - preventable routing, in other examples, data can be transmitted through the in - network topology 900 via a similar routing without chunking the data. Additionally, although the examples described herein generally follow the data moving downstream until it reaches the root node and reversing upstream (e.g., in one direction at a time), in other examples, as needed, the routing can include downstream and upstream movements (e.g., in both directions). For example, certain nodes may not be available for receiving and / or transmitting data (e.g., not having storage capacity and / or operating capabilities), such that the routing can include reversing into a layer, skipping a layer, etc. Additionally, Figures 9A to 9GAn example of a pod that illustrates two servers is shown. In other examples, the topology construction, routing, and / or chunking described herein can be for more and / or fewer pods of more and / or fewer servers and / or other devices.

[0085] Figure 10 FIG. 4 is a flowchart of an exemplary computer-implemented method 1000 for network collective offloading routing management. Figure 10 The steps shown in FIG. 4 can be performed by any suitable computer-executable code and / or computing system, including Figure 1 , Figure 2 , Figures 7A to 7C and / or Figures 9A to 9G the systems illustrated in FIG. 4. In one example, Figure 10 each step shown in FIG. 4 represents an algorithm, the structure of which includes multiple sub-steps and / or is represented by multiple sub-steps, examples of which are provided in more detail below.

[0086] As Figure 10 illustrated in FIG. 5, at step 1002, one or more of the systems described herein evaluate a plurality of communication cost models for a collective operation. For example, control circuit 112 can evaluate a plurality of communication cost models for a collective operation (e.g., for a full reduction operation Figure 8 ).

[0087] The systems described herein can perform step 1002 in various ways. In one example, evaluating the plurality of communication cost models includes optimizing the communication cost of the collective operation, and the values of the plurality of cost model parameters are determined based on the optimized communication cost. As described above, control circuit 112 can use an optimizer function (e.g., a circuit and / or instructions for evaluating and optimizing a cost function) on a cost model (e.g., having a cost function) to determine the values of the cost model parameters for the optimized communication cost. As further described herein, in some examples, the optimization can include using portions of different cost models in a segmented manner (e.g., using certain cost models for certain levels / nodes of a collective network).

[0088] At step 1004, one or more of the systems described herein determine the values of a plurality of cost model parameters based on the evaluation. For example, control circuit 112 can determine the values of the cost model parameters based on the evaluation.

[0089] At step 1006, one or more of the systems described herein configure the topology of a collective network for performing a collective operation using a plurality of topology parameters corresponding to the values of the plurality of cost model parameters. For example, control circuit 112 can use the topology parameters determined according to the cost model parameters to configure nodes according to the topology.

[0090] The system described herein can perform step 1006 in various ways. In one example, the value of a cost model parameter can be used as the value of a corresponding topology parameter for constructing a topology. For example, the values of cost model parameters that result in an optimized communication cost (such as the number of upstream ports, the number of downstream ports, the number of processors (e.g., nodes having a job profile that matches the processing needs), the number of ports per processor, the tree depth, and / or the step value) can thus be used as the values of topology parameters (e.g., using similar values to construct the topology). Figures 9A to 9G An example of a topology using the topology parameters as described above is illustrated.

[0091] In some examples, configuring the topology includes: configuring communication connections between nodes of a level of the collective network and nodes of an adjacent level of the collective network based on the plurality of topology parameters. For example, the control circuit 112 can establish links (e.g., upstream and downstream links on corresponding upstream and downstream ports) between nodes of various levels such as Figures 9A to 9G the GPU level and the switch level, the switch level and the NIC level, etc.

[0092] In some examples, the control circuit 112 can configure a part of the topology based on the plurality of topology parameters and configure a second part of the topology based on a second communication cost model. For example, Figures 9A to 9G an example of how to configure certain levels differently is illustrated (e.g., the NIC level, which can be further configured based on port availability; and / or the node level, which can be configured as needed based on node availability and / or collective operation processing needs). Additionally, the control circuit 112 can further configure and / or reconfigure the topology based on any other factors described herein for establishing a topology (e.g., dynamically).

[0093] Figure 11 is a flowchart of an exemplary computer-implemented method 1100 for network collective offload message chunk management. Figure 11 The steps shown in Figure 1 can be performed by any suitable computer-executable code and / or computing system (including Figure 2 ), Figures 7A to 7C ), Figures 9A to 9G and / or the system illustrated in Figure 11 ). In one example, each step shown in

[0094] As shown in Figure 11As illustrated, at step 1102, one or more systems in the systems described herein receive multiple data sets. For example, nodes such as computing device 102A or computing device 102B (e.g., Figure 9C the corresponding switch_0_0 in or any other node / process described herein) may receive multiple data sets (e.g., message blocks).

[0095] The systems described herein may perform step 1102 in various ways. In one example, according to a chunking scheme, the number of multiple data sets corresponds to the number of input ports. For example, in Figure 9C switch_0_0 may receive a first block (0_0_0) from a first input port (connected to Gpu_0_0) and a second block (0_1_0) from a second input port (connected to Gpu_0_1).

[0096] At step 1104, one or more systems in the systems described herein perform a collective operation of a collective network on the multiple data sets to generate a result data set. For example, switch_0_0 may perform a collective operation on the received blocks (e.g., see Figure 9D ).

[0097] At step 1106, one or more systems in the systems described herein split the result data set into multiple blocks. For example, in Figure 9D switch_0_0 may split the result into sub-blocks.

[0098] The systems described herein may perform step 1106 in various ways. In one example, according to a chunking scheme, the number of multiple data sets corresponds to the number of input ports, and the number of multiple blocks corresponds to the number of output ports. For example, in Figure 9D switch_0_0 with two downstream ports may split the result into two sub-blocks, as further described above.

[0099] At step 1108, one or more systems in the systems described herein transmit the multiple blocks to multiple destinations based on a topology corresponding to a collective operation that defines a path for the blocks to pass through the collective network, such that the corresponding blocks reach the same destination. For example, in Figure 9E switch_0_0 may transmit sub-block 0_∑_0_0 to Nic_0_0_0 and sub-block 0_∑_0_1 to Nic_0_1_0 according to the topology.

[0100] The system described herein can perform step 1108 in various ways. In one example, the topology defines connections between nodes of a collective network based on collective operations such that multiple data sets are received from a previous tier of the collective network and multiple chunks are transmitted to a subsequent tier of the collective network, which in some embodiments may be established by a collective engine as described herein (e.g., control circuit 112). Additionally, control circuit 112 may further configure and / or reconfigure the topology (e.g., dynamically) based on any other factors described herein for establishing the topology. In some examples, the path of the chunks through the collective network ensures that chunks matching vector indices reach the same destination for collective operations (see, e.g., as further described above Figures 9F to 9G ).

[0101] Figure 12 is a flowchart of an exemplary computer-implemented method 1200 for network collective offload routing and deadlock prevention management. Figure 12 The steps shown in Figure 1 can be performed by any suitable computer-executable code and / or computing system (including Figure 2 、 Figures 7A to 7C and / or Figures 9A to 9G the systems illustrated in Figure 12 ). In one example,

[0102] each of the steps shown in Figure 12 represents an algorithm, the structure of which includes multiple sub-steps and / or is represented by multiple sub-steps, examples of which will be provided in more detail below.

[0103] As illustrated in Figures 9A to 9G

[0104] at step 1202, one or more of the systems described herein receive a job profile and a ready state regarding initializing a collective network for collective operations from each of multiple nodes in the collective network. For example, control circuit 112 may receive a job profile and / or a ready state from nodes in the collective network (e.g., one or more iterations of computing devices 102A to computing devices 102B).

[0103] The system described herein can perform step 1202 in various ways. In one example, as part of initializing the collective network and more specifically initializing the collective network for collective operations (e.g., Figures 9A to 9G the allreduce operation in

[0104] At step 1204, one or more of the systems described herein select nodes that can handle the collective operations required according to the tier for each tier of the collective network based on the job profile and the ready state. For example, the control circuit 112 can select appropriate nodes for each tier of the collective network that can handle the collective operations (and / or portions thereof) as needed.

[0105] The systems described herein can perform step 1204 in various ways. In one example, the control circuit 112 can determine what processing is needed at each approximate tier based on the collective operations to determine the processing requirements, and select nodes (e.g., based on the job profile and / or the ready state) that can perform the processing (such as the reduction of blocks at the switch tier in Figures 9A to 9G and the reduction of sub-blocks at the node tier).

[0106] At step 1206, one or more of the systems described herein determine a topology for data routing for the collective operation based on the nodes selected for each tier. For example, the control circuit 112 can determine a topology that includes data routing (e.g., based on the links between the nodes).

[0107] The systems described herein can perform step 1206 in various ways. In one example, determining the topology includes determining the links of the selected nodes of the tier to the nodes of the previous tier and the nodes of the subsequent tier. For example, the control circuit 112 can determine the number of tiers of reduction (e.g., at the switch tier and the node tier in Figures 9A to 9G ) required for a full reduction operation, and further determine the intermediate tier (e.g., the NIC tier for forwarding data). In some embodiments, the control circuit 112 can further use topology parameters to determine the links between ports / nodes (e.g., as described above, using the step value for connecting the NIC tier to the node tier). Additionally, the control circuit 112 can further configure and / or reconfigure the topology (e.g., dynamically) based on any other factors described herein for establishing the topology.

[0108] At step 1208, one or more of the systems described herein configure the plurality of nodes based on the determined topology. For example, the control circuit 112 configures the iterations of the computing devices 102A to 102B based on the topology (e.g., see Figures 9A to 9G ).

[0109] The systems described herein can perform step 1208 in various ways. In one example, configuring multiple nodes includes establishing links between the selected nodes at each level based on the determined topology. For example, the control circuit 112 can instruct the nodes to establish links to the appropriate nodes in the adjacent levels according to the topology. In some examples, the links can be established dynamically when performing the corresponding collective operations. For example, using segment routing, each node can forward data to the appropriate node via the appropriate port, such that each node can perform different collective operations by processing / forwarding data according to the topology. In some examples, the control circuit 112 can re - establish and / or dynamically change the topology (e.g., the links) in response to changes in the collective network, such as changes in the ready state and / or job profile (e.g., processing capacity and / or memory capacity).

[0110] In addition, although Figure 10 、 Figure 11 and Figure 12 illustrate various steps, in some examples, Figure 10 、 Figure 11 and / or Figure 12 one or more of the steps illustrated in

[0111] As described above, the computing devices and systems described and / or illustrated herein broadly represent any type or form of computing device or system capable of executing computer - readable instructions (such as those included in the modules described herein). In their most basic configuration, these computing devices each include at least one storage device and at least one physical processor.

[0112] In some examples, the term "memory device" generally refers to any type or form of volatile or non - volatile storage device or medium capable of storing data and / or computer - readable instructions. In one example, the memory device stores, loads, and / or holds one or more of the modules and / or circuits described herein. Examples of storage devices include, but are not limited to, random access memory (RAM), read - only memory (ROM), flash memory, hard disk drive (HDD), solid - state drive (SSD), optical disk drive, cache, variations or combinations of one or more of the above components, or any other suitable memory.

[0113] In some examples, the term "physical processor" generally refers to any type or form of hardware-implemented processing unit capable of interpreting and / or executing computer-readable instructions. In one example, the physical processor accesses and / or modifies one or more modules stored in the aforementioned memory device. Examples of physical processors include, but are not limited to, microprocessors, microcontrollers, central processing units (CPUs), field-programmable gate arrays (FPGAs) implementing soft-core processors, application-specific integrated circuits (ASICs), systems-on-a-chip (SoCs), digital signal processors (DSPs), neural network engines (NNEs), accelerators, graphics processing units (GPUs), parts of one or more of the above, variants or combinations of one or more of the above, or any other suitable physical processor.

[0114] Although illustrated as separate elements, the modules described and / or illustrated herein may represent parts of a single module or application. Additionally, in certain specific implementations, one or more of these modules may represent one or more software applications or programs that, when executed by a computing device, cause the computing device to perform one or more tasks. For example, one or more of the modules described and / or illustrated herein represent modules that are stored and configured to run on one or more of the computing devices or systems described and / or illustrated herein. In some specific implementations, a module may be implemented as a circuit or circuitry. One or more of these modules may also represent all or part of one or more dedicated computers configured to perform one or more tasks.

[0115] Furthermore, one or more of the modules described herein transform data, physical devices, and / or representations of physical devices from one form to another. For example, one or more of the modules described herein may receive one or more data vectors to be transformed, transform the data, output the transformation result for transmission to other modules, use the transformation result to perform collective operations, and store the transformation result to perform the collective operations. Additionally or alternatively, one or more of the modules described herein may transform any other part of a processor, volatile memory, non-volatile memory, and / or physical computing device from one form to another by executing on the computing device, storing data on the computing device, and / or otherwise interacting with the computing device.

[0116] In some specific implementations, the term "computer-readable medium" generally refers to any form of device, carrier, or medium that can store or carry computer-readable instructions. Examples of computer-readable media include, but are not limited to, transmission media such as carrier waves, and non-transitory media such as magnetic storage media (e.g., hard disk drives, tape drives, and floppy disks), optical storage media (e.g., compact discs (CDs), digital video discs (DVDs), and Blu-ray discs), electronic storage media (e.g., solid-state drives and flash media), and other distribution systems.

[0117] The order of process parameters and steps described and / or illustrated herein is given by way of example only and may vary as needed. For example, although the steps illustrated and / or described herein are shown or discussed in a particular order, these steps need not necessarily be performed in the order shown or discussed. The various exemplary methods described and / or illustrated herein may also omit one or more steps described or illustrated herein, or include additional steps other than those disclosed.

[0118] The foregoing description has been provided to enable other technicians in the art to best utilize the various aspects of the exemplary specific implementations disclosed herein. This exemplary description is not intended to be exhaustive or limited to any precise form. Many modifications and variations are possible without departing from the spirit and scope of the present disclosure. The specific implementations disclosed herein should be considered illustrative rather than restrictive in all respects. When determining the scope of the present disclosure, reference should be made to the appended claims and their equivalents.

[0119] Unless otherwise indicated, the terms "connected to" and "coupled to" (and their derivatives) as used in the specification and claims will be considered to permit both direct and indirect (i.e., via other elements or components) connections. Additionally, the term "a" or "an" as used in the specification and claims will be considered to mean "at least one". Finally, for ease of use, the terms "comprising" and "having" (and their derivatives) used in the specification and claims may be interchanged with the word "including" and have the same meaning.

Claims

1. An apparatus, the apparatus comprising: a control circuit configured to: select a communication cost model for a collective operation; and configure a topology of a collective network for performing the collective operation based on the selected communication cost model.

2. The apparatus according to claim 1, wherein the control circuit is configured to select the communication cost model by: optimizing a communication cost of the collective operation by evaluating a plurality of communication cost models for the collective operation; and selecting, from the plurality of communication cost models, a cost model corresponding to the optimized communication cost.

3. The apparatus according to claim 2, wherein the communication cost model includes a plurality of parameters, and optimizing the communication cost includes determining optimized parameters of the plurality of parameters.

4. The apparatus according to claim 3, wherein configuring the topology is based on using the optimized parameters as topology parameters.

5. The apparatus according to claim 3, wherein the plurality of parameters includes at least one of the following: the number of upstream ports; the number of downstream ports; the number of processors; the number of ports per processor; the tree depth; and the step value.

6. The apparatus according to claim 5, wherein configuring the topology comprises: Configure communication connections between nodes of a level of the collective network and nodes of an adjacent level of the collective network based at least on an optimized step value corresponding to the number of nodes connected in a next level.

7. The device according to claim 2, wherein optimizing the communication cost comprises: Flatten a tree associated with the cost model.

8. The apparatus according to claim 1, wherein the control circuit is further configured to: configure a part of the topology based on the selected communication cost model.

9. The apparatus according to claim 8, wherein the control circuit is further configured to: configure a second part of the topology based on a second communication cost model.

10. A system, the system comprising: a memory; a processor; and a control circuit configured to: evaluate cost parameters of a communication cost model for a collective operation; and configure a topology of a collective network for performing the collective operation based on the cost parameters.

11. The system according to claim 10, wherein the control circuit is configured to: evaluate the cost parameters of the communication cost model by optimizing a communication cost of the collective operation.

12. The system according to claim 11, wherein optimizing the communication cost includes determining optimized parameter values of the cost parameters, and configuring the topology is based on using the optimized parameter values as topology parameters.

13. The system according to claim 12, wherein the cost parameters includes at least one of the following: the number of upstream ports; the number of downstream ports; the number of processors; the number of ports per processor; the tree depth; and the step value.

14. The system according to claim 13, wherein configuring the topology comprises: Configure communication connections between nodes of a level of the collective network and nodes of an adjacent level of the collective network based at least on an optimized step value corresponding to the number of nodes connected in a next level.

15. The system according to claim 11, wherein optimizing the communication cost includes: Flatten a tree associated with the cost model.

16. The system according to claim 10, wherein the control circuit is further configured to: configure a part of the topology based on the evaluated communication cost model and configure a second part of the topology based on a second communication cost model.

17. A method, the method comprising: evaluating a plurality of communication cost models for a collective operation; determining values of a plurality of cost model parameters based on the evaluation; and configuring a topology of a collective network for performing the collective operation using a plurality of topology parameters corresponding to the values of the plurality of cost model parameters.

18. The method according to claim 17, wherein evaluating the plurality of communication cost models includes optimizing the communication cost of the collective operation, and the values of the plurality of cost model parameters are determined based on the optimized communication cost.

19. The method according to claim 17, wherein configuring the topology comprises: Configuring communication connections between nodes of a level of the collective network and nodes of an adjacent level of the collective network based on the plurality of topology parameters.

20. The method according to claim 17, the method further comprising: Configuring a part of the topology based on the plurality of topology parameters and configuring a second part of the topology based on a second communication cost model.