Repeatable random rounding for intra-network computing
By configuring a seed value generation and storage mechanism in the switch, the problem of inconsistent seed value propagation and resource allocation in network computation is solved, and repeatable random rounding is achieved, ensuring the consistency and reliability of computation results.
Patent Information
- Application Number
- CN202510643867.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-20
- Filing Date
- 2025-05-19
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies lack effective mechanisms for propagating and coordinating seed values in network-based computation, leading to problems such as accumulated numerical deviations, repeatability failures, race conditions, and inconsistent resource allocation, which affect the reliability and consistency of computation results.
By configuring a seed value generation and storage mechanism in the switch, functionally interdependent derived seed values are derived from the basic seed value, and resource allocation is managed through an encryption mechanism to ensure repeatable random rounding operations without revealing the network topology.
It enables repeatable random rounding in a network-based computing environment, reduces numerical bias, ensures the consistency and reliability of calculation results, and supports multi-tenancy and independent operation of applications.
Smart Images

Figure CN120994161A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to reproducible random rounding for in-network computation operations, such as reductions and / or arithmetic operations. BACKGROUND
[0002] Switches and similar network devices are central components of many communication, security, and computing networks. Switches are often used to connect multiple devices to form a network. In some cases, switches have an in-network computation mode that allows certain computation functions, such as data reduction operations, to be performed by the switch itself. SUMMARY
[0003] In an illustrative example, a system includes at least one processing node to perform one or more computation processes as part of a distributed workload to generate an output. The at least one processing node is configured with a derived seed value generated from a base seed value. The system also includes a rounding circuit to perform a rounding operation on the at least one processing node in accordance with the derived seed value.
[0004] In another illustrative example, a processing node includes a computation circuit to perform one or more computation processes as part of a distributed workload to generate an output, and a rounding circuit to perform a random rounding operation on the output in accordance with a seed value in cooperation with the computation circuit.
[0005] In another illustrative example, a method includes configuring a port of a processing node with a plurality of seed values generated from a base seed value, and providing reproducible random rounding operations for the processing node based on the plurality of seed values.
[0006] The rounding methods described and set forth herein can be applied to switches, routers, or any other known or yet to be developed suitable type of networking device or general purpose computing device. Other features and advantages are described and set forth herein, which will be apparent from the following description and drawings. BRIEF DESCRIPTION OF DRAWINGS
[0007] The present disclosure is described in connection with the appended drawings that are not necessarily drawn to scale:
[0008] Figure 1 is a block diagram depicting an illustrative configuration of a system in accordance with at least some embodiments of the present disclosure;
[0009] Figure 2 is a block diagram depicting an example structure of a switch in accordance with at least some embodiments of the present disclosure;
[0010] Figure 3 is a flow diagram depicting a method in accordance with at least some embodiments of the present disclosure;
[0011] Figure 4 FIG. 8 is a flow diagram depicting another method in accordance with at least some embodiments of the present disclosure;
[0012] Figure 5 FIG. 9 is a flow diagram depicting yet another method in accordance with at least some embodiments of the present disclosure. DETAILED DESCRIPTION
[0013] The following description is provided for the purpose of illustrating certain embodiments and is not intended to limit the scope of the claims, applicability, or configuration in any way. Rather, the following description provides a description of possible implementations and configurations, along with some of the numerous possible components, arrangements, and criteria that can be used for practicing those implementations and configurations. It should be understood that numerous other implementations can be created, and that these implementations can be implemented in a variety of ways.
[0014] As can be appreciated from the following description, the components of the system can be arranged in any suitable locations within a network of distributed components for computational efficiency reasons, without affecting the operation of the system.
[0015] Furthermore, it should be understood that the various links connecting the elements can be wired, trace, or wireless links, or any suitable combination thereof, or any other suitable known or later developed element capable of providing and / or transmitting data to the connected elements. For example, the transmission medium used as the link can be any suitable electrical signal carrier, including coaxial cable, copper wire, and optical fiber, electrical traces on a printed circuit board (PCB), etc.
[0016] As used herein, the phrases “at least one,” “one or more,” “or,” and “and / or” are open-ended expressions that are both conjunctive and disjunctive in operation. For example, each of the expressions “at least one of A and B;” “one or more of A or B;” “A, B, or C;” “A, B, and C;” and “A, B, or C” means A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together.
[0017] As used herein, the term “automatically” and variants thereof mean any suitable process or operation that is performed without requiring substantial human input when performing the process or operation. However, a process or operation can be automatic even if performance of the process or operation uses substantial or non-substantial human input, so long as the input is received prior to performance of the process or operation. Human input is considered substantial if it influences the manner in which the process or operation will be performed. Human input that agrees to perform the process or operation is not considered “substantial.”
[0018] The terms “determine,” “calculate,” “estimate,” and variations thereof as used herein are used interchangeably and include any form of obtaining, assessing, or evaluating a value, including any appropriate type of method, process, operation, or technique.
[0019] Various aspects of the disclosure will be described with reference to the drawings, which are as schematic representations of idealized configurations.
[0020] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure.
[0021] As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” “includes,” and / or “including,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0022] Devices, including but not limited to personal computers, servers, central processing units (CPUs), graphics processing units (GPUs), and other types of computing devices, can be interconnected using network devices such as switches. These interconnected entities can form a network, enabling data communication and resource sharing among nodes. Switches and other computing devices can provide computing services, such as reduction and / or aggregation computations, on behalf of host devices.
[0023] For example, in the training process of a machine learning network that uses an in-network algorithm for reduction and aggregation, the addition or multiplication of a vector of floating point operands has higher precision than that of the operands sent by the host. After the computation is finished, a rounding operation can be needed. Random rounding is crucial for the training process to introduce controlled randomness, reduce bias and variance, improve generalization ability, and enhance robustness.
[0024] In standard rounding techniques, a number is typically rounded to the nearest representable number within a certain precision. For example, when rounding to the nearest integer, 2.3 is rounded to 2 and 2.5 is rounded to 3. This can introduce a consistent bias in one direction, especially when dealing with a large number of computations. Other rounding techniques, such as “round to the nearest even number” (RNE), in which a number is rounded to the nearest even number, also introduce bias, which is unacceptable for certain types of applications.
[0025] With stochastic rounding, numbers are rounded randomly up or down, rather than always rounding to the nearest number. The probability of a number rounding up or down can be proportional to the distance of that number from the two nearest representable numbers. For example, the number 2.3 can have a 30% chance of rounding up to 3 and a 70% chance of rounding down to 2. One advantage of stochastic rounding is that it can reduce systematic bias that results from multiple rounding operations. While each individual rounding operation can introduce error, these errors are not systematically biased up or down. Over a large number of operations, such errors tend to average out, which makes stochastic rounding particularly useful in iterative processes such as numerical optimization and machine learning.
[0026] The present disclosure provides solutions to the following problems in the context of stochastic rounding for in-network computation: 1) seed propagation and coordination: in-network devices propagate and coordinate seed values among themselves to ensure configuration consistency while avoiding undesirable numerical effects (e.g., numerical bias accumulation); 2) seed storage and retrieval: isolate computation streams between different applications and / or users - e.g., so that operations by user A do not change the numerical results for user B; 3) handle race conditions between endpoints or hosts, switches, or even individual ports in the network; and 4) enable users to receive an allocation of in-network computation resources that can be numerically different from a previously used allocation of in-network resources but is isomorphic to the previously used allocation (e.g., the two allocations have equivalent reduction trees formed from different sets of resources) - this feature can be implemented without revealing the underlying physical network topology to users of the network.
[0027] Solutions that lack mechanisms to handle the above problems will fail in any number of the following ways: 1) numerical bias accumulation will interfere with computation results; 2) reproducibility will fail because in-network computation is non-isomorphic (i.e., in-network operations are performed in different orders, so the numerical results will likely be different); 3) reproducibility will fail because of race conditions between packet arrival times; and 4) multiple tenants / applications / streams will interfere with each other and result in non-reproducible results.
[0028] In machine learning, and especially when training deep neural networks, stochastic rounding is very useful when dealing with low-precision arithmetic (e.g., 16-bit or 8-bit floating point numbers). Stochastic rounding helps maintain the accuracy of a model when precision is reduced, because it can prevent the accumulation of rounding errors that might otherwise cause severe bias or convergence problems.
[0029] According to one or more embodiments described herein, switches can enable a wide variety of nodes, such as other switches, servers, personal computers, and other computing devices, to communicate across a network. Ports of the switches can be used as communication endpoints, enabling the system to manage multiple simultaneous network connections with one or more nodes. Computing systems can perform one or more methods involving random rounding of a computation result. Such random rounding can be performed in a repeatable manner by the systems and methods described herein.
[0030] Repeatability, the ability to consistently reproduce the results of an experiment or computation, is a key aspect of computing processes, such as artificial intelligence (AI) model training performed by a host using a switch or other computing device that performs the computation. Repeatability has a variety of advantages in such scenarios. For example, in AI and machine learning, verifying results helps ensure that the model is accurate and reliable. When the results of rounding are repeatable, developers can more quickly resolve errors that arise during training. Repeatability helps identify and correct errors in AI computations. For example, if results can be consistently reproduced, it is easier to pinpoint where and why an error occurred, whether in the data, the algorithm, or the implementation.
[0031] Conventional methods of random rounding do not provide repeatability. There is a need for repeatability of random rounding to allow users to maintain snapshots and perform debugging using the exact same training process and ensure that the same sequence of random decisions is generated each time.
[0032] The present disclosure describes systems and methods for enabling a switch or other computing system to perform a computation (e.g., a reduction computation) on numbers received from, for example, one or more hosts. In some examples, the system generates different seed values from a base seed value such that the seed values are functionally dependent on one another. Each switch in a network of switches can be configured with a different seed value and used to perform random rounding of a computation result (e.g., a data reduction computation for in-network computation operations). Notably, example embodiments support repeatable random rounding in in-network computation environments (e.g., as implemented by NVIDIA’s Scalable Hierarchical Aggregation Protocol (SHARP) technology).
[0033] While examples provided herein refer to FP16 and FP32, it should be understood that implementations described herein can be used for any number format, including, for example, IEEE half-precision and / or single-precision floating-point numbers. For example, the present disclosure can also be applicable to non-IEEE floating-point formats, such as Bfloat16 (BF16), and the like. In certain embodiments, a host can send IEEE half-precision floating-point numbers, while a switch can perform computations in IEEE single-precision floating-point numbers.
[0034] Referring now to the drawings, various systems and methods for providing repeatable random rounding will be described. The rounding concepts described and set forth herein can be applied to rounding of numbers resulting from reduction operations as well as rounding of any other numbers. The implementations described below involve a particular example in which a host device uses a switch for computation and the switch returns a rounded result of the computation. However, it should be understood that the same or similar systems and methods can be used for various other uses, including any scenario in which a computing device seeks to round a number.
[0035] The term "data" as used herein should be understood to refer to any suitable discrete quantity of digitized information. The data received by a switch or other device can be in the form of packetized data or non-packetized data, without departing from the scope of the present disclosure. Moreover, certain embodiments will be described in connection with systems configured to receive data from a host and to reduce the received data. However, it should be understood that a host can not be required in certain implementations of the disclosed systems and methods. It should be understood that features and functionality of the systems and methods described herein can be used in a centralized architecture, a distributed architecture, or a single computing device.
[0036] As described in greater detail below, the inventive concept provides repeatable random rounding operations in the context of in-network computation (e.g., using SHARP technology) by using a pseudo-random algorithm based on seeds. At least one embodiment involves propagating seed values in nodes (e.g., switches) of a network. The seed values can be derived from an initial or base seed value such that all seed values are functionally dependent on one another. At least one embodiment involves how the seed values are stored (e.g., in dedicated switch memory) and retrieved, such as in response to a trigger such as a user command or in response to a node entering an in-network computation operation. At least one further embodiment involves allocating computing resources to users of in-network computation operations without revealing the topology of the network, while adhering to strict and complex topological restrictions.
[0037] Figure 1 A system 100 is shown that includes one or more switches 103, a central manager 104, and one or more hosts 203a-203d. Figure 2 An example structure of a switch 103 (also referred to herein as a processing node) is shown.
[0038] With reference to Figure 1 and Figure 2 The switch 103 can be part of a network in which multiple switches 103 communicate with one another, communicate with multiple hosts 203a-203d through ports 106a-106d, and / or communicate with a central manager 104. Such a network of switch-connected hosts can be used in a variety of environments, from data centers and cloud computing infrastructures to artificial intelligence systems.
[0039] Switches 103 can be or include, for example, network switches, network interface controllers (NICs), or other devices capable of receiving data and routing it to other nodes in a network. Switches 103 can be connected in a suitable topology (e.g., a fat tree topology) that includes, for example, top-of-rack (TOR) or core switches, spine switches, and / or leaf switches. Switches 103 can be capable of receiving, processing data (e.g., packets), and forwarding it to appropriate destinations in the network, such as other switches 103 and / or hosts 203. In some implementations, switches 103 can be contained in a switch pod, platform, or chassis that can contain one or more switches 103 as well as one or more power devices and other components.
[0040] Each host 203 can be a computing unit, such as a personal computer, a server, or other computing device, and can be responsible for executing applications and performing data processing tasks. As an example, the range of hosts 203 described herein can span from servers in a data center to desktop computers in a network, or to devices such as Internet of Things (IoT) sensors and smart devices. Hosts 203 can be or include host channel adapters (HCAs). Each host 203 can include one or more processing circuits, such as GPUs, CPUs, ASICs, FPGAs, or other circuits capable of performing computations, as well as memory and storage resources for running software applications, manipulating data processing, and performing specific tasks as needed. In some implementations, hosts 203 can also or alternatively include hardware such as GPUs for manipulating intensive tasks of machine learning, artificial intelligence (AI) workloads, or other complex processes. For example, hosts 203a-203d can utilize the computing power of switches 103 to aggregate data to arrive at a single result, such as by summing, finding a minimum or maximum value, or combining data sets. The data sent from hosts 203a-203d to switches 103 can be raw data that switches 103 can reduce.
[0041] The central manager 104 can manage one or more aspects on behalf of the system 100. The central manager 104 can have processing capabilities and be implemented by a server or other suitable computing device. Alternatively, the central manager 104 can be implemented internally within one or more of the switches 103, or by one or more of the switches 103. In some examples, the central manager 104 is responsible for generating or deriving seed values (for use in rounding operations within the switches 103) from a base seed value. The base seed value can be provided by the host 203 (e.g., by a user of an application running on the host 203). In at least one embodiment, the central manager 104 is responsible for encrypting messages indicating allocated computing resources that are available to an application for performing a distributed workload. The encrypted messages can also contain a description of allocated characteristics that must be replicated for a particular application, such as reduction topology criteria that define a topology of a reduction tree for the application (where a “reduction tree” refers to nodes that perform reduction operations on a distributed workload). The encrypted messages can be sent to the application or a user of the application and remain encrypted so as to not reveal the topology of the reduction tree to the application or the user of the application. Although not explicitly shown, the central manager 104 can communicate with other, non-illustrated elements of the system 100 (e.g., a job scheduler).
[0042] In some examples, the hosts 203 and switches 103 operate as a high-performance computing (HPC) cluster. The cluster of hosts 203 can contain numerous interconnected servers, each equipped with CPUs and / or GPUs. The hosts 203 can provide computational horsepower, e.g., for training large-scale AI models or running complex scientific simulations. For AI and machine learning tasks, the hosts 203 can contain one or more GPUs or other processing circuitry capable of handling the parallel processing demands of neural networks and other applications. The hosts 203 can participate in AI-related, research-related, and other processor-intensive tasks and leverage the network of switches 103 and other hosts 203 to handle distributed computing loads. Such hosts 203 can include workstations and personal computers used by researchers, data scientists, and professionals to develop, test, and run AI models and research simulations, for example.
[0043] In some implementations, the switches 103 are capable of providing computational power and performing computations on behalf of one or more hosts 203. For example, the switches 103 can perform one or more in-network computation processes as part of a distributed workload to generate an output. The distributed workload can correspond to a machine learning operation or other computational operation involving multiple computing resources (e.g., hosts) processing a large workload in parallel.
[0044] Data can flow through the network of switches 103 and hosts 203 using one or more protocols, such as the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP), or the Internet Protocol (IP). Upon receiving data from a host 203 or another switch 103, a switch 103 can inspect the data to determine a computation required for the data, perform the computation, round the result of the computation, and route the rounded result of the computation as data through the network.
[0045] With reference to Figure 2 A switch 103 can include a plurality of ports 106a-106d, buses 121a-121d, switching hardware 109, a buffer 112, one or more computing circuits 115, a processor 118, and a memory 124. The ports 106a-106d of a switch 103 can be capable of facilitating the transmission of data packets or non-packetized data to, from, and through the switch 103. These ports 106a-106d can serve as interface points for connecting network cables, connecting the switch 103 with other switches 103 and / or hosts 203.
[0046] Each port 106 can be capable of receiving incoming data packets from other devices and / or sending outgoing data packets to other devices. In some implementations, a port 106 can be configured to operate as a dedicated ingress or egress port 106, or can be enabled to operate in dual functionality capable of performing both ingress and egress functions. For example, an egress port 106 can be dedicated for sending data from the switch 103, while an ingress port 106 can be used only for receiving incoming data into the switch 103.
[0047] The switching hardware 109 of a switch 103 can be capable of manipulating received packets by performing ingress processing, reduction computations, generating numbers based on seed values, rounding the results of the reduction computations using the generated numbers, and performing egress processing on the rounded results of the reduction computations. Using the systems or methods described herein, the switching hardware 109 can be capable of providing reduction computation capabilities to one or more hosts 203 in a repeatable manner using random rounding.
[0048] Each port 106a-106d of a switch 103 can be associated with one or more buses 121a-121d. When data (such as vectors, streams of numbers, or data in any format) is received through a port 106a-106d, the data can be stored in the respective bus 121a-121d associated with the port 106a-106d. Data in numerical form on the buses can be used both for reduction computations and for generating numbers used to round the results of the reduction computations.
[0049] One or more computing circuits 115 can cause the switches 103 to perform computational tasks. These tasks can range from simple arithmetic computations to more complex logical decision processes. The computing circuits 115 described herein can be capable of performing various arithmetic operations, such as addition and / or subtraction, as well as logical operations (such as AND, OR, NOT, etc.). The computing circuits 115 can include one or more arithmetic logic units (ALUs), central processing units (CPUs), graphics processing units (GPUs), application-specific integrated circuits (ASICs), and / or field-programmable gate arrays (FPGAs) for processing computational tasks for the hosts 203.
[0050] According to embodiments of the present disclosure, the hosts 203 can offload reduction tasks with the switches 103 to minimize computational load and process data more efficiently. The reduction tasks described herein can include operations such as summing values, finding minimum or maximum values, or merging data sets. Example reduction operations include Reduce, AllReduce, ReduceScatter, BlockedReducedScatter, etc. In other words, the switches 103 can operate in an in-network compute mode, for example, according to NVIDIA’s SHARP technology, which has been introduced to greatly reduce latency of reduction operations. Specifically, SHARP defines a protocol for reduction operations performed on data as it traverses a reduction tree in a network. This enables manipulation of data as it is transmitted within a data center network without waiting for the data to reach a central CPU. Each switch 103 can utilize its computing circuits 115 to perform one or more of the reduction operations described above as it receives data from one or more hosts 203. As described above, the results of the operations can utilize rounding performed by one or more rounding circuits 116. In this scenario, the switches 103 can be configured to perform rounding operations (such as random rounding of the results of the computational operations) using one or more rounding circuits 116 and return the rounded results to one or more other nodes of the network, such as the hosts 203. The rounded results can be reproduced in later iterations, reducing rounding bias that can conflict with the results of computationally intensive tasks, such as training of AI models.
[0051] The operations performed by the compute circuit 115 can be floating point operations. For example, the bus 121 can receive one or more vectors containing a plurality of floating point numbers. To perform the operations, the compute circuit 115 can convert each floating point number to a floating point number having higher precision. This conversion can enable the compute circuit 115 to accurately perform the operations and minimize errors in the course of the operations. Once the numbers have the higher precision floating point format, the compute circuit 115 can perform floating point operations, such as addition or multiplication operations. As an example, the compute circuit 115 can iteratively add each higher precision floating point number to an accumulator.
[0052] After performing the floating point operations, the compute circuit 115 can output the results to the rounding circuit 116 to round the higher precision floating point results back to lower precision floating point numbers. The rounding circuit 116 can perform a random rounding using a seed value that is generated from a base seed value. Random rounding based on a seed will be described in more detail below, but generally should be understood as a form of random rounding augmented by an initial value called a seed.
[0053] Although the compute circuit 115 and the rounding circuit 116 are shown as separate circuits, these elements can be processes performed by a single unit, such as an ALU.
[0054] The one or more processors 118 can be configured to control various aspects of the switch hardware 109. In some implementations, the processors 118 can include CPUs, ASICs, and / or other processing circuitry that can be capable of handling computational, decisional, and management functions for the operation of the switch 103. The processors 118 can be configured to handle management and control functions of the switch 103, such as setting routing tables, configuring ports, and otherwise managing the operation of the switch 103. The processors 118 of the switch can execute software and / or firmware to configure and manage the switch 103, such as operating systems and management tools.
[0055] The memory 124 of the switch 103 described herein can include one or more storage elements that can store configuration settings, application data, operating system data, and other data. These storage elements can include, for example, random access memory (RAM), dynamic RAM (DRAM), flash memory, non-volatile RAM (NVRAM), ternary content-addressable memory (TCAM), static RAM (SRAM), and / or other formats of storage elements.
[0056] Example embodiments will now be described with reference to various methods that enable repeatable random rounding to be performed in the context of network- within-compute operations by the switch 103. The random rounding discussed herein is referred to as repeatable in that random rounding can be performed on the same value (e.g., a floating point value) at different times while obtaining the same result each time.
[0057] Example embodiments are directed to implementing random rounding in in-network computing environments to create a reliable, useful, and easy-to-use and integrate system. The present disclosure provides a solution to the following problems in the context of random rounding for in-network computing: 1) Seed propagation and coordination: in-network devices propagate and coordinate seed values among themselves to ensure configuration consistency while avoiding or reducing undesirable numerical effects (e.g., numerical bias accumulation); 2) Seed storage and retrieval: isolating computation streams between different applications and / or users— e.g., so that operations by user A do not change the numerical results of user B; 3) handling race conditions between endpoints, switches, and even individual ports in a processing network; 4) enabling users to receive an allocation of in-network computing resources that can be different in resource utilization from a previously used allocation of in-network resources but is isomorphic to the previously used allocation (e.g., both allocations have equivalent reduction trees formed from different sets of resources)— this feature can be implemented without revealing the underlying physical network topology to users of the network.
[0058] Referring to problem 1 (seed propagation and coordination) above, while all participating endpoints can send their initial seed to the in-network compute tree, only the seed of a designated endpoint is used as the base seed value. All other seed values are derived from this base seed value and are internally propagated to the in-network compute tree as part of the configuration step. Thus, the seed values are functionally interdependent to achieve repeatability while allowing each port to have a different seed value to overcome numerical bias accumulation issues.
[0059] Figure 3 A method 300 for seed propagation and coordination is shown in accordance with at least one embodiment. The method 300 can be performed by one or more elements described herein, such as the switch 103, the central manager 104, and / or the host 203.
[0060] Operation 304 includes receiving a base seed value. The base seed value can be provided by a user of the system 100, such as by a user of an application at the host 203. The base seed value can be requested from the host 203 in response to a request by the host to the system 100 to process a distributed workload to generate an output on which random rounding is performed. The base seed value can include a user-defined number, such as a real number (e.g., an integer), which can be represented in bits (FP16, FP32, etc.).
[0061] Operation 308 includes generating or deriving seed values from the base seed value. In some examples, a seed value can be generated or derived from the base seed value by using the base seed value as an input to an algorithm. The output of the algorithm can include a first derived seed value that is different from the base seed value. The first derived seed value can then be used as an input to the same algorithm to generate a second derived seed value that is different from the first derived seed value and the base seed value. This sequence of generating another derived seed value using a previously derived seed value as an input to the algorithm can continue until a sufficient number of derived seed values are generated. It can be appreciated that each derived seed value is different from the other derived seed values, but all of the derived seed values are functionally related to each other, which makes the reduction operation at the switches 103 repeatable while avoiding the numerical bias problem. The output of the algorithm that generates the derived seed values from the base seed value can also include an indicator that indicates whether each derived seed value is associated with an up-rounding operation or a down-rounding operation in the random rounding operation. In some examples, the derived seed values indicate a probability of up-rounding or down-rounding in the random rounding operation.
[0062] Operation 312 includes configuring the switches 103 using the derived seed values. For example, the seed values derived from operation 308 can be propagated throughout the network of the entire switches 103 such that each port 106 of each switch 103 is assigned or associated with one of the derived seed values. After the configuration is complete, each port 106 of each switch 103 that participates in the reduction operation can have a different derived seed value. It can be appreciated that operation 308 can include calculating the number of switch ports available for use by the application in order to generate the correct number of derived seed values.
[0063] Here, it should be appreciated that the method 300 can be executed separately for different applications such that a different set of derived seed values is generated for each application, which can avoid interference between the applications.
[0064] With reference to problem 2 (storage and retrieval of seed values) above, there can be two different solutions, each suitable for different switch hardware. The first solution uses a dedicated lookup mechanism within the firmware memory of the switch storing the derived seed values. Whenever the in-network compute function is used for a given workload, the derived seed value is fetched from the memory onto the relevant arithmetic unit of the switch 103. This can be done explicitly by the user (host 203) issuing a "load-seed" command. Only after the command is successful, the compute packets are sent to the switch 103. Alternatively, the switch firmware automatically loads the correct seed value by searching its memory. The second solution merges the seed values into the "start-up" mechanism of the in-network compute device (switch 103). For example, upon starting up or preparing the in-network compute resources for use by a first workload, as part of the configuration process, the seed value for the first workload is propagated onto the ports 106 (e.g., the seed value is generated according to method 300 and / or the stored seed value for the first workload is sent to the ports 106 by the host 203). If the first workload is temporarily paused and the in-network compute resources are available for use by other workloads, the seed value for the paused first workload is returned to the endpoint (host 203) and maintained or stored for use in case the first workload is resumed when the seed value will be propagated onto the ports 106 during the configuration process.
[0065] Figure 4 A method 400 for seed storage and retrieval is shown in accordance with at least one embodiment. The method 400 can be performed by one or more elements described herein, such as the switch 103, the central manager 104, and / or the host 203.
[0066] Operation 404 can include storing the derived seed values generated during the execution of the above-described method 300. For example, the specific progression of the derived seed values obtained by the algorithm mentioned in operation 308 can be stored for a particular workload (or stream) in order to be able to retrieve these derived seed values in case of a pause or interruption of the workload. Storing the progression of the derived seed values can involve maintaining an ordered list of the derived seed values in a memory (e.g., the buffer 112, the memory 124) to ensure that the same derived seed values are configured to the ports 106 before the workload stops and after the workload resumes. The ordered list of the derived seed values can be stored in a way that also indicates that the derived seed values are for a particular workload, to avoid interfering with seed values for other workloads.
[0067] Operation 408 includes determining that a derived seed value should be retrieved. Operation 408 can include, for example, detecting whether the stopped workload is resuming, e.g., if the workload has been queued for resumption, the resumption operation can occur automatically. In certain examples, detecting whether the stopped workload is resuming can occur in response to a request by the application to resume the workload.
[0068] Operation 412 includes retrieving the derived seed value stored in operation 404. Operation 412 can occur automatically in response to operation 408. That is, retrieving the derived seed value can occur in response to determining that the stopped workload is resuming. The derived seed value can be retrieved according to the capabilities of the system. One solution involves using a switch 103 firmware, such as a low-level communication library, that retrieves the derived seed value for the particular workload from memory, unloads the seed value from the port 106 used for the previous workload, and loads the derived seed value for the resuming workload to the port 106. Here, the system waits for a response indicating that the seed value propagation is complete before resuming the workload. Another solution involves building the retrieval functionality into the specialized hardware (e.g., compute circuitry 115) used for the reduction operation. In this case, operation 412 can include automatically loading the derived seed value belonging to the resuming workload from a scratchpad memory as part of the configuration process for resuming the workload.
[0069] With reference to problem 3 (race condition) described above, the way the in-network computation is executed is numerically equivalent to using a fixed order of summation for all operands, regardless of their individual arrival times. This can be achieved mathematically, such as by using a higher-precision internal representation.
[0070] With reference to problem 4 (in-network computation homogeneity) described above, when a user requests an in-network computation allocation, the user receives an encrypted string from the system (e.g., central manager 104). This encrypted string can only be decrypted by the central manager 104 of the in-network computation. If the user sends this string at some future date, the central manager 104 will be able to determine the topology of the in-network computation allocation (height of the tree, in-degree of each vertex, etc.) and assign resources to the user accordingly. Since the string has been encrypted, the user does not get access to the exact topology of the tree or the underlying network.
[0071] Figure 5 A method 500 for maintaining in-network computation homogeneity is shown in accordance with at least one embodiment. The method 500 can be performed by one or more elements described herein, such as the central manager 104.
[0072] Operation 504 includes receiving a request for in-network computing resources provided by the set of switches 103. The request can be sent by a user of a particular application to the central manager 104. The request for in-network computing resources can be a request for resources for performing in-network computing operations of the particular application. In response to the request, the central manager 104 can allocate in-network computing resources for the particular application. The allocation can be defined by characteristics that must be replicated for the workload of the application to be processed. Thus, the encrypted message can contain a description of the topology of the reduction tree, which is a logical tree used to perform reduction operations at the switches 103. For example, the encrypted message can contain a description of reduction topology criteria associated with the allocation, such as the height of the tree, the in-degree of each vertex, etc.
[0073] Operation 508 includes generating and sending an encrypted message including the allocation of in-network computing resources in response to the request and at a first point in time. The encrypted message can include a description of the topology of the allocated in-network computing resources, such as the height of the tree, the in-degree of each vertex, etc. The message can be encrypted by the central manager 104 and sent to the requesting application that stores or maintains the encrypted message. Notably, the application is unable to decrypt the encrypted message. Rather, the central manager 104 can be the only entity capable of decrypting the encrypted message. Thus, the topology of the reduction tree is not disclosed to the user / application, which is a significant feature of the method.
[0074] Operation 512 includes receiving the encrypted message at a second point in time later than the first point in time. For example, the central manager 104 can receive the encrypted message from the application when the application wishes to use the previously allocated in-network computing resources and send the encrypted message back to the central manager 104.
[0075] Subsequently, operation 516 includes enabling access to selected switches of the set of switches 103 according to the allocation. For example, in operation 516, the central manager 104 decrypts the encrypted message received from the application and determines the assigned reduction topology from the decrypted message. The central manager 104 can then select a subset of available switches 103 that satisfy the reduction topology criteria for the requesting application and enable the application to use the selected switches 103 for in-network computation operations. If the central manager 104 determines that the available switches 103 do not satisfy the reduction topology criteria, the central manager can notify the application of this condition. At this point, the user can instruct the system to delay in-network computation operations until enough switches 103 are available to satisfy the reduction topology criteria, or, alternatively, instruct the system to proceed with in-network computation operations despite the known failure to satisfy the reduction topology criteria. Notably, the selection of switches 103 for in-network computation operations is very flexible, as the central manager 104 need not select the same subset of switches 103 for a particular application in future iterations of the method 500. Rather, the central manager 104 can select a different subset of switches 104 in future iterations as long as the selected switches are able to satisfy the reduction topology criteria.
[0076] In light of the foregoing, it should be appreciated that example embodiments provide a system 100 that includes at least one processing node that executes one or more computational processes as part of a distributed workload to generate an output. The at least one processing node can correspond to or include a switch 103, and the one or more computational processes can be executed as part of in-network computation operations of the distributed workload. For example, as described above with reference to Figure 3 For example, the at least one processing node is configured with a derived seed value that is generated from a base seed value. As described with reference to Figure 4 In some examples, the derived seed value is retrieved from memory of the at least one processing node in response to a trigger. In at least one embodiment, the trigger includes a user command to load the derived seed value. In some embodiments, the trigger includes activation of in-network computation functionality of the at least one processing node. As described above, the at least one processing node can include a plurality of ports 106, and each port can be configured with a different seed value. As described above with reference to Figure 3 For example, the base seed value is used as an input to an algorithm that generates the subsequent derived seed value.
[0077] In at least one embodiment, system 100 further includes a rounding circuit 116 to perform a rounding operation on at least one processing node in accordance with a derived seed value. As described herein, the rounding operation can include a random rounding operation performed on a floating point value output from the computing circuit 115.
[0078] In at least one embodiment, system 100 further includes a central manager 104 in communication with the at least one processing node. The central manager 104 can perform operations described with reference to Figure 5 For example, the central manager 104 can receive a request for in-network computing resources provided by a set of processing nodes including the at least one processing node, and in response to the request and at a first point in time, send an encrypted message containing an allocation of the in-network computing resources. As described herein, the encrypted message can contain a description of a reduction topology standard that should be satisfied to ensure reproducibility. Thereafter, the central manager 104 can receive the encrypted message at a second point in time later than the first point in time, and enable access to selected nodes in the set of processing nodes in accordance with the allocation.
[0079] In light of the foregoing, at least one embodiment is directed to a processing node (e.g., switch 103) that contains a computing circuit 115 to perform one or more computing processes as part of a distributed workload to generate an output. As can be appreciated, the one or more computing processes can include a reduction operation. The processing node can also include a rounding circuit 116 that cooperates with the computing circuit 115 to perform a random rounding operation on the output in accordance with a seed value. In accordance with embodiments of the present disclosure, the output can include a floating point value on which the random rounding operation is performed. Further, the seed value can be generated from a base seed value in accordance with the discussion of Figure 3 The base seed value can be defined by a user (e.g., provided by a user of an application). The processing node can also include a plurality of ports each configured with a different seed value for performing the random rounding operation.
[0080] In light of the foregoing, at least one embodiment is directed to a method that includes configuring ports of a processing node using a plurality of seed values generated from a base seed value. Here, the processing node can correspond to a switch 103 having ports 106. Configuring the ports 106 using the plurality of seed values can be performed in accordance with the operations described with reference to Figure 3 The method can also include providing the processing node with reproducible random rounding operations based on the plurality of seed values. The reproducible random rounding operations can be performed on an output of the computing circuit 115 to ensure the same results are obtained at different times.
[0081] It should be understood that any of the features described herein can be claimed in combination with any of the other features described herein, whether those features come from the same described embodiment or not.
[0082] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments can be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been described in detail in order to avoid obscuring the embodiments.
[0083] While example embodiments of the present disclosure have been described herein, it should be understood that the inventive concept can be embodied in other ways and that the appended claims are intended to be construed as including other such embodiments, unless the existing technology limits the application.
Claims
1. A system comprising: at least one processing node to perform one or more computational processes as part of a distributed workload to generate an output, the at least one processing node configured with a derived seed value generated from a base seed value; and a rounding circuit to perform a rounding operation on the at least one processing node in accordance with the derived seed value.
2. The system of claim 1, wherein, The rounding operation is performed on a floating point value.
3. The system of claim 2, wherein, The rounding operation comprises a random rounding operation.
4. The system of claim 1, wherein, The at least one processing node comprises a plurality of ports, and wherein each port is configured with a different seed value.
5. The system of claim 4, wherein, The different seed values are generated using the base seed value.
6. The system of claim 1, wherein, The derived seed value is retrieved from a memory of the at least one processing node.
7. The system of claim 6, wherein, The derived seed value is retrieved from the memory in response to a trigger.
8. The system of claim 7, wherein, The trigger comprises a user command to load the derived seed value.
9. The system of claim 7, wherein, The trigger comprises activation of an in-network computing function for the at least one processing node.
10. The system of claim 1, further comprising: a central manager in communication with the at least one processing node, the central manager configured to: receive a request for in-network computing resources provided by a set of processing nodes including the at least one processing node; and in response to the request and at a first point in time, send an encrypted message comprising an allocation of the in-network computing resources, the encrypted message including a description of a reduction topology criterion associated with the allocation.
11. The system of claim 10, wherein the central manager is configured to: receive the encrypted message at a second point in time later than the first point in time; and enable access to selected processing nodes of the set of processing nodes in accordance with the allocation.
12. The system of claim 1, wherein, The at least one processing node comprises a network switch.
13. The system of claim 1, wherein, The one or more computational processes are performed as part of an in-network computing operation of the distributed workload.
14. A processing node comprising: computing circuitry to perform one or more computational processes as part of a distributed workload to generate an output; and a rounding circuit cooperating with the computing circuitry to perform a random rounding operation on the output in accordance with a seed value.
15. The processing node of claim 14, wherein, The output comprises a floating point value on which the random rounding operation is performed.
16. The processing node of claim 14, wherein, The seed value is generated in accordance with a base seed value.
17. The processing node of claim 16, wherein, The base seed value is user-defined.
18. The processing node of claim 14, further comprising: a plurality of ports, wherein each port is configured with a different seed value for performing the random rounding operation.
19. The processing node of claim 14, wherein, The one or more computational processes comprise a reduction operation.
20. A method comprising: configuring a port of a processing node with a plurality of seed values generated in accordance with a base seed value; and providing the processing node with a repeatable random rounding operation based on the plurality of seed values.