Elastic optimizer state for fully fragmented data parallelism
By sharing and replicating optimizer shards in distributed training, the problem of model parameter loss caused by node failure is solved, achieving fault resilience and training continuity, and improving the efficiency and reliability of large-scale ML training.
Patent Information
- Application Number
- CN202510094262.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-26
- Filing Date
- 2025-01-21
- Publication Date
- 2025-10-28
AI Technical Summary
Existing distributed training algorithms cannot provide resilience when nodes fail, leading to the loss of model parameters. In particular, in large-scale ML training, the failure of a single node may cause the entire learning process to be interrupted and resources to be wasted.
Failure resilience is achieved by sharing portions of optimizer shards across compute nodes. Specifically, this involves replicating optimizer shards on each node and dividing them into multiple parts, distributing these parts during the forward and backward propagation phases to ensure continuous updates and maintenance of the optimizer state in the event of node failure.
It enables continuous updating and maintenance of model parameters in the event of node failure, avoiding model restarts and resource waste, and ensuring the continuity and efficiency of distributed training.
Smart Images

Figure CN120849181A_ABST
Abstract
Description
Background Technology
[0001] Machine learning (ML) generally involves computer-implemented processes that use sample data (e.g., training data) to build models that make predictions or decisions without being explicitly programmed to do so. ML processes are used in a wide variety of applications, especially where developing conventional algorithms to perform various computational tasks is difficult or infeasible.
[0002] Distributed training is a subfield of machine learning in which multiple decentralized entities collaboratively train a common ML model using parallel execution on a subset of data or parameters held locally at each entity. Distributed training methods contrast with traditional centralized ML techniques that execute training sequentially on a single computing node. Attached Figure Description
[0003] The present disclosure is described in detail with reference to the following accompanying drawings, which illustrate one or more various embodiments. The drawings are provided for illustrative purposes only and depict merely typical or exemplary embodiments.
[0004] Figure 1 The illustration shows an example system for distributed training implemented according to an example of this disclosure.
[0005] Figure 2 This is a schematic block diagram depicting the forward propagation phase in distributed training implemented according to an example of this disclosure.
[0006] Figure 3 This is a schematic block diagram depicting the backpropagation phase in distributed training implemented according to an example of this disclosure.
[0007] Figure 4 This is a schematic block diagram depicting an example implementation of a processing flow for providing resilience to failures in distributed training, based on the present disclosure.
[0008] Figure 5 This is a schematic block diagram depicting a process implemented according to an example of this disclosure to ensure continuous resilience against node failures during distributed training.
[0009] Figure 6 This is an example computational component based on the implementation disclosed in this paper, which can be used to achieve various features of distributed training fault resilience.
[0010] Figure 7 This is another example computational component that can be used to achieve various features of distributed training fault resilience, based on the implementation disclosed in this paper.
[0011] Figure 8 These are example computer systems that can be used to implement various features of the distributed training faults disclosed herein.
[0012] The accompanying drawings are not exhaustive and do not limit this disclosure to the precise form disclosed. Detailed Implementation
[0013] Training large-scale ML models can be a challenging task, requiring significant computational power, resources, and time. For example, large language models (LLMs) can consist of numerous parameters being trained, and the number of such parameters has increased from 110 million in mid-2018 to one trillion in mid-2021. To efficiently train these large-scale LLMs, and other large-scale ML models, distributed training algorithms have been proposed, which partition training jobs across multiple computational resources (also known as “computing nodes”, “nodes”, or “network nodes”) in a distributed training network. For example, data parallelism (DP) replicates the entire model on each computing node and splits the training dataset into segments for each computing node. Another example is pipelined parallelism (PP), which divides the ML model into stages and distributes those stages across multiple computing nodes. Tensor parallelism (TP) is yet another approach, slicing tensors into multiple chunks, each of which can be executed on a computing node. However, these techniques are generally coupled to specific model architectures and are difficult to generalize to other models. For example, tensor parallelism may require tight coupling of computation nodes because parallelized matrix multiplication can be communication-intensive, so it only works across computational elements (e.g., GPUs) within a single server. The efficiency of pipeline parallelism may depend on the model itself and may generate inefficient execution (so-called "pipeline bubbles") when the model is irregular.
[0014] Fully Sharded Data Parallelism (FSDP) is another parallelization technique for distributed machine learning. FSDP shards (e.g., partitions or otherwise divides) model parameters (such as weight states and optimizer states) across a network of nodes. Each compute node can operate on a different training dataset (referred to as "local training data" or "local data" in this paper) held locally at that compute node. FSDP consists of at least two training phases: a forward propagation phase (sometimes referred to as "forward pass") and a back propagation phase (sometimes referred to as "back pass").
[0015] The forward propagation phase can be performed to obtain the output of a common ML model from local inputs. For example, an ML model may include multiple layers, each configured to perform a different transformation on one or more inputs. The transformation performed by each layer may depend on the weight states (sometimes referred to as "weights" in this paper) learned through training for each layer. The first layer of the model may be called the input layer, and the last layer may be called the output layer. Training data can be fed as input to the input layer. Multiple intermediate layers (sometimes called "hidden layers") can be provided sequentially between the input and output layers. The output from one layer is fed as input to the next layer in sequence. The forward propagation phase is used to compute the output from each layer in the order from the input layer to the output layer.
[0016] In the case of FSDP, during the forward propagation phase, each compute node operates to compute the local output of each layer from the common ML model by applying inputs locally held by the respective compute node. As mentioned above, each compute node holds a shard of model parameters for each layer, such as a shard of weight states for each layer (referred to herein as a "weight shard") and a shard of optimizer states (referred to herein as an "optimizer shard"). To perform the transformation from input to output, each compute node performs an all-gather operation for each layer to collect the weight shards held at other compute nodes and completely reconstruct the layer. Once the complete layer is obtained, each compute node performs a forward computation operation to obtain the output for that layer by feeding the locally held inputs to the reconstructed layer. Each compute node then discards the weight shards associated with other nodes (e.g., those received from other nodes) to free up space, thereby using the obtained outputs as input to repeat the process for the next layer in sequence. Repeat this process for each layer until the compute node obtains the output for the final layer.
[0017] The backpropagation phase can be performed to update the model parameters of the ML model by obtaining the loss function from previous layers. The backpropagation phase computes one or more gradients of the loss function relative to the weight states of a given layer. The backpropagation phase performs a backward propagation computation to obtain gradients one layer at a time, iterating backward from the output layer to the input layer. The backpropagation phase can utilize gradient descent, or variants such as stochastic gradient descent, to perform the backward computation. The backpropagation phase can utilize an optimization algorithm that can be defined by the optimizer state to compute updated weights relative to the obtained gradients. The term "optimizer state" can refer to the momentum vector of the optimization algorithm or a similar history-tracking characteristic. For example, the optimizer state for a gradient descent optimization algorithm can track the gradient and the moving average of the squared gradient.
[0018] In the case of FSDP, during backpropagation, each node operates to update the weight state of its weight slice by applying the optimizer slice of that node relative to the gradient corresponding to its weight slice. For example, for each layer, each compute node performs a full collection operation to collect weight slices from other nodes and completely reconstruct the layer. Once the complete layer is obtained, each compute node performs a backpropagation computation operation to compute the weight gradient and input gradient for the current layer relative to the weights of the fully reconstructed layer. This weight gradient and input gradient can be referred to as the local weight gradient and local input gradient, respectively. Each node then discards the weight slices collected from other nodes to free up space for the next layer. At this stage, each compute node holds the local weight gradient corresponding to each weight slice relative to the corresponding local input. The local weight gradients can then be aggregated (e.g., averaged) across compute nodes to obtain the global weight gradient. Each node can then perform a reduce-scatter operation on the global weight gradient to obtain a portion of the global weight gradient corresponding to the corresponding weight slice of that node. Each node can update the weights of its corresponding weight slice relative to that part of the global weight gradient.
[0019] The standard implementation of FSDP does not provide resilience to recover shards of model parameters (such as, but not limited to, optimizer shards) in the event of node failure, malfunction, or other anomalous behavior that renders a node unusable for distributed training (referred to as functional failures in this paper). Therefore, when a compute node becomes unavailable, the corresponding model parameters held by that node may be lost, and relearning may require re-initializing the entire process, at least with respect to the lost parameters. Thus, the failure of a single node can significantly disrupt the learning process, which can be exacerbated in large-scale ML training involving hundreds or thousands of compute nodes.
[0020] This disclosure provides a fault-resilience scheme, which can be implemented in the FSDP framework, to protect against node failures by sharing portions of optimizer shards among compute nodes. In the examples herein, each compute node may hold (e.g., store) a local shard of optimizer state (sometimes referred to herein as a “trained optimizer shard”), and a copy of a portion of the optimizer state shard stored at other compute nodes in the distributed training network (referred herein as a “copied optimizer shard portion”). The training optimizer shard may contain optimizer state that is local to each compute node of the optimizer algorithm, which is common among compute nodes and corresponds to a common ML model. Each compute node may be responsible for updating its weight state using its training optimizer shard relative to the gradient during the backpropagation phase, and updating its training optimizer shard relative to the acquired gradient (e.g., a portion of the global weight gradient corresponding to the corresponding weight shard of that compute node) during the backpropagation phase. To achieve resilience against failures, each compute node can replicate its training optimizer shard, divide the replicated training optimizer shard into replicated optimizer shard portions, and distribute the replicated optimizer shard portions to other compute nodes at any point during the forward and / or backpropagation phases. For example, replicated optimizer shard portions can be distributed via an all-to-al operation performed at any point during the forward and / or backpropagation phases. In an illustrative example, replicated optimizer shard portions can be distributed via an all-to-all operation performed simultaneously or approximately simultaneously with the full collection operation of the backpropagation phase. Thus, each compute node can hold replicated optimizer shard portions received from each of the other nodes on the network. According to some examples, each training optimizer shard can be replicated and divided into multiple (N) replicated optimizer shard portions of equal size, where N is one less than the number of compute nodes used for training. When a failure of a compute node is detected, each remaining node (referred to herein as a "functional node") can update its corresponding training optimizer shard with a replica of the optimizer shard corresponding to the failed node (e.g., a replica of the optimizer shard initiated or received from the failed node during a previous iteration of the backpropagation phase). Therefore, the updated training optimizer shard may include the previous training optimizer shard merged with the replica of the optimizer shard from the failed node.
[0021] To provide further resilience against subsequent node failures, each functionally healthy node can replicate and partition its updated training optimizer shards. These updated shards can then be distributed to the remaining functionally healthy nodes during the next iteration of forward or backward propagation. In this way, the latest optimizer state for the failed node can be maintained and updated across the network without being lost due to the failure. Therefore, the remaining functionally healthy nodes can continue uninterrupted training in ML mode.
[0022] Therefore, the implementation of this disclosure can provide resilience against node failures by sharing portions of the optimizer state among compute nodes at any time during the FSDP framework. Furthermore, by uniformly distributing the optimizer state, which is included in the optimizer shards of the failed node, across functionally healthy nodes, the workload can be shared equally among each functionally healthy node, thereby minimizing computational overhead in terms of power and resources.
[0023] It should be noted that the terms “optimized,” “best,” etc., used herein can be used to mean making or achieving the most efficient or perfect performance possible. However, as those skilled in the art who read this document will recognize, perfection cannot always be achieved. Therefore, these terms can also encompass making or achieving the best possible, most efficient, or most practical performance in a given scenario, or making or achieving better performance than could be achieved using other settings or parameters.
[0024] Figure 1 An example system 100 for distributed training, implemented according to an example of the present disclosure, is illustrated. The example system 100 includes a distributed training network 110 having multiple computing nodes 10A-10G (also collectively referred to as node 10, or individually referred to as node 10A-10G) in a cluster or group of computing nodes.
[0025] Each node 10 can be coupled to other nodes 10 via a network, which may include, for example, any one or more of the Internet, intranet, PAN (Personal Area Network), LAN (Local Area Network), WAN (Wide Area Network), SAN (Storage Area Network), MAN (Metropolitan Area Network), wireless network, cellular communication network, public switched telephone network, and / or other networks. Furthermore, depending on the implementation, the components described herein may be implemented in hardware and / or software configuring the hardware.
[0026] The cluster of distributed training network 110 can include any number, configuration, and connections between nodes 10. Thus, Figure 1The arrangement of nodes 10 shown is for illustrative purposes only. Node 10 can be a fixed computing device or a mobile computing device. Although in Figure 1 The diagram illustrates node 10A in detail, but each node 10 within node 10 can be configured in the manner shown. Figure 1 In the example, node 10A includes one or more processors 20 (for convenience, they are interchangeably referred to herein as multiple processors 20, (multiple) processors 20, or processor 20) and one or more storage devices 40 (for convenience, they are interchangeably referred to herein as multiple storage devices 40, (multiple) storage devices 40, or storage device 40), as well as other components. In the example, one or more compute nodes in the compute node may be implemented as graphics processing units (GPUs). The (multiple) storage devices 40 may hold (e.g., store) data 48 that is locally accessible to node 10A (referred to herein as local data). Other nodes 10 in the distributed training network 110 (e.g., nodes 10B-10G in this example) may not have access to the local data 48.
[0027] In some examples, storage devices 40 may store a distributed ledger 42, one or more models 44 (for convenience, they are interchangeably referred to herein as multiple models 44, multiple models 44, or model 44), and / or multiple rules 46. The distributed ledger 42 may include a series of data blocks referencing at least one other block (such as a previous block). In this way, data blocks can be chained together as the distributed ledger 42. In some examples, the distributed ledger 42 may store blocks indicating the state of node 10A in relation to machine learning during iteration. Thus, the distributed ledger 42 may store immutable records of state transitions of node 10A. In this way, the distributed ledger 42 may store the current model state and historical model state of each model 44. However, it should be noted that in some embodiments, a collection of records, models, and smart contracts from one or more other nodes (e.g., multiple nodes 10B-10G) may be stored in the distributed ledger 42.
[0028] As described herein, model 44 can be locally trained at node 10 based on local data 48, and subsequently updated based on model parameters learned at other nodes 10. The properties of model 44 will be based on the specific implementation of node 10 itself. For example, model 44 can be defined by learning parameters relating to: features of self-driving vehicles (such as sensor information, as it is relevant to object detection), network configuration features for network configuration, security features related to network security (such as intrusion detection), healthcare features related to patients' medical records and health-related information, social science features related to human behavior in social and cultural semantic aspects, and / or other context-based models.
[0029] Model 44 can be stored as a local instance of the ML algorithm and model parameters determined by training the ML algorithm. Model parameters can be stored as various model states, such as, but not limited to, weights, biases, optimizers, gradients, etc., that can constrain a specific instance of model 44. Each model 44 can include multiple layers, each defined by a set of model parameters used to perform different transformations on the input. The first layer of the model can be an input layer and the last layer can be an output layer, to which local data 48 can be fed. Multiple intermediate or hidden layers can be presented sequentially between the input and output layers, where the output from one layer can be fed as input to the next layer. The transformations performed by each layer can depend on the model parameters learned for that layer. The (multiple) models 44 can include any model of the general class of ML algorithms, including but not limited to many statistical and classical ML algorithms used in vertical domains (such as regression-based decision trees (DT), support vector machines (SVM), etc.). Training methods can include, but are not limited to, standard batch training.
[0030] In the examples presented herein, model 44 can be stored as a shard of model state that defines a local instance of the ML algorithm. For example, node 10A is illustratively shown as holding a weight shard 52A and an optimizer shard 54 of model 44, which defines a local instance of the ML algorithm. Each node 10 can hold other shards that collectively define the entire common layer. Model 44 can hold one or more shards of model state for each layer. For example, weight shard 52A can hold weights that define the transformations for one layer of the ML model, and model 44 can hold one or more other weight shards that define the transformations for one or more other layers of model 44. Similarly, optimizer shard 54A can hold optimizer state that defines a local instance of the optimization algorithm for one layer of the ML model (e.g., the same layer as weight shard 52A), and model 44 can hold one or more other optimizer shards for one or more other layers of model 44.
[0031] Rule 46 may include smart contracts or computer-readable rules that configure nodes to behave in some way related to distributed training and enable scatter control. For example, Rule 46 may specify deterministic state transitions, when to initiate machine learning iterations, whether nodes are allowed to register in iterations, the number of nodes required to agree on consensus decisions, the percentage of voting participants required to agree on consensus decisions, and / or other actions that node 10A may take for distributed machine learning.
[0032] Rule 46 can specify hyperparameters that define how the ML framework 24 and the flexible framework 28 are constructed. Hyperparameters can be considered mechanisms used to manage the training process, such as determining how many training iterations should be performed, how many nodes 10 should be used for training, setting training stopping criteria, setting data parallelism techniques, etc. Hyperparameters can be pre-set, adjustable parameters that can be tuned to obtain / generate an ML model / algorithm with optimal / tuned performance. In some examples, hyperparameters can be set by the operator via a front-end dashboard.
[0033] According to the examples disclosed herein, rule 46 may include one or more hyperparameters that configure the ML framework 24 to utilize the FSDP technique for distributed training. In this case, rule 46 may include one or more hyperparameters that specify the number of shards of model states to be created from the parameters of model 44. For example, rule 46 may specify the number of shards to be created from the set of weights for a given layer of model 44 by dividing (e.g., partitioning) the weights that define the transformation of the layer into a specified number of weight shards. Similarly, the set of optimizer states of model 44 may be divided into a specified number of optimizer shards. Rule 46 may then allocate the weight shards and optimizer shards to each node 10 for storage, for example, in storage devices 40(multiple) . According to the examples, the number of shards (e.g., the number of weight shards and the number of optimizer shards) may be specified as the number of registered nodes 10 used for training. In various examples, each shard may be equal in size, given the amount of memory required to store each shard. That is, for example, each weight partition can be equal in size, and each optimizer partition can be equal in size. However, weight partitions do not need to be equal in size relative to optimizer partitions.
[0034] Based on the examples disclosed herein, rule 46 may also include one or more hyperparameters that configure the resilient framework 28 to protect against node failures. In this case, rule 46 may include one or more hyperparameters that configure each node 10 to replicate its corresponding optimizer shard to create a copy of that optimizer shard. Rule 46 may also include one or more hyperparameters that configure each node 10 to divide the replicated optimizer shard into multiple replicated optimizer shard portions and distribute the replicated optimizer shard portions to other nodes 10. Each replicated optimizer shard portion may include data defining a different part of the optimizer shard (e.g., given the optimizer state contained therein), which does not overlap with another part of any other replicated optimizer shard portion originating from the same optimizer shard. Thus, the replicated optimizer shard portions can collectively recreate the entire optimizer shard. The number of replicated optimizer shard portions created may be based on the number of nodes 10 that are expected to operate as if performing distributed training or otherwise function. For example, the number of replicated optimizer shards can be one less than the number of nodes 10 registered for training.
[0035] In some examples, rule 46 may also include one or more hyperparameters that configure each node 10 to replicate its corresponding weight shard to create a copy of that weight shard. In this example, rule 46 may also include one or more hyperparameters that configure each node 10 to divide the replicated weight shard into multiple replicated weight shard portions and distribute the replicated weight shard portions of the weight shard to other nodes 10. Each replicated weight shard portion may include data that defines a different part of the weight shard (e.g., based on the weights contained therein), and this different part of the weight shard does not overlap with another part of any other replicated weight shard portion originating from the same weight shard. Thus, the replicated weight shard portions can collectively recreate the entire weight shard. As described above, the number of replicated weight shard portions created may be based on the number of nodes 10 that are expected to operate as if performing distributed training or otherwise function.
[0036] Rule 46 may also include one or more checks for detecting node failures (such as functional failures). A functional failure can refer to a situation where a computing node becomes unavailable for distributed training; for example, the node becomes inoperable or otherwise becomes a non-participant in the training. In some cases, node failure may cause damage to the storage device(s)40 of the failed node, potentially resulting in the loss of model parameters held by the failed node. In other cases, even if model parameters are not lost, a non-participating node may not update its model parameters along the training curve, meaning that when the node is reinserted into the training, it will have to catch up with the other nodes.
[0037] Due to, but not limited to, functional failures (also known as malfunctions) of nodes, a node may become functionally unusable for distributed training. Functional failures can include any failure of the compute node, including hardware failures and anomalous behavior demonstrated by the compute node. Hardware failures can refer to a situation where a hardware component of the compute node has failed or otherwise does not operate as expected. Anomalous behavior can include, but is not limited to, software failures (e.g., the compute node hardware functions, but the software hangs or otherwise fails to perform as expected), networking failures (e.g., the compute node operates as expected, but the compute node is unreachable because the switch port in the link between the compute node and other nodes is not working or is otherwise defective), and performance failures (e.g., the compute node fails to provide its shards to other compute nodes within a reasonable amount of time).
[0038] The examples in this paper can utilize any techniques used to detect such node failures. While illustrative examples are provided, they are not intended to be limiting, and any methods or techniques for detecting node failures or unavailability can be used in the examples disclosed herein. In some examples, a management system implemented by nodes 10 can be provided in the distributed training network 110, which monitors each participating node 10, detects when a node 10 has failed or other nodes are no longer participating, and generates alerts that can be provided to the remaining functional nodes 10 to notify of the failure. In another example, the ML framework 24 may include a watchdog process configured to monitor the learning process. As part of the ML framework 24, each node 10 may be expected to produce some data and exchange that data with other nodes 10 at certain times. The watchdog process can be configured to monitor the expected data exchanges, and if the expected communication is not received within the expected amount of time, the watchdog process can alert node 10 of a failure. Nodes that fail to provide the expected communication can be identified as failed nodes.
[0039] Multiple processors 20 may have access to local data 48 that is accessible locally by node 10A, but not necessarily by other nodes 10A. Such local data 48 may include, for example, private data not intended to be shared with other devices. Multiple processors 20 may be programmed by one or more computer program instructions. For example, processor 20 may be programmed to execute application layer 22, ML framework 24, interface layer 26, elastic framework 28, or other instructions to perform various operations, each of which is described in more detail herein. As used herein, for convenience, when various instructions actually program processor 20 (and therefore node 10A) to perform operations, the various instructions will be described as performing the operation.
[0040] Application layer 22 can execute applications on node 10A. For example, application layer 22 may include an agent (not shown) that programs node 10A to participate in distributed machine learning across distributed training network 110, as described herein. In the example, each node 10 may be programmed using the same agent, ensuring that each action is performed according to the same set rules (such as rules encoded using rule 46). For example, the agent can program each node 10 to act as a participant node based on hyperparameters specified by rule 46. For example, according to the following... Figures 2-6 As further described, the application layer 22 can perform machine learning through the ML framework 24 and the elastic framework 28.
[0041] The ML framework 24 can train a model based on local data 48 held at node 10A. For example, the ML framework 24 can generate one or more model parameters by applying local data 48 to a local instance of an ML algorithm (e.g., model 44). The ML framework 24 learns weights, biases, optimizers, and / or gradients as one or more model parameters (which are interchangeably referred to herein as "one or more local parameters" or "(multiple) local parameters"), which may constrain a particular model 44 and be stored in storage device 40. In this example, the ML framework 24 may use the FSDP framework, although other frameworks may also be used.
[0042] Based on various examples, the ML framework 24 can use FSDP to distribute training across nodes 10. For example, the ML framework 24 can perform multiple stages of training. In the case of distributed training via FSDP, the ML framework 24 operates on the local training data 48 and performs forward propagation and backpropagation stages to obtain model parameters. The ML framework 24 can perform a forward propagation stage to obtain the output of model 44 from the inputs iterated forward through each layer from the input layer to the output layer. The ML framework 24 can also perform a backpropagation stage to adjust the weights and optimizer of model 44 by computing the gradients as a loss function relative to the weights and local inputs of a given layer iterated backward through from the last layer to the first layer. The following section combines... Figure 2-Figure 3 Additional details are provided.
[0043] Application layer 22 can interact with and participate in distributed training network 110 using interface layer 26 for collaborative machine learning across multiple participant nodes 10. Interface layer 26 can communicate with other nodes, for example, by broadcasting transactions and writing blocks to distributed ledger 42 based on those transactions.
[0044] Interface layer 26 can share (multiple) local model parameters and inference with other participant nodes 10. Interface layer 26 may include messaging interfaces for communicating with other participant nodes 10 via the network. These messaging interfaces can be configured for Message Passing Interface (MPI) send / receive operations. Other types of messaging interfaces may also be used.
[0045] The elastic framework 28 ensures that distributed training is resilient to node failures on the distributed training network 110. For example, the elastic framework 28 can be executed by node 10A to replicate optimizer shard 54A and divide the replicated optimizer shard into multiple replicated optimizer shard portions according to rule 46. The elastic framework 28 can also be executed to distribute the replicated optimizer shard portions to other nodes 10 of the distributed training network 110, for example, by sending a different replicated optimizer shard portion to each of the other nodes 10. Each node 10 can similarly execute a corresponding elastic framework to create a replicated optimizer shard portion of the corresponding optimizer shard and distribute the replicated optimizer shard portion to other nodes 10. Node 10A can execute the elastic framework 28 to receive replicated optimizer shard portions from each of the other nodes on the network and hold the received replicated optimizer shard portions in storage devices(s) 40. Therefore, each node 10 can hold a different replica optimizer shard portion that is received (e.g., derived from) each of the other nodes 10.
[0046] According to rule 46, the elastic framework 28 can detect node failures. Upon detecting a node failure, each remaining functional node 10 (e.g., the remaining nodes 10 other than the failed node) can update its corresponding optimizer shard based on the replicated optimizer shard portion held at that functional node 10 and associated with the failed node. As an illustrative example, node 10A can execute the elastic framework 28 to update optimizer shard 54A using the replicated optimizer shard portion received from the failed node before the failure was detected. In the example, optimizer shard 54A can be updated by merging the replicated optimizer shard portion with the local optimizer shard 54A to produce an updated optimizer shard. The update of the local optimizer shard can be performed at each node 10 that remains active and involved in training. Therefore, the optimizer state contained in the optimizer shard of the failed node can be preserved and maintained in the updated optimizer shards of the remaining nodes 10. The updated local optimizer shard of node 10 can be used by the ML framework 24 to continue distributed training without interruption from detected faults.
[0047] Additionally, according to some examples, the elastic framework 28 may optionally be executed by node 10A to replicate weight shard 52A and divide the replicated weight shard into multiple replicated weight shard portions according to rule 46. The elastic framework 28 may also be executed to distribute the replicated weight shard portions to other nodes 10 of the distributed training network 110, for example, by sending a different replicated weight shard portion to each of the other nodes 10 in a manner similar to distributing replicated optimizer shard portions as described above. Each node 10 may similarly execute a corresponding elastic framework to create a replicated weight shard portion of the corresponding weight shard and distribute the replicated weight shard portion to other nodes 10. Node 10A may execute the elastic framework 28 to receive replicated weight shard portions from each of the other nodes in the network and store the received replicated weight shard portions in storage(s)40. Thus, each node 10 may hold a different replicated weight shard portion received (e.g., derived from) each of the other nodes 10.
[0048] Upon detection of a node failure, according to an optional example, each remaining functional node 10 can update its corresponding weight shard based on a replicated weight shard portion held at that functional node 10 and associated with the failed node. As an illustrative example, node 10A can execute resilient framework 28 to update weight shard 52A using the replicated weight shard portion received by node 10A from the failed node prior to the failure detection. In the example, weight shard 52A can be updated by merging the replicated weight shard portion with the local weight shard 52A to produce an updated weight shard. Updates to local weight shards can be performed at each node 10 that remains functional and involved in training. Thus, along with maintaining the optimizer shards of the failed node as described above, according to this example, the weights contained in the weight shards of the failed node can be preserved and maintained in the updated weight shards of the remaining nodes 10. The updated local weight shards of node 10 can be used by the ML framework 24 to continue distributed training without interruption from detected faults.
[0049] In some implementations, node 10A may include encapsulation and deployment 50, which can encapsulate and deploy model 44 as a containerized object. For example, encapsulation and deployment 50 can encapsulate (multiple) local model parameters and other inference into a containerized object, which can be shared with other participant nodes 10 via interface layer 26. For example, but not limited to, encapsulation and deployment 50 can use a Docker platform to generate Docker files including model 44. In another example, encapsulation and deployment 50 can use a Docker platform to generate Docker files including shard portions of weight shard 52A, optimizer shard 54A, and / or shard portions of weight shard 52A. Other containerization platforms may also be used. In this way, various applications at node 10 can access and use model 44 in a platform-independent manner. Thus, the model can not only be built based on common parameters from nodes in a distributed training network, but can also be encapsulated and deployed in various environments.
[0050] Figure 2 This is a schematic block diagram illustrating the processing flow of the forward propagation phase 200 in distributed training, implemented according to examples of this disclosure. Figure 2 In the example, the forward propagation phase 200 can be performed by multiple nodes 10, such as in combination. Figure 1 As described. Therefore, one or more operations in the forward propagation phase 200 can be performed by, for example, one or more of the application layer 22, ML framework 24, interface layer 26, and / or resilient framework 28, as performed by processor(s) 20. Figure 2 In the example shown, forward propagation phase 200 is illustratively depicted as being performed by four nodes 10A-10D. However, forward propagation phase 200 can be performed by any number of nodes required for a given machine learning application.
[0051] This can be applied to ML models (e.g., Figure 1 The forward propagation phase 200 process is executed iteratively for each layer of model 44 to obtain the output of each layer based on the input applied by each node 10. The output obtained from one iteration can be used as the input for the next iteration of the forward propagation phase 200. The forward propagation phase 200 includes multiple operations that... Figure 2 The process is illustrated as being grouped into steps 210 and 220. The forward propagation phase 200 may execute steps 210 and 220 for each layer of the ML model, starting from the first layer (e.g., the input layer) and iterating through multiple intermediate layers (e.g., hidden layers) to the last layer (e.g., the output layer) in the order of the layers.
[0052] Before the first iteration of the forward propagation phase 200 at the first layer, registration can occur, whereby each node 10A-10D can register or enroll itself for distributed learning. In one example, this can be a one-time process. In other examples, registration or enrollment can be performed as a type of verification process after a period of time. In the example, each node 10 can subsequently record its associated attributes in the learning contract, such as a Uniform Resource Locator (URL), from which the model parameter set local to node 10 can be downloaded by other nodes.
[0053] Additionally, prior to the first iteration of the forward propagation phase 200, hyperparameters can be loaded from storage device 40 into, for example, the ML framework 24 for each node. As mentioned above, the hyperparameters define how the ML framework 24 and the resilient framework 28 are constructed. Hyperparameters can be selected for use at each node, and training can be performed based on these hyperparameters. In the example, hyperparameters can govern the training process, for example by specifying how many nodes (e.g., nodes 10A-10D) will perform training; how many weight shards will be created; how many replicated shard portions will be created by each node; how node failures are detected or what process is considered a node failure, etc.
[0054] In the example, the trained ML model can be constrained by a common ML algorithm that includes various model parameters (such as weight states and optimizer states). Each layer can be constrained by a set of model parameters. As described above, the weights constraining each layer of the ML model can be partitioned into multiple weight slices. Figure 2In this example, the common weights are divided into four weight slices 52A-52D, which are assigned and stored at each node 10A-node 10D. For example, weight slice 52A is assigned to node 10A, weight slice 52B to node 10B, weight slice 52C to node 10C, and weight slice 52D to node 10D. Therefore, each node 10A-node 10D can hold weight slice 52A-weight slice 52D, and the corresponding node 10A-node 10D is responsible for performing a transformation on the input to compute the usage in the output for each given layer. In this example, weight slices 52A-weight slice 52D collectively comprise the weights that define the entire layer.
[0055] Similarly, in Figure 2 In this example, the common optimizer state is divided into four optimizer shards 54A-54D, which are allocated and stored at each node 10A-10D. For example, optimizer shard 54A is allocated to node 10A, optimizer shard 54B to node 10B, optimizer shard 54C to node 10C, and optimizer shard 54D to node 10D. In this example, optimizer shards 54A-54D collectively comprise the optimizer state of the entire layer.
[0056] In the example, at step 210, for a given layer, each node 10A-10D reconstructs the corresponding layer. For example, the ML framework 24 at each node 10A-10D performs a full collection operation 212 to obtain weight shards from other nodes and reconstructs the complete layer based on the collected weights. For example, node 10A performs a full collection operation 212 to obtain weight shards 52B-52D and stores weight shards 52A-52D in the storage(s) of node 10A. The ML framework 24 can access the storage(s) of the storage(s) to retrieve the weights of each shard and reconstruct the complete layer by applying the weights of the various weight shards to a common ML algorithm. Similarly, node 10B obtains weight slices 52A, 52C, and 52D; node 10C obtains weight slices 52A, 52B, and 52D; and node 10D obtains weight slices 52A-52C to construct a complete layer at each corresponding node. Due to the utilization of weight slices 52A-52D used for training, weight slices 52A-52D can be considered as training slices.
[0057] At step 220, the ML framework 24 of each node 10A-10D can perform operations 222A-222D to obtain the output for the fully constructed layer. For example, as part of operation 222A, node 10A executes ML framework 24 to perform forward computation operations. ML framework 24 can perform forward computation by feeding local inputs to the reconstruction layer of the common ML model and compute (e.g., obtain) the local output for that layer. The reconstruction layer can perform transformations on the local inputs according to the weights of that reconstruction layer. In the case of the first iteration of forward propagation phase 200, ML framework 24 reconstructs the first layer and applies the local training data 48 as input to the first layer to obtain the local output of the first layer. Node 10A can apply the local output obtained for a given reconstruction layer as the local input for the next reconstruction layer during subsequent iterations of forward propagation phase 200. Although the above example is provided with reference to node 10A, each node 10B-10D can perform similar operations to obtain the local output for each layer of the ML model based on its corresponding local inputs. In this way, the forward propagation phase 200 iterations pass through each layer of the ML model.
[0058] Once the output of a given reconstruction layer is obtained, weight shards obtained from other nodes 10A-10D can be discarded to free up space for the next iteration. For example, node 10A can execute ML framework 24 to discard, delete, or otherwise remove weight shards 52B-52D from its(multiple) storage devices 40. Similarly, node 10B can discard weight shards 52A, 52C, and 52D; node 10C can discard weight shards 52A, 52B, and 52D; and node 10D can discard weight shards 52A, 52B, and 52C.
[0059] Furthermore, depending on the hyperparameters, each node 10A-node 10D can be configured to create a copy of a portion of its corresponding weight shard, which can be distributed to other nodes 10A-node 10D. For example, the hyperparameters can configure the resilient framework of each node 10A-node 10D to replicate the corresponding optimizer shard, divide the replicated optimizer shard into multiple replicated optimizer shard portions according to the hyperparameters, and distribute the replicated optimizer shard portions of the weights to other nodes 10A-node 10D. Each replicated optimizer shard portion can include data defining a different portion of the corresponding optimizer shard (e.g., optimizer shard 54A of node 10A), which does not overlap with a portion of any other replicated optimizer shard portion originating from the same optimizer shard. Therefore, the replicated optimizer shard portions can collectively define the complete optimizer shard. The number of replicated optimizer shard portions created can be based on the number of nodes used to perform distributed training. For example, the number of replicated optimizer shard portions can be less than the total number of nodes.
[0060] exist Figure 2 In the example, node 10A can execute elastic framework 28 to generate three replicated optimizer shard portions 54A-1 to 54A-3. For example, elastic framework 28 can access storage device(s) 40 to obtain optimizer shard 54A and create a copy of optimizer shard 54A. Elastic framework 28 can divide the copy of optimizer shard 54A into portions by splitting optimizer shard 54A into three replicated optimizer shard portions 54A-1 to 54A-3. In various examples, the three segments can be substantially equal in size.
[0061] Node 10A can execute elastic framework 28 to distribute replicated optimizer shard portions 54A-1 to 54A-3 to nodes 10B-10D. Each node 10B-10D executes its corresponding elastic framework to receive one replicated optimizer shard portion from 54A-1 to 54A-3 and stores the received replicated optimizer shard portion in the corresponding storage(s) device(s). Figure 2In the example, node 10B holds replicated optimizer shard portion 54A-1, node 10C holds replicated optimizer shard portion 54A-2, and node 10D holds replicated optimizer shard portion 54A-3. Although the examples herein refer to replicated optimizer shard portions 54A-1 through 54A-3, the reference numerals are intended to indicate portions of the whole (e.g., 33% of optimizer shard 54A) and are not intended to assign a sequential order to the replicated optimizer shard portions. Thus, for example, replicated optimizer shard portion 54A-2 could be the first sequential portion of optimizer shard 54A, replicated optimizer shard portion 54A-3 could be the second sequential portion, and replicated optimizer shard portion 54A-1 could be the last sequential portion, or other arrangements as desired.
[0062] In some examples, during the initial distribution of the replicated optimizer shard portion, nodes 10A-10D may distribute the replicated optimizer shard portion before or as part of step 210, for example, using an all-to-all operation performed by the appropriate elastic framework. In some examples, the distribution of the replicated optimizer shard portion may be performed concurrently with operation 212 (e.g., simultaneously or nearly simultaneously). However, these are merely examples, and the initial distribution of the replicated optimizer shard portion can be performed at any point during or before the forward propagation phase, as desired.
[0063] In the example, each node 10A-node 10D can execute its elastic framework to update the replicated optimizer shard portion held on it based on the updated optimizer state obtained during the backpropagation phase. For example, in conjunction with the following... Figure 3 As described, each node 10A-node 10D updates its optimizer fragment 54A-optimizer fragment 54D relative to the global gradient. Then, each node 10A-node 10D can also execute its elastic framework to update the replicated optimizer fragment it holds by receiving updated optimizer fragments from other nodes 10A-node 10D. In the example, compute nodes 10A-node 10D can each execute their elastic framework to... Figure 2 At any desired point during the forward propagation phase shown, perform an all-to-all operation (e.g., as combined with the following text). Figure 3The described all-to-all operation 314). For example, the all-to-all operation can be performed during step 210, such as, for example, with operation 212 serially grounded (e.g., simultaneously or nearly simultaneously). In another example, the all-to-all operation can be performed during step 220, such as, for example, with operations 222A-222D serially grounded (e.g., simultaneously or nearly simultaneously).
[0064] An all-to-all operation performed by a specific compute node can be used to update the replicated optimizer shard portion corresponding to that specific compute node's copy held at other compute nodes. For example, compute node 10A can update optimizer shard 54A during the backpropagation phase. Compute node 10A can replicate the updated optimizer shard 54A, partition the updated optimizer shard 54A, and perform an all-to-all operation that updates the replicated optimizer shard portions 54A-1 to 54A-3 at each compute node 10B-10D. The all-to-all operation may include sending only a portion of the updated optimizer shard 54A to a given compute node holding the corresponding replicated optimizer shard portion. For example, compute node 10A can send an updated instance of the replicated optimizer shard portion 54A-1 to compute node 10B, an updated instance of the replicated optimizer shard portion 54A-2 to compute node 10C, and an updated instance of the replicated optimizer shard portion 54A-3 to compute node 10D. Compute nodes 10B-10D can similarly send updated instances of the corresponding replicated optimizer shard portions to compute nodes 10A-10D.
[0065] In some examples, the hyperparameter can also optionally configure the resilient framework of each node 10A-node 10D to replicate the corresponding weight shards, divide the replicated weight shards into multiple replicated weight shard portions according to the hyperparameter, and distribute the replicated weight shard portions of the weights to other nodes 10A-node 10D. Each replicated weight shard portion may include data defining a different part of the corresponding weight shard (e.g., weight shard 52A of node 10A), which does not overlap with any part of any other replicated weight shard portion originating from the same weight shard. Thus, the replicated weight shard portions can collectively define the complete weight shard. The number of replicated weight shard portions created can be based on the number of nodes used to perform distributed training. For example, the number of replicated weight shard portions can be less than the total number of nodes.
[0066] exist Figure 2In the example, node 10A can execute elastic framework 28 to generate three replicated weight shard portions 52A-1 to 52A-3. For example, elastic framework 28 can access storage device(s) 40 to obtain weight shard 52A and create copies of weight shard 52A. Elastic framework 28 can divide the copy of weight shard 52A into multiple parts by splitting weight shard 52A into three replicated weight shard portions 52A-1 to 52A-3. In various examples, the three segments can be substantially equal in size.
[0067] According to this example, node 10A can execute elastic framework 28 to distribute replicated weight shard portions 52A-1 to 52A-3 to nodes 10B-10D. Each node 10B-10D executes its corresponding elastic framework to receive one replicated weight shard portion from replicated weight shard portions 52A-1 to 52A-3 and stores the received replicated weight shard portion in the corresponding storage(s) device(s). Figure 2 In the example, node 10B holds the replicated weight shard portion 52A-1, node 10C holds the replicated weight shard portion 52A-2, and node 10D holds the replicated weight shard portion 52A-3. Although the examples herein refer to replicated weight shard portions 52A-1 through 52A-3, the reference numerals are intended to indicate portions of the whole (e.g., 33% of weight shard 52A) and are not intended to assign a sequential order to the replicated weight shard portions. Thus, for example, replicated weight shard portion 52A-2 could be the first part of the sequence of weight shard 52A, replicated weight shard portion 52A-3 could be the second part of the sequence, and replicated weight shard portion 52A-1 could be the last part of the sequence, or other arrangements as desired.
[0068] In some examples, as part of operation 212, nodes 10A-10D can distribute replicated weight slices, for example, as part of a full collection operation performed by the corresponding ML framework. In this example, by distributing the replicated weight slices using the same operation performed as part of the training process, no additional communication overhead is required. In another example, the replicated weight slices can be distributed before the first iteration of forward propagation phase 200.
[0069] In the example where the replicated weight slices are distributed among the computation nodes, each node 10A-node 10D can execute its resilient framework to update the replicated weight slices it holds based on the fully reconstructed layer during step 210. For example, as described above in conjunction with step 210, each node 10A-node 10D collects weight slices held at other nodes to reconstruct the layer. Each node 10A-node 10D can also execute (e.g., during step 220) its resilient framework to update the replicated weight slices using the weight slices received during step 210. For example, the replicated weight slices may contain weights learned during previous iterations of training the current layer that may need to be updated to the most recently learned weights. The weight slices received during step 210 may contain the most recent weights for that layer, which can be used to update the replicated weight slices. For example, node 10A can acquire weight shards 52B-52D during full collection operation 212 and execute elastic framework 28 to update the replicated weight shard portion 52B-1, replicated weight shard portion 52C-1, and replicated weight shard portion 52D-1 respectively using the corresponding weights contained in weight shards 52B-52D. Nodes 10B-10C can similarly update their held replicated weight shard portions respectively during step 220.
[0070] Figure 3 This is a schematic block diagram illustrating the processing flow for the backpropagation phase 300 in distributed training, implemented according to an example of this disclosure. Figure 3 In the example, backpropagation phase 300 can be executed by multiple nodes 10, such as in combination. Figure 1 As described. Therefore, one or more operations in the backpropagation phase 300 can be performed by, for example, one or more of the application layer 22, ML framework 24, interface layer 26, and / or elastic framework 28 (as performed by processor(s) 20). In Figure 3 In the example shown, backpropagation phase 300 is illustratively depicted as being performed by nodes 10A-10D, which can be used to perform... Figure 2 The same node in the forward propagation phase 200. Although Figure 3 The example illustration shows four nodes, but backpropagation phase 300 can be performed by any number of nodes.
[0071] This can be applied to ML models (e.g., Figure 1 Each layer of model 44) performs Figure 3The backpropagation stage 300 shown obtains the gradient of a given layer relative to its weights from the loss function through layer-by-layer backward iteration from the last layer to the first layer of the ML model. The backpropagation stage 300 then uses the gradient to update the weight state and optimizer state based on it. The backpropagation stage 300 includes several operations that... Figure 3 The process is illustratively described as being grouped into steps 310, 320, 330, and 340, which can be performed iteratively for each layer. In the example, it can be... Figure 2 After the forward propagation phase 200, the backward propagation phase 300 is executed.
[0072] In the example, at step 310, for a given layer, each node 10A-node 10D reconstructs the corresponding layer of the common ML model. For example, the ML framework at each node 10A-node 10D can be executed to perform a full collection operation 312 to collect the weight shards held at other nodes and reconstruct the complete layer based on the collected weights. In various examples, step 310 can be substantially similar to... Figure 2 The procedure is executed in the manner described in step 210.
[0073] At step 320, each node 10A-node 10D can execute its corresponding ML framework to perform operations 322A-operation 322D. Operations 322A-operation 322D may include performing a backward computation operation to obtain the input gradient and weight gradient of the fully reconstructed layer relative to the local input at each node 10A-node 10D. For example, as part of operation 322A, node 10A can execute ML framework 24 to perform a backward computation operation to obtain the input gradient and weight gradient relative to the input locally held at node 10A, which was utilized during forward propagation phase 200 (and stored in storage device 40). The backward computation operation may utilize, for example, but not limited to, gradient descent or variants (such as stochastic gradient descent), to obtain the input gradient and weight gradient based on the local input of the current layer relative to the loss function between the current layer and layers sequentially following the current layer in the ML model. Similarly, as part of operations 322B-322D, nodes 10B-10D each execute their respective ML frameworks to perform inverse computation operations to obtain the input gradients and weight gradients. The inverse computation operations performed during each of operations 322A-322D can utilize the same algorithm (e.g., gradient descent or its variants, such as stochastic gradient descent) or different algorithms, depending on the desired application.
[0074] Once the input gradients and weight gradients for a given reconstruction layer are obtained, weight fragments obtained from other nodes 10A-10D can be discarded to free up space for the next iteration. For example, node 10A can execute ML framework 24 to discard, delete, or otherwise remove weight fragments 52B-52D from its(multiple) storage devices 40. Similarly, node 10B can discard weight fragments 52A, 52C, and 52D; node 10C can discard weight fragments 52A, 52B, and 52D; and node 10D can discard weight fragments 52A, 52B, and 52C.
[0075] At step 330, each node 10A-node 10D can execute its respective ML framework to distribute global weight gradients among nodes 10A-node 10D via operation 332. Operation 332 may include a reduction scattering operation. For example, operation 332 performs a reduction operation that collects the set of local weight gradients obtained by each node 10A-node 10D and aggregates these local weight gradient sets together to produce a global weight gradient set. For example, during step 320, each node 10A-node 10D computes weight gradients relative to the local inputs of each node for the fully reconstructed layer. Thus, each node 10A-node 10D obtains local weight gradients relative to different inputs for each weight slice 52A-weight slice 52D, which are local to each node 10A-node 10D. Each set of local weight gradients obtained at a given node may include one or more local weight gradients corresponding to each weight slice 52A-weight slice 52D. The reduction operation of operation 322 collects the local weight gradient set from each node 10A-10D and aggregates (e.g., sums) the local weight gradient sets based on the weight sets. For example, the local weight gradients corresponding to weight piece 52A can be obtained from each node 10A-10D, and these local weight gradients can be summed together (or by other aggregation functions, such as, but not limited to, average, minimum, maximum, etc.) to obtain the global weight gradient for weight piece 52A. Similarly, the global weight gradient can be obtained for each weight piece 52B-52D, and weight pieces 52B-52D together with weight piece 52A can constitute the global weight gradient set.
[0076] Then, the global weight gradient set can be scattered across each node from node 10A to node 10D. For example, the global weight gradient for each weight slice can be scattered to the node associated with the corresponding slice. As an illustrative example, the global weight gradient for weight slice 52A can be scattered to node 10A, the global weight gradient for weight slice 52B can be scattered to node 10B, the global weight gradient for weight slice 52C can be scattered to node 10C, and the global weight gradient for weight slice 52D can be scattered to node 10D.
[0077] At step 340, each node 10A-10D can execute its respective ML framework to update the weights of the corresponding weight slice relative to the global weight gradient obtained at step 330. For example, node 10A can execute ML framework 24 to perform operation 342, which updates the weights of weight slice 52A relative to the global weight gradient obtained from step 330 for weight slice 52A. Similarly, node 10B can update the weights of weight slice 52B, node 10C can update the weights of weight slice 52C, and node 10D can update the weights of weight slice 52D. In this example, operation 342 performed by each node 10A-10D may include an optimization algorithm that uses the optimizer states contained in optimizer slices 54A-54D held at the respective node to compute the updated weights relative to the global weight gradient. For example, node 10A can hold an optimizer slice of the optimizer state, which can be applied as a local instance of the optimization algorithm (e.g., local relative to node 10A) to update weight slice 52A relative to the global weight gradient. Similarly, nodes 10B-10D can each hold an optimizer slice that can be used to update weight slice 10B-10D, respectively.
[0078] At step 340, each node 10A-10D can also execute its corresponding ML framework to update the optimizer state of the corresponding optimizer shard relative to the global weight gradient obtained at step 330. For example, node 10A can execute ML framework 24 to perform operation 342, which may include updating the optimizer state of optimizer shard 54A relative to the weight gradient obtained from step 330 (e.g., a portion of the global weight gradient corresponding to the corresponding weight gradient of node 10A). Similarly, node 10B can update the optimizer state of optimizer shard 54B, node 10C can update the weights of optimizer shard 54C, and node 10D can update the optimizer state of optimizer shard 54D. In this example, optimizer shards may contain local optimizer states of optimizer algorithms that are common to computational nodes 10A-10D corresponding to a common ML model.
[0079] In the examples disclosed in this article, such as Figure 3 As shown, each node 10A-node 10D can hold the weighted shard portion and the optimizer shard portion of the replication associated with other nodes. As described above, each node 10A-node 10D can execute its elastic framework to create and distribute the weighted shard portion and the optimizer shard portion of the replication. Additionally, as described above... Figure 2 As described, during step 320, each node 10A-node 10D can execute its elastic framework to perform operations that update the replicated weight shard portion held on it based on the weight shards obtained from other nodes during step 310.
[0080] Furthermore, each node 10A-node 10D can execute its elastic framework to update the replicated optimizer shard portion held thereon based on the updated optimizer state obtained at step 340 of the previous iteration of backpropagation. For example, each node 10A-node 10D can execute its elastic framework to perform operation 314, which may include receiving the updated optimizer shard portion from each of the other nodes 10A-node 10D, and including updating the replicated optimizer shard portion held at each corresponding compute node 10A-node 10D. As an illustrative example, compute node 10A can execute its elastic framework 28 to perform operation 314, which obtains updated optimizer states for optimizer shard replicas 54B-1, 54C-1, and 54D-1 from each compute node 10B-10D. The elastic framework 28 can then use the corresponding optimizer states to update each replicated optimizer shard replica 54B-1, 54C-1, and 54D-1. Similarly, the elastic framework can be executed for each compute node 10B-10D to update the respective replicated optimizer shard replicas.
[0081] According to various examples, operation 314 can be an all-to-all operation performed by each compute node 10A-10D. For example, compute node 10A can update optimizer shard 54A at step 340. Compute node 10A can copy the updated optimizer shard 54A, partition the updated optimizer shard 54A, and perform the all-to-all operation as operation 314, which updates the copied optimizer shard portions 54A-1 to 54A-3 at each compute node 10B-10D. The all-to-all operation may include sending only a portion of the updated optimizer shard 54A (e.g., optimizer state) to a specific compute node that holds the corresponding copied optimizer shard portion. For example, compute node 10A can transmit the updated optimizer state for the replicated optimizer shard portion 54A-1 to compute node 10B, the updated optimizer state for the replicated optimizer shard portion 54A-2 to compute node 10C, and the updated optimizer state for the replicated optimizer shard portion 54A-3 to compute node 10B. Compute nodes 10B-10D can similarly transmit the updated optimizer state for the corresponding replicated optimizer shard portion to compute nodes 10A-10D.
[0082] exist Figure 3 In the example, operation 314 is executed in conjunction with operation 312 (e.g., simultaneously or nearly simultaneously). However, the implementation disclosed herein is not intended to be limited to this example. Operation 314 can be performed in other ways. Figure 3 It is executed at any desired point during the backpropagation phase, and at any point during the forward propagation phase, as described above. Figure 2 As described. For example, operation 314 may be performed during step 320, such as, for example, being performed simultaneously or nearly simultaneously with operations 222A-222D. As another example, operation 314 may be performed during step 330, such as, for example, being performed simultaneously or nearly simultaneously with operations 322.
[0083] Figure 4 This is a schematic block diagram depicting a handling flow 400 for providing resilience to node failures in distributed training, implemented according to examples of this disclosure. Figure 4 In the example, process 400 can be executed by multiple nodes 10, such as in combination. Figure 1 As described. Therefore. Figure 4One or more of the operations shown may be performed by, for example, one or more of the application layer 22, ML framework 24, interface layer 26, and / or elastic framework 28 (such as by processor(s) 20). Figure 4 The processing flow 400 is illustratively described as being executed by four nodes 10A-10D. However, by Figure 4 The flexibility provided by the processing flow 400 shown can be executed by any number of nodes.
[0084] In the example, processing flow 400 can be performed at any point during distributed training by utilizing replicated optimizer shard portions held at each of the nodes from node 10A to node 10D, for example, during forward propagation phase 200 and / or backpropagation phase 300. By distributing the optimizer shard held at one node (e.g., optimizer shard 54A held by node 10A) across other nodes (e.g., nodes 10B to 10D) as replicated optimizer shard portions (e.g., replicated optimizer shard portions 54A-1 to replicated optimizer shard portions 54A-3), the example in this document can, for example, be performed at the failure of node 10A (e.g., as combined above). Figure 1 This provides resilience in the event of functional failures (such as those described). Similarly, the examples in this paper can be resilient to failures of other nodes by distributing the corresponding replicated optimizer shard portions.
[0085] For example, at step 410, a functional failure of a node can be detected based on an inspection (e.g., rule 46), as described above. Figure 1 As described. In the illustrative example, node 10A may fail for any reason or become unavailable for distributed training, thus rendering node 10A non-participatory and / or inactive (as confirmed by the "X"). The remaining functional nodes 10B to 10D can be notified of the failure by any desired notification technique, as described above, which can be considered as detecting the failure of node 10A.
[0086] Based on the detected functional failure (e.g., in response to the detection that node 10A has failed), each remaining functionally healthy node (e.g., nodes 10B-10D) can update its optimizer shard to include a portion of the optimizer shard corresponding to the failed node. For example, at step 420, each functionally healthy node can execute its elastic framework to move the portion of the optimizedr shard corresponding to the failed node (e.g., node 10A) to the corresponding optimizer shard of that functionally healthy node. At step 430, each functionally healthy node can execute its elastic framework to update the corresponding optimizer shard of that functionally healthy node by merging the portion of the optimizedr shard corresponding to the failed node with the corresponding optimizer shard of that functionally healthy node. The resulting updated optimizer shard can then be used for training (e.g., as an updated training shard).
[0087] In the illustrative implementation, each remaining functional node (e.g., node 10B-node 10D) can perform step 420 by executing the appropriate resilient framework to locate the replicated optimizer shard portion corresponding to the failed node. In some examples, the replicated optimizer shard portion may be stored at each node in association with a unique identifier of the compute node (e.g., MAC address, IP address, or any other unique identifier), and the replicated optimizer shard portion originates from (e.g., is received from) that compute node. That is, for example, replicated optimizer shard portion 54A-1 may be stored in one or more storage devices at node 10B, and this replicated optimizer shard portion 54A-1 is tagged or otherwise associated with the unique identifier of node 10A; replicated optimizer shard portion 54A-2 may be stored in one or more storage devices at node 10C, and this replicated optimizer shard portion 54A-2 is tagged or otherwise associated with the unique identifier of node 10A; and replicated optimizer shard portion 54A-3 may be stored in one or more storage devices at node 10D, and this replicated optimizer shard portion 54A-3 is tagged or otherwise associated with the unique identifier of node 10A. Each replicated optimizer shard portion originating from other nodes may similarly be tagged or associated with the identifier of the source node. Therefore, when a node failure is detected, the resilient framework of the remaining functional nodes can locate the replicated optimizer shard portion corresponding to the failed node in the corresponding storage devices(s).
[0088] Once each remaining functional node (e.g., nodes 10B-10D in this example) locates a portion of the replicated optimizer shard received from the failed node (e.g., node 10A in this example) (e.g., replicated optimizer shard portions 54A-1 to 54A-3), the remaining functional nodes can execute their respective elastic frames to move the located replicated optimizer shard portion to the corresponding optimizer shard. The elastic frame of each remaining functional node can then operate to merge the replicated optimizer shard portion with the corresponding optimizer shard of that functional node, thereby generating an updated optimizer shard. For example, as... Figure 4 As shown in step 430, node 10B can execute its elastic framework to merge the replicated optimizer shard portion 54A-1 with optimizer shard 54B to generate an updated optimizer shard 64B. Similarly, node 10C can merge the replicated optimizer shard portion 54A-2 with optimizer shard 54C to generate an updated optimizer shard 64C, and node 10D can merge the replicated optimizer shard portion 54A-3 with optimizer shard 54D to generate an updated optimizer shard 64D.
[0089] Therefore, optimizer shard 54A held by the failed node 10A can be preserved by merging it with other optimizer shards of the remaining functional nodes. Thus, the optimizer state of optimizer shard 54A can be maintained and updated by the remaining functional nodes 10B-10D, and used to update weights. Therefore, distributed machine learning performed by nodes 10A-10D can be resilient to the failure of node 10A, and can be performed uninterruptedly via nodes 10B-10D using updated optimizer shards 64B-64D as updated training shards.
[0090] Based on various examples, for instance, since each replicated optimizer shard portion 54A-1 to replicated optimizer shard portion 54A-3 is substantially equal in size, the updated weight shards 64B to 64D can be increased in size by the same amount. Therefore, the workload and communication distributed to the remaining functional nodes can be shared equally, and individual nodes can remain unloaded.
[0091] In some examples, additional resilience to node failures can be provided by distributing the weight shard held at a node (e.g., weight shard 52A held by node 10A) into replicated weight shard portions (e.g., replicated weight shard portions 52A-1 to replicated weight shard portions 52A-3).
[0092] exist Figure 4In the illustrative example shown, based on the functional failure detected at step 410, the weight slice includes a portion of the weight slice corresponding to the failed node. For example, at step 420, each functional node can execute its resilient framework to move the portion of the weight slice corresponding to the failed node (e.g., node 10A) to the corresponding weight slice of the functional node. In this example, step 430 can also update the corresponding weight slice of the functional node by merging the portion of the weight slice corresponding to the failed node with the corresponding weight slice of the functional node. The resulting updated weight slice can then be used for training (e.g., as an updated training slice).
[0093] According to the illustrative implementation of this example, step 420 may optionally include each remaining functional node (e.g., nodes 10B-10D) executing a corresponding resilient framework to locate the replicated weight shard portion corresponding to the failed node. In some examples, the replicated weight shard portion may be stored at each node in association with a unique identifier of the compute node, the replicated weight shard portion originating from that compute node. That is, for example, replicated weight shard portion 52A-1 may be stored in one or more storage devices at node 10B, replicated weight shard portion 52A-1 being tagged or otherwise associated with a unique identifier of node 10A; replicated weight shard portion 52A-2 may be stored in one or more storage devices at node 10C, replicated weight shard portion 52A-2 being tagged or otherwise associated with a unique identifier of node 10A; and replicated weight shard portion 52A-3 may be stored in one or more storage devices at node 10D, replicated weight shard portion 52A-3 being tagged or otherwise associated with a unique identifier of node 10A. Each replicated weighted shard portion originating from other nodes can be similarly tagged or associated with the source node's identifier. Therefore, when a node failure is detected, the resilient framework of the remaining functional nodes can locate the weighted shard portion corresponding to the failed node's replicate in the corresponding storage(s) devices.
[0094] Once each remaining functional node (e.g., nodes 10B-10D in this example) locates a replicated weight shard portion (e.g., replicated weight shard portions 52A-1 to 52A-3) received from the failed node (e.g., node 10A in this example), step 420 may include each of the remaining functional nodes executing its respective elastic framework to move the located replicated weight shard portion to its corresponding weight shard. Then, step 430 may include each remaining functional node executing its respective elastic framework to merge the replicated weight shard portion with its corresponding weight shard, thereby generating an updated weight shard. Figure 4 In the example shown, in step 430, node 10B can execute its resilient framework to merge the replicated weight shard portion 52A-1 with weight shard 52B to generate an updated weight shard 62B. Similarly, node 10C can merge the replicated weight shard portion 52A-2 with weight shard 52C to generate an updated weight shard 62C, and node 10D can merge the replicated weight shard portion 52A-3 with weight shard 52D to generate an updated weight shard 62D.
[0095] Therefore, weight shard 52A held by the failed node 10A can be preserved by merging it with other weight shards of the remaining functional nodes. Thus, the weights of weight shard 52A can be maintained and updated by the remaining functional nodes 10B-10D. Therefore, the distributed machine learning performed by nodes 10A-10D can be further resilient to the failure of node 10A, and can be performed uninterruptedly via nodes 10B-10D using the updated weight shards 62B-62D as updated training shards.
[0096] According to various examples, for instance, since the weighted shard portions 52A-1 to 52A-3 of each replication are substantially equal in size, the workload and communication distributed to the remaining functional nodes can be shared equally, and individual nodes can be kept from being overloaded.
[0097] To ensure that distributed training is not only affected by the failure of the current node (e.g., combined with...) Figure 4 The described node 10A is resilient and also resilient to subsequent node failures (e.g., Figure 4 The implementation of this disclosure is resilient to the example of node 10B-node 10D, and the replicated optimizer shard portion can be updated based on the updated optimizer shard. More specifically, the replicated optimizer shard portion can be, for example, in combination with... Figure 3The backpropagation phase 300 discussed during and / or in combination Figure 2 The operations 314 performed during the forward propagation phase 200 discussed are updated. That is, for example, the optimized state of the updated optimizer shards 64B-64D can be copied, partitioned, and distributed to nodes using an all-to-all operation. As described above, the optimized state contained in the distributed optimizer shards can be used to update the corresponding replicated optimizer shard portion at each node.
[0098] Additionally, in some examples, the replicated weight shard portion can be updated based on the updated weight shard. For example, the replicated weight shard portion can be updated as part of, for instance, a full collection operation performed during forward propagation phase 200 and / or backpropagation phase 300. That is, updated weight shards 62B-62D can be replicated, partitioned, and distributed during steps 210 and / or 310 as weight shards are distributed among nodes to reconstruct a given layer. As described above, the weights contained in the distributed weight shards can then be used to update the corresponding replicated weight shard portion at each node.
[0099] For example, Figure 5 A schematic block diagram illustrating a process flow that ensures continuous resilience against node failures, according to an example of this disclosure, is shown. Figure 5 The processing flow shown can be followed by 500. Figure 4 The optimizer partition update at step 430. Thus, in Figure 5 In the example, process 500 is executed by the remaining functional nodes 10B to 10D. Therefore, Figure 5 One or more of the operations shown can be performed by, for example, one or more of the application layer 22, ML framework 24, interface layer 26, and / or elastic framework 28, as performed by processor(s) 20. Although the process flow 500 is illustratively depicted as being performed by three nodes 10B-10D, the elasticity provided by the process flow 500 can be performed by any number of nodes.
[0100] As above combined Figure 4As described, node 10A may have failed, and nodes 10B-10D may have generated corresponding updated optimizer shards 64B-64D based on the replicated optimizer shard portions received from node 10A. However, at the beginning of step 510, the replicated optimizer shard portions held at each remaining functional node (except for the replicated optimizer shard portion corresponding to node 10A) may remain unchanged. That is, for example, node 10B may hold replicated optimizer shard portions 54C-2 and 54D-2, node 10C may hold replicated optimizer shard portions 54B-2 and 54D-3, and node 10D may hold replicated optimizer shard portions 54B-3 and 54C-3. Therefore, replicated optimizer shard portion 54B-1, replicated optimizer shard portion 54C-1, and replicated optimizer shard portion 54D-1 may be lost due to a failure of node 10A, and the replicated optimizer shard portions for optimizer shards 64B-64D may be incomplete or nonexistent (for example, the portions of optimizer shards 64B-64D corresponding to replicated optimizer shard portions 54A-1 to replicated optimizer shard portions 54A-3 may not be represented in the replicated optimizer shards currently held at the node).
[0101] To ensure that the distributed training performed by the remaining functional nodes 10B-10D is resilient to future node failures, each remaining functional node 10B-10D can execute its resilience framework to update the replicated optimizer shard portion via steps 510 and 520. Steps 510 and 520 can be executed as part of forward propagation phase 200 or backpropagation phase 300 following a node failure (e.g., node 10A). In each case, forward propagation phase 200 or backpropagation phase 300 can be executed as described above, except that only the remaining functional nodes (e.g., nodes 10B-10D, since node 10A no longer exists in the process) are utilized, and forward propagation phase 200 or backpropagation phase 300 is modified as described below to update the replicated optimizer shard portion.
[0102] For example, at step 510, each node 10B-node 10D can execute its elastic framework to perform a full collection operation 512 to collect the weight shards held at other nodes and reconstruct the layer based on the collected weights. The full collection operation 512 can be either the full collection operation 212 or the full collection operation 312 described above. Therefore, after operation 512, each node 10B-node 10D holds weight shard 62B-weight shard 62D.
[0103] exist Figure 5 In the example, at step 510, each node 10B-10D can execute its elastic framework to perform operation 514, which distributes a portion of that node's updated optimizer shard 64B-update optimizer shard 64D to other nodes. For example, compute node 10B can replicate the updated optimizer shard 64B, partition the updated optimizer shard 64B, and perform operation 514 (e.g., an all-to-all operation) to supply portions of the updated optimizer shard 64B (e.g., portions of the optimizer state) to compute nodes 10C and 10D. Compute nodes 10C and 10D can similarly supply portions of the updated optimizer shard 64C and updated optimizer shard 64D to compute nodes 10B-10D.
[0104] For example, each node 10B-10D utilizes a portion of the updated optimizer shard 64B-updated optimizer shard 64D obtained during step 510 to update the replicated optimizer shard portion held at each node 10B-10D. The updated optimizer shard 64B-updated optimizer shard 64D received at step 510 may contain the latest optimizer state for that layer, which can be used during step 520 to update the replicated optimizer shard portion and adjust its size. For example, during operation 514, node 10B may obtain portions of updated optimizer shard 64C and updated optimizer shard 64D, and execute node 10B's elastic framework to update replicated optimizer shard portions 54C-2 and 54D-2 to include portions of the optimizer state of updated optimizer shards 64C and 64D, thereby creating updated replicated optimizer shard portions 64C-1 and 64D-1. Similarly, node 10C can update the replicated optimizer shard portion 54B-2 and the replicated optimizer shard portion 54D-3 to create updated replicated optimizer shard portions 64B-1 and 64D-2, and node 10D can update the replicated optimizer shard portion 54B-3 and the replicated optimizer shard portion 54C-3 to create updated replicated optimizer shard portions 64B-2 and 64C-2.
[0105] In the example, updating the replicated optimizer shard portion may include any technique for updating the replicated optimizer shard portion using portions of updated optimizer shards 64B-64D obtained from the compute node. In one example, updating the replicated optimizer shard portion may include replacing previously replicated optimizer shard portions (e.g., replicated optimizer shard portion 54C-2, replicated optimizer shard portion 54D-2) with portions of the updated optimizer shards (e.g., portions of optimizer shards 64C and 64D). In the example, compute node 10B may delete replicated optimizer shard portions 54C-2 and 54D-2, and store portions of optimizer shards 64C and 64D as replicated optimizer shard portions 64C-1 and 64D-1. Compute nodes 10C and 10D can perform similar operations to produce replicated optimizer shard portions 64B-1, 64D-2, 64B-2, and 64C-2. In another example, updating a replicated optimizer shard portion may include merging the obtained updated optimizer shard portion with the replicated optimizer shard portion stored thereon. In any case, the updated replicated optimizer shard portions may collectively form the entire updated optimizer shard. That is, for example, replicated optimizer shard portion 64B-1 and replicated optimizer shard portion 64B-2 together represent optimizer shard 64B, replicated optimizer shard portion 64C-1 and replicated optimizer shard portion 64C-2 together represent optimizer shard 64C, and replicated optimizer shard portion 64D-1 and replicated optimizer shard portion 64D-2 together represent optimizer shard 64D.
[0106] Additionally, in the example of the optimizer shard portion 64B-64D of the replicated optimizer shard portion generated from compute node 10B to compute node 10D, as described above... Figure 4As described in the example, at the beginning of step 510, the replicated weight shard portion held at each remaining functional node (except for the replicated weight shard portion corresponding to node 10A) may remain unchanged. That is, for example, node 10B may hold replicated weight shard portions 52C-2 and 52D-2, node 10C may hold replicated weight shard portions 52B-2 and 52D-3, and node 10D may hold replicated weight shard portions 52B-3 and 52C-3. Therefore, replicated weight shard portion 52B-1, replicated weight shard portion 52C-1, and replicated weight shard portion 52D-1 may be lost due to a failure of node 10A, and the replicated weight shard portions for weight shards 62B-62D may be incomplete or non-existent (for example, the portions of weight shards 62B-62D corresponding to replicated weight shard portions 52A-1 to replicated weight shard portions 52A-3 may not be represented in the replicated weight shards currently held at the node).
[0107] In this example, each remaining functional node 10B-10D can execute its resilient framework to update the replicated weight shard portion via steps 510 and 520. For example, as discussed above, at step 510, each node 10B-10D performs a full collection operation 512 to collect the weight shards held at other nodes and reconstruct the layer based on the collected weights. Thus, after operation 512, each node 10B-10D holds weight shard 62B-weight shard 62D. At step 520, according to this example, each node 10B-10D can execute its resilient framework to perform operation 522B-operation 522D. Operation 522B-operation 522D may include updating the replicated weight shard portion held at each node 10B-10D using the weight shards received during step 510. Operations 522B-522D can be performed as part of either Operations 222B-222D or Operations 322B-322D as described above, which may include combinations as described above. Figure 2 and Figure 3 The described update replication weighted shard portion.
[0108] For example, each node 10B-10D can perform a corresponding operation 522B-operation 522D to update the replicated weight shard portion held at each node 10B-10D using the updated weight shard 62B-weight shard 62D obtained during step 510. The updated weight shard 62B-update received at step 510 may contain the latest weights for that layer, which can be used during step 520 to update the replicated weight shard portion and adjust its size. For example, node 10B can acquire weight shards 62C and 62D during the full collection operation 512, and execute node 10B's elastic framework to update the replicated weight shard portion 52C-2 and replicated weight shard portion 52D-2 to include the weights of the updated weight shards 62C and 62D, thereby creating updated replicated weight shard portions 62C-1 and 62D-1. Similarly, node 10C can update the replicated weighted shard portion 52B-2 and the replicated weighted shard portion 52D-3 to create updated replicated weighted shard portions 62B-1 and 62D-2, and node 10D can update the replicated weighted shard portion 52B-3 and the replicated weighted shard portion 52C-3 to create updated replicated weighted shard portions 62B-2 and 62C-2.
[0109] Figure 6 The illustration shows example computing components, according to various embodiments, that can be used to implement node failure resilience in distributed training. Reference is now made to... Figure 6 The computing component 600 can be, for example, a server computer, a controller, or any other similar computing component capable of processing data. Figure 6 In the example implementation, computing component 600 includes a hardware processor 602 and a machine-readable storage medium 604. In the example, computing component 600 may be... Figure 1 An example of node 10 in node 10. In another example, hardware processor 602 may be a plurality of hardware processors coupled to a plurality of machine-readable storage media, which may represent Figure 1 Multiple nodes 10.
[0110] Hardware processor 602 may be one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for retrieving and executing instructions stored in machine-readable storage medium 604. Hardware processor 602 may retrieve, decode, and execute instructions (such as instructions 606-612) to control processes or operations for fail-safe resilience. Alternatively to, or in addition to, retrieving and executing instructions, hardware processor 602 may include one or more electronic circuits comprising electronic components for the function of executing one or more instructions, such as, but not limited to, graphics processing units (GPUs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), or other electronic circuits.
[0111] Machine-readable storage media (such as machine-readable storage media 604) can be any electronic, magnetic, optical, or other physical storage device that contains or stores executable instructions. Therefore, machine-readable storage media 604 can be, for example, random access memory (RAM), non-volatile RAM (NVRAM), electrically erasable programmable read-only memory (EEPROM), storage devices, optical discs, etc. In some embodiments, machine-readable storage media 604 can be a non-transitory storage medium, wherein the term "non-transitory" does not cover transient propagation signals. As described in detail below, machine-readable storage media 604 can be encoded using executable instructions (e.g., instructions 606-612).
[0112] Hardware processor 602 can execute instructions 606 to store a first optimizer slice of the optimizer state of the common ML model at the first compute node. For example, the first optimizer slice may include a subset or segment of optimizer state local to the first compute node, which can be used to define local instances of the common optimization algorithm, as described above. Figure 1-Figure 5 described.
[0113] Hardware processor 602 can execute instructions 608 to receive, by the first compute node, first plurality of optimizer shard portions from the first plurality of compute nodes. For example, the first compute node can be included as a distributed training network (such as, in conjunction with the above). Figure 1This describes a portion of a cluster of computational nodes in a distributed training network. In this example, the cluster of computational nodes may include a first computational node and a first plurality of computational nodes. Each optimizer shard portion may be received by the first computational node from a corresponding computational node among the first plurality of computational nodes. That is, for example, each computational node among the first plurality of computational nodes may provide an optimizer shard portion to the first computational node, which may collectively constitute the first plurality of optimizer shard portions. Each optimizer shard portion may be a copy of a portion of optimizer shards of a common optimization algorithm stored at each corresponding computational node of the plurality of computational nodes. For example, as described above... Figure 1-Figure 5 As described, each of the first plurality of compute nodes may store optimizer shards of optimizer state (e.g., different segments of optimizer state for local instances of a common optimization algorithm), each optimizer shard in which may be replicated and divided into optimizer shard portions and shared with the first compute nodes.
[0114] In the example, the first multiple optimizer sharding portion can be received during one of the following periods: forward propagation and back propagation of training the public ML model, as combined above. Figures 2-5 described.
[0115] In the example, the first compute node can be configured to provide a second plurality of optimizer shard portions to a first plurality of compute nodes. For example, hardware processor 602 can execute instructions to cause the first compute node to copy the first optimizer shard, divide the copied first optimizer shard into a second plurality of optimizer shard portions, and transfer the second plurality of optimizer shard portions to the first plurality of compute nodes. The transfer of the second plurality of optimizer shard portions can be performed during a full collection operation, which is performed during one of the following: forward propagation and back propagation of training a common ML model.
[0116] In response to a failure of a compute node among the first plurality of compute nodes, the hardware processor 602 may execute instruction 610 to update the first optimizer shard by merging the optimizer shard portion corresponding to the failed compute node with the first optimizer shard. For example, a failure of at least one compute node among the first plurality of compute nodes may be detected, as described above. Figure 1 and Figure 4 As described. In response to the detection of a fault (e.g., receiving a notification or warning), the first compute node can update the first optimizer shard by merging the optimizer shard portion of the first plurality of optimizer shard portions corresponding to the faulty node among the first plurality of shard nodes (e.g., a shard portion received from the faulty node or otherwise derived from the faulty node) with the first optimizer shard.
[0117] In the example, each of the second plurality of compute nodes can be configured to update the corresponding optimizer shard in the first plurality of optimizer shards with the optimizer shard portion corresponding to the failed compute node in response to a detected fault. In this example, the second plurality of compute nodes could be the first plurality of compute nodes in which the failed compute node was removed.
[0118] In the example, hardware processor 602 can execute instructions to receive third plurality of optimizer shard portions from a second plurality of compute nodes by a first compute node. Each optimizer shard portion of the third plurality of shard portions may be a copy of a portion of a corresponding updated optimizer shard stored at the corresponding compute node in the second plurality of compute nodes. For example, during a full collection operation, the first compute node may divide the updated first shard into fourth plurality of optimizer shard portions and transmit the fourth plurality of optimizer shard portions to the second plurality of compute nodes, the full collection operation being performed during one of the following: forward propagation and back propagation of training a common ML model.
[0119] Hardware processor 602 can execute instructions 612 to update the weights of the common ML model based on the updated first optimizer slice. For example, as described above... Figure 4 and Figure 5 In view of Figure 2 and Figure 3 As described, the updated first optimizer slice can be used as a training optimizer slice for future iterations of the backpropagation phase of the distributed training process, for example, in updating the weights of the public ML model. In the example, the updated first optimizer slice, along with updated first multiple optimizer slices, can be used to train the public ML model, as in combination. Figure 5 As described. Therefore, distributed training performed by the first computing node and the first plurality of computing nodes in a distributed training network can be resilient to node failures, allowing learning to continue uninterrupted.
[0120] Figure 7 The illustration shows another example computational component, according to various embodiments, that can be used to implement node failure resilience in distributed training. Now refer to... Figure 7 The computing component 700 can be, for example, a server computer, a controller, or any other similar computing component capable of processing data. Figure 7 In the example implementation, computing component 700 includes a hardware processor 702 and a machine-readable storage medium 704. In the example, computing component 700 may be... Figure 1 An example of node 10 in node 10.
[0121] The hardware processor 702 may be one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for retrieving and executing instructions stored in the machine-readable storage medium 704. The hardware processor 702 may retrieve, decode, and execute instructions (such as instructions 706-712) to control processes or operations for fail-safe resilience. Alternatively to, or in addition to, retrieving and executing instructions, the hardware processor 702 may include one or more electronic circuits comprising electronic components for the function of executing one or more instructions, such as, but not limited to, a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other electronic circuits.
[0122] Machine-readable storage media (such as machine-readable storage media 704) can be any electronic, magnetic, optical, or other physical storage device that contains or stores executable instructions. Therefore, machine-readable storage media 704 can be, for example, random access memory (RAM), non-volatile RAM (NVRAM), electrically erasable programmable read-only memory (EEPROM), storage devices, optical discs, etc. In some embodiments, machine-readable storage media 704 can be a non-transitory storage medium, wherein the term "non-transitory" does not cover transient propagation signals. As described in detail below, machine-readable storage media 704 can be encoded using executable instructions (e.g., instructions 706-712).
[0123] Hardware processor 702 can execute instructions 706 to receive each of a first plurality of optimizer shard portions from a respective compute node in a first plurality of compute nodes, wherein each optimizer shard portion is a part of a corresponding optimizer shard of optimizer state associated with the respective compute node. In the example, the optimizer state defines the optimization algorithm for the ML model. For example, each optimizer shard associated with the respective compute node may include a subset or segment of optimizer state, which may define a local instance of the optimization algorithm, as described above. Figure 1-Figure 5 As described above, the compute component 700 can receive portions of these optimizer shards from respective compute nodes in the first plurality of compute nodes, which can copy (e.g., duplicate) their respective optimizer shards and divide the copied optimizer shards into portions. Each compute node in the first plurality of compute nodes can transmit a portion of its copied optimizer shard to the compute component 700 for storage in machine-readable storage medium 704, as described above. Figure 1-Figure 5 described.
[0124] Hardware processor 702 can execute instructions 708 to recover the optimizer shard corresponding to the failed compute node by updating a first optimizer shard using a portion of the optimizer shard associated with the failed compute node, based on the detection of a failure in at least one of the first plurality of compute nodes. The first optimizer shard, which may be associated with compute component 700 (e.g., local to compute component 700), may be a segment of optimizer state, similar to each of the first plurality of optimizer shards, as described above. Instruction 708 can be executed to locate a portion of the optimizer shard associated with the failed node stored in machine-readable storage medium 704 and merge the located portion with the first optimizer shard of compute component 700, for example, as described above. Figure 3 and Figure 4 As described. Therefore, at least a portion of the optimizer shard associated with the failed node can be recovered, which would otherwise be lost due to the failure. In the example, each of the first plurality of compute nodes (e.g., a subset of the first plurality of compute nodes excluding the failed node) can similarly hold a portion of the optimizer shard associated with the failed node, which can be located and merged with the optimizer shard associated with each compute node.
[0125] The hardware processor 702 can execute instructions 710 to transfer portions of a second plurality of optimizedr shards, updated from the first plurality of compute nodes, to a subset of the first plurality of compute nodes during one of the following: forward propagation and back propagation of training the ML model. For example, as in combination Figure 3 and Figure 4 As described, the second plurality of optimizer shards can be passed to a subset of the first plurality of compute nodes during an all-to-all operation performed during one of the following: forward propagation and back propagation for training the ML model. In the example, the ML model can be trained partly based on the first optimizer shards during a first iteration of fully sharded data parallelism and partly based on the updated first optimizer shards during a second (e.g., subsequent) iteration of fully sharded data parallelism, for example, as described above. Figure 1-Figure 5 described.
[0126] While the examples disclosed herein are described with reference to providing resilience relative to a single compute node failure, malfunction, or otherwise becoming functionally unavailable for distributed training, the examples are not intended to be limited to a single functional failure. The examples disclosed herein can be extended to multiple simultaneous functional failures, thereby rendering more than one compute node functionally unavailable for distributed training. In this case, compute nodes (e.g., Figure 1Compute nodes 10A through 10G can store multiple replicated optimizer shards for each of the other compute nodes. For example, refer to... Figure 2 and Figure 3 Compute node 10A can store a first replica of optimizer shard portions corresponding to compute node 10B, a second replica of optimizer shard portions corresponding to compute node 10C, and a third replica of optimizer shard portions corresponding to compute node 10D. Similarly, compute nodes 10B-10D can store multiple replica shard portions, each corresponding to a specific compute node. Therefore, for example, in the event of a functional failure of compute nodes 10A and 10B, the optimizer shard portions stored on compute nodes 10C and 10D, respectively, for the replicas of compute nodes 10A and 10B, can be merged with optimizer shards 54C and 54D to generate updated optimizer shards, which in this example provide resilience for optimizer shards 54A and 54B. Figures 1-7 The functions discussed can be operated in a basically similar manner, except that multiple compute nodes may become functionally unavailable and their corresponding weight shards can be restored by merging with the weight shards of the remaining functional compute nodes.
[0127] In some examples, compute nodes (e.g., Figure 1 Compute nodes 10A through 10G can also store multiple replicated weight shard portions for each of the other compute nodes. For example, compute node 10A can also store a first replicated set of weight shard portions corresponding to compute node 10B, a second replicated set of weight shard portions corresponding to compute node 10C, and a third replicated set of weight shard portions corresponding to compute node 10D. In this example, compute nodes 10B through 10D can also store multiple replicated sets of weight shard portions, each set corresponding to a specific compute node. Therefore, for example, in the event of a functional failure of compute nodes 10A and 10B, the weight shard portions stored on compute nodes 10C and 10D, respectively replicated for compute nodes 10A and 10B, can be merged with weight shards 52C and 52D to generate updated weight shards, which in this example provide resilience for weight shards 52A and 52B.
[0128] Figure 8A block diagram of an example computer system 800 in which various embodiments described herein may be implemented is depicted. The computer system 800 includes a bus 802 or other communication mechanism for communicating information, and one or more hardware processors 804 coupled to the bus 802 for processing information. The hardware processors 804(s) may, for example, be one or more general-purpose microprocessors. The computer system 800 may be implemented, for example, Figure 1 Computation node 10.
[0129] Computer system 800 also includes main memory 806 (such as random access memory (RAM), cache, and / or other dynamic storage devices) coupled to bus 802 for storing information and instructions to be executed by processor 804. Main memory 806 can also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 804. When stored in a storage medium accessible to processor 804, such instructions cause computer system 800 to function as a special-purpose machine customized to perform the operations specified in the instructions. Main memory 806 and other memory components can store instructions that, when executed by processor 804, cause processor 804 to perform actions in conjunction with… Figure 1-Figure 5 One or more operations as described.
[0130] The computer system 800 also includes a read-only memory (ROM) 808 or other static storage device coupled to the bus 802 for storing static information and instructions for the processor 804. Storage devices 810 (such as disks, optical discs, or USB thumb drives (flash drives)) are provided and coupled to the bus 802 for storing information and instructions.
[0131] The computing system 800 may include a user interface module to implement a GUI, which may be stored in a mass storage device as executable software code to be executed by the computing device(s). For example, this module and other modules may include (by way of example) components (such as software components, object-oriented software components, class components, and task components), procedures, functions, properties, programs, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.
[0132] In general, the terms “component,” “engine,” “system,” “database,” and “data storage” used in this document can refer to logic embodied in hardware or firmware, or to a collection of software instructions that may have entry and exit points and are written in a programming language such as, for example, Java, C, or C++. Software components can be compiled and linked into an executable program installed in a dynamic link library, or can be written in an interpreted programming language such as BASIC, Perl, or Python. It should be understood that software components can be invoked from other components or from themselves, and / or can be invoked in response to detected events or interrupts. Software components configured to execute on a computing device can be provided on computer-readable media (such as optical discs, digital video discs, flash drives, hard disks, or any other tangible media) or as digital downloads (and can initially be stored in a compressed or installable format that requires installation, decompression, or decryption before execution). Such software code can be stored, partially or entirely, on a memory device executing the computing device for execution by the computing device. Software instructions can be embedded in firmware (such as EPROM). It should also be understood that hardware components may include connected logic units (such as gates and flip-flops) and / or may include programmable units (such as programmable gate arrays or processors).
[0133] Computer system 800 may implement the techniques described herein using custom hard-wired logic, one or more ASICs or FPGAs, firmware, and / or program logic. This custom hard-wired logic, one or more ASICs or FPGAs, firmware, and / or program logic, combined with the computer system, enables or allows computer system 800 to become a special-purpose machine. According to one embodiment, the techniques described herein are executed by computer system 800 in response to processor(s) 804 executing one or more sequences of one or more instructions contained in main memory 806. Such instructions may be read into main memory 806 from another storage medium (such as storage device 810). Execution of the instruction sequence contained in main memory 806 causes processor(s) 804 to perform the processing steps described herein. In alternative embodiments, hard-wired circuitry may be used in place of software instructions, or in combination with software instructions.
[0134] As used herein, the term "non-transitory medium" and similar terms refer to any medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such non-transitory media can include non-volatile media and / or volatile media. Non-volatile media include, for example, optical discs or magnetic disks (such as storage device 810). Volatile media include dynamic memory (such as main memory 806). Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape, or any other magnetic data storage medium, CD-ROMs, any other optical data storage media, any physical medium with a perforated pattern, RAM, PROMs, and EPROMs, FLASH-EPROMs, NVRAMs, any other memory chips or cassette tapes, and networked versions of the foregoing.
[0135] Non-transitory media differ from transmission media, but can be used in conjunction with transmission media. Transmission media participate in the transfer of information between non-transitory media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including lines containing bus 802. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.
[0136] Computer system 800 also includes a network interface 818 (also referred to as a communication interface) coupled to bus 802. Network interface 818 provides bidirectional data communication coupled to one or more network links connected to one or more local networks. For example, network interface 818 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem providing data communication connectivity to a corresponding type of telephone line. As another example, network interface 818 may be a Local Area Network (LAN) card to provide data communication connectivity to a LAN-compatible network (or a WAN component communicating with a WAN). Wireless links may also be implemented. In any such implementation, network interface 818 transmits and receives electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.
[0137] A network link typically provides data communication to other data devices via one or more networks. For example, a network link can provide connectivity to a host computer or to data equipment operated by an Internet Service Provider (ISP) via a local network. The ISP, in turn, provides data communication services through a global packet data communication network now commonly referred to as the "Internet." Both local networks and the Internet use electrical, electromagnetic, or optical signals that carry digital data streams. Signals through various networks, as well as signals on network links and through network interface 818, are example forms of transmission media carrying digital data to and from computer system 800.
[0138] Computer system 800 can transmit messages and receive data (including program code) through (multiple) networks, network links, and network interface 818. In the Internet example, the server can transmit request codes for the application through the Internet, ISP, local network, and network interface 818.
[0139] The received code may be executed by processor 804 when it is received, and / or stored in storage device 810 or other non-volatile memory for later execution.
[0140] Each of the processes, methods, and algorithms described in the foregoing sections may be embodied in a code component executed by one or more computer systems or computer processors including computer hardware, and may be fully or partially automated by that code component. One or more computer systems or computer processors may also operate to support the performance of related operations in a “cloud computing” environment or as “Software as a Service” (SaaS). Processes and algorithms may be implemented, partially or entirely, in dedicated circuitry. The various features and processes described above may be used independently of each other or may be combined in various ways. Different combinations and sub-combinations are intended to fall within the scope of this disclosure, and certain methods or processing blocks may be omitted in some implementations. The methods and processes described herein are not limited to any particular order, and the blocks or states associated with them may be executed in other suitable orders, or may be executed in parallel, or in some other manner. Blocks or states may be added to or removed from the disclosed example embodiments. The execution of certain operations or processes may be distributed among computer systems or computer processors, residing not only within a single machine but also deployed across multiple machines.
[0141] As used herein, circuits can be implemented using any form of hardware, software, or a combination thereof. For example, one or more processors, controllers, ASICs, PLAs, PALs, CPLDs, FPGAs, logic components, software routines, or other mechanisms can be implemented to constitute the circuit. In implementation, the various circuits described herein can be implemented as discrete circuits, or the described functions and features can be shared partially or wholly among one or more circuits. Even if various features or functional elements can be described or declared separately as separate circuits, such features and functions can be shared among one or more common circuits, and such description should not require or imply the need for separate circuits to implement such features or functions. When using software to implement circuits wholly or partially, such software can be implemented to operate in conjunction with a computing or processing system (such as a computer system 800) capable of performing the functions described herein.
[0142] As used herein, the term "or" may be interpreted in an inclusive or exclusive sense. Furthermore, descriptions of resources, operations, or structures in the singular form should not be construed as excluding the plural form. Unless otherwise expressly stated or otherwise understood within the context in which they are used, conditional language (such as "can," "may," "may," or "may," among other things) is generally intended to convey that certain embodiments include, while other embodiments do not, certain features, elements, and / or steps.
[0143] Unless otherwise expressly stated, the terms and phrases used in this document, and their variations thereof, should be interpreted as open-ended rather than restrictive. Adjectives such as “regular,” “traditional,” “normal,” “standard,” “known,” and similar terms should not be interpreted as limiting the described items to a given time period or items available at a given time, but should be interpreted as encompassing regular, traditional, normal, or standard techniques available or known at any time now or in the future. In some instances, the presence of extended terms and phrases such as “one or more,” “at least,” “but not limited to,” or other similar phrases should not be interpreted as implying an intention or need for a narrower meaning in instances where such extended phrases might not exist.
Claims
1. A system comprising: The first plurality of compute nodes, each of the first plurality of compute nodes storing a shard from the first plurality of optimizer shards of a common machine learning ML model; as well as The first computing node stores the first optimizer shard of the public ML model, and the first computing node is configured as follows: The first plurality of optimizer shard portions are stored, each optimizer shard portion being received from a corresponding compute node in the first plurality of compute nodes and being a copy of a portion of the corresponding optimizer shard in the first plurality of optimizer shards stored at the corresponding compute node; as well as In response to a failure of one of the first plurality of compute nodes, the first optimizer shard is updated with the optimizer shard portion corresponding to the failed compute node. The weights of the public ML model are updated using the updated first optimizer shards.
2. The system according to claim 1, wherein the first computing node is further configured to: Copy the first optimizer fragment; The copied first optimizer shard is divided into a second plurality of optimizer shard portions; and The second plurality of optimizer shards are transmitted to the first plurality of computing nodes.
3. The system of claim 2, wherein the first computing node transmits the second plurality of optimizer shards to the first plurality of computing nodes during an all-to-all operation, the all-to-all operation being performed during one of the following: forward propagation and back propagation of training the common ML model.
4. The system of claim 2, wherein the second plurality of optimizer shards comprises a plurality of equal-sized portions of the first optimizer shard, wherein each equal-sized portion of the first optimizer shard is transmitted to a compute node among the first plurality of compute nodes.
5. The system of claim 4, wherein the number of equal optimizer shards is equal to the number of computing nodes in the first plurality of computing nodes.
6. The system of claim 1, wherein in response to the failure of one of the first plurality of computing nodes, each of the second plurality of computing nodes is configured to: The optimizer shard in the first plurality of optimizer shards is updated using the optimizer shard portion corresponding to the faulty compute node. The second plurality of computing nodes are the first plurality of computing nodes excluding the computing nodes that are faulty.
7. The system of claim 6, wherein the first computing node is further configured to: Receive a third plurality of optimizer shard portions from the second plurality of compute nodes, each of the third plurality of optimizer shard portions being a copy of a portion of a corresponding updated optimizer shard stored at the corresponding compute node in the second plurality of compute nodes; The updated first optimizer shard is divided into a fourth plurality of optimizer shard parts; as well as The fourth plurality of optimizer shards are transmitted to the second plurality of computing nodes.
8. The system of claim 1, wherein the first computing node is configured as: A set of storage optimizer shard portions, each set comprising a plurality of optimizer shard portions received from a corresponding compute node among the first plurality of compute nodes, and each set of optimizer shard portions being a copy of a portion of the first plurality of optimizer shards corresponding to the corresponding compute node; and In response to the failure of two or more of the first plurality of compute nodes, the first optimizer shard is updated with the optimizer shard subset corresponding to the two or more compute nodes.
9. The system of claim 1, wherein the first computing node is further configured to: In response to the failure of one of the first plurality of compute nodes, a weight shard is updated with a weight shard portion corresponding to the failed compute node, wherein the weight shard includes the weights of the public ML model local to the first compute node, and wherein the weight shard portion is stored at the first compute node and was received from the failed compute node prior to the failure.
10. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a method, the method comprising: The first optimizer shard of the optimizer state of the public machine learning ML model is stored at the first computing node; The first computing node receives a first plurality of optimizer shards from a first plurality of computing nodes. Each optimizer shard is received from a corresponding computing node among the first plurality of computing nodes and is a copy of a portion of the optimizer state of the common ML model stored at the corresponding computing node. In response to a failure of a compute node among the first plurality of compute nodes, the first optimizer shard is updated by merging the optimizer shard portion corresponding to the failed compute node with the first shard. as well as The weights of the public ML model are updated based on the updated first optimizer shard.
11. The non-transitory computer-readable medium of claim 10, wherein the first computing node receives the first plurality of optimizer shard portions during an all-to-all operation, the all-to-all operation being performed during forward and backward propagation of training the common ML model.
12. The non-transitory computer-readable medium of claim 10, wherein the method further comprises: Copy the first optimizer fragment; The copied first optimizer shard is divided into a second plurality of optimizer shard parts; as well as The second plurality of optimizer shards are transmitted to the first plurality of computing nodes.
13. The non-transitory computer-readable medium of claim 12, wherein the method further comprises: During the all-to-all operation, the first computing node transmits the second plurality of optimizer shards to the first plurality of computing nodes, the all-to-all operation being performed during one of the following: forward propagation and back propagation of training the common ML model.
14. The non-transitory computer-readable medium of claim 12, wherein dividing the first optimizer fragment into a second plurality of optimizer fragment portions comprises: The copied first optimizer shard is divided into multiple equal-sized portions, wherein each equal-sized portion of the first optimizer shard is transmitted to a compute node among the first plurality of compute nodes.
15. The non-transitory computer-readable medium of claim 14, wherein the number of equal optimizer shard portions is equal to the number of computing nodes in the first plurality of computing nodes.
16. The non-transitory computer-readable medium of claim 10, wherein the method further comprises: In response to the failure of one of the first plurality of compute nodes, each of the second plurality of compute nodes updates the corresponding optimizer shard of the first plurality of optimizer shards with the optimizer shard portion corresponding to the failed compute node, wherein the second plurality of compute nodes are the first plurality of compute nodes excluding the failed compute node.
17. The non-transitory computer-readable medium of claim 16, wherein the method further comprises: The first compute node receives a third plurality of optimizer shard portions from the second plurality of compute nodes, each of the third plurality of optimizer shard portions being a copy of a portion of a corresponding updated optimizer shard stored at the corresponding compute node in the second plurality of compute nodes; The updated first optimizer shard is divided into a fourth plurality of optimizer shard parts; as well as The fourth plurality of optimizer shards are transmitted to the second plurality of computing nodes.
18. A computing node, comprising: The memory stores instructions and the first optimizer slice, which stores the optimizer state of the machine learning (ML) model. as well as A processor, operably connected to the memory and configured to execute the instructions to: Each optimizer shard portion is received from the respective compute node of the first plurality of compute nodes, wherein each optimizer shard portion is a part of a corresponding optimizer shard of the ML model associated with the respective compute node; Based on detecting a fault in at least one of the first plurality of compute nodes, the optimizer shard corresponding to the at least one compute node is restored by updating the first optimizer shard with the optimizer shard portion associated with the at least one compute node. as well as During backpropagation training of the ML model, a second plurality of optimized shard portions of the first optimized shard are transmitted to a subset of the first plurality of computing nodes.
19. The compute node of claim 18, wherein the second plurality of optimizer shard portions are transferred to the subset of the first plurality of compute nodes during the all-to-all operation.
20. The computing node of claim 18, wherein the processor is further configured to execute the instructions to: During the fully sharded data parallel iteration, the ML model is trained in part based on the first optimizer shards; and During subsequent iterations of fully sharded data parallelism, the ML model is trained in part based on the updated first optimizer shards.
Citation Information
Patent Citations
Elastic training method of large language model, cluster system, product and medium
CN117786412A
Elastically managing workers of multi-worker workloads on accelerator devices
US20230236837A1
Scaling deep graph learning in distributed setting
WO2022250910A1