Resilient optimization status for fully shaded data parallel

The fail-safe approach in FSDP framework addresses node failures by sharing and replicating optimizer and weight shards, enabling continuous and efficient distributed machine learning.

DE102024137356A1Pending Publication Date: 2025-10-30HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024137356
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-26
Filing Date
2024-12-12
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Conventional Fully Constrained Data Parallel (FSDP) implementations in distributed machine learning lack fail-safes to restore model parameters in case of node failures, leading to significant interruptions and re-initialization requirements.

Method used

Implement a fail-safe approach in the FSDP framework by sharing and replicating optimizer states among compute nodes, distributing replicated optimizer and weight shards to maintain updated states across the network, ensuring continuous training even in the event of node failures.

Benefits of technology

Ensures uninterrupted distributed training by maintaining and updating optimizer and weight states across functional nodes, minimizing computational load and resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Systems and procedures for fault tolerance in the distributed training of machine learning (ML) models are provided. Examples include a plurality of compute nodes that store optimizer portions of a plurality of optimizer portions, and a first compute node that stores a first optimizer portion of optimizer states. The first compute node can store portions of optimizer shards, each of which can be received by a corresponding compute node of the plurality of compute nodes and can be a replica of a portion of a corresponding optimizer shard of the plurality of optimizer shards stored at the corresponding compute node.In response to the failure of one of the multitude of compute nodes, the first compute node can update the first optimizer shard with an optimizer shard part corresponding to the failed compute node, and the ML model can be trained based on the updated first optimizer shard.
Need to check novelty before this filing date? Find Prior Art

Description

background

[0001] Machine learning (ML) is generally a computer-implemented process that uses example data (e.g., training data) to build a model for making predictions or decisions without being explicitly programmed. ML processes are used in a wide variety of applications, especially where it is difficult or impractical to develop traditional algorithms to perform various computational tasks.

[0002] Distributed training is a subfield of machine learning (ML) where multiple decentralized units jointly train a common ML model by executing subsets of data or parameters that are locally available at each unit in parallel. Distributed training approaches contrast with traditional centralized ML techniques, where training is performed sequentially on a single compute node. Brief description of the drawings

[0003] The present disclosure is described in detail in accordance with one or more different embodiments with reference to the following figures. The figures serve only for illustration and represent only typical or exemplary embodiments. Fig. Figure 1 shows an example system for distributed training according to an example implementation of the present disclosure. Fig. Figure 2 is a schematic block diagram representing a forward propagation phase in distributed training according to the example implementations of the present disclosure. Fig. Figure 3 is a schematic block diagram representing a backward propagation phase in distributed training according to the example implementations of the present disclosure. Fig. Figure 4 is a schematic block diagram that illustrates a process flow for providing fault tolerance in distributed training according to example implementations of the present disclosure. Fig. Figure 5 is a schematic block diagram illustrating a process flow to ensure continuous resilience to node failures during distributed training according to example implementations of this disclosure. Fig. 6 is an example of a computer component that can be used to implement various features of distributed training fault tolerance in accordance with the implementations disclosed here. Fig. 7 is another example of a computer component that can be used to implement various features of distributed training fault tolerance in accordance with the implementations disclosed here. Fig. Figure 8 is an example of a computer system that can be used to implement various features of the distributed training failure described in the present disclosure.

[0004] The illustrations are not exhaustive and do not limit the present disclosure to the exact form that is disclosed. Detailed description

[0005] Training machine learning (ML) models at scale can be a demanding task requiring significant computing power, resources, and time. For example, large language models (LLMs) can consist of a vast number of parameters to be trained, and the number of such parameters has increased from 110 million in 2018 to one trillion in 2021. To effectively train these large-scale LLMs, as well as other large-scale ML models, distributed training algorithms have been proposed. These algorithms distribute the training task across a range of computing resources (also known as "compute nodes" or "nodes") within a distributed training network. For instance, data parallelism (DP) can replicate an entire model on each compute node and split the training datasets into multiple segments for each node.Another example is pipeline parallelism (PP), where a machine learning model is divided into stages and these stages are distributed across multiple compute nodes. Tensor parallelism (TP), another approach, divides tensors into multiple chunks, each of which can be executed on a single compute node. However, these techniques are generally tied to a specific model architecture and are difficult to generalize to other models. For example, tensor parallelism may require tight coupling of compute nodes because parallelizing matrix multiplication can be communication-intensive, making it feasible only across compute elements (e.g., GPUs) within a single server. The efficiency of pipeline parallelism can depend on the model itself and lead to inefficient execution (so-called "pipeline bubbles") if the model is irregular.

[0006] Fully sharded data parallel (FSDP) is another parallelization technique for distributed machine learning. In FSDP, model parameters, such as weighting and optimization states, are distributed across a network of nodes (e.g., shared or otherwise partitioned). Each compute node can work with a different set of training data stored locally on that node (referred to here as "local training data" or "local data"). FSDP includes at least two training phases: a forward propagation phase (sometimes called a "forward pass") and a backward propagation phase (sometimes called a "backward pass").

[0007] The forward propagation phase can be performed to obtain outputs from local inputs of a common machine learning (ML) model. For example, an ML model might comprise multiple layers, each configured to perform different transformations on one or more inputs. The transformations performed by each layer can depend on weighting states (sometimes referred to here as "weights") learned for each layer through training. The first layer of the model can be called the input layer, and the last layer the output layer. Training data can be fed to the input layer. Several intermediate layers (sometimes called "hidden layers") can be arranged sequentially between the input and output layers. The outputs of one layer are passed on as inputs to the next layer.The forward propagation phase is used to calculate the outputs of each layer in order from the input layer to the output layer.

[0008] In the case of FSDP, during the forward propagation phase, each compute node computes local outputs from each layer of the shared machine learning model by applying inputs stored locally within that node. As mentioned earlier, each compute node has n shards of model parameters for each layer, such as n shards of weight states (referred to here as "weight shards") and n shards of optimizer states (referred to here as "optimizer shards") for each layer. To perform the input-to-output conversion, each compute node performs an "all-gather" operation for each layer to collect weight shards held at the other compute nodes and reconstruct the complete layer. Once the complete layer is obtained, each compute node performs a forward computation operation by feeding locally held inputs into the reconstructed layer to obtain outputs for that layer.Each compute node then discards the weight splitters assigned to the other nodes (e.g., received from the other nodes) to make room for repeating the process for the next sequential layer, using the received outputs as inputs. The process is repeated for each layer until the compute nodes receive the outputs for the last layer.

[0009] The backward propagation phase can be performed to update the model parameters of the machine learning model by obtaining loss functions from a previous layer. In the backward propagation phase, one or more gradients of a loss function are computed with respect to the weight states of a given layer. The backward propagation phase performs a backward-pass computation to obtain the gradients layer by layer by iterating backward from the output layer to the input layer. The backward propagation phase can use gradient descent or variants such as stochastic gradient descent to perform the backward computation. The backward propagation phase can employ an optimization algorithm, which can be defined by the optimization states, to compute updated weights with respect to the obtained gradients.The term "optimization state" can refer to momentum vectors or similar historically traceable properties of an optimization algorithm. For example, an optimization state for a gradient descent optimization algorithm might track moving averages of the gradient and the squared gradient.

[0010] In the case of FSDP, during backward propagation, each node updates the weight states of its weight shards by applying its optimization shards with respect to the gradients corresponding to its weight shard. For example, each compute node performs an all-collect operation for each layer to gather the weight shards from the other nodes and reconstruct the complete layer. Once the complete layer is obtained, each compute node performs a backward-pass computation to calculate the weight gradients and input gradients for the current layer with respect to the weights of the fully reconstructed layer. These weight gradients and input gradients can be referred to as local weight gradients and local input gradients, respectively. Each node then discards the weight shards collected from other nodes to make room for the next layer.At this stage, each compute node has local weight gradients corresponding to each weight fraction with respect to its respective local inputs. These local weight gradients can then be aggregated (e.g., averaged) across the compute nodes to obtain global weight gradients. Each node can then perform a reduction-dispersion operation on the global weight gradients to obtain a portion of the global weight gradients corresponding to its respective weight shard. Each node can then update the weights of its respective weight shards with respect to this portion of the global weight gradients.

[0011] Traditional FSDP implementations do not provide fault tolerance to restore portions of model parameters, such as, but not limited to, optimizer parameters, in the event of a node failure, malfunction, or other anomalous behavior that renders the node unavailable for distributed training (referred to here as a functional failure). Therefore, if a compute node becomes unavailable, the corresponding model parameters held by that node may be lost and may need to be relearned, at least with respect to the lost parameters, through a complete reinitialization of the entire process. Consequently, the failure of a single node can significantly disrupt the learning process, a situation that can be exacerbated in the case of large-scale machine learning training involving hundreds or thousands of compute nodes.

[0012] This disclosure presents a fault tolerance approach that can be implemented in the FSDP framework to protect against node failures by sharing portions of the optimizer shards among the compute nodes. In these examples, each compute node can contain local shards of optimizer states (here sometimes referred to as a "training optimizer shard") and copies of portions of shards of optimizer states (here referred to as "replicated optimizer shard portions") that are stored on the other compute nodes of a distributed training network. The training optimizer shards can contain optimizer states that are local to each compute node for an optimization algorithm that is common to the compute nodes and corresponds to a common machine learning model.Each compute node can be responsible for updating the weight states using its training optimizer shard with respect to the gradients during the backward propagation phase and for updating its training optimizer shard with respect to the received gradients (e.g., a portion of the global weight gradients corresponding to its respective weight shard) during the backward propagation phase. To achieve fault tolerance, each compute node can replicate its training optimizer shard, partition the replicated training optimizer shard into replicated optimizer shard parts, and distribute the replicated optimizer shard parts to the other compute nodes at any time during the forward and / or backward propagation phases.For example, replicated optimizer shard parts can be distributed by an all-to-all operation performed at any point during the forward propagation and / or backward propagation phases. In an intuitive example, replicated optimizer shard parts can be distributed by an all-to-all operation performed concurrently or nearly concurrently with an all-gatherer operation during the backward propagation phase. In this way, each compute node can contain replicated optimizer shard parts that it has received from each of the other nodes in the network. According to some examples, each training optimization shard can be replicated and divided into a number (N) of equally sized replicated optimization shards, where N is one less than the number of compute nodes used for training.When a compute node fails, each remaining node (referred to here as a "functional node") can update its respective training optimization shard with a replicated optimization shard section corresponding to the failed node (e.g., a replicated optimization shard section originating from or received by the failed node during the previous iteration of the backward propagation phase). Consequently, the updated training optimization shard may include the previous training optimization shard, merged with the replicated optimization shard section of the failed node.

[0013] To ensure further fault tolerance in the event of a subsequent node failure, each functional node can replicate and partition its updated training optimization splitter, which can then be distributed to the remaining functional nodes during the next iteration of forward or backward propagation. This way, the latest optimization states for the failed node are maintained and updated throughout the network and are not lost due to the failure. Thus, the remaining functional nodes can continue training the machine learning mode without interruption.

[0014] Accordingly, implementations of this disclosure can ensure fault tolerance by sharing portions of the optimization states between compute nodes at any time during the FSDP framework. By evenly distributing the optimization states contained in the failed node's optimization splitter across the functional nodes, the workload can be evenly distributed across each functional node, thereby minimizing computational overhead in terms of processing power and resources.

[0015] It should be noted that the terms "optimize," "optimal," and the like, as used here, can be interpreted as meaning to make or achieve performance as effective or perfect as possible. However, as any professional reading this document will recognize, perfection cannot always be achieved. Accordingly, these terms can also mean that performance is as good or effective as possible, or practical under the given circumstances, or that performance is better than what can be achieved with other settings or parameters.

[0016] Fig. Figure 1 shows an example system 100 for distributed training according to an example implementation of the present disclosure. The example system 100 comprises a distributed training network 110 with a plurality of compute nodes 10A-10G in a cluster or group of compute nodes (also referred to collectively as node 10 or individually as node 10A-10G).

[0017] Each node 10 can be connected to other node 10 via a network, which may include, for example, the internet, an intranet, a PAN (Personal Area Network), a LAN (Local Area Network), a WAN (Wide Area Network), a SAN (Storage Area Network), a MAN (Metropolitan Area Network), a wireless network, a cellular communications network, a public telephone network, and / or another type of network. Furthermore, the components described here can be implemented in various hardware and / or software configurations, configuring the hardware.

[0018] The multitude of nodes 10 in the cluster of the distributed training network 110 can comprise any number, configuration, and connections between the nodes 10. The in Fig. The arrangement of nodes 10 shown in Figure 1 is therefore for illustrative purposes only. Node 10 can be a fixed or mobile computer device. While node 10A is in Fig. As shown in detail in section 1, each of the 10 nodes can be configured in the manner shown. In the example of Fig. 1. Node 10A comprises one or more processors 20 (referred to herein for simplicity as processors 20, processor(s) 20, or processor 20) and one or more storage devices 40 (referred to herein for simplicity as storage devices 40, storage device(s) 40, or storage device 40), as well as other components. In examples, one or more of the compute nodes may be implemented as a graphics processing unit (GPU). The storage device(s) 40 may contain (e.g., store) data 48 that node 10A can access locally (referred to hereafter as local data). The local data 48 may not be accessible to other nodes 10 in the distributed training network 110 (e.g., nodes 10B-10G in this example).

[0019] In some examples, the storage device(s) 40 can store a distributed ledger 42, one or more models 44 (referred to here for simplicity as models 44, model(s) 44, or model 44), and / or rule(s) 46. The distributed ledger 42 can contain a series of data blocks that refer to at least one other block, such as a previous block. In this way, the data blocks can be concatenated to form the distributed ledger 42. In some examples, the distributed ledger 42 can store blocks that indicate a state of node 10A with respect to its machine learning during an iteration. Thus, the distributed ledger 42 can store an immutable record of the state transitions of a node 10A. In this way, the distributed ledger 42 can store a current and historical model state of each model 44.However, it should be noted that in some embodiments a collection of records, models and smart contracts can be stored in the distributed ledger 42 by one or more other nodes (e.g. nodes 10B-10G).

[0020] As described here, the model 44 can be trained locally at a node 10 based on local data 48 and then updated based on model parameters learned at other nodes 10. The nature of the model 44 depends on the specific implementation of the node 10 itself. For example, the model 44 can be defined by learned parameters relating to the following aspects: features of the self-driving vehicle, such as sensor information related to object detection; network configuration features; security features related to network security, such as intruder detection; healthcare features related to medical records and patient health information; social science features related to human behavior in social and cultural semantic aspects; and / or other context-based models.

[0021] The model 44 can be stored as a local instance of a machine learning algorithm as well as model parameters determined by training the machine learning algorithm. Model parameters can be stored as different model states, such as weights, biases, optimizers, gradients, or similar properties, which can define a specific instance of the model 44. Each model 44 can consist of multiple layers, with each layer defined by a set of model parameters for performing various transformations on an input. The first layer of the model can be an input layer to which local data 48 can be fed, and the last layer can be an output layer. Several intermediate or hidden layers can exist between the input and output layers in a sequential order, with the outputs of one layer being fed as inputs to the next layer.The transformations performed by each layer may depend on the model parameters learned for that layer. Model(s) 44 can comprise any model from a general class of machine learning algorithms, including, but not limited to, many statistical and classical machine learning algorithms used by vertical industries, such as regression-based algorithms, decision trees (DT), support vector machines (SVM), etc. Training methods may include, among others, standard batch training.

[0022] In the examples presented, the model 44 can be stored as shards of model states, each defining a local instance of a machine learning algorithm. For instance, node 10A is shown to contain a weight shard 52A and an optimization shard 54 of the model 44, defining a local instance of the machine learning algorithm. Each node 10 can contain other shards that together define the entire common layer. The model 44 can contain one or more shards of model states for each layer. For example, the weight layer 52A can contain weights that define a transformation for one layer of the machine learning model, and the model 44 can contain one or more other weight layers that define a transformation for one or more other layers of the model 44.Similarly, the optimization shard 54A can contain optimization states that define a local instance of an optimization algorithm for a layer of the ML model (e.g., the same layer as the weighting shard 52A), and the model 44 can contain one or more other optimization shards for one or more other layers of the model 44.

[0023] Rules 46 can include smart contracts or machine-readable rules that configure nodes to behave in specific ways with respect to distributed training and enable decentralized control. For example, rules 46 can define deterministic state transitions, when to initiate a machine learning iteration, whether a node should be allowed to enroll in an iteration, the number of nodes that must agree to a consensus decision, the percentage of voting nodes that must agree to a consensus decision, and / or other actions that node 10A can perform for distributed machine learning.

[0024] Rule 46 allows for the definition of hyperparameters, which define how the ML framework 24 and the resilience framework 28 are structured. Hyperparameters can be viewed as a mechanism for controlling the training process, such as deciding how many training iterations to perform, how many nodes 10 to use for training, defining criteria for ending training, defining data parallelism techniques, and so on. Hyperparameters can be customizable, predefined parameters that can be tuned to obtain / generate an ML model / algorithm with optimal / tuned performance. In some examples, the hyperparameters can be set by an operator via a front-end dashboard.

[0025] According to the examples disclosed here, the rules 46 can contain one or more hyperparameters that configure the ML framework 24 to use FSDP techniques for distributed training. In this case, the rules 46 can contain one or more hyperparameters that specify a number of shards of model states to be created from parameters of a model 44. For example, rule 46 can specify a number of shards to be created from a collection of weights for a particular layer of the model 44 by partitioning the weights that define the transformations of that layer into the specified number of weight shards. Similarly, a collection of optimization states of the model 44 can be partitioned into the specified number of optimization shards. The rules 46 can then assign 10 weight shards and optimization shards to each node, for example,40 are stored in one or more storage devices. In some examples, the number of shards (e.g., the number of weight shards and the number of optimizer shards) can be specified as the number of nodes registered for training. In various examples, each shard can be the same size in terms of the storage required to store each shard. That is, for example, each weight shard can be the same size and each optimizer shard can be the same size. However, the weight shards do not have to be the same size as the optimization shards.

[0026] According to the examples disclosed here, the rules 46 can also contain one or more hyperparameters that configure the Resiliency Framework 28 to protect against node failures. In this case, the rules 46 can contain one or more hyperparameters that configure each node 10 to replicate its respective optimization shard to create a copy of it. Rules 46 can also contain one or more hyperparameters that configure each node 10 to partition the replicated optimizer shard into a number of replicated optimizer shard parts and distribute the replicated optimizer shard parts of optimizer shards to other nodes 10. Each replicated optimizer shard part can include data that defines a specific part of an optimizer shard (e.g.,(with respect to the optimizer states they contain) that do not overlap with any other part of any of the other replicated optimizer shard parts originating from the same optimizer shard. In this way, the replicated optimizer shard parts can collectively replicate the entire optimizer shard. The number of replicated optimizer shard parts created can be based on the number of nodes 10 that are operational or otherwise functioning as expected to perform distributed training. For example, the number of replicated optimizer shard parts can be one less than the number of nodes 10 registered for training.

[0027] In some examples, Rule 46 may also include one or more hyperparameters that configure each node 10 to replicate its respective weight sliver to create a copy of it. In this example, Rule 46 may also include one or more hyperparameters that configure each node 10 to divide the replicated weight sliver into a number of replicated weight sliver sections and distribute the replicated weight sliver sections to other nodes 10. Each replicated weight sliver section can include data that defines a specific section of a weight sliver (e.g., with respect to the weights it contains) that does not overlap with any other section of any of the other replicated weight sliver sections that originated from the same weight sliver. In this way, the replicated weight sliver parts can collectively reconstruct the entire weight sliver.The number of replicated weight shard sections created can be based on the number of nodes 10 that are operational as described above or otherwise functioning as expected to perform distributed training.

[0028] Rule 46 may also include one or more checks for detecting node failures, such as a node malfunction or similar. A malfunction can refer to a situation where a compute node is no longer available for distributed training, for example, because the node is no longer functioning or is otherwise no longer participating in the training. In some cases, a node failure may result in damage to the storage device(s) 40 of the failed node, which may lead to the loss of model parameters stored by the failed node. In other cases, even if the model parameters are not lost, a non-participating node may not update its model parameters during training, meaning that when the node is reinserted into the training, it will have to catch up with the other nodes.

[0029] A node may become unavailable for distributed training due to functional failures (also referred to as malfunctions). A functional failure can encompass any failure of the compute node, including hardware failures and anomalous behavior. A hardware failure can refer to a situation where hardware components of a compute node have failed or are otherwise not functioning as intended. Anomalous behaviors include, but are not limited to, software errors (e.g., the compute node's hardware is functional, but the software hangs or does not execute as expected), network errors (e.g., the compute node functions as expected but is unreachable because a switch port is dead or there is another defect in the connection between the compute node and other nodes), and performance errors (e.g.,(The compute node cannot transfer its shards to other compute nodes within a reasonable timeframe).

[0030] In the examples described here, any technique for detecting such node failures can be used. Some illustrative examples are given here, but these should not be considered a limitation. Any method or technique for detecting node failures or unavailability can be used in the examples disclosed here. In some examples, a management system implemented by a node 10 can be deployed in the distributed training network 110. This system monitors each participating node 10, detects when a node 10 has failed or others have stopped participating, and generates alerts that can be forwarded to the remaining functioning nodes 10 to inform them of the failure. In another example, the ML framework 24 can include a watchdog process configured to monitor the learning process.Each node 10 can be expected to produce specific data at certain times as part of the ML framework 24 and exchange this data with other nodes 10. The watchdog process can be configured to monitor the expected data exchange, and if an expected communication is not received within the expected timeframe, the watchdog process can alert node 10 to an error. The node that failed to provide the expected communication can be identified as a failed node.

[0031] The processor(s) 20 can receive local data 48 that is accessible locally to node 10A, but not necessarily to other nodes 10A. Such local data 48 may, for example, contain private data that should not be shared with other devices. The processor(s) 20 can be programmed by one or more computer program instructions. For example, the processors 20 can be programmed to execute the application layer 22, the machine learning framework 24, the interface layer 26, the resilience framework 28, or other instructions to perform various operations, each of which is described in more detail below. For simplicity, the various instructions are described here as the execution of an operation when the various instructions actually program the processors 20 (and thus node 10A) to perform the operation.

[0032] Application layer 22 can run applications on node 10A. For example, application layer 22 can contain an agent (not shown) that programs node 10A to participate in distributed machine learning over a distributed training network 110, as described here. In examples, each node 10 can be programmed with the same agent, ensuring that each acts according to the same rules, which can be encoded, for example, with rules 46. For instance, the agent can program each node 10, according to the hyperparameters specified by rules 46, to act as a participant node. Application layer 22 can perform machine learning over the ML framework 24 and the resilience framework 28, for example, as described below in conjunction with the Fig. 2-6 described procedures.

[0033] The ML framework 24 can train a model based on local data 48 stored at node 10A. For example, the ML framework 24 can generate one or more model parameters by applying the local data 48 to a local instance of an ML algorithm (e.g., model 44). The ML framework 24 learns weights, biases, optimizers, and / or gradients as one or more model parameters (here interchangeably referred to as "one or more local parameters" or "local parameter(s)") that define a particular model 44 and can be stored in the storage device 40. In one example, the ML framework 24 can use the FSDP framework, although other frameworks can also be used.

[0034] According to various examples, the ML framework 24 can use FSDP to distribute the training across the nodes 10. The ML framework 24 can, for example, perform multiple training phases. In the case of distributed training via FSDP, the ML framework 24 works with local training data 48 and performs a forward propagation phase and a backward propagation phase to obtain model parameters. The ML framework 24 can perform the forward propagation phase to obtain the outputs of a model 44 from the inputs by iterating through each layer from the input layer to the output layer. The ML framework 24 can also perform the backward propagation phase to adjust the weights and optimizers of the model 44 by calculating gradients as loss functions with respect to the weights of a given layer and local inputs by iterating backward from the last layer to the first layer.Further details will be provided below in connection with the . Fig. 2-3 described.

[0035] The application layer 22 can use the interface layer 26 to interact with and participate in the distributed training network 110 for cooperative machine learning across multiple participating nodes 10. The interface layer 26 can communicate with other nodes, for example by sending transactions and writing blocks to the distributed ledger 42 based on these transactions.

[0036] Interface layer 26 can share local model parameters and conclusions with other participating nodes 10. Interface layer 26 can include a message passing interface used for communication across a network with other participating nodes 10. The messaging interface can be configured as a Message Passing Interface (MPI) send / receive operation. Other types of message passing interfaces can also be used.

[0037] Resilience framework 28 can ensure that distributed training is resilient to node failures in the distributed training network 110. For example, resilience framework 28 can be executed by node 10A to replicate optimizer shard 54A and partition the replicated optimizer shard into a multitude of replicated optimizer shard parts according to rules 46. Resilience framework 28 can also be executed to distribute the replicated optimizer shard parts to the other nodes 10 of the distributed training network 110, for example, by transferring a specific replicated optimizer shard part to each of the other nodes 10. Each node 10 can similarly execute a corresponding resilience framework to create replicated optimizer shard parts from corresponding optimizer shards and distribute the replicated optimizer shard parts to other nodes 10.Node 10A can execute the resilience framework 28 to receive a replicated optimizer splitter portion from each of the other nodes in the network and store the received replicated optimizer splitter portions in storage device(s) 40. As a result, each node 10 can hold different replicated optimizer splitter portions received from (e.g., originating from) each of the other nodes 10.

[0038] Resiliency Framework 28 can detect a node failure according to Rule 46. After detecting a node failure, each remaining functional node 10 (e.g., the remaining nodes 10 except the failed node) can update its respective optimizer shard based on a replicated optimizer shard portion held in that functional node 10 and associated with the failed node. As an illustrative example, node 10A can execute Resiliency Framework 28 to update optimization shard 54A using a replicated optimization shard portion that node 10A received from the failed node prior to the detected failure. In other examples, optimization shard 54A can be updated by merging the replicated optimization shard portion with the local optimization shard 54A to create an updated optimization shard.The update of local optimizer shards can be performed on any node 10 that is still functioning and participating in the training. As a result, the optimizer states contained in the optimizer shard of the failed node may not be lost and can be retained in the updated optimizer shards of the remaining node 10. The updated local optimizer shards of the remaining node 10 can be used by the ML framework 24 to continue the distributed training without interruption due to the detected failure.

[0039] Furthermore, according to some examples, the resilience framework 28 can optionally be executed by node 10A to replicate the weight splitters 52A and partition the replicated weight splitters into a multitude of replicated weight splitter parts according to rules 46. The resilience framework 28 can also be executed to distribute the replicated weight splitter parts to the other nodes 10 of the distributed training network 110, for example, by transferring a specific replicated weight splitter part to each of the other nodes 10 in a manner similar to the distribution of replicated optimizer splitter parts described above. Each node 10 can similarly execute a corresponding resilience framework to create replicated weight splitter parts of the corresponding weight splitters and distribute the replicated weight splitter parts to other nodes 10.Node 10A can execute the resilience framework 28 to receive a replicated weight splitter portion from each of the other nodes in the network and store the received replicated weight splitter portions in the storage device(s) 40. As a result, each node 10 can hold different replicated weight splitter portions received from (e.g., originating from) each of the other nodes 10.

[0040] After node failure is detected, according to an optional example, each remaining functional node 10 can update its respective weight shard based on a replicated weight shard portion held within that functional node 10 and connected to the failed node. As an illustrative example, node 10A can execute the resilience framework 28 to update weight shard 52A using a replicated weight shard portion that node 10A received from the failed node before the failure was detected. In other examples, weight shard 52A can be updated by merging the replicated weight shard portion with the local weight shard 52A to create an updated weight shard. Updating the local weight shards can be performed on any node 10 that remains functional and participating in training.In addition to retaining the optimizer shard of the failed node, as described above, the weights contained in the failed node's weight shard are not lost in this example and can be retained in the updated weight shards of the remaining 10 nodes. The updated local weight shards of the remaining 10 nodes can be used by the ML framework 24 to continue distributed training without interruption due to the detected failure.

[0041] In some implementations, node 10A can include a packaging and deployment function 50 that can package and deploy a model 44 as a containerized object. For example, the packaging and deployment function 50 can package local model parameters and other inferences into a containerized object that can be shared with other participating nodes 10 via interface layer 26. For example, and without limitation, the packaging and deployment function 50 can use the Docker platform to generate Dockerfiles containing models 44. In another example, the packaging and deployment function 50 can use the Docker platform to generate Dockerfiles containing weight splitters 52A, replicated splitter parts of the optimizer splitter 54A, and / or replicated splitter parts of the weight splitter 52A. Other containerization platforms can also be used.In this way, various applications at node 10 can access and use model 44 regardless of the platform. This allows models to be created not only based on collective parameters from nodes in a distributed training network, but also packaged and deployed in different environments.

[0042] Fig. Figure 2 is a schematic block diagram of a process flow for a forward propagation phase 200 in distributed training, according to example implementations of the present disclosure. In the example of Fig. 2. The forward propagation phase 200 can be carried out by a multitude of nodes 10, as in conjunction with Fig. 1 described. Accordingly, one or more of the operations of the forward propagation phase 200 can, for example, be executed by one or more of the application layer 22, the ML framework 24, the interface layer 26 and / or the resilience framework 28, as executed by the processor(s) 20. In the Fig. The example shown illustrates the forward propagation phase 200, which is performed by four nodes 10A-10D. However, the forward propagation phase 200 can be performed by any number of nodes required for a specific machine learning application.

[0043] The forward propagation phase 200 process can be iteratively performed for each layer of an ML model (e.g., model 44 of Fig. 1) are performed to obtain the outputs of each layer based on the inputs applied by each node 10. The outputs obtained in one iteration can be used as inputs for the next iteration of the forward propagation phase 200. The forward propagation phase 200 comprises several operations that are described in Fig. 2 are divided into steps 210 and 220. In the forward propagation phase 200, steps 210 and 220 can be executed for each layer of the ML model, starting with the first layer (e.g., input layer) and iterating through a number of middle layers (e.g., hidden layers) to the last layer (e.g., output layer) according to the sequential order of the layers.

[0044] Before performing the first iteration of the forward propagation phase 200 on the first layer, registration can take place, allowing each node 10A-10D to register for use in distributed learning. In one example, this might be a one-time process. In other examples, registration or sign-up can be performed after some time as a kind of verification process. Subsequently, each node 10 can record its relevant attributes in a learning contract, such as the Uniform Resource Locator (URL), from which its local set of model parameters can be downloaded by other nodes.

[0045] Furthermore, before performing the first iteration of the forward propagation phase, 200 hyperparameters can be loaded from the storage device(s) 40 into the ML frame 24 of each node. As mentioned above, hyperparameters define how the ML frame 24 and the resilience frame 28 are structured. The hyperparameters can be selected for use at each node, and training can be performed according to these hyperparameters. In examples, the hyperparameters can govern the training process, such as specifying how many nodes (e.g., nodes 10A-10D) perform the training, how many weight splitters should be created, how many replicated splitter parts should be created from each node, how a node failure should be detected, or which processes are considered a node failure, etc.

[0046] In examples, the machine learning (ML) model being trained can be defined by a general ML algorithm that includes various model parameters, such as weight and optimization status. Each layer can be defined by a set of model parameters. As described above, the weights defining each layer of the ML model can be divided into a number of weight splitters. In the example of Fig. The collective weights are divided into four weight blocks 52A-52D, which are assigned to and stored at each node 10A-10D. For example, weight block 52A is assigned to node 10A, weight block 52B to node 10B, weight block 52C to node 10C, and weight block 52D to node 10D. Thus, each node 10A-10D can contain a weight block 52A-52D, for the use of which that node 10A-10D is responsible when performing transformations on inputs to calculate outputs for each specific layer. In total, the weight blocks 52A-52D in this example comprise the weights that define a complete layer.

[0047] Similarly, in the example of Fig. 2. The collective optimizer states are divided into four optimizer shards, 54A-54D, which are assigned to and stored at each node, 10A-10D. For example, optimizer shard 54A is assigned to node 10A, optimizer shard 54B to node 10B, optimizer shard 54C to node 10C, and optimizer shard 54D to node 10D. In this example, the optimizer shards 54A-54D collectively comprise the optimizer states of an entire layer.

[0048] In one example, in step 210, for a given layer, each node 10A-10D reconstructs the respective layer. For instance, the ML frame 24 performs an all-collect operation 212 at each node 10A-10D to obtain weight shards from the other nodes and reconstructs the complete layer from the collected weights. For example, node 10A performs an all-collect operation 212 to obtain weight shards 52B-52D and stores weight shards 52A-52D in its storage device(s) 40. The ML frame 24 can access the storage device(s) 40 to retrieve the weights of each shard and reconstruct the complete layer by applying the weights of the different weight shards to the common ML algorithm.Similarly, node 10B receives weight splitters 52A, 52C, and 52D; node 10C receives weight splitters 52A, 52B, and 52D; and node 10D receives weight splitters 52A-52C to construct the complete layer at each respective node. Weight splitters 52A-52D can be considered training chips due to their use for training.

[0049] In step 220, the ML frame 24 of each node 10A-10D can execute operations 222A-222D to obtain outputs for the fully constructed layer. For example, node 10A executes ML frame 24 to perform a forward computation operation as part of operations 222A. ML frame 24 can perform forward computation by feeding local inputs to the reconstructed layer of the shared ML model and computing (e.g., obtaining) local outputs for that layer. The reconstructed layer can perform transformations on the local inputs according to the weights of the reconstructed layer. In the case of a first iteration of the forward propagation phase 200, ML frame 24 reconstructs the first layer and applies local training data 48 as inputs to the first layer to obtain local outputs of the first layer.Node 10A can use the local outputs obtained for a given reconstructed layer as local inputs for the next reconstructed layer during a subsequent iteration of forward propagation phase 200. While the example above refers to node 10A, each node 10B-10D can perform similar operations to obtain local outputs for each layer of the ML model based on its respective local inputs. In this way, forward propagation phase 200 iterates through the various layers of the ML model.

[0050] Once the outputs of a particular reconstructed layer have been obtained, the weight shards received from other nodes 10A-10D can be discarded to make room for the next iteration. For example, node 10A can execute the ML framework 24 to discard, delete, or otherwise remove weight shards 52B-52D from its storage device(s) 40. Similarly, node 10B can discard weight shards 52A, 52C, and 52D; node 10C can discard weight shards 52A, 52B, and 52D; and node 10D can discard weight shards 52A, 52B, and 52C.

[0051] Furthermore, each node 10A-10D can be configured according to hyperparameters to create copies of portions of its respective weight shard, which can then be distributed to other nodes 10A-10D. For example, hyperparameters can configure the Resiliency Frameworks of each node 10A-10D to replicate a corresponding optimizer shard, partition the replicated optimizer shard into a number of replicated optimizer shard portions according to the hyperparameters, and distribute the replicated optimizer shard portions of weights to other nodes 10A-10D. Each replicated optimizer shard portion can contain data that defines a specific part of the respective optimizer shard (e.g., optimizer shard 54A from node 10A) that does not overlap with any part of any of the other replicated optimizer shard portions originating from the same optimizer shard.Thus, the replicated optimizer shard segments can collectively define a complete optimizer shard. The number of replicated optimizer shard segments created can be based on the number of nodes used to perform distributed training. For example, the number of replicated optimizer shard segments can be one less than the total number of nodes.

[0052] In the example of Fig. 2. Node 10A can execute Resiliency Framework 28 to create three replicated optimizer shard segments, 54A-1 to 54A-3. For example, Resiliency Framework 28 can access storage device(s) 40 to obtain optimizer shard 54A and create a copy of it. Resiliency Framework 28 can partition the copy of optimization shard 54A into segments by dividing it into three replicated optimization shard segments, 54A-1 to 54A-3. In various examples, the three segments can be essentially the same size.

[0053] Node 10A can execute the resilience framework 28 to distribute the replicated optimization splitter parts 54A-1 to 54A-3 to nodes 10B-10D. Each node 10B-10D executes its respective resilience framework to receive one of the replicated optimization splitter parts 54A-1 to 54A-3 and store the received replicated optimization splitter part in one or more storage devices. In the example of Fig. In node 2, node 10B contains the replicated optimizer shard part 54A-1, node 10C contains the replicated optimizer shard part 54A-2, and node 10D contains the replicated optimizer shard part 54A-3. While this example refers to replicated optimizer shard sections 54A-1 through 54A-3, the reference numbers are intended to represent parts of a whole (e.g., 33% of optimizer shard 54A) and are not meant to give the replicated optimizer shard sections a sequential order. Thus, for example, B. the replicated optimizer shard section 54A-2 can be a sequential first section of optimizer shard 54A, the replicated optimizer shard section 54A-3 can be a sequential second section, and the replicated optimizer shard section 54A-1 can be a sequential last section, or any other arrangement as desired.

[0054] In some examples, nodes 10A-10D can distribute the replicated optimizer shard parts during an initial distribution before or as part of step 210, for example, using an all-to-all operation performed by the respective resiliency framework. In some examples, the distribution of replicated optimizer shard parts can be performed in tandem (e.g., concurrently or nearly concurrently) with operation 212. However, this is just one example, and the initial distribution of replicated optimizer shard parts can be performed at any time before or during the forwarding phase.

[0055] In examples, each node 10A-10D can execute its resilience framework to update replicated optimizer shard parts stored on it, based on updated optimizer states obtained during a reverse propagation phase. As shown below in conjunction with Fig. As described in section 3, for example, each node 10A-10D updates its optimizer shard 54A-54D with respect to global gradients. Each node 10A-10D can then also execute its resilience framework to update replicated optimizer shard sections stored on it by receiving updated optimizer shard sections from the other nodes 10A-10D. In one example, each of the compute nodes 10A-10D can execute its resilience framework to perform an all-to-all operation (e.g., the all-to-all operation 314, as shown below in conjunction with Fig. 3 described) at any point during the in Fig. to perform the forward propagation phase shown in Figure 2. For example, the all-to-all operation can be performed during step 210, e.g., in tandem (e.g., simultaneously or almost simultaneously) with operation 212. In another example, the all-to-all operation can be performed during step 220, e.g., in tandem (e.g., simultaneously or almost simultaneously) with operations 222A-222D.

[0056] The all-to-all operation performed by a given compute node can be used to update replicated optimizer shard portions corresponding to that specific compute node and held on the other compute nodes. For example, compute node 10A can update optimization shard 54A during a backward propagation phase. Compute node 10A can replicate the updated optimization shard 54A, partition the updated optimization shard 54A, and perform an all-to-all operation that updates the replicated optimization shard portions 54A-1 through 54A-3 on each compute node 10B-10D. The all-to-all operation can involve transferring only the portion of the updated optimizer shard 54A to a given compute node that contains the corresponding replicated optimizer shard portion.For example, compute node 10A can send an updated instance of the replicated optimizer shard part 54A-1 to compute node 10B, an updated instance of the replicated optimizer shard part 54A-2 to compute node 10C, and an updated instance of the replicated optimizer shard part 54A-3 to compute node 10B. Compute nodes 10B-10D can similarly send updated instances of their respective replicated optimization shard sections to compute nodes 10A-10D.

[0057] In some examples, hyperparameters can optionally configure the resilience frameworks of each node 10A-10D to replicate a corresponding weight splitter group, divide the replicated weight splitter group into a number of replicated weight splitter group sections according to the hyperparameters, and distribute the replicated weight splitter group sections to other nodes 10A-10D. Each replicated weight splitter section can include data that defines a specific section of the respective weight splitter (e.g., weight splitter 52A of node 10A) that does not overlap with any section of any of the other replicated weight splitter sections originating from the same weight splitter. In this way, the replicated weight splitter sections can collectively define a complete weight splitter.The number of replicated weight splitter sections created can be based on the number of nodes used to perform the distributed training. For example, the number of replicated weight splitter sections can be one less than the total number of nodes.

[0058] In the example of Fig. 2. Node 10A can execute the Resiliency Framework 28 to create three replicated weight splitter parts 52A-1 to 52A-3. For example, the Resiliency Framework 28 can access the storage device(s) 40 to obtain weight splitter 52A and create a copy of weight splitter 52A. The Resiliency Framework 28 can partition the copy of weight splitter 52A into parts by dividing weight splitter 52A into three replicated weight splitter parts 52A-1 to 52A-3. In various examples, the three segments can be essentially the same size.

[0059] In this example, node 10A can execute Resiliency Framework 28 to distribute the replicated weight splitter parts 52A-1 to 52A-3 to nodes 10B-10D. Each node 10B-10D executes its respective Resiliency Framework to receive one of the replicated weight splitter parts 52A-1 to 52A-3 and store the received replicated weight splitter part in one or more storage devices. In the example of Fig. Node 10B holds the replicated weight splitter section 52A-1, node 10C the replicated weight splitter section 52A-2, and node 10D the replicated weight splitter section 52A-3. While this example refers to replicated weight splitter sections 52A-1 through 52A-3, the reference numbers are intended to represent parts of a whole (e.g., 33% of weight splitter 52A) and are not meant to give the replicated weight splitter sections a sequential order. For example, replicated weight splitter 52A-2 can be the first part of the 52A weight splitters, replicated weight splitter 52A-3 can be the second part, and replicated weight splitter 52A-1 can be the last part, or any other arrangement can be chosen.

[0060] In some examples, nodes 10A-10D can distribute the replicated weight splitter parts as part of operation 212, for example, as part of the all-collect operation performed by the respective ML frameworks. In this example, distributing the replicated weight splitter parts might not require any additional communication overhead, since it uses the same operation performed as part of the training process. In another example, replicated weight splitter parts can be distributed before the first iteration of the forward propagation phase 200.

[0061] In examples where replicated weight snippets are distributed across compute nodes, each node 10A-10D can execute its resilience framework to update replicated weight snippets stored on it based on the complete layer reconstructed in step 210. For example, as described above in conjunction with step 210, each node 10A-10D collects weight snippets held by the other nodes to reconstruct a layer. Each node 10A-10D can also execute its resilience framework, for example, during step 220, to update replicated weight snippets using the weight snippets received during step 210. For example, the replicated weight snippets might contain weights learned during a previous iteration of training a current layer, which might need to be updated to reflect the most recently learned weights.The weight shards received in step 210 can contain the latest weights for that layer, which can be used to update the replicated weight shard sections. For example, during an all-collect operation 212, node 10A can receive weight shards 52B-52D and execute the Resiliency Framework 28 to update the replicated weight shard sections 52B-1, 52C-1, and 52D-1 using the corresponding weights in weight shards 52B-52D. Similarly, nodes 10B-10C can update their respective replicated weight shard sections in step 220.

[0062] Fig. Figure 3 is a schematic block diagram of a process flow for a backward propagation phase 300 in distributed training, according to example implementations of the present disclosure. In the example of Fig. 3. The backward propagation phase 300 can be carried out by a multitude of nodes 10, as in connection with Fig. 1 described. Accordingly, one or more of the operations of the backward propagation phase 300 can, for example, be executed by one or more of the application layer 22, the ML framework 24, the interface layer 26 and / or the resilience framework 28, as executed by the processor(s) 20. In the Fig. The example shown in Figure 3 illustrates the backward propagation phase 300 as it is executed by nodes 10A-10D, which may be the same nodes used to execute the forward propagation phase 200. Fig. 2 are used. While the example in Fig. As 3 shows four nodes, the backward propagation phase 300 can be carried out by any number of nodes.

[0063] The in Fig. The reverse propagation phase 300 shown can be used for each layer of an ML model (e.g., model 44 of Fig. 1) This is performed to obtain gradients from a loss function for a given layer with respect to its weights by iterating backward layer by layer from the last layer to the first layer of the ML model. The backward propagation phase 300 can then use the gradients to update the weights and optimization states according to the gradients. The backward propagation phase 300 includes several operations that are described in Fig. The process is divided into steps 310, 320, 330, and 340, which can be executed iteratively for each layer. For example, the backward propagation phase 300 can follow the forward propagation phase 200. Fig. 2 will be carried out.

[0064] In one example, step 310 reconstructs the respective layer of the shared ML model for a given layer at each node 10A-10D. For instance, the ML framework can be executed at each node 10A-10D to perform an all-collect operation 312 to gather weight splitters held at the other nodes and reconstruct the complete layer from the collected weights. Step 310 can be executed in various examples in a manner essentially analogous to step 210 of Fig. 2 corresponds.

[0065] In step 320, each node 10A-10D can execute its respective ML framework to perform operations 322A-322D. Operations 322A-322D may include performing a backward computation operation to obtain input gradients and weight gradients for the fully reconstructed layer with respect to the local inputs at each node 10A-10D. For example, node 10A can execute ML framework 24 to perform a backward computation operation as part of operations 322A to obtain input and weight gradients with respect to inputs held locally at node 10A that were used (and stored in storage device(s) 4) during forward propagation phase 200.The backward computation operation can, for example, but not exclusively, use gradient descent or variants such as stochastic gradient descent to obtain input gradients and weight gradients based on the local inputs for the current layer with respect to a loss function between the current layer and a layer that sequentially follows the current layer in the ML model. Similarly, nodes 10B-10D each execute a corresponding ML frame to perform a backward computation operation as part of operations 322B-322D to obtain input gradients and weight gradients. The backward computation operation performed during each of operations 322A-322D can use the same algorithm (e.g., gradient descent or variants such as stochastic gradient descent) or different algorithms, depending on the desired application.

[0066] Once the input and weight gradients of a given reconstructed layer have been obtained, the weight shards received from other nodes 10A-10D can be discarded to make room for the next iteration. For example, node 10A can execute the ML framework 24 to discard, delete, or otherwise remove weight shards 52B-52D from its storage device(s) 40. Similarly, node 10B can discard weight shards 52A, 52C, and 52D; node 10C can discard weight shards 52A, 52B, and 52D; and node 10D can discard weight shards 52A, 52B, and 52C.

[0067] In step 330, each node 10A-10D can execute its respective machine learning framework to distribute the global weight gradient across nodes 10A-10D via operation 332. Operation 332 can include a reduction-dispersion operation. For example, operation 332 performs a reduction operation that collects sets of local weight gradients obtained from each node 10A-10D and aggregates these sets of local weight gradients to produce a set of global weight gradients. During step 320, for example, each node 10A-10D computed weight gradients for the fully reconstructed layer with respect to the local inputs of each node. As a result, each node 10A-10D receives local weight gradients for each weight splitter 52A-52D with respect to different inputs that are local to each node 10A-10D.Each set of local weight gradients obtained at a given node can include one or more local weight gradients corresponding to each weight shard 52A-52D. The reduction operation of operation 322 collects the sets of local weight gradients from each node 10A-10D and aggregates (e.g., summes) the sets of local weight gradients on a weight shard basis. For example, local weight gradients corresponding to weight class shard 52A can be obtained from each node 10A-10D, which can be summed (or any other aggregation function, such as, but not limited to, average, minimum, maximum, etc.) to obtain a global weight gradient for weight class 52A. Similarly, global weight gradients can be determined for each weight shard 52B-52D, which together with weight shard 52A can form a set of global weight gradients.

[0068] The configured global weight gradients can then be distributed across each of the nodes 10A-10D. For example, the global weight gradients for each weight class can be distributed to the node associated with that class. As a clear example, the global weight gradients for weight half 52A can be distributed to node 10A, the global weight gradients for weight half 52B to node 10B, the global weight gradients for weight half 52C to node 10C, and the global weight gradients for weight half 52D to node 10D.

[0069] In step 340, each node 10A-10D can execute its respective ML frame to update the weights of a given weight type with respect to the global weight gradients obtained in step 330. For example, node 10A can execute ML frame 24 to perform operation 342, which updates the weights of weight type 52A with respect to the global weight gradients for weight type 52A obtained in step 330. Similarly, node 10B can update the weights of weight type 52B, node 10C can update the weights of weight type 52C, and node 10D can update the weights of weight type 52D.In examples, operation 342, performed by each node 10A-10D, can include an optimization algorithm that uses optimization states contained in optimization shards 54A-54D, held at each node, to compute updated weights with respect to the global weight gradients. For example, node 10A can contain an optimization section with optimization states that can be applied as a local instance of an optimization algorithm (e.g., local with respect to node 10A) to update the weight section 52A with respect to the global weight gradients. Similarly, nodes 10B-10D can contain optimizer shards that can be used to update their respective weight shards 10B-10D.

[0070] In step 340, each node 10A-10D can also execute its respective ML framework to update the optimization states of a given optimization shard with respect to the global weight gradients obtained in step 330. For example, node 10A can execute ML framework 24 to perform operation 342, which may involve updating the optimizer states of optimizer shard 54A with respect to the weight gradients obtained in step 330 (e.g., a portion of the global weight gradient corresponding to the respective weight shard). Similarly, node 10B can update the optimization states of optimization shard 54B, node 10C can update the weights of optimization shard 54C, and node 10D can update the optimization states of optimization shard 54D.In examples, the optimizer shards can contain local optimizer states for an optimizer algorithm that is common to the compute nodes 10A-10D, which correspond to the common ML model.

[0071] In the examples disclosed here, as in Fig. As shown in Figure 3, each node 10A-10D can contain replicated weight and optimization splitter pieces that are assigned to other nodes. As described above, each node 10A-10D can execute its resilience framework to create and distribute the replicated weight and optimization splitter pieces. Furthermore, each node 10A-10D, as shown above in conjunction with Fig. 2 described, its Resiliency Framework during step 320 to perform operations that update replicated weight slivers stored on it based on the weight slivers received from other nodes during step 310.

[0072] Furthermore, each node 10A-10D can execute its resilience framework to update replicated optimizer shard portions stored on it, based on updated optimizer states obtained in step 340 of a previous iteration of backward propagation. For example, each node 10A-10D can execute its resilience framework to perform operation 314, which may involve receiving updated optimization shard portions from each of the other nodes 10A-10D and updating replicated optimization shard portions corresponding to and held at each compute node 10A-10D. As a clear example, compute node 10A can execute its resilience framework 28 to perform operation 314, which retrieves updated optimizer states for the replicated optimizer shard sections 54B-1, 54C-1 and 54D-1 of the optimizer shard from each compute node 10B-10D.The resilience framework 28 can then update each replicated optimizer shard section 54B-1, 54C-1, and 54D-1 using the respective optimizer states. The resilience framework for each compute node 10B-10D can be executed similarly to update the respective replicated optimizer shard sections.

[0073] Operation 314 can be an all-to-all operation performed by each compute node 10A-10D, as shown in various examples. For instance, in step 340, compute node 10A can update the optimization shard 54A. Compute node 10A can replicate the updated optimization shard 54A, partition it, and then perform an all-to-all operation (Operation 314) that updates the replicated optimization shard sections 54A-1 to 54A-3 on each compute node 10B-10D. The all-to-all operation can involve transferring only the portion of the updated optimizer shard 54A (e.g., the optimizer states) to a specific compute node that contains the corresponding replicated optimizer shard portion.For example, compute node 10A can send updated optimizer states for replicated optimizer shard part 54A-1 to compute node 10B, updated optimizer states for replicated optimizer shard part 54A-2 to compute node 10C, and updated optimizer states for replicated optimizer shard part 54A-3 to compute node 10B. Compute nodes 10B-10D can similarly send updated optimization states for their respective replicated optimization shard sections to compute nodes 10A-10D.

[0074] In the example of Fig. 3. Operation 314 is performed in tandem (e.g., simultaneously or nearly simultaneously) with operation 312. However, the implementations disclosed here are not intended to be limited to this example. Operation 314 can be performed at any point during the backward propagation phase of Fig. 3 and at any point during the forward propagation phase, as described above in connection with Fig. 2 described. For example, operation 314 can be performed during step 320, such as in tandem (e.g., simultaneously or almost simultaneously) with operations 222A-222D. As another example, operation 314 can be performed during step 330, such as in tandem (e.g., simultaneously or almost simultaneously) with operation 322.

[0075] Fig. Figure 4 is a schematic block diagram illustrating a process flow 400 for providing resilience to node failures during distributed training, as described in the example implementations of this disclosure. In the example of Fig. 4. The process flow 400 can be carried out by a multitude of nodes 10, as in conjunction with Fig. 1 described. Accordingly, one or more of the following may occur: Fig. The operations shown in Figure 4 are performed, for example, by one or more of the application layer 22, the ML framework 24, the interface layer 26, and / or the resilience framework 28, as executed by the processor(s) 20. The process flow 400 of Fig. Figure 4 is shown illustratively as being executed by four nodes 10A-10D. The one in Fig. The elasticity provided in the 400 process flow shown can, however, be executed by any number of nodes.

[0076] In examples, process flow 400 can be executed at any point during distributed training by using the replicated optimizer shard portions held at each of nodes 10A-10D, e.g., during the forward propagation phase 200 and / or the backward propagation phase 300. By distributing an optimization shard held at one node (e.g., optimization shard 54A held by node 10A) to the other nodes (e.g., nodes 10B-10D) as replicated optimization shard portions (e.g., replicated optimization shard portions 54A-1 to 54A-3), the examples contained herein can provide fault tolerance in the event of a failure of, for example, node 10A (e.g., functional failures and the like, as described above in conjunction with...). Fig. 1 described). Similarly, the examples described here can be resilient in the event of a failure of the other nodes due to the distribution of the respective replicated optimization splitter parts.

[0077] For example, in step 410, a node failure can be detected according to the checks (e.g., rules 46), as described above in conjunction with Fig. As described in section 1. In an illustrative example, node 10A may fail for any reason or become unavailable for distributed training, thus classifying node 10A as non-participating and / or non-functional (as indicated by the crossed-out "X"). The remaining functional nodes 10B-10D can be notified of the failure using any notification technique as described above, which can be considered detection of a failure of node 10A.

[0078] Based on the detected functional failure (e.g., in response to the detection that node 10A has failed), each remaining functional node (e.g., nodes 10B-10D) can update its optimizer shard to include a replicated optimizer shard portion corresponding to the failed node. For example, in step 420, each functional node can execute its resilience framework to move a replicated optimizer shard portion corresponding to the failed node (e.g., node 10A) into its respective optimizer shard. In step 430, each functional node can execute its resilience framework to update its respective optimizer shard by merging the replicated optimizer shard portion corresponding to the failed node with its respective optimizer shard. The resulting updated optimizer shard can then be used for training (e.g.,(as an updated training shard).

[0079] In an exemplary implementation, each remaining functional node (e.g., nodes 10B-10D) can perform step 420 by executing a suitable resilience framework to find a replicated optimization splitter section corresponding to the failed node. In some examples, replicated optimization splitter sections can be stored at each node in conjunction with a unique identifier of the compute node (e.g., a MAC address, IP address, or other unique identifier) ​​from which the replicated optimization splitter section originated (e.g., was received).This means, for example, that the replicated optimizer shard part 54A-1 can be stored in storage devices at node 10B that are labeled with the unique identifier of node 10A or otherwise associated with it; the replicated optimizer shard part 54A-2 can be stored in storage devices at node 10C that are labeled with the unique identifier of node 10A or otherwise associated with it; and the replicated optimizer shard part 54A-3 can be stored in storage devices at node 10D that are labeled with the unique identifier of node 10A or otherwise associated with it. Each replicated optimization shard originating from other nodes can be similarly labeled or associated with an identifier of the originating node.Therefore, if a node failure is detected, the resilience framework of the remaining functional nodes can locate under a replicated optimizer shard portion that corresponds to a failed node in the respective storage device(s).

[0080] Once each remaining functional node (e.g., nodes 10B-10D in this example) has located the replicated optimizer shard portion (e.g., replicated optimizer shard portions 54A-1 to 54A-3) received from the failed node (e.g., node 10A in this example), the remaining functional nodes can execute their respective resilience frameworks to move the located replicated optimizer shard portions into their respective optimizer shards. Each remaining functional node's resilience framework can then merge the replicated optimizer shard portion with its respective optimizer shard, thereby creating an updated optimizer shard. As shown in Fig. As shown in step 430, node 10B, for example, can execute its resilience framework to merge the replicated optimizer shard part 54A-1 with optimizer shard 54B to create an updated optimizer shard 64B. Similarly, node 10C can merge the replicated optimizer shard part 54A-2 with optimizer shard 54C to create an updated optimizer shard 64C, and node 10D can merge the replicated optimizer shard part 54A-3 with optimizer shard 54D to create an updated optimizer shard 62D.

[0081] Accordingly, the optimization shard 54A, which was owned by the failed node 10A, can be salvaged by merging it with the other optimization shards of the remaining functional nodes. In this way, the optimization states of optimization shard 54A can be maintained and updated by the remaining functional nodes 10B-10D and used to update the weights. As a result, the distributed machine learning performed by nodes 10A-10D can withstand the failure of node 10A and continue uninterrupted via nodes 10B-10D, using the updated optimizer shards 64B-64D as the updated training shards.

[0082] According to various examples, the updated weight splitter 64B-64D can be increased by the same amount, for example, because each replicated optimizer splitter 54A-1 to 54A-3 is essentially the same size. In this way, the work and communication overhead can be evenly distributed among the remaining functional nodes, and no single node is overburdened.

[0083] In some examples, additional fault tolerance in the event of node failures can be achieved by distributing a weight splitter held at one node (e.g., weight splitter 52A held by node 10A) to the other nodes (e.g., nodes 10B-10D) as replicated weight splitter parts (e.g., replicated weight splitter parts 52A-1 to 52A-3).

[0084] In the Fig. In the example shown, the weight shards contain replicated weight shard portions corresponding to a failed node, based on the functional failure detected in step 410. For example, each functional node can execute its resilience framework to move a replicated weight shard portion corresponding to the failed node (e.g., node 10A) into its respective weight shard in step 420. In this example, the respective weight shard can also be updated in step 430 by merging the replicated weight shard portion corresponding to the failed node with its respective weight shard. The resulting updated weight shardware can then be used for training (e.g., as updated training shardware).

[0085] According to one exemplary implementation of this example, step 420 can optionally include each remaining functional node (e.g., nodes 10B-10D) executing a corresponding resilience framework to locate a replicated weight splitter fragment corresponding to the failed node. In some examples, replicated weight splitter fragments can be stored at each node along with a unique identifier of the compute node from which the replicated weight splitter fragment originated.This means, for example, that the replicated weight splitter section 52A-1 can be stored in the storage device(s) at node 10B, which are identified with the unique identifier of node 10A or otherwise linked to it; the replicated weight splitter section 52A-2 can be stored in the storage device(s) at node 10C, which are identified with the unique identifier of node 10A or otherwise linked to it; and the replicated weight splitter section 52A-3 can be stored in the storage device(s) at node 10D, which are identified with the unique identifier of node 10A or otherwise linked to it. Each replicated weight splitter section originating from other nodes can be similarly identified or linked to an identifier of the originating node.Therefore, if a node failure is detected, the Resiliency Framework, which contains the remaining functional nodes, can locate a replicated weight splitter section corresponding to a failed node in the respective storage device(s).

[0086] Once each remaining functional node (e.g., nodes 10B-10D in this example) has located the replicated weight splitter portion (e.g., replicated weight splitter portion 52A-1 to 52A-3) received from the failed node (e.g., node 10A in this example), step 420 can involve each of the remaining functional nodes executing its respective resilience framework to move the located replicated weight splitter portion into its respective weight splitter. Then, step 430 can involve each remaining functional node executing its respective resilience framework to merge the replicated weight splitter portion with its respective weight splitter, creating an updated weight splitter. In the example in Fig. In the example shown, node 10B can execute its resilience framework in step 430 to merge the replicated weight splitter section 52A-1 with weight splitter 52B to create an updated weight splitter 62B. Similarly, node 10C can merge the replicated weight splitter section 52A-2 with weight splitter 52C to create an updated weight splitter 62C, and node 10D can merge the replicated weight splitter section 52A-3 with weight splitter 52D to create an updated weight splitter 62D.

[0087] Accordingly, the weight splitter group 52A, which was owned by the failed node 10A, can be stored by merging it with the other weight splitter groups of the remaining functional nodes. In this way, the weights of weight splitter 52A can be maintained and updated by the remaining functional nodes 10B-10D. As a result, the distributed machine learning performed by nodes 10A-10D can remain unaffected by the failure of node 10A and continue uninterrupted via nodes 10B-10D using the updated weight splitters 62B-62D as the updated training splitters.

[0088] According to various examples, the updated weight splitter 62B-62D can be increased by the same amount, e.g., because each replicated weight splitter section 52A-1 to 52A-3 is essentially the same size. In this way, the work and communication overhead can be evenly distributed among the remaining functional nodes, and no single node is overburdened.

[0089] To ensure that distributed training is not only effective against the failure of a current node (e.g., node 10A, as in connection with Fig. 4 described), but also against the failure of a subsequent node (e.g., one of nodes 10B-10D in the example of Fig. 4) If the implementation is resilient, it can update replicated optimizer shard sections based on the updated optimizer shards. In particular, replicated optimizer shard sections can be updated by operation 314, which, for example, occurs during backward propagation phase 300, as described in conjunction with Fig. 3 discussed, and / or during the forward propagation phase 200, as in connection with Fig. As discussed in section 2, this is implemented. This means, for example, that the optimizer states of the updated optimizer shards 64B-64D can be replicated, partitioned, and distributed to the nodes using an all-to-all operation. As described above, the optimization states contained in the distributed optimization shards can be used to update the respective replicated optimization shard portions at each node.

[0090] Additionally, in some examples, replicated weight snippets can be updated based on updated weight snippets. For instance, replicated weight snippets can be updated as part of an all-collect operation, which is performed, for example, during forward propagation stage 200 and / or backward propagation stage 300. That is, the updated weight snippets 62B-62D can be replicated, partitioned, and distributed during steps 210 and / or 310 because the weight snippets are distributed to the nodes to reconstruct a specific layer. As described above, the weights contained in the distributed weight snippets can then be used to update the respective replicated weight snippets at each node.

[0091] Fig. Figure 5, for example, shows a schematic block diagram illustrating a process flow for ensuring continuous fault tolerance in the event of node failures, as described in an example from this disclosure. The diagram in Fig. The process flow shown in step 500 can be based on the update of the optimization snippets in step 430 of Fig. 4 follow. In the example of Fig. In step 5, process flow 500 is carried out by the remaining function nodes 10B-10D. Accordingly, one or more of the nodes in Fig. The operations shown in Figure 5 can be performed, for example, by one or more of the application layer 22, the ML framework 24, the interface layer 26, and / or the resilience framework 28, as executed by the processor(s) 20. While the process flow 500 in the figure is executed by three nodes 10B-10D, the resilience provided by the process flow 500 can be executed by any number of nodes.

[0092] As above in connection with Fig. As described in section 4, node 10A may have failed, and nodes 10B-10D may each have generated updated optimization shards 64B-64D based on replicated optimization shard portions received from node 10A. However, at the beginning of step 510, the replicated optimizer shard portions in each remaining functional node—except those corresponding to node 10A—may be unchanged. That is, for example, node 10B may contain the replicated optimizer shard portions 54C-2 and 54D-2, node 10C may contain the replicated optimizer shard portions 54B-2 and 54D-3, and node 10D may contain the replicated optimizer shard portions 54B-3 and 54C-3.Thus, the replicated optimizer shard parts 54B-1, 54C-1 and 54D-1 may have been lost due to the failure of node 10A, and the replicated optimizer shard parts for optimizer shard parts 64B-64D may be incomplete or missing (e.g., the parts of optimizer shard parts 64B-64D corresponding to replicated optimizer shard parts 54A-1 to 54A-3 may not be represented in a replicated optimizer shard currently held at a node).

[0093] To ensure that the distributed training performed by the remaining functional nodes 10B-10D is resilient to future node failures, each remaining functional node 10B-10D can execute its resilience framework to update replicated optimizer shard segments through steps 510 and 520. Steps 510 and 520 can be performed as part of a forward propagation phase 200 or a backward propagation phase 300 following the failure of a node (e.g., node 10A). In either case, the forward propagation phase 200 or the backward propagation phase 300 can be performed as described above, but only with the remaining functional nodes (e.g., nodes 10B-10D, since node 10A is no longer present in the due process) and modified as described below to update replicated optimizer shard segments.

[0094] For example, in step 510, each node 10B-10D can execute its resilience framework to perform an all-collect operation 512 to gather weight shards held at the other nodes and reconstruct the layer from the collected weights. The all-collect operation 512 can be the all-collect operation 212 described above or the all-collect operation 312. Thus, after operation 512, each node 10B-10D holds weight shards 62B-62D.

[0095] In the example of Fig. In step 510, each node 10B-10D can execute its resilience framework to perform an operation 514 that distributes portions of its updated optimizer shards 64B-64D to the other nodes. For example, compute node 10B can replicate the updated optimization shard 64B, partition the updated optimization shard 64B, and perform an operation 514 (e.g., an all-to-all operation) to deliver portions (e.g., portions of the optimization states) of the updated optimization shard 64B to compute nodes 10C and 10D. Compute nodes 10C and 10D can similarly deliver portions of the updated optimization shards 64C and 64D to compute nodes 10B-10D.

[0096] For example, each node 10B-10D uses portions of the updated optimizer shards 64B-64D obtained in step 510 to update replicated optimizer shard portions held in each node 10B-10D. The updated optimizer shards 64B-64D obtained in step 510 may contain the latest optimizer states for that layer, which can be used in step 520 to update and resize the replicated optimizer shard portions. For example, during operation 514, node 10B can receive parts of the updated optimizer shards 64C and 64D and execute its resilience framework to update the replicated optimizer shard parts 54C-2, 54D-2 so that they contain some of the optimizer states of the updated optimizer shards 64C and 64D, resulting in updated replicated optimizer shard parts 64C-1, 64D-1.Similarly, node 10C can update the replicated optimizer shard sections 54B-2, 54D-3 to create updated replicated optimizer shard sections 64B-1 and 64D-2, and node 10D can update the replicated optimizer shard sections 54B-3, 54C-3 to create updated replicated weight shard sections 64B-2 and 64C-2.

[0097] In examples, updating replicated optimizer shard sections can include any technique for updating replicated optimizer shard sections using sections of the updated optimizer shard sections 64B-64D obtained from compute nodes. In one example, updating replicated optimizer splitter sections can include replacing previous replicated optimizer splitter sections (e.g., replicated optimizer splitter sections 54C-2, 54D-2) with sections of updated optimizer splitters (e.g., sections of optimizer splitters 64C and 64D). In one example, compute node 10B can delete the replicated optimizer shard sections 54C-2, 54D-2 and store parts of the optimizer shard sections 64C and 64D as replicated optimizer shard sections 64C-1 and 64D-1.Compute nodes 10C and 10D can perform similar operations to generate replicated optimizer shard parts 64B-1, 64D-2, 64B-2, and 64C-2. In another example, updating the replicated optimizer shard parts can involve merging the received part of the updated optimizer shard with the replicated optimizer shard part stored on it. In each case, the updated replicated optimizer shard parts can collectively form the entire updated optimizer shard. This means that, for example, the replicated parts of the optimization shards 64B-1 and 64B-2 together represent the optimization shard 64B, the replicated parts of the optimization shards 64C-1 and 64C-2 together represent the optimization shard 64C, and the replicated parts of the optimization shards 64D-1 and 64D-2 together represent the optimization shard 64D.

[0098] Additionally, in examples where the compute nodes 10B-10D have generated replicated weight splitter sections 64B-64D, as in the example above, Fig. As described in section 4, at the beginning of step 510, the replicated weight splitter sections at each remaining functional node—except for those corresponding to node 10A—will be unchanged. This means, for example, that node 10B may contain the replicated weight splitter parts 52C-2 and 52D-2, node 10C may contain the replicated weight splitter parts 52B-2 and 52D-3, and node 10D may contain the replicated weight splitter parts 52B-3 and 52C-3. Thus, the replicated weight card parts 52B-1, 52C-1 and 52D-1 may be lost due to the failure of node 10A, and the replicated weight card parts for weight cards 62B-62D may be incomplete or missing (e.g., the parts of weight cards 62B-62D corresponding to replicated weight card parts 52A-1 to 52A-3 may not be represented in a replicated weight card currently held at a node).

[0099] In this example, each remaining functional node 10B-10D can execute its resilience framework to update replicated weight splitter portions in steps 510 and 520. For example, as described above, each node 10B-10D performs a collection operation 512 to gather weight splitters held at the other nodes and reconstructs the layer from the gathered weight splitters in step 510. Thus, after operation 512, each node 10B-10D holds 62B-62D weight splitters. In step 520, in this example, each node 10B-10D can execute its resilience framework to perform operations 522B-522D. Operations 522B-522D can include updating replicated weight splitter portions held at each node 10B-10D using the weight splitters received in step 510.Operations 522B-522D can be performed as part of operations 222B-222D or 322B-322D described above, which may include updating the replicated weight splitter sections, as described above in conjunction with the . Fig. 2 and Fig. 3 described.

[0100] For example, each node 10B-10D can perform a corresponding operation 522B-522D to use updated weight shards 62B-62D obtained in step 510 to update replicated weight shard pieces held at each node 10B-10D. The updated weight shards 62B-62D obtained in step 510 can contain the latest weights for that layer, which can be used in step 520 to update and resize the replicated weight shard pieces. For example, during the collection process 512, node 10B can receive weight shards 62C and 62D and execute its resilience framework to update the replicated weight shard sections 52C-2, 52D-2 so that they contain the weights of the updated weight shards 62C and 62D, resulting in updated replicated weight shard sections 62C-1 and 62D-1.Similarly, node 10C can update the replicated weight splitter sections 52B-2, 52D-3 to create updated replicated weight splitter sections 62B-1 and 62D-2, and node 10D can update the replicated weight splitter sections 52B-3, 52C-3 to create updated replicated weight splitter sections 62B-2 and 62C-2.

[0101] Fig. Figure 6 shows an example of a computer component that can be used to implement node fault tolerance in distributed training environments in accordance with various implementations. As shown in Fig. As shown in Figure 6, the computer component 600 can be, for example, a server computer, a controller, or another similar computer component capable of processing data ( ). In the example implementation of Fig. In Figure 6, the computer component 600 comprises a hardware processor 602 and a machine-readable storage medium 604. In an example, the computer component 600 can be an example of one of the nodes 10 of Fig. 1. In another example, the hardware processor 602 can be a multitude of hardware processors coupled with a multitude of machine-readable storage media, which are a multitude of nodes 10 of Fig. 1 can be represented.

[0102] The Hardware Processor 602 may consist of one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices capable of retrieving and executing instructions stored on the machine-readable memory medium 604. The Hardware Processor 602 can retrieve, decode, and execute instructions, such as instructions 606-612, to control processes or operations for fault tolerance. Alternatively or in addition to retrieving and executing instructions, the Hardware Processor 602 may include one or more electronic circuits comprising electronic components for performing the functionality of one or more instructions, such as, but not limited to, graphics processing units (GPUs), programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), or other electronic circuits.

[0103] A machine-readable storage medium, such as the machine-readable storage medium 604, can be any electronic, magnetic, optical, or other physical storage device that contains or stores executable instructions. The machine-readable storage medium 604 can be, for example, random access memory (RAM), non-volatile random access memory (NVRAM), electrically erasable programmable solid-state memory (EEPROM), a storage device, an optical disk, or the like. In some embodiments, the machine-readable storage medium 604 can be a non-transient storage medium, the term "non-transient" excluding the transitive transmission signals. As described in detail below, the machine-readable storage medium 604 can be encoded with executable instructions, for example, instructions 606-612.

[0104] The 602 hardware processor can execute instruction 606 and store a first optimizer subset of optimizer states of a common ML model on a first compute node. The first optimization subset comprises, for example, a subset or segment of optimization states that pertain to the first compute node and can be used to define a local instance of a common optimization algorithm, as described above in the context of the Fig. 1-5 described.

[0105] The hardware processor 602 can execute instruction 608 to receive a first plurality of optimizer splitter pieces from a first plurality of compute nodes through the first compute node. For example, the first compute node can be part of a cluster of compute nodes in a distributed training network, as shown above in conjunction with Fig. 1. A distributed training network is described. The cluster of compute nodes in this example can include the first compute node and the first plurality of compute nodes. Each optimization shard can be received by the first compute node from a corresponding compute node of the first plurality of compute nodes. That is, for example, each of the first plurality of compute nodes can provide an optimization shard segment to the first compute node, which together can form the first plurality of optimization shard segments. Each optimizer shard segment can be a replica of a portion of an optimizer shard containing the optimizer states of the common optimization algorithm, which is stored at the respective compute node of the plurality of compute nodes. For example, as above in conjunction with the Fig. As described in 1 to 5, each of the first plurality of compute nodes stores an optimizer shard of the optimizer states (e.g., a specific segment of the optimizer states that define a local instance of the common optimization algorithm), each of which can be replicated and partitioned into optimizer shard sections and shared with the first compute node.

[0106] In examples, the first multitude of optimization snippets can be received during one of the following phases: forward propagation or backward propagation of the training of the joint ML model, as above in conjunction with the Fig. 2-5 described.

[0107] In some examples, the first compute node can be configured to provide a second set of optimizer shard parts to the first set of compute nodes. For instance, the 602 hardware processor can execute instructions that cause the first compute node to replicate the first optimization shard, partition the replicated first optimization shard into the second set of optimization shard parts, and transfer the second set of optimization shard parts to the first set of compute nodes. The transfer of the second set of optimizer shard parts can be performed during an all-collect operation, which is carried out either during forward or backward propagation of the joint machine learning model training.

[0108] The 602 hardware processor can execute instruction 610 to update the first optimizer shard in response to the failure of a compute node from the first compute plurality by merging an optimizer shard portion corresponding to the failed compute node with the first optimizer shard. For example, a failure of at least one of the first compute plurality can be detected, as described above in conjunction with the Fig. 1 and Fig. 4 described. In response to the detection of the failure (e.g., by receiving a notification or alarm), the first compute node can update the first optimization shard by merging an optimization shard section from the first plurality of optimization shard sections that corresponds to the failed node of the first plurality of shards nodes (e.g., the shard section received from or otherwise originating from the failed node) with the first optimization shard.

[0109] In examples, each of the second set of compute nodes can be configured to update a corresponding optimizer shard of the first set of optimizer shards with an optimizer shard portion that corresponds to the failed compute node in response to the detection of the failure. In this example, the second set of compute nodes can be the first set of compute nodes with the failed compute node removed.

[0110] In examples, the hardware processor can execute 602 instructions to receive a third plurality of optimization splitter parts from the second plurality of compute nodes via the first compute node. Each optimizer splitter part of the third plurality of splitter parts can be a replica of a part of a respective updated optimizer splitter stored at the respective compute node of the second plurality of compute nodes. The first compute node can partition the updated first splitter into a fourth plurality of optimizer splitter parts and transfer the fourth plurality of optimizer splitter parts to the second plurality of compute nodes, for example, during an all-collect operation performed either during forward or backward propagation of the joint ML model training.

[0111] Hardware processor 602 can execute instruction 612 to update the weights of the joint ML model based on the updated first optimization split. As mentioned above in the context of the Fig. 4 and Fig. 5 with regard to the Fig. 2 and Fig. As described in section 3, the updated first optimizer shard can, for example, be used as a training optimizer shard for future iterations of a backward propagation phase of a distributed training process, such as when updating the weights of the joint ML model. In examples, training the joint ML model can be performed using the updated first optimizer shard as well as the updated first plurality of optimizer shards, as described in conjunction with Fig. 5 described. In this way, the distributed training, which is carried out by the distributed training network of the first compute node and the first plurality of compute nodes, can withstand a node failure, so that the learning can continue without interruption.

[0112] Fig. Figure 7 shows another example of a computer component that can be used to implement node fault tolerance in distributed training according to various implementations. As in Fig. As shown in Figure 7, computer component 700 can be, for example, a server computer, a controller, or another similar computer component capable of processing data. In the example implementation of Fig. In section 7, the computer component 700 comprises a hardware processor 702 and a machine-readable storage medium 704. In an example, the computer component 700 can be an example of one of the nodes 10 from Fig. Be 1.

[0113] The Hardware Processor 702 may consist of one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices capable of retrieving and executing instructions stored in the machine-readable memory medium 704. The Hardware Processor 702 can retrieve, decode, and execute instructions, such as instructions 706-712, to control processes or operations for fault tolerance. Alternatively or in addition to retrieving and executing instructions, the Hardware Processor 702 may include one or more electronic circuits containing electronic components for performing the functionality of one or more instructions, such as, but not limited to, graphics processing units (GPUs), a field-programmable gate array (FPGA), application-specific integrated circuits (ASICs), or other electronic circuits.

[0114] A machine-readable storage medium, such as the machine-readable storage medium 704, can be any electronic, magnetic, optical, or other physical storage device that contains or stores executable instructions. The machine-readable storage medium 704 can be, for example, a working memory (RAM), a non-volatile working memory (NVRAM), an electrically erasable programmable solid-state memory (EEPROM), a storage device, an optical disk, and the like. In some embodiments, the machine-readable storage medium 704 can be a non-transitory storage medium, the term "non-transitory" excluding the transitive transmission signals. As described in detail below, the machine-readable storage medium 704 can be encoded with executable instructions, for example, instructions 706-712.

[0115] The 702 hardware processor can execute instruction 706 to receive each of a first plurality of optimization splitter sections from a respective compute node of a first plurality of compute nodes, where each optimization splitter section is a segment of a respective optimization splitter of optimization states associated with that compute node. In examples, the optimization states define an optimization algorithm of a machine learning model. For instance, each optimization section associated with a respective compute node can comprise a subset or segment of the optimization states that can define a local instance of the optimization algorithm, as described above in the context of the Fig. 1-5 described. The computer component 700 can receive parts of these optimization splitters from the respective compute nodes of the first plurality of compute nodes, which may have replicated (e.g., copied) their respective optimization splitters and may have divided the replicated optimization splitters into parts. Each of the first plurality of compute nodes can transfer a part of its replicated optimization splitters to the computer component 700 for storage in a machine-readable storage medium 704, as described above in connection with the Fig. 1-5 described.

[0116] The hardware processor 702 can execute instruction 708 to restore an optimizer shard corresponding to the failed compute node based on the detection of a failure of at least one of the compute nodes of the first plurality of compute nodes. This is done by updating a first optimizer shard with the optimizer shard portion associated with the failed compute node. The first optimizer shard, which can be connected to the compute component 700 (e.g., locally within the compute component 700), can be a segment of optimizer states, similar to each optimizer shard of the first plurality of optimizer shards, as described above.Instruction 708 can be executed to locate a part of the optimization splitter associated with the failed node stored in the machine-readable storage medium 704, and to merge the located part with the first optimization splitter of computer component 700, for example, as above in conjunction with the . Fig. 3 and Fig. 4 described. In this way, at least the portion of the optimization shard associated with the failed node, which would otherwise have been lost due to the failure, can be recovered. In examples, each compute node of the first plurality of compute nodes that is not the failed node (e.g., a subset of the first plurality of compute nodes that does not include the failed node) can similarly contain a portion of the optimization shard associated with the failed node, which can be located and merged with an optimization shard associated with each compute node.

[0117] The 702 hardware processor can execute instruction 710 to propagate a second set of optimization splitter sections of the updated first optimization splitter to a subset of the first set of compute nodes, either during forward or backward propagation of the ML model training. For example, as in conjunction with Fig. 3 and Fig. As described in section 4, the second set of optimizer shard parts are transferred to the subset of the first set of compute nodes during an all-to-all operation performed during forward or backward propagation of the ML model training. For example, the ML model can be trained during a first iteration of fully split data parallelism partly based on the first optimizer shard, and during a second (e.g., subsequent) iteration of fully split data parallelism partly based on the updated first optimizer shard, as described above in conjunction with the Fig. 1-5 described

[0118] While the examples described herein relate to providing fault tolerance in the event of the failure, malfunction, or other unavailability of a distributed training compute node, they are not limited to a single failure. The examples disclosed here can be extended to multiple simultaneous failures, resulting in more than one distributed training compute node becoming unavailable. In this case, the compute nodes (e.g., compute nodes 10A-10G in Fig. 1) Store multiple replicated optimization splitters for each of the other compute nodes. For example, compute node 10A, as in the Fig. 2 and Fig. Figure 3 shows a first set of replicated optimization shard parts corresponding to compute node 10B, a second set of replicated optimization shard parts corresponding to compute node 10C, and a third set of replicated optimization shard parts corresponding to compute node 10D. Similarly, compute nodes 10B-10D can store multiple sets of replicated shard sections, with each set corresponding to a specific compute node. For example, in the event of a failure of compute nodes 10A and 10B, replicated optimization shard sections for each compute node 10A and 10B, stored on compute nodes 10C and 10D, can be merged with optimization shards 54C and 54D, respectively, to generate updated optimization shards that, in this example, ensure the fault tolerance of optimization shards 54A and 54B. The related Fig. The functions discussed in points 1 to 7 can essentially function in a similar way, except that several compute nodes become functionally unavailable and their respective weight shards can be recovered by merging them with the weight shards of the remaining functional compute nodes.

[0119] In some examples, compute nodes (e.g., compute nodes 10A-10G of Fig. 1) also store multiple replicated weight splitter parts for each of the other compute nodes. For example, compute node 10A can also store a first set of replicated weight splitter parts corresponding to compute node 10B, a second set of replicated weight splitter parts corresponding to compute node 10C, and a third set of replicated weight splitter parts corresponding to compute node 10D. Compute nodes 10B-10D can also store multiple sets of replicated weight splitter parts in this example, with each set corresponding to a specific compute node. Thus, for example, in the event of a failure of compute nodes 10A and 10B, replicated weight splitter parts for each compute node 10A and 10B, stored on compute nodes 10C and 10D, can be replaced with weight splitters 52C and 52D, respectively.52D are merged to create updated weight splitters that, in this example, provide failover for weight splitters 52A and 52B.

[0120] Fig. Figure 8 shows a block diagram of an example computer system 800, in which various embodiments of the configurations described here can be implemented. The computer system 800 comprises a bus 802 or other communication mechanism for transmitting information, and one or more hardware processors 804 coupled to the bus 802 for processing information. The hardware processor(s) 804 can be, for example, one or more general-purpose microprocessors. The computer system 800 can, for example, be configured as a compute node 10 of Fig. 1 will be implemented.

[0121] The Computer System 800 also includes a main memory 806, such as random access memory (RAM), a cache, and / or other dynamic memory devices connected to the 802 bus to store information and instructions to be executed by the 804 processor. The main memory 806 can also be used to store temporary variables or other intermediate information during the execution of instructions to be carried out by the 804 processor. When such instructions are stored in memory media accessible to the 804 processor, the Computer System 800 becomes a specialized machine adapted to perform the operations specified in the instructions. The main memory 806, as well as other memory components, can store instructions which, when executed by the 804 processor, cause the 804 process to perform one or more operations in conjunction with the Fig.to perform the operations described in 1-5.

[0122] The Computer System 800 also includes a read-only memory (ROM) 808 or other static storage device connected to the 802 bus to store static information and instructions for the 804 processor. A storage device 810, such as a magnetic disk, an optical disk, or a USB flash drive, etc., is provided and connected to the 802 bus to store information and instructions.

[0123] The Computer System 800 can include a user interface module for implementing a graphical user interface, which can be stored on a mass storage device as executable software code that is executed by the computer device(s). This and other modules can include components such as software components, object-oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.

[0124] In general, the terms "component," "engine," "system," "database," "data store," and the like, as used here, can refer to logic embodied in hardware or firmware, or to a collection of software instructions that may have entry and exit points and are written in a programming language such as Java, C, or C++. A software component may be compiled and linked into an executable program, installed in a dynamic link library, or written in an interpreted programming language such as BASIC, Perl, or Python. It is understood that software components may be invoked by other components or by themselves, and / or may be invoked in response to detected events or interruptions. Software components configured to run on computer devices may be stored on a computer-readable medium, such as...Software code may be provided on a compact disc, digital video disc, flash drive, magnetic disk, or other tangible medium, or as a digital download (and may initially be stored in a compressed or installable format that requires installation, decompression, or decryption before execution). Such software code may be stored partially or entirely in memory on the executing computer device for execution by that device. Software instructions may be embedded in firmware, such as an EPROM. Hardware components may consist of interconnected logic units, such as gates and flip-flops, and / or programmable units, such as programmable gate arrays or processors.

[0125] The Computer System 800 can implement the techniques described herein using custom hard-wired logic, one or more ASICs or FPGAs, firmware, and / or program logic which, in combination with the Computer System, causes or programs the Computer System 800 to be a special-purpose machine. According to one embodiment, the techniques described herein are executed by the Computer System 800 in response to the Processor(s) 804 executing one or more sequences of one or more instructions contained in the main memory 806. Such instructions may be read into the main memory 806 from another storage medium, such as the Storage Device 810. The execution of the instruction sequences contained in the main memory 806 causes the Processor(s) 804 to perform the process steps described herein.In alternative embodiments, hard-wired circuits can be used instead of, or in combination with, software instructions.

[0126] The term "non-volatile media" and similar terms as used here refer to all media that store data and / or instructions that make a machine operate in a particular way. Such non-volatile media can include both non-volatile and volatile media. Non-volatile media include, for example, optical or magnetic disks, such as the Storage Device 810. Volatile media include dynamic memory, such as the Main Memory 806. Common forms of non-volatile media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tapes or other magnetic data storage media, CD-ROMs, other optical data storage media, physical media with hole patterns, RAM, PROM and EPROM, FLASH-EPROM, NVRAM, other memory chips or cartridges, and their networked versions.

[0127] Non-transmittable media differ from transmission media but can be used in conjunction with them. Transmission media are involved in the transfer of information between non-transmittable media. Examples of transmission media include coaxial cable, copper and fiber optic cables, including the wires that make up the 802 bus. Transmission media can also take the form of sound or light waves, such as those generated in radio and infrared data communication.

[0128] The Computer System 800 also includes a network interface 818 (also called a communication interface) connected to bus 802. The network interface 818 provides a two-way data communication connection to one or more network connections that are connected to one or more local area networks (LANs). For example, the network interface 818 could be an ISDN (Integrated Services Digital Network) card, a cable modem, a satellite modem, or a modem to establish a data communication connection to a corresponding type of telephone line. Alternatively, the network interface 818 could be a LAN (Local Area Network) card to establish a data communication connection to a compatible LAN (or a WAN component for communication with a WAN). Wireless connections could also be implemented.In each of these implementations, the 818 network interface sends and receives electrical, electromagnetic, or optical signals that transmit digital data streams with various types of information.

[0129] A network connection typically enables data communication over one or more networks to other data devices. For example, a network connection might establish a connection over a local area network to a host computer or to data devices operated by an Internet service provider (ISP). The ISP, in turn, provides data communication services over the worldwide packet data communication network, commonly known today as the "Internet." Both the local area network and the Internet use electrical, electromagnetic, or optical signals to transmit digital data streams. The signals in the various networks and the signals on the network link and across the 818 network interface, which transmit digital data to and from the Computer System 800, are examples of transmission media.

[0130] The Computer System 800 can send messages and receive data, including program code, over the network(s), network connection, and network interface 818. In the Internet example, a server could transmit requested code for an application program over the Internet, the ISP, the local network, and network interface 818.

[0131] The received code can be executed by the 804 processor upon receipt and / or stored in the 810 memory device or other non-volatile memory for later execution.

[0132] Each of the processes, methods, and algorithms described in the preceding sections can be embodied in code components and fully or partially automated by one or more computer systems or processors, including computer hardware. These computer systems or processors can also be operated in a cloud computing environment or as Software as a Service (SaaS). The processes and algorithms can be partially or fully implemented in application-specific circuits. The various features and procedures described above can be used independently or combined in various ways.Various combinations and subcombinations are said to fall within the scope of this disclosure, and certain procedural or process blocks may be omitted in some implementations. The methods and processes described herein are also not restricted to a particular order, and the associated blocks or states may be executed in other suitable orders, in parallel, or otherwise. Blocks or states may be added to or removed from the disclosed examples. The execution of certain operations or processes may be distributed across computer systems or computer processors, not just within a single machine, but distributed across a number of machines.

[0133] A circuit can be implemented in any form of hardware, software, or a combination thereof. For example, one or more processors, controllers, ASICs, PLAs, PALs, CPLDs, FPGAs, logic components, software routines, or other mechanisms can be implemented to form a circuit. In implementation, the various circuits described here can be implemented as discrete circuits, or the described functions and features can be used partially or entirely by one or more circuits. Even if various features or functional elements are individually described or claimed as separate circuits, these features and functions can be shared by one or more common circuits, and such a description is not intended to require or imply that separate circuits are necessary to implement these features or functions.If a circuit is wholly or partially implemented in software, this software can be implemented to work with a computer or processing system capable of performing the functionality described in this context, such as the Computer System 800.

[0134] The term "or" used here can be understood in both an inclusive and an exclusive sense. Furthermore, the singular description of resources, processes, or structures is not to be understood as excluding the plural. Conditional expressions such as "may," "could," "might," or "can" are, unless expressly stated otherwise or understood differently in context, generally to be understood as meaning that certain embodiments include certain features, elements, and / or steps, while other embodiments do not.

[0135] Unless explicitly stated otherwise, the terms and expressions used in this document, as well as their variations, are to be understood as open rather than restrictive. Adjectives such as "conventional," "traditional," "normal," "standard," "known," and terms with similar meanings are not to be understood as limiting the described subject matter to a specific period or to an item available at a particular time, but should be understood as encompassing conventional, traditional, normal, or standard technologies that may be available or known now or at any time in the future.The presence of expansive words and phrases such as "one or more", "at least", "but not limited to" or similar phrases in some cases is not to be understood as implying that the narrower case is intended or required when such expansive phrases are not present.

Claims

[1] A system that includes the following: a first plurality of compute nodes, wherein each of the first plurality of compute nodes stores a sliver of a first plurality of optimization slivers of a common machine learning (ML) model; and a first compute node that stores a first optimization sliver of the joint ML model, wherein the first compute node is configured such that it: a first plurality of optimizer splitter parts, wherein each optimizer splitter part is received by a respective compute node of the first plurality of compute nodes and is a replica of a part of a respective optimizer splitter of the first plurality of optimizer splitters that is stored at the respective compute node; and In response to the failure of a compute node of the first plurality of compute nodes, the first optimizer shard is updated with an optimizer shard portion corresponding to the failed compute node. where the weights of the joint ML model are updated using the updated first optimization split. [2] System according to claim 1, wherein the first compute node is further configured such that it: replicate the first optimizer shard; Partitioning the replicated first optimizer shard into a second plurality of optimizer shard parts; and Transferring the second set of optimizer slivers to the first set of computation nodes. [3] System according to claim 2, wherein the first compute node transfers the second plurality of optimizer splitter sections to the first plurality of compute nodes during an all-to-all operation performed during one of the following operations: a forward propagation and a backward propagation of the training of the joint ML model. [4] System according to claim 2, wherein the second plurality of optimization splitter parts comprises a number of equal parts of the first optimization splitter, wherein each equal part of the first optimization splitter is transferred to a computation node of the first plurality of computation nodes. [5] System according to claim 4, wherein the number of identical optimizer splitter sections is equal to the number of computation nodes of the first plurality of computation nodes. [6] System according to claim 1, wherein, in response to the failure of the compute node of the first plurality of compute nodes, each of a second plurality of compute nodes is configured such that it: Updating each optimizer shard from the first set of optimizer shards with an optimizer shard portion corresponding to the failed compute node, where the second set of computation nodes is the first set of computation nodes excluding the failed computation node. [7] System according to claim 6, wherein the first compute node is further configured such that it: a third plurality of optimization splitter sections are received from the second plurality of compute nodes, each optimization splitter section of the third plurality of optimization splitter sections being a replica of a section of a respective updated optimization splitter stored at the respective compute node of the second plurality of compute nodes; Partitioning the updated first optimizer shard into a fourth set of optimizer shard parts; and Transferring the fourth set of optimizer sliver pieces to the second set of computation nodes. [8] System according to claim 1, wherein the first compute node is configured such that it: Storing sets of optimizer splitter sections, wherein each set of optimizer splitter sections comprises multiple optimizer splitter sections received by a respective compute node of the first plurality of compute nodes and are replicas of sections of an optimizer splitter of the first plurality of optimizer splitters corresponding to the respective compute node; and In response to a failure of two or more compute nodes of the first plurality of compute nodes, the first optimizer shard is updated with the sets of optimizer shard parts corresponding to the two or more compute nodes. [9] System according to claim 1, wherein the first compute node is further configured such that it: In response to the failure of a compute node of the first plurality of compute nodes, a weight splitter group is updated with a weight splitter group fraction corresponding to the failed compute node, wherein the weight splitter group comprises weights of the common ML model that is local to the first compute node, and wherein the weight splitter group fraction is stored in the first compute node that was received from the failed compute node prior to the failure. [10] A non-transitory, computer-readable medium containing instructions which, when executed by one or more processors, cause the one or more processors to perform a procedure comprising: Storage of an initial optimization sliver of optimization states of a common machine learning (ML) model on a first compute node; Receiving a first plurality of optimizer slivers from a first plurality of compute nodes by the first compute node, wherein each optimizer sliver is received by a respective compute node of the first plurality of compute nodes and is a replica of a part of a respective optimizer sliver of the optimizer states of the common ML model stored at the respective compute node; in response to the failure of a compute node of the first plurality of compute nodes, updating the first optimizer shard by merging an optimizer shard portion corresponding to the failed compute node with the first shard; and Update of the weights of the joint ML model based on the updated first optimization split. [11] Non-transitory, computer-readable medium according to claim 10, wherein the first computation node receives the first plurality of optimizer splitter sections during an all-to-all operation performed either during forward propagation or backward propagation of the training of the joint ML model. [12] Non-transitory, computer-readable medium according to claim 10, wherein the method further comprises: the replication of the first optimizer shard; Partitioning the replicated first optimizer shard into a second plurality of optimizer shard sections; and Transferring the second plurality of optimizer sliver pieces to the first plurality of computation nodes. [13] Non-transitory, computer-readable medium according to claim 12, wherein the method further comprises: Transfer of the second set of optimizer splitter sections by the first compute node to the first set of compute nodes during an all-to-all operation, which is performed either during forward propagation or backward propagation of the training of the joint ML model. [14] Non-transitory, computer-readable medium according to claim 12, wherein partitioning the first optimizer splitter into a second plurality of optimizer splitter parts comprises: Partitioning the replicated first optimizer shard into a number of equal parts, with each equal part of the first optimizer shard being transferred to a compute node of the first plurality of compute nodes. [15] Non-transitory, computer-readable medium according to claim 14, wherein the number of identical optimizer splitter sections is equal to the number of computation nodes of the first plurality of computation nodes. [16] Non-transitory, computer-readable medium according to claim 10, wherein the method further comprises: In response to the failure of the compute node from the first plurality of compute nodes, each optimizer shard from the first plurality of optimizer shards is updated with an optimizer shard portion corresponding to the failed compute node by each from a second plurality of compute nodes, wherein the second plurality of compute nodes is the first plurality of compute nodes excluding the failed compute node. [17] Non-transitory, computer-readable medium according to claim 16, wherein the method further comprises: Received by the first compute node from the second plurality of compute nodes, a third plurality of optimization splitter sections, each splitter section of the third plurality of optimization splitter sections being a replica of a section of a respective updated optimization splitter stored at the respective compute node of the second plurality of compute nodes; Partitioning the updated first optimizer shard into a fourth set of optimizer shard sections; and Transferring the fourth plural of optimizer sliver pieces to the second plural of compute nodes. [18] A compute node with: a memory that stores instructions and a first optimizer splitter of optimizer states of a machine learning (ML) model; and a processor that is operationally connected to the memory and configured to execute instructions to: Receiving each of a first plurality of optimizer slivers from a respective compute node of a first plurality of compute nodes, wherein each optimizer sliver is a part of a respective optimizer sliver of the ML model that is connected to the respective compute node; based on detecting a failure of at least one compute node of the first plurality of compute nodes, restoring an optimizer shard corresponding to the at least one compute node by updating the first optimizer shard with the optimizer shard portion connected to the at least one compute node; and Transferring a second set of optimization splitter parts of the updated first optimization splitter to a subset of the first set of computation nodes during a backpropagation of the ML model training. [19] Computing node according to claim 18, wherein the second plurality of optimizer splitter sections is transferred to the subset of the first plurality of computing nodes during an all-to-all operation. [20] Computing node according to claim 18, wherein the processor is further configured to execute the instructions to: Train the ML model during an iteration of fully split data parallelism partly based on the first optimization splitter; and The ML model is trained during a subsequent iteration of the fully split data parallelism, partially based on the updated first optimization splitter.