Error Recovery System and Method
By distributing redundant model copies across worker groups, the system facilitates rapid error recovery in multiprocessor AI systems, minimizing downtime and maintaining throughput.
Patent Information
- Application Number
- JP2022541236
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-27
- Filing Date
- 2020-12-16
- Publication Date
- 2025-07-09
- Estimated Expiration
- 2040-12-16
AI Technical Summary
In multiprocessor AI systems, when one node experiences an error, a significant amount of data needs to be recomputed, leading to slow recovery and reduced throughput due to the reliance on global checkpoints.
Implementing a system where redundant copies of the AI model are stored across worker groups, allowing for rapid recovery by reloading the previous iteration's model at the affected node, eliminating the need for abrupt returns to global checkpoints.
Enables fast error recovery within seconds, reducing the impact of failures on performance by allowing affected nodes to quickly resume processing without reprocessing previous iterations.
Smart Images

Figure 0007705225000002 
Figure 0007705225000003 
Figure 0007705225000004
Abstract
Description
Technical Field
[0001]
[0001] This disclosure relates to computing. More particularly, this disclosure relates to techniques for error recovery in artificial intelligence processing.
[0002]
[0002] Artificial intelligence (AI) processing typically includes loading part or all of an AI model (e.g., a neural network model) onto one or more processors. A data set is supplied as input to the AI model, and an output is generated. By way of example, the output may correspond to the classification or recognition of certain features of the input data set. For training, the output is compared against known outputs for the input data, errors are backpropagated across the model, and the model's parameters are adjusted. In the case of large models and data sets, processing may be split across multiple processors to obtain results more quickly.
[0003]
[0003] One problem that arises with such systems is when one node of a multiprocessor system is hit by an error. Often, it is most likely that a large amount of data needs to be recomputed in order to resume the computation.
[0004]
[0004] In the figures of the accompanying drawings, various embodiments of the disclosure are illustrated by way of example and not limitation.
Brief Description of the Drawings
[0005]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Mode for Carrying Out the Invention
[0006]
[0019] In the following description, for purposes of explanation only, numerous examples and specific details are set forth in order to provide a thorough understanding of the present disclosure. Such examples and details should not be construed as unduly limiting the elements of the claims or the claimed subject matter as a whole. It will be apparent to those skilled in the art that, based on the language of the different claims, the claimed subject matter may include some or all of the features in these examples, alone or in combination, and may further include modifications and equivalents of the features and techniques described herein.
[0007]
[0020] Artificial intelligence (AI) processing systems are often required to process large amounts of data. Distributed processing increases the processing speed. For example, distributed training in deep learning that uses synchronous or hybrid data parallelism is an effective way to converge a model across many AI processors with high throughput and high accuracy.
[0008]
[0021] An example of a technique used in an AI network (e.g., for training) is what is called data parallelism. Data parallelism subdivides a training data set into fragments and loads a model for processing this data in parallel onto AI processors. For example, in one embodiment of data parallelism, the training data can be split into fragments (also known as shards), and each shard can be distributed for processing across multiple AI processors (also known as workers or target processors). On the other hand, the shards are split into minibatches, and the minibatches are iteratively processed by multiple AI processors in successive iterations. Between each iteration, the AI processors receive a minibatch (e.g., a minibatch of training data) and determine a change (also known as a "gradient" or "delta") in the model parameters. At the end of each iteration, the AI processors can combine and synchronize their model parameters and update the model with the new parameter values.
[0009]
[0022] The features and advantages of the present disclosure include a process of recovering from a failure. FIG. 1 shows a plurality of AI processors configured to process data using model M in parallel. In this example, the AI processors are configured into a plurality N of worker groups 101-103, where N is an integer. A worker group may include, for example, one or more AI processors. The AI processors can include a graphics processor (GPU), an AI accelerator, or other digital processors optimized for AI processing (e.g., a von Neumann architecture processor such as a matrix multiplication pair x86 processor). For example, examples of AI processors can include a GPU (e.g., an Nvidia Volta (registered trademark) having 800 cores and 64 multi-accumulators), or a tensor processor unit (TPU: Tensor Processor Unit) (e.g., four cores that execute 16k processing in parallel).
[0010]
[0023] In the iteration shown in this example, each worker group receives input data (e.g., a mini-batch) and processes this data using models 105 - 107. In this example, the iteration starts at 100 and each worker group starts with substantially the same model. For example, as part of a previous iteration, models 105 - 107 in each of the worker groups may synchronize their parameters (e.g., by performing an All-Reduce). In one embodiment, one or more copies of the model can also be saved as model 104, e.g., between each iteration cycle. At 111, each worker group processes different input data, e.g., a mini-batch from the same training data set. However, in this example, an error (e.g., a hardware or software failure) occurs in one of the worker groups 102. Advantageously, at 112, the model 104 that was used and saved at the start of the iteration can be loaded into worker group 102, and worker group 102 can quickly resume processing and generate results. At 112, the results of all of the worker groups 101 - 103 can be combined to generate an updated model, and the resulting model can be saved again, e.g., for the next iteration. In various embodiments described in more detail below, the worker group that was hit by the error can receive a new model 104 from, e.g., a controller (shown below), another worker group, or local memory of that worker group.
[0011]
[0024] Figure 2 shows an error recovery method according to an embodiment. At 201, a computing error is detected. For example, the computing error may be a software error or a hardware error in an AI processor of a plurality of artificial intelligence processors that process a data set. At 202, the AI processor can eliminate this error. For example, in an embodiment, some or all elements (e.g., hardware or software components) of the AI processor can be restarted. As shown in the following various embodiment examples, for example, the AI processor can be restarted by a controller coupled to the AI processor. At 203, a model is loaded into the AI processor. Here, the model corresponds to the same model that was processed by a plurality of AI processors during previous iterations of processing data from the data set.
[0012]
[0025] Features and advantages of the present disclosure include that a worker group can access the model used at the start of each iteration of processing in order to resume quickly. Previously, an AI system could not reach a global checkpoint where the state information of the system was saved until after many iterations. When an error occurred, one of the systems had to go back through many iterations to the global checkpoint, which took time. Advantageously, an AI processor affected by a failure can return to the start of the current iteration, while other processors can wait until the results of the current iteration are generated. Once the failed AI processor is reset and the error is cleared, the current iteration model can be reloaded and resumed. As described herein, the model can be stored in a number of different locations, as long as it is accessible to an AI processor affected by an error state. Examples of AI models include combinations of AI parameters such as weights or biases for a particular AI topology. Processing these models can include generating a gradient during each iteration. The gradient can include a deviation (delta) from the current parameter value (e.g., a delta value for a particular weight in a neural network). The gradient is generated as a result of processing by each AI processor and combined (e.g., aggregated by average, mean, etc.) and then can be applied to the value of the model at the start of the iteration. For example, calculate the average delta for all weights in a neural network model and apply this average delta to generate a subsequent model to be used in the next iteration.
[0013]
[0026] Figure 3 shows error recovery according to an exemplary embodiment in a computer processing system that performs training. In this example, training data 301 is used to train parameters of an AI model, such as weights of a neural network. The training data set 301 can be divided into fragments (here called slices or "shards") 302a - 302N. On the other hand, the shards are transferred to different worker groups for processing. Each of the shards can be further divided into smaller fragments 303a - 303N (here called "mini - batches" or sometimes simply "batches"). The mini - batches 303a - 303N of each shard are then successively combined, one at a time, with worker groups 304a - 304N. The worker groups can receive mini - batches of training data, perform AI training on model 320, and generate training results. Then, at 350, the training results from each worker group can be combined to generate an updated model 321. The updated model 321 can be loaded into each worker group, for example, to process the next mini - batch. As used herein, an "epoch" occurs when every worker group has processed all of their shards and one complete training data set 301 has been processed once. The training data set 301 is processed over multiple epochs and may, for example, reach a final set of trained model parameters. It should be understood that other ways of partitioning the data set 301 across multiple worker groups for processing can also use the error recovery techniques described herein.
[0014]
[0027] In this example, one iteration includes the step of receiving a mini-batch by worker groups 304a to 304N, the step of processing the mini-batch to generate results, and the step of combining these results to generate an updated model. The iteration further includes the step of loading the updated model into the worker groups (e.g., at the start or end of the iteration). FIG. 3 shows the i-th iteration, where N i-th mini-batches (minibatch_i) are loaded into N worker groups 304a to 304N for processing (e.g., N is an integer). The i-th model M_i 320 generated in the previous (i - 1)-th iteration is loaded into each of the worker groups 304a to 304N. The results obtained by processing each mini-batch_i are combined at 350 to generate the subsequent (or next) model model_i+1 321. Then, the model model_i+1 321 is loaded into each of the worker groups 304a to 304N to process the (i + 1)-th mini-batch in the next i + 1-th iteration. As described above, if one of the worker groups is affected by an error (e.g., a hard or soft failure), each of the worker groups only needs to be able to access the model for each iteration. Thus, for example, if a failure occurs while worker group 304b is processing mini-batch i, the i-th model (M_i) can be reloaded to complete its processing. Another system can also detect that worker group 304b is affected by an error and wait. When worker group 304b clears the error and generates results, the results from the worker group are combined to continue the computing.
[0015]
[0028] Figure 4 shows a computing architecture for processing AI data. In this example, multiple worker groups can be coupled to a controller, and the controller can further be coupled to a network. For example, worker groups 410a - 410N are coupled to controller 401, and worker groups 411a - 411N are coupled to controller 402. Through network 450, controllers 401 and 402 can be coupled together (e.g., via an Ethernet connection and one or more network switches - not shown). Also, worker groups can be coupled together via communication links such as links 451 and 452 (e.g., PCIe). For example, to train data, as described above, multiple such controller / worker groups can be used to process AI data in parallel. In various embodiments, combinations of the processing results (e.g., delta_parameters) described above can also be executed, for example, by a controller, between worker groups (e.g., via an all - reduce), or using combinations thereof.
[0016]
[0029] As described above, each worker group may include one or more workers, and each worker may be one or more of, for example, a GPU, TPU, or other AI processor optimized to perform multiplication and addition (multiply - accumulate, "MAC"), matrix multiplication ("MatMul"), and other operations. The controller may also be referred to as a host or gateway. The controller may be, for example, a conventional CPU, FPGA, system - on - chip (SoC), application - specific integrated circuit (ASIC), or embedded ARM controller, or other processor that can execute software and communicate with the worker groups based on instructions in this software. The system can include a driver that causes software to schedule and control tasks that need to be executed on a target device.
[0017]
[0030] The upper-level representation of a typical synchronous data parallelism flow is shown in Figure 5. In this example, every iteration ends with a synchronization of the model across the entire worker group (WG). Previously, global checkpoints were taken periodically to recover from any error or failure in the worker group. Checkpoints by previous systems, if frequent, could significantly reduce throughput, so in many cases global checkpoints were spread out (e.g., once an hour). However, as shown in Figure 6, one potential problem is that recovery from such errors is also slow. Due to the failure shown in Figure 6, all worker groups are interrupted and abruptly return to the global checkpoint. If errors or failures occur frequently enough (such as poisoning in a large cluster), this can have a serious impact on performance.
[0018]
[0031] The features and advantages of the present disclosure are that, without the need to abruptly return the entire group to the global checkpoint, access to the model from the previous iteration allows recovery from errors and certain failures occurring within a large cluster, and the recovery is much faster (e.g., within seconds rather than hours). As shown in Figure 7 and as described previously, errors occurring in one or more worker groups can be resolved during the current iteration, and based on the model used at the start of that iteration, local recalculation is performed. Thus, the worker group experiencing the error can recover quickly, and all worker groups can proceed to subsequent iterations without the need to reprocess data from multiple previous iterations, for example.
[0019]
[0032] Embodiment examples of the present disclosure can utilize the observation that as long as there is a fast redundant copy accessible for recovery, a certain state (e.g., a model) can be recalculated from the previous state. Thus, in one embodiment, the "master copy" of the current model (e.g., parameters such as the weights of a neural network used at the start of iteration by a worker group) may be stored at a location accessible by each worker group (e.g., on a controller). Note that the master copy only needs to be the minimum state information required for recalculation, and thus the copy of the model from the current iteration may not have any recalculable state information (e.g., activation, for example). Alternatively, when a worker group is hit by an error, the master copy may be directly resident on the worker group for local access by a specific worker group (e.g., in local memory protected by an error correction code (ECC)). In yet other embodiments, each worker group maintains an extra copy of the model during the current iteration. Since this extra copy is not updated during processing, it is also available to other worker groups that may be hit by an error state. Advantageously, when the model of the current iteration is maintained by each worker group, different parts of this model (different subsets of the entire model) can be sent simultaneously from a plurality of different worker groups to the worker group that has encountered a failure, which in some architectures can be much faster than sending the model from the controller to the failed worker group, for example.
[0020]
[0033] In one embodiment, redundant copies of the model can be spread across worker groups such that each worker group obtains two different sections of the two copies (e.g., if a worker group holds the same section of the two copies, a failure in that worker group would result in an irrecoverable loss). The master copy can be updated frequently at the end of each iteration. Also, in certain forms of data parallelism that allow local updates, it can be updated even more frequently. Finally, in some example embodiments, workers in a worker group can notify the controller of any irrecoverable errors (such as parity errors). Also, if a local timeout is set, this will be much smaller than the value obtained by subtracting the recovery time estimated from the global timeout, but, for example, large enough to recognize an error. Instead of a timeout, a worker can send a heartbeat to the controller so that the controller can determine when a worker has been hit by an error.
[0021]
[0034] In various embodiments, the recovery method most appropriately depends on the circumstances of the failure. In the case of a parity error (poisoning): The controller can reset the worker group and re - execute from the master copy of the model using the same mini - batch data again. In the case of a local timeout (or heartbeat misses), the controller can force the failing worker to reset (e.g., by sideband operation). If this is successful, the recovery proceeds as in the previous case of parity error or poisoning. If this is not successful after repeated attempts, the controller can re - compile the less efficient model on the same worker group or, for example, employ a dedicated spare worker group. If none of these options are successful or available, there is a risk that the controller itself will fail.
[0022]
[0035] Regarding controller failures, all controllers can have the same master copy of the model at the end of each iteration. That is, even if a global timeout occurs due to a controller failure, the need to roll back to the global checkpoint can be eliminated. For example, after the software adjusts the worker groups and data shards so that they can operate with the new cluster size, the controller can continue from the current iteration point.
[0023]
[0036] In various embodiments, there may be multiple ways to restore the redundant copy from the end of the previous iteration. In one embodiment, the controller provides a copy from its own copy in memory. In other embodiments, the worker group that experienced the failure can have the master copy in local memory (e.g., in directly implemented ECC-protected memory). In still other embodiments, the worker group that experienced the failure collects copies from one or more operational worker groups (e.g., in parallel).
[0024]
[0037] Figure 8 shows an example of an error recovery method, where the controller side interacts with the target device in the worker group. Here, the "device" 891 is a worker or a worker group (for example, a group of devices sharing one copy of the model). In this example, only one device is shown for simplicity of illustration, but the controller 890 can have a process for each device 891. The arrows 850 - 857 indicate the flow of data and / or control information between the controller and, for example, an artificial intelligence processor. The example method shown in Figure 8 shows multiple iterations. At 801, the controller 890 can initialize the device 891. Thus, the device 891 may perform, for example, a soft reset at 820. At 802, the model is initialized and each device 891 can receive a copy of the initial model. At 803 and 804, each controller 890 and the associated device 891 perform an initial synchronization. At 804, the iteration begins, indicating the first iteration (for example, iter = 1). At 805, the controller 890 causes a mini - batch of data to be sent to each device 891 (for example, DevId is the device identifier). Each device 891 receives the data at 823 and runs the data against the model at 824. At 825, the device 891 may or may not detect an error. Similarly, the controller 890 investigates the device for errors during processing. If no error is detected by the device at 825 or by the controller at 806, the controller and the device synchronize at 809 and 829, and the controller 890 can start, for example, the next iteration. However, if an error is detected by the device 891 at 825, at 826, the device having this error can wait for the controller 890 to reset it. In this example, the controller 890 can detect the device ID ("devID") of the device having the error at 806 and perform a soft reset of this device at 807.At 808, the controller can send a copy of the model used during the current iteration to the device having an error. At 827, the device performs a soft reset, and at 828, the device receives and loads the model. The "RecoverModel" box can correspond to, for example, one of the embodiments described above for the recovery technique. Then, the device reloads the data at 823 and communicates the data in comparison with the reloaded model at 824. Other devices not affected by the error enter a waiting state and can resume after the device affected by the error has completed the iteration process. In certain embodiments described herein, other devices may receive a portion of the model and a portion of the data, and the step of loading and reprocessing the data for the device affected by the error may be performed by multiple devices, for example, to reduce the recovery time.
[0025]
[0038] In one embodiment of the fault (controller - recovery) in FIG. 9, the controller can have a master copy. In this example, "n" controllers 901 - 905 may be coupled to "k" worker groups (e.g., "k" groups of one or more artificial intelligence processors) and memories 911 - 914 respectively. In each iteration, the master copy of the model can be stored, for example, in the memories 911 - 914 coupled to the controllers 901 - 905. These memories may be protected with an error - correcting code (ECC), for example, to ensure the integrity of the stored model. Recovery can be initiated by a controller that has detected either a local timeout (or lost heartbeat) or poisoning. In either case, assume that the faulting worker can come back to life. If the worker itself has died, all the controller can do is to notify an error that can only be repaired by going back to the global checkpoint and readjusting the cluster. In other scenarios where the faulting worker has died, the controller can re - balance the same minibatch across the remaining workers (if possible). However, the state shown in FIG. 9 is fully recoverable. Because the worker may only reports poisoning of detectable soft errors.
[0026]
[0039] As described above, in the second embodiment (self - recovery), an ECC - protected memory is attached to each worker. When a worker detects poisoning, it attempts to self - recover. The worker resumes and reloads the model / graph / data from the attached ECC memory to retry the same mini - batch. Further, to speed up the recovery, poisoning can be segmented by category. For example, the worker specifies where the poisoning occurred (by address range), and then, before resuming, uses a recovery code to repair only that segment. In the case of self - recovery, a soft - hung worker may still be recoverable if it incorporates a watchdog timer interrupt (self - heartbeat), and this is possible if there is one dedicated core provided for this purpose.
[0027]
[0040] In the third embodiment (neighbor - recovery), the worker group may or may not have a controller, and consists of k workers (e.g., T1 to Tk). Even in the case of a hard failure, this worker group can recover by reorganizing into a small group that continues to process the same mini - batch. To achieve this, the group can incorporate model redundancy. This is made possible, in particular, by model partitioning (model parallelism). In this case, one worker group divides one model across multiple workers (e.g., different workers process different parts of the model). In this partitioning, a part of each worker's memory holds (carries) a redundant copy of the model state of other workers (e.g., just the minimum model state necessary for recovery) in a mutually exclusive way. For example, whenever worker T1 updates its segment Seg(l), the redundant state in worker Tk is also updated. This can be performed as a hardware - assisted mirrored write, a software write, or during model updates after all - reduces. For example, it is as follows.
[0028]
Table 1
[0029]
[0041] Thus, in various embodiments, using the distribution of redundant copies, two or more copies can be distributed in mutually exclusive partitions (i.e., the same target does not hold the same segment of different copies) in such a way that any new (or resumed) target can collect an in-tact copy from no other member. By having two copies, one recovery from failure is ensured, and with three copies, two recoveries from failure are ensured, and so on. However, even for a large cluster, two copies may be used to recover from soft errors or resumption.
[0030]
[0042] Thus, in various embodiments, using the master copy of the current iteration model stored in the controller, stored locally on the worker group, or stored for multiple workers in the worker group, recovery can be made local, and the master copy need only be exclusively partitioned across multiple workers on the same worker group (e.g., it may be exclusively partitioned according to the original copy so that workers do not have overlapping sections of the model).
[0031]
[0043] That is, when multiple workers are within one worker group, the master copy need only be mutually and exclusively partitioned from the running copy across the same worker group. As an alternative example, two or more copies can be mutually and exclusively partitioned across all workers in such a way that any failure can be recovered by collecting one of the in-tact copies within the target that is the resumed target or replacement target. In other embodiments, this copy may be a redundant copy from a particular target.
[0032]
[0044] Figure 10 shows recovery according to one example embodiment when an error occurs during the result aggregation phase. In some embodiments, it may be advantageous to recover from errors that occur during the result aggregation phase of each iteration. For example, as shown in FIG. 10, an iteration may include synchronization of the model across all worker groups at 1001. At 1002, a mini-batch is received and applied to the model to generate a result. At 1003, which is the start of the result aggregation phase, post data synchronization can be performed. In some cases, an error in one of the artificial intelligence processors may occur after the data has been applied to the model. Typically, each worker group can generate a unique vector of delta values (e.g., gradients) that indicate changes in the model parameters, for example, after each mini-batch of data has been processed.
[0033]
[0045] FIG. 11 shows an example of result generation according to an embodiment. Here, N worker groups WG0 to WGN (N is an integer) generate N vectors of length M (for example, M is an integer equal to the number of weights of the neural network in the model). The delta value Δij in each vector may be, for example, a floating-point number. When the system is operating (for example, when operating without errors), the vectors generated by each worker group are passed to all other worker groups, and each worker group aggregates one subset of the fields from each vector. It is natural that there are N partitions among the N worker groups, and each worker group aggregates the results for a specific field of the vectors received from other worker groups. For example, worker group WG0 can receive vectors from other worker groups and aggregate the Δ1j to Δij fields to generate, for example, result array R0. The aggregation can include, for example, the average of the weights or other functions known to those skilled in AI processing. However, if one of the worker groups is hit by an error during the processing of the results, that worker group can send an invalid result indicator to the other worker groups. In this example, WG0 sends a result vector of length M that includes a garbage bit (shown here as "xxxx"). During the processing of the results, when another worker group receives an invalid result indicator from another worker group, it can trigger that worker group to enter a waiting state. Therefore, these worker groups can wait while the worker group hit by the error eliminates the error and processes valid results.
[0034]
[0046] Figure 12 shows an example of result aggregation according to an embodiment. In one embodiment, the worker groups can also be configured in a ring shape, and the worker groups can pass the gradient vectors (described above) and then the results (e.g., the aggregated gradients) to other worker groups. In this example, each worker group can receive the result array of the aggregated gradients. When all worker groups have all the results from all other worker groups, each worker group will have the full set of aggregated results, and using this aggregated result, the worker group can modify those versions of the model. In this example, since all worker groups start with the same model, as a result of each update of the model, the model remains substantially the same (e.g., the AI parameters change together, so each worker group has substantially the same model over all iterations). Also in this case, if a worker group is hit by an error, this worker group can output an invalid result flag, and the other worker groups can wait until the worker group hit by the error recovers and sends a valid result.
[0035]
[0047] Figure 13 shows error recovery according to an embodiment in a multiprocessor computing environment. In this example, worker group 1301 is hit by an error and outputs x 1311, which is an invalid result flag. Other worker groups (e.g., 1300, 1302) can generate valid gradient vectors Δ (e.g., 1310, 1312). In this example, each of the other worker groups can wait until worker group 1301 eliminates its error and generates a valid result. Then, the system passes the valid gradient vectors, calculates the aggregated results, and transfers this result to other worker groups during the result aggregation phase, so that each worker group has, for example, an updated model.
[0036]
[0048] FIG. 14 illustrates the process of distributing the calculations of a failed processor across multiple processors, according to an embodiment. In one embodiment, when a worker group detects and eliminates an error, different parts of the model can be loaded across the entire worker group, including the worker group affected by the error. Thus, the time to recalculate results during a particular iteration can be reduced for the worker group affected by the error. As shown in FIG. 14, worker groups 1410-1413 may be processing mini-batches, for example, using the same model. Here, worker group 1412 is affected by an error. However, in this example, the model for the current iteration is partitioned across the entire plurality of worker groups including worker group 1412. Referring to FIG. 14, when worker group 1412 has finished eliminating the error, worker group 1412 can prompt the loading of model 1450 across all of worker groups 1410-1413. Thus, the portion of the training data intended for processing by worker group 1412 in the current iteration is processed across the plurality of worker groups 1410-1413, and for example, the recovery time can be reduced.
[0037] Still other exemplary embodiments
[0049] In various embodiments, the present disclosure includes an error recovery method. This method can be embodied in a non-transitory computer-readable storage medium. A program code executable by a computer system is stored on the non-transitory computer-readable storage medium, and this program code causes the computer system to execute the techniques described herein. In one embodiment, the computer system may include a plurality of artificial intelligence processors and one or more controllers. The non-transitory computer-readable storage medium may be, for example, a memory and may be coupled to, for example, one or more controllers or one or more artificial intelligence processors.
[0038]
[0050] The following techniques can be implemented alone or in different combinations, and can also be implemented together with other techniques described in this specification.
[0039]
[0051] For example, in one embodiment, the present disclosure includes a method. This method includes detecting a computing error in a first artificial intelligence processor among a plurality of artificial intelligence processors during a first processing iteration of data from a data set, eliminating the error from the first artificial intelligence processor, and loading a model in one or more of the artificial intelligence processors including the first artificial intelligence processor, where this model corresponds to the same model processed by the plurality of artificial intelligence processors during the first processing iteration of data from the data set.
[0040]
[0052] In one embodiment, while the first artificial intelligence processor eliminates the error, a plurality of artificial intelligence processors other than the first artificial intelligence processor wait, and the plurality of processors use the same second identical model generated from the same model used in the first processing iteration to simultaneously process data from the data set in the next processing iteration.
[0041]
[0053] In one embodiment, the computing error is detected during the result aggregation phase of the first processing iteration, and at least some of the plurality of artificial intelligence processors complete the result aggregation phase after waiting for the first artificial intelligence processor to generate a valid result during the aggregation phase.
[0042]
[0054] In one embodiment, the first artificial intelligence processor sends an invalid result flag to at least some of the plurality of artificial intelligence processors to prompt waiting.
[0043]
[0055] In one embodiment, the result aggregation phase is an All-Reduce.
[0044]
[0056] In one embodiment, the step of loading the model includes loading different parts of the model in one or more of the artificial intelligence processors including the first artificial intelligence processor, and the method further includes processing a first portion of the data received by the first artificial intelligence processor in a first processing iteration in one or more of the artificial intelligence processors including the first artificial intelligence processor.
[0045]
[0057] In one embodiment, the step of loading the model includes loading the model into the first artificial intelligence processor, and the method further includes processing a first portion of the data received by the first artificial intelligence processor in a first processing iteration in the first artificial intelligence processor.
[0046]
[0058] In one embodiment, the model is received by the first artificial intelligence processor from a controller.
[0047]
[0059] In one embodiment, the model is received by the first artificial intelligence processor from one or more other processors of the plurality of artificial intelligence processors.
[0048]
[0060] In one embodiment, the model is received by the first artificial intelligence processor from the local memory of the first artificial intelligence processor.
[0049]
[0061] In one embodiment, the model includes artificial intelligence parameters.
[0050]
[0062] In one embodiment, the model includes the weights of a neural network.
[0051]
[0063] In one embodiment, the data set is a training data set.
[0052]
[0064] The above description has shown various embodiments of the present disclosure, together with examples of how aspects of a particular embodiment can be implemented. The above examples have been presented to illustrate the flexibility and advantages of the particular embodiments defined by the following claims and should not be regarded as the only embodiments. Based on the above disclosure and the following claims, other arrangements, embodiments, implementations, and equivalents can also be adopted without departing from the scope of the present disclosure defined by the claims.
Claims
1. A system comprising: a plurality of artificial intelligence processors; one or more controllers; a memory storing program code executable by the one or more controllers and the plurality of artificial intelligence processors; wherein the program code causes the system to: during a first time period of a first processing iteration of data from a data set, detect a computing error in a first artificial intelligence processor of the plurality of artificial intelligence processors; eliminate the computing error from the first artificial intelligence processor; during a second time period of the first processing iteration, cause the first artificial intelligence processor to load a model, the model corresponding to the same model that a second artificial intelligence processor of the plurality of artificial intelligence processors successfully processed at least partially during the first time period.
2. The system of claim 1, wherein the plurality of artificial intelligence processors other than the first artificial intelligence processor wait while the first artificial intelligence processor eliminates the computing error, and the plurality of artificial intelligence processors other than the first artificial intelligence processor use a second identical model generated from the same model used in the first processing iteration to simultaneously process data from the data set in a next processing iteration.
3. The system of claim 1, wherein the computing error is detected during a result aggregation phase of the first processing iteration, and at least some of the plurality of artificial intelligence processors wait to generate valid results during the result aggregation phase of the first processing iteration and then complete the result aggregation phase of the first processing iteration.
4. The system of claim 3, wherein the first artificial intelligence processor sends an invalid result flag to at least some of the plurality of artificial intelligence processors to prompt the waiting.
5. The system of claim 3, wherein the result aggregation phase is an All-Reduce.
6. The system of claim 1, wherein the loading of the model includes loading different parts of the model in a plurality of artificial intelligence processors including the first artificial intelligence processor. A system in which the program code causes the system to process a first portion of data received by the first artificial intelligence processor in the first processing iteration in the plurality of artificial intelligence processors including the first artificial intelligence processor.
7. The system according to claim 1, wherein loading the model includes loading the model into the first artificial intelligence processor, and the system further processes a first portion of data received by the first artificial intelligence processor in the first processing iteration in the first artificial intelligence processor.
8. The system according to claim 1, wherein the model is received by the first artificial intelligence processor from a controller in the first artificial intelligence processor.
9. The system according to claim 1, wherein the model is received by the first artificial intelligence processor from one or more other processors among the plurality of artificial intelligence processors in the first artificial intelligence processor.
10. The system according to claim 1, wherein the model is received by the first artificial intelligence processor from the local memory of the first artificial intelligence processor in the first artificial intelligence processor.
11. The system according to claim 1, wherein the model includes artificial intelligence parameters.
12. The system according to claim 1, wherein the model includes weights of a neural network.
13. The system according to claim 1, wherein the data set is a training data set.
14. A computer-readable storage medium storing program code executable by a computer system, the program code causing the computer system to detect a computing error in a first artificial intelligence processor among a plurality of artificial intelligence processors during a first time period of a first processing iteration of data from a data set, eliminate the computing error from the first artificial intelligence processor, load a model in the first artificial intelligence processor during a second time period of the first processing iteration, the model corresponding to the same model that has succeeded in being at least partially processed by a second artificial intelligence processor among the plurality of artificial intelligence processors during the first time period.
15. An error recovery method, comprising: detecting, during a first time period of a first processing iteration of data from a data set, a computing error in a first artificial intelligence processor among a plurality of artificial intelligence processors; eliminating the computing error from the first artificial intelligence processor; loading, during a second time period of the first processing iteration, a model in the first artificial intelligence processor; wherein the model corresponds to the same model that has been successfully processed at least partially by a second artificial intelligence processor among the plurality of artificial intelligence processors during the first time period.
Citation Information
Patent Citations
Method and system for distributed deep machine learning
US20170220949A1
Communication optimizations for distributed machine learning
US20190205745A1