Fault Recovery in Distributed Training via Dynamic Participant Lists

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for fault recovery in distributed deep neural network training, such as manual worker process recovery and disk checkpoint restoration, are inefficient due to prolonged downtime and high re-computation costs, especially with large models and datasets, as they require constant monitoring and are affected by slower disk read/write speeds.

Innovation Solution

A fault recovery system with a master node that detects faults and adjusts the collective communication participant list, using a dynamic rendezvous and in-memory checkpoint method to quickly recover worker processes and prevent training losses by storing model parameters in memory, allowing for faster state restoration and reduced downtime.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual worker process recovery and disk checkpoint restoration are used, then fault recovery can be achieved, but downtime is prolonged and re-computation costs increase

Engineering Contradiction:
Improvefault recovery capabilityVSAvoiddowntime
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by maintaining in-memory checkpoints of worker process states during normal operation. When a fault occurs, these pre-prepared memory checkpoints enable immediate state restoration without disk I/O, significantly reducing downtime compared to traditional disk checkpoint methods

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the mechanical disk I/O system with a memory-based checkpoint system. By substituting slow disk read/write operations with fast memory access, the system achieves quicker fault recovery while maintaining the same fault tolerance functionality

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If disk checkpoint method is used for state restoration, then fault recovery is possible, but re-computation costs are high due to slower disk read/write speeds

Engineering Contradiction:
Improvestate restoration capabilityVSAvoidre-computation cost
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent substitutes disk-based checkpoint storage with memory-based storage. This replacement eliminates the bottleneck of disk read/write speeds, enabling rapid loading of model parameters and reducing re-computation costs when restoring worker processes after faults

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system pre-loads and maintains checkpoint data in memory during normal operation. When faults occur, the pre-positioned memory checkpoints can be immediately utilized without expensive disk I/O operations, thereby reducing re-computation overhead and improving productivity

Inventive Principle:
Principle #10Preliminary action

3Productivity

If collective communication participant list is adjusted dynamically, then training continuity is maintained after fault, but system complexity increases

Engineering Contradiction:
Improvetraining continuityVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The master node implements feedback mechanisms by monitoring worker node health status and dynamically adjusting the collective communication participant list. This feedback loop enables automatic adaptation to faults, maintaining training continuity without requiring complex manual intervention or system redesign

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system transitions from a static participant list to a dynamic one that automatically adapts to changing system conditions. The master node dynamically adds or removes worker nodes from the participant list based on their operational status, enabling continuous training while managing complexity through automated rules

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12197278B2System, apparatus and method with fault recovery
Publication Date: 2025.01.14 SAMSUNG ELECTRONICS CO LTD
  • US12197278B2 patent drawing
  • US12197278B2 patent drawing
  • US12197278B2 patent drawing

AI summary

A system with fault recovery includes: a plurality of worker nodes configured to perform distributed training; and a master node configured to control the plurality of worker nodes, wherein the master node is configured to: detect a fault of the plurality of worker nodes based on a predetermined period; adjust a collective communication participant list in response to the detecting of the fault; and transmit the adjusted participant list to one or more worker nodes in the adjusted participant list.