Fault Recovery in Distributed Training via Dynamic Participant Lists
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for fault recovery in distributed deep neural network training, such as manual worker process recovery and disk checkpoint restoration, are inefficient due to prolonged downtime and high re-computation costs, especially with large models and datasets, as they require constant monitoring and are affected by slower disk read/write speeds.
Innovation Solution
A fault recovery system with a master node that detects faults and adjusts the collective communication participant list, using a dynamic rendezvous and in-memory checkpoint method to quickly recover worker processes and prevent training losses by storing model parameters in memory, allowing for faster state restoration and reduced downtime.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual worker process recovery and disk checkpoint restoration are used, then fault recovery can be achieved, but downtime is prolonged and re-computation costs increase
Solution Approach 1:
The system performs preliminary actions by maintaining in-memory checkpoints of worker process states during normal operation. When a fault occurs, these pre-prepared memory checkpoints enable immediate state restoration without disk I/O, significantly reducing downtime compared to traditional disk checkpoint methods
Solution Approach 2:
The patent replaces the mechanical disk I/O system with a memory-based checkpoint system. By substituting slow disk read/write operations with fast memory access, the system achieves quicker fault recovery while maintaining the same fault tolerance functionality
2Reliability
If disk checkpoint method is used for state restoration, then fault recovery is possible, but re-computation costs are high due to slower disk read/write speeds
Solution Approach 1:
The patent substitutes disk-based checkpoint storage with memory-based storage. This replacement eliminates the bottleneck of disk read/write speeds, enabling rapid loading of model parameters and reducing re-computation costs when restoring worker processes after faults
Solution Approach 2:
The system pre-loads and maintains checkpoint data in memory during normal operation. When faults occur, the pre-positioned memory checkpoints can be immediately utilized without expensive disk I/O operations, thereby reducing re-computation overhead and improving productivity
3Productivity
If collective communication participant list is adjusted dynamically, then training continuity is maintained after fault, but system complexity increases
Solution Approach 1:
The master node implements feedback mechanisms by monitoring worker node health status and dynamically adjusting the collective communication participant list. This feedback loop enables automatic adaptation to faults, maintaining training continuity without requiring complex manual intervention or system redesign
Solution Approach 2:
The system transitions from a static participant list to a dynamic one that automatically adapts to changing system conditions. The master node dynamically adds or removes worker nodes from the participant list based on their operational status, enabling continuous training while managing complexity through automated rules
Data Source
AI summary
A system with fault recovery includes: a plurality of worker nodes configured to perform distributed training; and a master node configured to control the plurality of worker nodes, wherein the master node is configured to: detect a fault of the plurality of worker nodes based on a predetermined period; adjust a collective communication participant list in response to the detecting of the fault; and transmit the adjusted participant list to one or more worker nodes in the adjusted participant list.


