Node Failure Recovery via Storage Link Redirection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Distributed applications in node networks lack effective internal failure-management solutions, leading to loss of local backup data and calculation steps due to physical failures, with existing backup levels being inefficient in terms of cost and complexity.
Innovation Solution
A method that redirects the link between a storage medium and its node to another node upon failure, allowing for backup without the need for proactive copying across all nodes, maintaining efficiency similar to intermediate-level backups at the cost and complexity of local backups, using a PCIe switch for redirection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If intermediate-level backup (L2) is performed by duplication on a partner node, then backup robustness is improved, but device complexity and cost increase
Solution Approach 1:
The patent introduces a storage medium as an intermediary component that decouples the backup process from direct node-to-node duplication. The storage medium acts as a mediator that can be redirected to any node in the network, providing robust backup capabilities without requiring complex node-to-node connection management or partner node coordination
Solution Approach 2:
The patent separates the backup function from the compute nodes by introducing dedicated storage media. This segmentation allows the backup system to be independently managed and redirected without affecting the computational nodes, reducing the complexity of coordinated backup operations across multiple nodes
2Loss of time
If local backup (L1) is performed frequently, then computation time loss during failure is minimized, but backup robustness decreases
Solution Approach 1:
The storage medium is designed to serve multiple functions: it can be rapidly accessed by the original node for frequent local backups, and simultaneously be redirectable to any other node in the network for robust backup recovery. This multi-functionality allows the system to achieve both fast recovery time and high backup robustness through the same infrastructure
3Reliability
If global backup (L4) is performed, then backup robustness is maximized, but computation time loss during failure increases significantly
Solution Approach 1:
The patent implements a dynamic backup system where the storage medium can be flexibly redirected to different nodes based on failure conditions. This dynamic approach allows the system to achieve global backup robustness on-demand without maintaining continuous global backup operations, thereby avoiding the significant computation time losses associated with traditional global backup methods
Data Source
AI summary
Disclosed is a failure management method in a network of nodes, including, for each considered node: first, a step of locally saving the state of this considered node, to a storage medium for this node in question. Then, if the considered node has failed, retrieving the local backup of the state of this considered node, by redirecting the link between the considered node and its storage medium to connect this storage medium to an operational node other than the considered node, this operational node already in the process of carrying out this calculation, the local backups of these considered nodes, used for the retrieving steps being coherent with each other so as to correspond to the same state of calculation. If a considered node failed, returning this local backup for this considered node to a new additional node added to the network at the time of the failure.


