Node Failure Recovery via Storage Link Redirection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Distributed applications in node networks lack effective internal failure-management solutions, leading to loss of local backup data and calculation steps due to physical failures, with existing backup levels being inefficient in terms of cost and complexity.

Innovation Solution

A method that redirects the link between a storage medium and its node to another node upon failure, allowing for backup without the need for proactive copying across all nodes, maintaining efficiency similar to intermediate-level backups at the cost and complexity of local backups, using a PCIe switch for redirection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If intermediate-level backup (L2) is performed by duplication on a partner node, then backup robustness is improved, but device complexity and cost increase

Engineering Contradiction:
Improvebackup robustnessVSAvoidbackup complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces a storage medium as an intermediary component that decouples the backup process from direct node-to-node duplication. The storage medium acts as a mediator that can be redirected to any node in the network, providing robust backup capabilities without requiring complex node-to-node connection management or partner node coordination

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent separates the backup function from the compute nodes by introducing dedicated storage media. This segmentation allows the backup system to be independently managed and redirected without affecting the computational nodes, reducing the complexity of coordinated backup operations across multiple nodes

Inventive Principle:
Principle #1Segmentation

2Loss of time

If local backup (L1) is performed frequently, then computation time loss during failure is minimized, but backup robustness decreases

Engineering Contradiction:
Improvecomputation time lossVSAvoidbackup robustness
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The storage medium is designed to serve multiple functions: it can be rapidly accessed by the original node for frequent local backups, and simultaneously be redirectable to any other node in the network for robust backup recovery. This multi-functionality allows the system to achieve both fast recovery time and high backup robustness through the same infrastructure

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If global backup (L4) is performed, then backup robustness is maximized, but computation time loss during failure increases significantly

Engineering Contradiction:
Improvebackup robustnessVSAvoidcomputation time loss
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements a dynamic backup system where the storage medium can be flexibly redirected to different nodes based on failure conditions. This dynamic approach allows the system to achieve global backup robustness on-demand without maintaining continuous global backup operations, thereby avoiding the significant computation time losses associated with traditional global backup methods

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11477073B2Procedure for managing a failure in a network of nodes based on a local strategy
Publication Date: 2022.10.18 BULL SA
  • US11477073B2 patent drawing
  • US11477073B2 patent drawing
  • US11477073B2 patent drawing

AI summary

Disclosed is a failure management method in a network of nodes, including, for each considered node: first, a step of locally saving the state of this considered node, to a storage medium for this node in question. Then, if the considered node has failed, retrieving the local backup of the state of this considered node, by redirecting the link between the considered node and its storage medium to connect this storage medium to an operational node other than the considered node, this operational node already in the process of carrying out this calculation, the local backups of these considered nodes, used for the retrieving steps being coherent with each other so as to correspond to the same state of calculation. If a considered node failed, returning this local backup for this considered node to a new additional node added to the network at the time of the failure.