Node Fault Recovery via Storage Link Redirection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Distributed applications in networks lack effective internal fault management, leading to loss of local backup data and calculation progress due to physical faults, with existing backup solutions being either inefficient or excessively complex and costly.

Innovation Solution

A method that redirects the link between a storage medium and its node to another node in the event of a fault, allowing for backup and recovery with similar efficiency to intermediate backups at the cost and complexity of local backups, thereby optimizing data backup and resilience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If intermediate backup (level L2) is performed by duplication on a partner node, then fault recovery capability is improved, but backup cost and complexity increase

Engineering Contradiction:
Improvefault recovery capabilityVSAvoidbackup complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces a storage medium as an intermediary component that decouples the backup process from direct node-to-node duplication. The storage medium receives data from the defective node and makes it accessible to operational nodes without requiring complex inter-node communication protocols or partner node coordination, thereby simplifying the backup mechanism while maintaining recovery capability

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent extracts the backup and recovery function from the node pair relationship and relocates it to a dedicated storage medium. This separation allows the backup system to operate independently of node availability and simplifies the overall system architecture by removing the need for nodes to directly manage each other's backup states

Inventive Principle:
Principle #2Taking out (Extraction)

2Speed

If local backup is performed frequently, then recovery speed is improved, but data loss in case of node failure increases

Engineering Contradiction:
Improverecovery speedVSAvoiddata loss
Core Design Contradiction:
SpeedVSLoss of information

Solution Approach 1:

The patent applies different backup strategies to different scenarios: for frequent failures, local backups provide rapid recovery, while for catastrophic node failures, the storage medium provides persistent data retention. This localized optimization of backup quality for different failure modes allows the system to achieve both fast recovery and data protection without requiring frequent expensive global backups

Inventive Principle:
Principle #3Local quality

3Reliability

If global backup at file system level is performed, then backup robustness is improved, but backup time and cost increase significantly

Engineering Contradiction:
Improvebackup robustnessVSAvoidbackup time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the backup system into multiple levels: local node backups for rapid recovery and a centralized storage medium for persistent data retention. This segmentation allows the system to achieve robustness through the storage medium without requiring frequent full-system global backups, thereby reducing backup time and computational overhead while maintaining data safety

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11249868B2Method of fault management in a network of nodes and associated part of network of nodes
Publication Date: 2022.02.15 THE FRENCH ALTERNATIVE ENERGIES & ATOMIC ENERGY COMMISSION
  • US11249868B2 patent drawing

AI summary

The invention relates to a method of fault management in a network of nodes (2), comprising, for each node considered (2) of all or part of the nodes (2) of the network performing one and the same calculation: firstly, a step of local backup of the state of this node considered (21), at the level of a storage medium (31) for this node considered (21), the link (6) between this storage medium (31) and this node considered (21) being able to be redirected from this storage medium (31) to another node (23), thereafter, a step of relaunching: either of the node considered (21) if the latter is not defective, on the basis of the local backup of the state of this node considered (21), or of an operational node (23) different from the node considered (21), if the node considered (21) is defective, on the basis of the recovery of the local backup of the state of this node considered (21), by redirecting said link (6) between the node considered (21) and its storage medium (31) so as to connect said storage medium (31) to said operational node (23), the local backups of these nodes considered (2), used for the relaunching steps, are mutually consistent so as to correspond to one and the same state of this calculation.