Shared Failure Domain Database for Network Remediation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In three-tier Clos networks, combined hardware or link failures lead to complex capacity degradation and traffic congestion due to the lack of available neighboring device information for remediation, making it difficult to calculate and address failures effectively.

Innovation Solution

The implementation of a Shared Failure Domain (SFD) database that groups devices and links with a shared failure model, allowing for coordinated remedial actions across neighboring devices, such as deactivating related ports or devices, based on a central repository of failure domain data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional network failure management is used without shared failure domain information, then device complexity is reduced, but network reliability deteriorates due to inability to effectively remediate combined failures

Engineering Contradiction:
Improvenetwork reliabilityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

A centralized controller acts as an intermediary between network devices and failure management functions. The controller maintains a shared failure domain database that maps relationships between devices and links, enabling coordinated remediation actions across multiple devices when failures occur, thereby improving network reliability without increasing individual device complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary actions by pre-establishing failure domain relationships and populating the shared failure domain database before failures occur. When a failure is detected, the controller can immediately query the database and execute pre-planned remediation actions across related devices, improving response time and reliability without adding complex real-time decision-making to individual devices

Inventive Principle:
Principle #10Preliminary action

2Productivity

If ECMP routing is used with single-hop decision limited to single router, then device complexity is reduced, but productivity deteriorates due to inefficient packet forwarding when routers fail

Engineering Contradiction:
Improvepacket forwarding efficiencyVSAvoidrouting decision complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system extends ECMP routing from single-hop local decisions to multi-hop coordinated decisions by adding a new dimension of controller-mediated communication. When a router failure is detected, the controller queries the shared failure domain database and coordinates alternative path selection across multiple routers, improving packet forwarding efficiency without requiring each individual router to perform complex global routing calculations

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10999127B1Configuring network devices using a shared failure domain
Publication Date: 2021.05.04 AMAZON TECH INC
  • US10999127B1 patent drawing
  • US10999127B1 patent drawing
  • US10999127B1 patent drawing

AI summary

Remediation of network devices that are failing is accomplished using a Shared Failure Domain (SFD) database that provides neighboring device/link information to remediation tools. SFD refers to a group of objects (links/devices) that share a same failure model. A state change of one or multiple of the objects results in a corresponding action on other devices linked together through the SFD. Moreover, the SFD data is available in a central repository and software tools consult the central repository for failure domain data before taking remedial actions. SFD data is generated using configuration generation and device state. Software tools lookup SFD data during operational events (device/link down) and take appropriate actions on the neighboring devices.