FCoE Network Failure Detection via Out-of-Band Management Controller

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Fibre Channel over Ethernet (FCoE) networks experience significant delays in detecting node or interconnect failures, leading to traffic black-holing, which is unacceptable in critical deployments.

Innovation Solution

Incorporating a management controller within each network device to detect failures and communicate through an out-of-band management network, allowing for rapid identification and mitigation of failures by stopping input/output processes across the network.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If traditional FCoE failure detection methods are used, then the system maintains simplicity, but failure detection time is delayed (20-225 seconds) causing traffic black-holing

Engineering Contradiction:
Improvefailure detection timeVSAvoidsystem complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent introduces a management controller as an intermediary component that operates independently from the main FCoE protocol stack. This management controller monitors network device status and communicates through an out-of-band management network, enabling rapid failure detection without interfering with normal FCoE operations. The intermediary architecture allows parallel failure detection while maintaining system simplicity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent adds a separate out-of-band management network dimension for failure detection, independent from the primary FCoE data plane. This dimensional separation allows failure detection traffic to flow through a dedicated channel with different performance characteristics, achieving sub-second detection without congesting or complicating the main FCoE network path.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If rapid failure detection is implemented, then traffic black-holing is prevented, but system complexity increases due to additional management infrastructure

Engineering Contradiction:
Improvenetwork reliabilityVSAvoidmanagement infrastructure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The management controller operates autonomously to detect and report failures without requiring complex external management systems. It self-monitors the status of network devices, self-communicates failure information through the out-of-band network, and enables other devices to self-adjust by stopping I/O processes, thereby improving reliability while minimizing the complexity of external management infrastructure.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary failure detection and notification actions before traffic black-holing occurs. The management controller continuously monitors device status and proactively communicates failures to the management system and peer devices, enabling preventive action (stopping I/O processes) before data loss occurs, thus improving reliability without requiring complex reactive measures.

Inventive Principle:
Principle #10Preliminary action

3Loss of substance

If sub-second failure detection is achieved, then data loss is prevented, but the system requires additional out-of-band management network infrastructure

Engineering Contradiction:
Improvedata lossVSAvoidnetwork infrastructure complexity
Core Design Contradiction:
Loss of substanceVSDevice complexity

Solution Approach 1:

The patent extracts the failure detection function from the main FCoE protocol stack and places it in a separate management controller that communicates through an out-of-band network. This extraction isolates the failure detection mechanism from data plane operations, preventing data loss through rapid detection while confining the additional infrastructure requirements to a separate management network rather than complicating the primary FCoE network.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS9762432B2Systems and methods for rapid failure detection in fibre channel over ethernet networks
Publication Date: 2017.09.12 DELL PROD LP
  • US9762432B2 patent drawing
  • US9762432B2 patent drawing
  • US9762432B2 patent drawing

AI summary

An information handling system is provided herein. The information handling system includes a central processor in communication with a network processor, a plurality of ports coupled to the network processor for sending and receiving Fiber Channel over Ethernet (FCoE) frames, and an Ethernet controller in communication with a physical connector and with the central processor. The information handling system further includes a management controller configured to communicate with a management system through the Ethernet controller to report a failure to be mitigated by temporarily stopping inputs and outputs on a coupled network device. Associated methods and computer-readable media having associated instructions are also provided herein.