I/O Fabric Error Routing to Affected Host Nodes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current PCI Express bus systems do not allow sharing of adapters among multiple host computer systems, leading to increased costs and space constraints in blade systems, and existing error reporting methods notify all host nodes regardless of whether the error affects all or only some, causing unnecessary system downtime.

Innovation Solution

A method for routing error messages in a multi-root environment to only those host computer systems affected by the error, using routing tables to identify affected nodes and suspending traffic until all affected nodes acknowledge the error before clearing it.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If error reporting methods notify all host nodes regardless of whether the error affects all or only some, then all host nodes are informed of the error, but unnecessary system downtime occurs for unaffected nodes

Engineering Contradiction:
Improveerror notification accuracyVSAvoidsystem downtime
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the error notification process by routing error messages to only those host nodes affected by the error, rather than notifying all host nodes. This is achieved through the I/O fabric switch which uses routing tables to forward error messages selectively to specific host nodes based on the error's scope and type.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by tailoring the error notification to the specific local context of each host node. The system determines which host nodes are affected by the error and notifies only those, allowing unaffected nodes to continue operations without interruption.

Inventive Principle:
Principle #3Local quality

2Reliability

If adapters are dedicated to individual blade systems, then each system has reliable access to its adapter, but costs and space constraints increase

Engineering Contradiction:
Improveadapter access reliabilityVSAvoidnumber of adapters required
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent implements universality by enabling multiple blade systems to share common I/O adapters through the I/O fabric. The I/O fabric switch acts as a mediator that allows different blade systems to access the same adapter resources, making the adapters multi-functional and reducing the total number of adapters required in the system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of information

If all host nodes are notified of errors, then complete error awareness is achieved, but resource suspension increases

Engineering Contradiction:
Improveerror information deliveryVSAvoidresource utilization
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent segments the error notification process by routing error messages to only those host nodes affected by the error, rather than notifying all host nodes. This is achieved through the I/O fabric switch which uses routing tables to forward error messages selectively to specific host nodes based on the error's scope and type.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by providing error notification only to the extent necessary - i.e., only to host nodes that are actually affected by the error. This avoids the excessive action of notifying all host nodes when only a subset is impacted, thereby preserving resource utilization for unaffected nodes.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS7707465B2Routing of shared I/O fabric error messages in a multi-host environment to a master control root node
Publication Date: 2010.04.27 LENOVO GLOBAL TECHNOLOGIES SWITZERLAND INTERNATIONAL GMBH
  • US7707465B2 patent drawing
  • US7707465B2 patent drawing
  • US7707465B2 patent drawing

AI summary

A computer-implemented method, apparatus, and computer program product are disclosed for routing error messages in a multiple host computer system environment to only those host computer systems that are affected by the error. The environment includes multiple host computer systems that share multiple devices utilizing a switched fabric. An error is detected in one of the devices. Routing tables that are stored in fabric devices in the fabric are used to identify ones of the host computer systems that are affected by the error. An error message that identifies the error is routed to only the identified ones of the host computer systems.