I/O Fabric Error Routing to Affected Host Nodes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current PCI Express bus systems do not allow sharing of adapters among multiple host computer systems, leading to increased costs and space constraints in blade systems, and existing error reporting methods notify all host nodes regardless of whether the error affects all or only some, causing unnecessary system downtime.
Innovation Solution
A method for routing error messages in a multi-root environment to only those host computer systems affected by the error, using routing tables to identify affected nodes and suspending traffic until all affected nodes acknowledge the error before clearing it.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If error reporting methods notify all host nodes regardless of whether the error affects all or only some, then all host nodes are informed of the error, but unnecessary system downtime occurs for unaffected nodes
Solution Approach 1:
The patent segments the error notification process by routing error messages to only those host nodes affected by the error, rather than notifying all host nodes. This is achieved through the I/O fabric switch which uses routing tables to forward error messages selectively to specific host nodes based on the error's scope and type.
Solution Approach 2:
The patent applies local quality by tailoring the error notification to the specific local context of each host node. The system determines which host nodes are affected by the error and notifies only those, allowing unaffected nodes to continue operations without interruption.
2Reliability
If adapters are dedicated to individual blade systems, then each system has reliable access to its adapter, but costs and space constraints increase
Solution Approach 1:
The patent implements universality by enabling multiple blade systems to share common I/O adapters through the I/O fabric. The I/O fabric switch acts as a mediator that allows different blade systems to access the same adapter resources, making the adapters multi-functional and reducing the total number of adapters required in the system.
3Loss of information
If all host nodes are notified of errors, then complete error awareness is achieved, but resource suspension increases
Solution Approach 1:
The patent segments the error notification process by routing error messages to only those host nodes affected by the error, rather than notifying all host nodes. This is achieved through the I/O fabric switch which uses routing tables to forward error messages selectively to specific host nodes based on the error's scope and type.
Solution Approach 2:
The patent applies partial action by providing error notification only to the extent necessary - i.e., only to host nodes that are actually affected by the error. This avoids the excessive action of notifying all host nodes when only a subset is impacted, thereby preserving resource utilization for unaffected nodes.
Data Source
AI summary
A computer-implemented method, apparatus, and computer program product are disclosed for routing error messages in a multiple host computer system environment to only those host computer systems that are affected by the error. The environment includes multiple host computer systems that share multiple devices utilizing a switched fabric. An error is detected in one of the devices. Routing tables that are stored in fabric devices in the fabric are used to identify ones of the host computer systems that are affected by the error. An error message that identifies the error is routed to only the identified ones of the host computer systems.


