Peripheral Bus Arbitration Entity for Device Failure Isolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current proprietary support systems for managing peripheral bus device failures in information handling systems are cumbersome, costly, and often limited to specific operating systems or types of devices, failing to effectively address failures and requiring resource-intensive management, especially when dealing with complex devices like storage controllers and GPGPUs.

Innovation Solution

A method and system that utilize an arbitration entity and management module to detect, isolate, and manage failures of peripheral bus devices by suspending communication, identifying redundant devices, and assigning a secondary device to maintain system availability, thereby isolating the failure from the root complex and allowing for high availability across multiple systems without requiring operating system awareness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If proprietary support systems are used to manage bus device failures, then device availability can be maintained through failover mechanisms, but the system becomes cumbersome, costly, and limited to specific operating systems or device types

Engineering Contradiction:
Improvebus device availabilityVSAvoidmanagement system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an arbitration entity as an intermediary component that sits between the bus devices and the operating system. This arbitration entity manages failover operations and failure detection, acting as a mediator that eliminates the need for complex proprietary support systems while maintaining device availability. The arbitration entity handles the complexity internally, presenting a simplified interface to both the operating system and bus devices.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The arbitration entity is designed with multi-functional capabilities that allow it to manage multiple types of bus devices (storage controllers, GPGPUs, network cards) across different operating systems through a unified mechanism. This universal approach replaces the need for device-specific or OS-specific proprietary support systems, reducing overall system complexity while maintaining broad applicability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If proprietary support systems modify individual bus device drivers, then specific device failures can be managed, but bus devices become burdened with increasing vendor and operating system-specific features

Engineering Contradiction:
Improvedevice failure managementVSAvoiddevice driver compatibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The arbitration entity serves as an intermediary layer between the operating system and bus device drivers, handling failure management and failover operations without requiring modifications to the device drivers themselves. This preserves driver compatibility across different vendors and operating systems while providing robust failure management capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent extracts the failure management functionality from the bus device drivers and relocates it to the arbitration entity. This separation allows the drivers to remain simple and compatible while concentrating the complex failure management logic in the arbitration entity, thereby maintaining driver versatility without sacrificing reliability.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If management of bus device availability is performed at the operating system level, then device failures can be detected and managed, but system resources are consumed

Engineering Contradiction:
Improvedevice failure detectionVSAvoidoperating system resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The arbitration entity acts as an intermediary that assumes the responsibility of failure detection and management, transferring this workload from the operating system to the arbitration entity. This reduces operating system resource consumption while maintaining reliable failure detection and management capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The arbitration entity provides self-service functionality by autonomously detecting failures, initiating failover operations, and managing bus device availability without requiring continuous operating system intervention. This autonomous operation reduces the computational resources consumed by the operating system while maintaining high reliability.

Inventive Principle:
Principle #25Self-service

4Reliability

If proprietary support systems are used for specific device types like storage controllers, then those device failures can be managed, but it becomes difficult or impractical to extend support to other device types

Engineering Contradiction:
Improvespecific device failure managementVSAvoidsupport system applicability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The arbitration entity is designed as a universal platform that can manage multiple types of bus devices including storage controllers, GPGPUs, and network cards through a single unified mechanism. This eliminates the need for separate proprietary support systems for each device type while maintaining comprehensive failure management capabilities across all device types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10114688B2System and method for peripheral bus device failure management
Publication Date: 2018.10.30 DELL PROD LP
  • US10114688B2 patent drawing
  • US10114688B2 patent drawing
  • US10114688B2 patent drawing

AI summary

A system and method for managing peripheral device failures is disclosed. The method includes detecting, at a processor of a peripheral bus, a failure of a first bus device at a downstream port from the processor. The downstream port is populated by the first bus device and the processor is communicatively coupled at an upstream port to a root complex. The processor is configured to isolate the failure of the first bus device from the root complex. The method also includes, responsive to detecting the failure, suspending communication of data to the first bus device, receiving information regarding a second bus device selected from a cluster of a plurality of bus devices, and assigning the second bus device to the downstream port.