Self-Virtualizing IO Error Recovery via Adjunct Partition Coordination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In logically partitioned computers, error handling for self-virtualizing IO resources, such as SRIOV Ethernet adapters, is complex due to the need to coordinate error recovery across multiple physical and virtual functions and operating systems, leading to inefficiencies and potential errors in Extended Error Handling (EEH) recovery.

Innovation Solution

Simplified error handling is achieved by utilizing a physical function adjunct partition to coordinate error recovery for self-virtualizing IO resources, where each virtual function adjunct partition is restarted, eliminating the need for coordination within logical partitions and ensuring that virtual function adjunct partitions execute through their normal initialization paths to prevent stale data or untested recovery paths.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If error recovery coordination is implemented across multiple logical partitions and operating systems for self-virtualizing IO resources, then error handling completeness is improved, but system complexity and processing overhead increase significantly

Engineering Contradiction:
Improveerror handling completenessVSAvoiderror recovery coordination complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary error recovery coordination mechanism that acts as a mediator between multiple logical partitions and the self-virtualizing IO resource. This intermediary layer centralizes error recovery coordination, allowing comprehensive error handling across all LPARs without requiring complex peer-to-peer coordination between partitions. The intermediary receives error notifications from the IO resource and distributes recovery actions to affected LPARs, simplifying the overall error handling architecture while maintaining reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The error recovery process is segmented into distinct phases: error detection by the self-virtualizing IO resource, error notification to the intermediary, coordination of recovery actions, and individual LPAR recovery execution. This segmentation allows each component to focus on specific tasks, reducing the complexity that would arise from requiring all LPARs to simultaneously coordinate with each other for error recovery.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If traditional device driver error handling is used in logically partitioned systems, then compatibility with conventional systems is maintained, but error recovery efficiency and speed deteriorate

Engineering Contradiction:
Improvecompatibility with conventional systemsVSAvoiderror recovery efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The self-virtualizing IO resource performs self-service error detection and notification functions, autonomously identifying errors and initiating the error recovery process without requiring traditional device driver intervention in each LPAR. The IO resource notifies the intermediary of errors and participates in its own recovery process, eliminating the inefficiencies of traditional driver-based error handling while maintaining system compatibility through the intermediary layer.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Error recovery actions are prepared and coordinated in advance by the intermediary upon error detection, rather than waiting for traditional device driver error handling sequences to unfold. The intermediary pre-coordinates recovery actions with affected LPARs, allowing faster recovery execution compared to traditional sequential error handling through multiple device drivers.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If error recovery involves multiple operating systems and logical partitions coordinating simultaneously, then comprehensive error coverage is achieved, but processing overhead and time consumption increase

Engineering Contradiction:
Improveerror coverageVSAvoiderror recovery time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The intermediary serves as a centralized coordination point that manages error recovery across multiple LPARs and operating systems simultaneously. Instead of requiring parallel coordination between all LPARs (which would be time-consuming), the intermediary sequentially manages recovery for each affected LPAR, achieving comprehensive error coverage while minimizing total recovery time through efficient resource orchestration.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The error recovery process is made dynamic and adaptive, with the intermediary adjusting the recovery sequence and resource allocation based on real-time system state and LPAR priorities. This dynamic approach allows the system to optimize recovery time by recovering critical LPARs first while maintaining comprehensive error coverage across all partitions.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8645755B2Enhanced error handling for self-virtualizing input/output device in logically-partitioned data processing system
Publication Date: 2014.02.04 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8645755B2 patent drawing
  • US8645755B2 patent drawing
  • US8645755B2 patent drawing

AI summary

Error handling is simplified for a self-virtualizing IO resource that utilizes a physical function adjunct partition for a physical function in the self-virtualizing IO resource to coordinate error recovery for the self-virtualizing IO resource, by restarting each virtual function adjunct partition associated with that physical function to avoid the need to coordinate error recovery within the logical partitions to which such virtual function adjunct partitions are assigned.