Dynamic Resource Allocation for I/O Recovery Events

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Computing systems face degraded performance due to I/O recovery events, such as storage system failures, which require substantial time and resources for recovery, leading to increased workload and potential service outages, and existing solutions do not effectively mitigate these issues without incurring additional billing costs for users.

Innovation Solution

A resource management module dynamically allocates additional computing resources, including CPU cores, memory, and I/O devices, to affected partitions during and after I/O recovery events to shorten recovery times and complete workload backlogs, while ensuring that these additional resources are not billed to users as internal operational costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If additional computing resources are allocated to partitions during I/O recovery events, then recovery time and performance are improved, but billing costs for users increase

Engineering Contradiction:
ImproveI/O recovery performanceVSAvoidbilling costs
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system converts the harmful effect of I/O recovery events (which cause performance degradation and service outages) into a beneficial opportunity by automatically allocating additional computing resources to affected partitions. This resource allocation accelerates the recovery process and completes workload backlogs, transforming the negative impact of I/O failures into improved service reliability without requiring user payment for the additional resources used during recovery operations

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

2Loss of time

If additional computing resources are allocated dynamically during I/O recovery events, then recovery time is reduced, but system complexity increases

Engineering Contradiction:
Improverecovery timeVSAvoidresource allocation complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The system implements a feedback mechanism where the resource management module continuously monitors I/O recovery events and automatically triggers additional resource allocation when recovery events are detected. This closed-loop approach ensures that resources are dynamically adjusted based on actual system conditions, reducing recovery time while managing complexity through automated decision-making rather than manual intervention

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The resource management module operates autonomously to detect I/O recovery events and allocate additional computing resources without requiring user input or manual configuration. The system self-manages the complex task of resource allocation, balancing act between service level agreement compliance and cost management, thereby reducing recovery time while containing operational complexity within the automated system

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11327767B2Increasing resources for partition to compensate for input/output (I/O) recovery event
Publication Date: 2022.05.10 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11327767B2 patent drawing
  • US11327767B2 patent drawing
  • US11327767B2 patent drawing

AI summary

Embodiments of dynamically increasing the resources for a partition to compensate for an input/output (I/O) recovery event are provided. An aspect includes allocating a first set of resources to a partition that is hosted on a data processing system. Another aspect includes operating the partition on the data processing system using the first set of resources. Another aspect includes, based on detection of an input/output (I/O) recovery event associated with operation of the partition, determining a compensation for the I/O recovery event. Another aspect includes allocating a second set of resources in addition to the first set of resources to the partition, the second set of resources corresponding to the compensation for the I/O recovery event. Another aspect includes operating the partition on the data processing system using the first set of resources and the second set of resources.