Automated LPAR Recovery via HMC Error Code Evaluation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In computing storage environments, manual processes for creating or rebuilding logical partitions (LPARs) are costly and resource-intensive, requiring on-site travel and human interaction with code, leading to potential errors and prolonged maintenance windows due to the need for removable media and optical drives.

Innovation Solution

An automated recovery mechanism using a Hardware Management Console (HMC) that evaluates failure scenarios, performs cleanup operations, and retries network connections to create or repair LPARs without human intervention, eliminating the need for removable media and optical drives, and enabling remote service operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual processes are used for creating or rebuilding LPARs, then human operators can handle complex scenarios with code evaluation, but labor costs and travel expenses increase significantly

Engineering Contradiction:
Improveerror handling capabilityVSAvoidlabor cost efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs automated error recovery and LPAR creation without requiring human intervention. The hardware management console automatically evaluates failure scenarios, identifies error codes, executes appropriate recovery actions, and recreates LPARs, enabling the system to serve itself rather than requiring manual operator intervention for each failure scenario

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes (physical travel to remote sites, hands-on code evaluation and execution) with an automated electronic system. The hardware management console uses software-based error code evaluation and automated script execution to substitute for human operators, eliminating the need for physical presence at remote locations

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual code evaluation is performed to identify errors, then accurate error diagnosis is achieved, but maintenance time and resource expenditure increase

Engineering Contradiction:
Improveerror diagnosis accuracyVSAvoidmaintenance window duration
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system pre-evaluates multiple failure scenarios and has predetermined recovery actions ready for each scenario. When an error occurs, the hardware management console matches the error code against pre-defined failure scenarios and executes the corresponding pre-prepared recovery script, eliminating the need for real-time code evaluation and reducing maintenance time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements automated feedback loops where error codes are continuously monitored, evaluated against known failure scenarios, and trigger automatic recovery actions. This closed-loop feedback mechanism enables rapid error diagnosis and response without requiring manual intervention, significantly reducing maintenance window duration while maintaining high diagnostic accuracy

Inventive Principle:
Principle #23Feedback

3Reliability

If removable media and optical drives are used for LPAR creation, then system compatibility and reliability are maintained, but device complexity and operational difficulty increase

Engineering Contradiction:
Improvesystem compatibilityVSAvoidhardware requirement complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and removes the requirement for removable media and optical drives from the LPAR creation process. By using automated network-based image transfer and hardware management console functionality, the system eliminates the need for physical media, reducing device complexity while maintaining reliability through alternative automated delivery mechanisms

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces mechanical media insertion and optical drive operations with electronic network-based image transfer. The hardware management console automatically transfers LPAR images over the network and provisions them to the target system, substituting mechanical processes with electronic automation to reduce hardware complexity and operational difficulty

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Productivity

If automated recovery mechanisms are implemented, then service reliability and speed are improved, but system complexity and initial resource requirements increase

Engineering Contradiction:
Improverecovery speedVSAvoidautomation system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The hardware management console is designed to perform multiple functions: it manages LPAR creation, monitors system health, evaluates error codes, executes recovery scripts, and provisions images. This multi-functional approach consolidates automation capabilities into an existing management infrastructure, reducing the need for separate dedicated automation components and minimizing additional system complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8458510B2LPAR creation and repair for automated error recovery
Publication Date: 2013.06.04 MAPLEBEAR INC
  • US8458510B2 patent drawing
  • US8458510B2 patent drawing
  • US8458510B2 patent drawing

AI summary

Various embodiments for automated error recovery in a computing storage environment by a processor device are provided. In one embodiment, pursuant to performing one of creating a new and rebuilding an existing logical partition (LPAR) operable in the computing storage environment by a hardware management console (HMC) in communication with the LPAR, at least one failure scenario is evaluated by identifying error code. If a failure is caused by an operation of the HMC and a malfunction of a current network connection, a cleanup operation is performed on at least a portion of a current HMC configuration, an alternative network connection to the current network connection is made, and a retry operation is performed.