Spare Partition Unit Replacement in Partitionable Computing Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multiple-core microprocessor servers face challenges in reliability and failover mechanisms, as they can become single points of failure when consolidating multiple servers, potentially leading to increased downtime and maintenance costs.
Innovation Solution
A method and apparatus for managing spare partition units in a partitionable computing device using a global management entity and local operating systems, allowing for the addition and replacement of spare partition units without requiring recompilation of computer-executable instructions, enabling controlled and orderly failover processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If multiple servers are consolidated into a single multiple-core microprocessor server, then cost and energy consumption are reduced, but reliability deteriorates due to single point of failure
Solution Approach 1:
The system is segmented into multiple partition units (PUs), each capable of running an independent operating system instance. These PUs are distributed across multiple physical processors or cores, allowing the system to maintain multiple logical servers within a single physical server. This segmentation enables failover capabilities while consolidating hardware resources.
Solution Approach 2:
The system dynamically changes the operational state of partition units by transitioning them between active and standby states. When a failure is detected, the system changes parameters such as processor assignment, memory allocation, and OS instance activation to activate backup partition units, thereby maintaining reliability while keeping hardware consolidated.
2Reliability
If partition units are replaced to improve reliability, then system reliability is enhanced, but system downtime increases during the replacement process
Solution Approach 1:
Backup partition units are pre-configured and maintained in a standby state before failures occur. The system prepares replacement partition units in advance by allocating resources and loading operating system images, so that when a failure occurs, the replacement can occur rapidly without significant downtime.
Solution Approach 2:
The system uses an intermediary mechanism (such as a hypervisor or management entity) that coordinates the replacement process. This intermediary manages the transition between active and standby partition units, handling resource allocation and OS instance activation to minimize disruption and reduce downtime during replacements.
3Reliability
If spare partition units are added to enable failover, then reliability is improved, but device complexity increases
Solution Approach 1:
The system designs partition units to be universal and multi-functional, where each PU can serve multiple purposes: running active OS instances, serving as backup units, and being dynamically reassigned to different processors or cores. This universality reduces the need for dedicated backup hardware, thereby limiting complexity growth while maintaining failover capabilities.
Solution Approach 2:
The system implements self-service mechanisms where the management entity automatically monitors system health, detects failures, and triggers failover processes without requiring manual intervention. This automation reduces operational complexity by handling failover management internally, allowing spare partition units to be effectively utilized without proportionally increasing system complexity.
Data Source
AI summary
A method and apparatus for managing spare partition units in a partitionable computing device is disclosed. The method comprises detecting if a spare partition unit is required for addition or replacement in a local operating system and if a spare partition unit is required for addition, initiating an addition of a spare partition unit. If a spare partition unit is required for replacement, a replacement of a failing partition unit with a spare partition unit is initiated; part of the memory of the failing partition unit is passively migrated into the memory of the spare partition unit's partition; part of the memory of the failing partition unit is also actively migrated into the memory of the spare partition unit's partition; and the partitionable computing device is cleaned up. Partition units are replaced without requiring that computer-executable instructions be recompiled.


