Hypervisor Workload Migration for Hardware Repair
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Enterprise-level software has not kept pace with processor improvements in live error recovery and hardware component maintenance, requiring spare systems and time for workload migration during repairs.
Innovation Solution
A hypervisor migrates workload from a failing hardware resource to another within the computing system, allowing for the repair of the first resource while maintaining system operation, and may also attempt to drive operational parameters back within acceptable ranges or fail over to another system if necessary.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If workload is migrated between entire computing systems, then hardware component reliability is improved, but system complexity and time required for repair increase
Solution Approach 1:
The patent segments the computing system into virtual machines and hardware resources, allowing independent migration of virtual machines between hardware resources. This enables repair of individual hardware components without migrating entire computing systems, thereby improving reliability while reducing system complexity.
Solution Approach 2:
The patent introduces a intermediary layer (virtual machine migration mechanism) that allows workload to be transferred between hardware resources without requiring complete system migration. This intermediary approach enables targeted hardware repair while maintaining overall system operation.
2Reliability
If workload is migrated between entire computing systems, then hardware component reliability is improved, but time required for repair increases
Solution Approach 1:
The patent enables segmentation of workload migration at the virtual machine level rather than system level, allowing rapid migration of individual virtual machines to available hardware resources while leaving other virtual machines operational. This significantly reduces repair time compared to migrating entire computing systems.
Solution Approach 2:
The patent implements preliminary actions by maintaining standby hardware resources and pre-configuring virtual machine migration capabilities, enabling rapid failover when hardware failures occur. This reduces repair time by having readiness measures in place before failures occur.
3Reliability
If spare computing systems are used for workload migration, then hardware component reliability is improved, but device complexity and resource requirements increase
Solution Approach 1:
The patent makes hardware resources universal by enabling any hardware resource to host any virtual machine through standardized migration mechanisms. This eliminates the need for dedicated spare computing systems, as any available hardware resource can serve as a failover target, thereby improving reliability without increasing device complexity.
4Ease of repair
If hardware component repair requires system downtime, then repair simplicity is improved, but productivity decreases
Solution Approach 1:
The patent ensures continuity of useful action by migrating virtual machines to available hardware resources before halting the failed hardware component for repair. This allows repair activities to proceed without system downtime, maintaining productivity while simplifying the repair process through controlled virtual machine migration and hardware replacement.
Data Source
AI summary
Hardware component repair in a computing system while workload continues to execute on the computing system includes receiving an indication that an operational parameter of a first hardware resource of said computing system does not meet operational acceptability criteria; migrating workload of the computing system from said first hardware resource to a second hardware resource within the computing system; and halting operation of said first hardware resource for repair.


