VM State Validation After Physical Server Crash
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In private cloud computing infrastructures, physical server failures leading to hypervisor crashes result in unpredictable VM server migrations, providing minimal and general notification to stakeholders, lacking detail on the impact on services.
Innovation Solution
A system that includes a physical server crash register, a physical to virtual device mapping transformer, and feedback/notification modules to identify impacted VM servers, validate their states, and communicate real-time service interruption notifications to stakeholders.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If dynamic VM server placement is implemented for load balancing, then resource utilization is improved, but predictability of placement and timing is worsened
Solution Approach 1:
The system performs preliminary actions by pre-identifying impacted VM servers and pre-preparing notification content before the actual failure occurs. The crash response assistance system proactively monitors physical server health and pre-maps the relationship between physical servers and their hosted VM servers, enabling rapid response when failures do occur.
Solution Approach 2:
The system implements feedback mechanisms by continuously monitoring the state of physical servers and VM servers, tracking migration status, and providing real-time updates to stakeholders. The feedback loop includes monitoring whether VM servers have recovered to powered-on state after migration and notifying stakeholders of the current status.
2Device complexity
If general notification is provided for server failures, then system complexity is reduced, but information completeness for stakeholders is worsened
Solution Approach 1:
The notification system is segmented into multiple functional components: a physical server crash register for capturing failure details, a physical to virtual device mapping transformer for identifying impacted VM servers, and feedback/notification modules for delivering tailored information to different stakeholder groups. This segmentation allows complex information processing while maintaining manageable system architecture.
Solution Approach 2:
The physical to virtual device mapping transformer acts as an intermediary that translates physical server failure information into impacted VM server lists. This intermediary component processes and enriches the failure information with relevant VM and application details before notification, bridging the gap between simple failure detection and comprehensive stakeholder information needs.
3Loss of time
If immediate notification is provided to stakeholders, then response time is improved, but system complexity for tracking and validation is worsened
Solution Approach 1:
The system performs preliminary mapping between physical servers and VM servers, and pre-identifies stakeholders before failures occur. This preliminary setup enables immediate notification when failures happen, as the relationships and contact information are already established and stored in the crash response assistance system.
Solution Approach 2:
The system performs self-service validation by automatically checking whether impacted VM servers have recovered to powered-on state without requiring manual intervention. The feedback module autonomously tracks migration status and validates recovery conditions, reducing the need for complex external tracking mechanisms.
Data Source
AI summary
Crash response assistance in the event of a physical server experiencing a hardware and/or software failure (i.e., crashing) while hosting Virtual Machine (VM) servers. Notification of a physical server hardware and/or software failure triggers acquisition of details about the failure. Based on the crash details, impacted VM servers are identified and validations are performed to determine whether impacted VM servers have recovered to a powered-on state or remain in a powered-off state. Once stakeholders/users and/or support groups associated with the physical server, impacted VM servers and/or applications executing on the impacted VM servers are identified and, service interruption notifications are communicated to the identified stakeholders/users and support groups, which identify at least one impacted VM server and at least one application executing on the identified VM server(s) and whether the VM server has returned to a powered-on state or remains in a powered-off state.


