Virtual Machine Memory Management via BMC-RAC Failover
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current information handling systems, particularly in data centers, face challenges with memory management due to the inefficiencies in firmware features storage and customization, leading to unnecessary memory usage and limited scalability, as well as complexities in managing devices across different hardware capabilities.
Innovation Solution
The method involves assigning memory modules between compute and storage nodes, detecting connectivity losses, and establishing secondary communication channels to transfer data, allowing for dynamic memory allocation and management, including the use of non-volatile and volatile memory modules, to ensure continuous operation of virtual machines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If memory modules are assigned to virtual machines over a communication channel between compute and storage nodes, then memory management flexibility and scalability are improved, but system complexity and connectivity failure risks increase
Solution Approach 1:
The system separates memory resources into distinct modules located at storage nodes, with each virtual machine assigned specific memory modules. This segmentation allows independent management and allocation of memory resources without affecting the entire system, resolving the contradiction between flexibility and complexity.
Solution Approach 2:
A management controller is introduced as an intermediary component that handles memory allocation, monitoring, and failover operations. This intermediary abstracts the complexity of distributed memory management, providing simplified control while maintaining the flexibility of modular memory assignment.
2Adaptability or versatility
If memory modules are located at storage nodes rather than compute nodes, then memory resource sharing and scalability are improved, but access speed and connectivity reliability deteriorate
Solution Approach 1:
The system performs preliminary actions by assigning memory modules to virtual machines before connectivity issues occur. When connectivity is lost, pre-configured failover mechanisms immediately activate, transferring memory access to alternative locations without interruption to the virtual machine operations, thus maintaining speed despite distributed architecture.
Solution Approach 2:
Redundant memory modules are pre-positioned at compute nodes as cushioning resources. When storage node connectivity fails, these pre-positioned resources immediately compensate for the loss, maintaining access speed while preserving the scalability benefits of distributed memory architecture.
3Reliability
If connectivity loss detection and failover mechanisms are implemented, then system reliability is improved, but device complexity and management overhead increase
Solution Approach 1:
The management controller implements continuous feedback monitoring of connectivity between compute and storage nodes. When connectivity loss is detected, the controller automatically triggers failover procedures, providing reliable operation without requiring complex manual intervention or management overhead.
Solution Approach 2:
The system performs self-service failover operations where the management controller automatically detects connectivity issues and reallocates memory resources without external intervention. This automation improves reliability while reducing management complexity by eliminating the need for manual fault response procedures.
Data Source
AI summary
Virtual machine memory management, including assigning, for each virtual machine operating at a compute node, respective first memory modules to the virtual machine operating at a storage node; accessing, by the virtual machines and over a first communication channel between the compute node and the storage node, the respective first memory modules; detecting, for a particular virtual machine, a loss of connectivity of the particular virtual machine with respective first memory modules assigned to the particular virtual machine, and in response: assigning, for the particular virtual machine, second memory modules to the particular virtual machine operating at the compute node; establishing a second communication channel between a BMC of the storage node and a RAC of the computing node; transferring, over the second communication channel, data stored at the respective first memory modules assigned to the particular virtual machine to the second memory modules assigned to the particular virtual machine.


