Failover Support for Shared Resources in Multi-Computer Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current server systems with shared resources lack effective failover support mechanisms, leading to increased costs and limitations, as they require additional system management controllers or restrict access, which is impractical for large-scale rack-mount server systems.
Innovation Solution
A failover support mechanism is implemented using priority detection signals and GPIO pins to determine the highest priority system management controller, allowing it to monitor and control shared resources, with a simple handshake protocol among controllers to ensure seamless resource management and failover.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If an additional system management controller is provided to monitor shared resources, then the reliability of resource management is improved, but the device complexity and cost increase
Solution Approach 1:
Multiple system management controllers are merged into a single functional unit that shares common resources. The controllers collectively manage shared resources through priority-based selection rather than each having dedicated control, reducing overall system complexity while maintaining reliable resource monitoring and management.
Solution Approach 2:
The system management controllers are designed with multi-functionality to handle both resource monitoring and failover selection. Each controller can assume the role of managing shared resources when selected as highest priority, eliminating the need for dedicated separate controllers and reducing device complexity.
2Device complexity
If access to shared resources is restricted to one specific system management controller, then the device complexity is reduced, but the adaptability and failover support ability are limited
Solution Approach 1:
The system implements dynamic access control where the eligibility to manage shared resources is not fixed but determined through real-time priority detection. Controllers can dynamically assume management roles based on their priority status and operational state, enabling flexible adaptability while maintaining simple access mechanisms.
Solution Approach 2:
The system employs feedback mechanisms where controllers continuously monitor their own priority status and the operational state of other controllers. This feedback loop enables automatic failover and adaptation when primary controllers fail or priority changes occur, providing versatility without increasing access complexity.
3Adaptability or versatility
If multiple system management controllers are provided without additional cost, then the adaptability is improved, but the device complexity increases
Solution Approach 1:
The system management controllers perform self-service by autonomously detecting their own priority status and determining eligibility to manage shared resources. The controllers independently monitor system conditions and execute failover decisions without requiring external coordination, reducing system architecture complexity while maintaining high adaptability.
Data Source
AI summary
Managing shared resources in a multi-computer system with failover support, including: reading priority detection signals from a computer inserted into the multiple-computer system, the priority detection signals representing a priority of the inserted computer; reading planar detection signals from the computer, the planar detection signals representing an insertion state of all computers currently inserted into the multiple-computer system; determining if the computer has the highest priority among all the computers inserted into the multiple-computer system in accordance with the priority detection signals and the planar detection signals; and, in response to determining that the computer has the highest priority, monitoring shared resources and outputting a specific output signal associated with the highest priority computer, the specific output signal providing an identification of the highest priority computer to other computers currently inserted into the multiple-computer system and representing control, by the highest priority computer, of the shared resources.


