Distributed Blade Server Automatic Failover Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The reliability of blade servers in high-density deployments is lower compared to tower and rack servers, and in distributed systems, the reliability of nodes is lower than that of minicomputers, leading to system unreliability and slow spare blade replacement.
Innovation Solution
A distributed blade server system with a management server that determines a standby blade when a blade is in abnormal operation, delivers configuration and startup commands to enable the standby blade to access storage partitions, allowing for quick service recovery and reduced service interruption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If blade servers are deployed in high-density environments to save space and reduce power consumption, then space utilization and energy efficiency are improved, but system reliability deteriorates
Solution Approach 1:
The patent establishes standby blades in advance that are pre-configured with backup operating systems and application programs. When a working blade fails, the management server automatically activates the corresponding standby blade to take over its services, eliminating the need for manual intervention and reducing service interruption time despite high-density deployment constraints
Solution Approach 2:
The patent dynamically changes the operational state of blades from working to standby and vice versa based on real-time monitoring of blade health status. The management server adjusts the configuration commands to storage systems and startup commands to blades, transforming static high-density architecture into a dynamic fault-tolerant system that maintains reliability while preserving space efficiency
2Reliability
If spare blades are kept ready for replacement to improve reliability, then system availability is improved, but replacement speed deteriorates due to manual intervention requirements
Solution Approach 1:
The patent implements automatic failover where standby blades self-activate when their corresponding working blades fail. The management server automatically sends configuration commands to storage systems and startup commands to standby blades, enabling them to load backup operating systems and application programs without human intervention, thus achieving rapid service recovery
Solution Approach 2:
The management server continuously monitors blade operational status and receives feedback on working blade failures. Based on this real-time feedback, the server automatically triggers the activation sequence for standby blades, creating a closed-loop system that responds immediately to failures and minimizes service interruption time
3Manufacturing precision
If manual configuration and startup procedures are used for blade switching to ensure proper system state, then configuration accuracy is improved, but switching speed deteriorates
Solution Approach 1:
The patent replaces manual mechanical configuration procedures with automated electronic command transmission. The management server electronically sends configuration commands to storage systems and startup commands to blades, substituting human operators with automated software control that maintains configuration accuracy while dramatically increasing switching speed
Solution Approach 2:
The management server acts as an intermediary between the failed working blade and the standby blade. It receives failure notifications, determines the appropriate standby blade, and coordinates the activation process by sending commands to both storage systems and blades, ensuring proper configuration state while enabling rapid automated switching
Data Source
AI summary
A distributed blade server system, a management server and a switching method are provided. The method includes: determining a standby blade of a first blade when it is determined that the first blade is in abnormal operation; delivering, based on an access relationship between a startup card of the first blade and a first storage partition, a first configuration command to a storage system, the first configuration command including information of an access relationship between a startup card of the standby blade and the first storage partition, so that the storage system configures the access relationship between the startup card of the standby blade and the first storage partition; and delivering a startup command to the standby blade.


