Distributed Blade Server Automatic Failover Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The reliability of blade servers in high-density deployments is lower compared to tower and rack servers, and in distributed systems, the reliability of nodes is lower than that of minicomputers, leading to system unreliability and slow spare blade replacement.

Innovation Solution

A distributed blade server system with a management server that determines a standby blade when a blade is in abnormal operation, delivers configuration and startup commands to enable the standby blade to access storage partitions, allowing for quick service recovery and reduced service interruption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If blade servers are deployed in high-density environments to save space and reduce power consumption, then space utilization and energy efficiency are improved, but system reliability deteriorates

Engineering Contradiction:
Improvespace utilizationVSAvoidsystem reliability
Core Design Contradiction:
Area of stationary objectVSReliability

Solution Approach 1:

The patent establishes standby blades in advance that are pre-configured with backup operating systems and application programs. When a working blade fails, the management server automatically activates the corresponding standby blade to take over its services, eliminating the need for manual intervention and reducing service interruption time despite high-density deployment constraints

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent dynamically changes the operational state of blades from working to standby and vice versa based on real-time monitoring of blade health status. The management server adjusts the configuration commands to storage systems and startup commands to blades, transforming static high-density architecture into a dynamic fault-tolerant system that maintains reliability while preserving space efficiency

Inventive Principle:
Principle #35Parameter changes

2Reliability

If spare blades are kept ready for replacement to improve reliability, then system availability is improved, but replacement speed deteriorates due to manual intervention requirements

Engineering Contradiction:
Improvesystem availabilityVSAvoidreplacement time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements automatic failover where standby blades self-activate when their corresponding working blades fail. The management server automatically sends configuration commands to storage systems and startup commands to standby blades, enabling them to load backup operating systems and application programs without human intervention, thus achieving rapid service recovery

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The management server continuously monitors blade operational status and receives feedback on working blade failures. Based on this real-time feedback, the server automatically triggers the activation sequence for standby blades, creating a closed-loop system that responds immediately to failures and minimizes service interruption time

Inventive Principle:
Principle #23Feedback

3Manufacturing precision

If manual configuration and startup procedures are used for blade switching to ensure proper system state, then configuration accuracy is improved, but switching speed deteriorates

Engineering Contradiction:
Improveconfiguration accuracyVSAvoidswitching speed
Core Design Contradiction:
Manufacturing precisionVSSpeed

Solution Approach 1:

The patent replaces manual mechanical configuration procedures with automated electronic command transmission. The management server electronically sends configuration commands to storage systems and startup commands to blades, substituting human operators with automated software control that maintains configuration accuracy while dramatically increasing switching speed

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The management server acts as an intermediary between the failed working blade and the standby blade. It receives failure notifications, determines the appropriate standby blade, and coordinates the activation process by sending commands to both storage systems and blades, ensuring proper configuration state while enabling rapid automated switching

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9189349B2Distributed blade server system, management server and switching method
Publication Date: 2015.11.17 HUAWEI TECH CO LTD
  • US9189349B2 patent drawing
  • US9189349B2 patent drawing
  • US9189349B2 patent drawing

AI summary

A distributed blade server system, a management server and a switching method are provided. The method includes: determining a standby blade of a first blade when it is determined that the first blade is in abnormal operation; delivering, based on an access relationship between a startup card of the first blade and a first storage partition, a first configuration command to a storage system, the first configuration command including information of an access relationship between a startup card of the standby blade and the first storage partition, so that the storage system configures the access relationship between the startup card of the standby blade and the first storage partition; and delivering a startup command to the standby blade.