Multi-Vendor Server Management via Redfish API Automation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current server management systems face challenges in efficiently managing multiple vendors' servers, particularly in small-scale companies where specialized skills are lacking, leading to difficulties in promptly addressing server failures and resulting in significant losses due to interruptions and high recovery costs.
Innovation Solution
A server management system that includes an administrator terminal, a client terminal, and a management server capable of collecting and analyzing data from multiple vendors' servers, providing proactive failure handling, firmware updates, and standardization using Redfish API, enabling efficient multi-vendor support and automation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a specialized server manager is employed, then server management capability is improved, but hiring cost increases
Solution Approach 1:
The server management system enables self-service through automated monitoring, alerting, and management functions. The system automatically detects server failures, sends notifications to administrators, and provides management capabilities without requiring specialized external personnel, thus reducing hiring costs while maintaining reliable server management
Solution Approach 2:
The server management system is designed to universally manage servers from multiple vendors through a single platform. It provides multi-vendor support with standardized interfaces, allowing one system to perform multiple management functions across different server types, eliminating the need for specialized managers for each vendor
2Reliability
If post-failure recovery method is used, then recovery is achieved, but server operation stops during recovery period causing significant loss
Solution Approach 1:
The system performs preliminary actions by proactively detecting potential server failures before they occur and notifying administrators in advance. This allows preventive maintenance and preparation of recovery measures before actual failures happen, minimizing downtime when issues do occur
Solution Approach 2:
The system implements continuous monitoring and feedback mechanisms that detect server status changes and immediately notify administrators. This real-time feedback enables rapid response to failures, reducing the time servers remain non-operational by quickly alerting personnel to intervene
3Adaptability or versatility
If multiple vendors' servers are managed, then system versatility is improved, but management complexity increases
Solution Approach 1:
The server management system achieves multi-vendor support through universal standardized interfaces and protocols. It provides a single management platform that can handle servers from multiple vendors without requiring separate management systems, thus maintaining versatility while avoiding increased complexity through standardization
Data Source
AI summary
Disclosed is an extended high-performance server integrated management system based on remote server management standard to efficiently manage existing high-performance servers and ultra-high-performance servers capable of supporting extended BMC functions. This system includes multiple modules responsible for each major function, considering a possibility of large-scale expansion, and each module can be run on an independent server depending on the size of the management system. In addition, this system utilizes the Intelligent Platform Management Interface (IPMI) standard, which is widely used for efficient remote hardware management of ultra-high-performance servers, and supports function for collecting hardware monitoring information linked to BMC for various functions such as temperature, power, and fan speed and collectively controlling the hardware on the basis of the collected data. In this case, an out-of-band based agent communication method was used to extend the BMC functionality. In addition, it also supports group management of each server, boot option change, Serial Over Lan (SOL)-based remote console access function, and the function of delivering event-specific warnings and alarms based on collected hardware monitoring information. Furthermore, the collected hardware status information of each server is continuously stored in the database and managed to be used as basic data for future failure prediction and prevention.


