Dual Chassis Management Controller Status Reporting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data center rack systems face operational disruptions when a single Chassis Management Controller (CMC) fails, as they lose the ability to monitor and report status, leading to potential undetected issues with network devices, and existing backup solutions are not fully effective in maintaining continuous monitoring.
Innovation Solution
Implementing a dual chassis system with a master and partner CMCs, where the partner CMC takes over status reporting if the master CMC fails, ensuring continuous monitoring and operation by establishing communication through system buses and network interfaces, using protocols like I2C or UART for data exchange.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single CMC is used to manage network devices, then device complexity is reduced, but reliability deteriorates when the CMC fails
Solution Approach 1:
The patent assigns different functional roles to different CMCs: the first CMC manages first network devices while the second CMC manages second network devices. This local specialization allows the system to maintain management functionality even when one CMC fails, as the other CMC continues to operate its designated devices without interference.
Solution Approach 2:
The patent establishes communication pathways between CMCs and with external systems in advance. The first CMC is pre-configured to communicate with both the first network devices and the second network devices, creating redundant communication routes before failure occurs. This preliminary setup ensures that status reporting can continue through alternative paths when the primary CMC fails.
2Loss of information
If the first CMC monitors both first and second network devices, then monitoring coverage is improved, but reliability deteriorates due to single point of failure
Solution Approach 1:
The patent divides the monitoring function into two separate CMCs, each responsible for specific network devices. The first CMC monitors first network devices while the second CMC monitors second network devices. This segmentation eliminates the single point of failure problem, as the failure of one CMC does not result in loss of status information for all devices, only for its designated subset.
Solution Approach 2:
The patent introduces the second CMC as an intermediary that can assume the monitoring role when the first CMC fails. The second CMC is positioned to receive status information from network devices and relay it to external systems, serving as a backup mediator that maintains information flow continuity when the primary monitoring path is disrupted.
3Reliability
If communication paths are established between CMCs and network devices, then status reporting capability is improved, but device complexity increases
Solution Approach 1:
The patent designs the first CMC with multi-functionality, enabling it to communicate with both first network devices and second network devices. This universal communication capability allows the first CMC to perform its primary monitoring function while also serving as a backup for the second CMC, reducing the need for dedicated communication paths and simplifying the overall infrastructure.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A dual CMC structure is disclosed for a rack mounted structure. The structure has a first chassis with power supplies, a first chassis management controller, and a first set of network devices. A second chassis includes power supplies, a second chassis management controller and a second set of network devices. The respective chassis management controllers obtain status data of the power supplies, as well as status data from the other chassis management controllers. The first chassis management controller is designated as the master controller and reports the status data from both the first and second chassis. The structure is operable to change communication of the status data to the second chassis management controller, in the event the first chassis management controller fails.