Manageability Redundancy for Micro Server SoC Deployments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current manageability solutions for System-on-a-Chip (SoC) systems, especially in micro server and clustered deployments, face reliability, availability, and serviceability issues due to single-point failures in manageability access points, leading to potential failures of entire components when one SoC fails.
Innovation Solution
Implementing a system with dynamically reconfigurable integrated circuit blocks that can perform management functions and task functions across multiple SoC blocks, utilizing an all-to-all communication infrastructure to enable redundancy and failover capabilities, allowing other blocks to take over management and task responsibilities if a block fails.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If each SoC has its own manageability access point (MAP), then manageability functionality is provided for each node, but a failure of a single MAP leads to failure of the entire component, reducing reliability
Solution Approach 1:
The patent implements a universal management controller that can perform management functions for multiple SoC nodes through a single MAP. This multi-functional approach allows one MAP to manage multiple nodes, eliminating the single-point-failure issue where each node had its own dedicated MAP. The management controller can dynamically allocate its capabilities to different nodes as needed.
Solution Approach 2:
The patent combines multiple individual MAP functions into a single unified management controller. Instead of having separate MAP instances for each SoC node, the system merges these functions into one centralized controller that can service multiple nodes, thereby improving reliability by eliminating single-point failures while reducing the overall number of manageability access points.
2Productivity
If multiple SoC systems are integrated within a single silicon package, then processing density is increased, but current manageability capabilities become inadequate, causing RAS issues
Solution Approach 1:
The management controller is designed with universal capabilities to manage multiple SoC nodes within a single silicon package. It can dynamically allocate its management functions to different nodes, providing adequate RAS capabilities for high-density integrations without requiring a separate dedicated MAP for each node, thus maintaining reliability while supporting high processing density.
Solution Approach 2:
The management controller implements dynamic functionality allocation, where its management capabilities can be dynamically assigned to different SoC nodes based on operational needs. This dynamic approach allows the system to scale from single-node to multi-node configurations within the same silicon package, adapting to varying processing density requirements while maintaining adequate manageability and RAS capabilities.
3Reliability
If a single MAP is used to manage multiple IC blocks, then redundancy and failover capabilities are enabled, but the system complexity increases
Solution Approach 1:
The management controller is designed as a universal resource that can service multiple IC blocks, enabling redundancy and failover capabilities. Instead of each block having its own dedicated MAP, the universal controller can dynamically allocate its management functions to different blocks, providing fault tolerance while avoiding the complexity of multiple redundant MAP instances.
Solution Approach 2:
The patent implements a virtualization approach where the universal management controller creates virtual management instances for different IC blocks. Rather than duplicating physical MAP hardware for each block, the system uses software-based virtualization to provide copies of management functions, achieving redundancy and failover capabilities with reduced hardware complexity.
Data Source
AI summary
Technologies for providing manageability redundancy for micro server and clustered System-on-a-Chip (SoC) deployments are presented. A configurable multi-processor apparatus may include multiple integrated circuit (IC) blocks where each IC block includes a task block to perform one or more assignable task functions and a management block to perform management functions with respect to the corresponding IC block. Each task block and each management block may include one or more instruction processors and corresponding memory. Each IC block may be controllable to perform a function of one or more other IC blocks. The IC blocks may communicate with each other via a management communication infrastructure that may include a communication path from each of the management blocks to each of the other management blocks. Via the management communication infrastructure, the management blocks may bridge communication paths between pairs of management blocks.


