Controller Chip Cluster Redundancy for Fault Tolerance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing controller chips lack sufficient redundancy in power supply and clock generation, leading to insufficient functionality during failures in the main application cluster, as previous architectures rely on isolated power management systems that are insufficient to maintain chip functionality during faults.
Innovation Solution
A microcontroller chip with multiple clusters, each having its own power supply and clock tree, along with a monitoring cluster that detects failures and redistributes tasks between clusters to maintain functionality, allowing for seamless operation even when one cluster fails by having a secondary cluster take over the functions of the primary cluster.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If completely redundant controllers or system-on-chip on separate dies are used, then reliability is improved, but device complexity increases
Solution Approach 1:
The controller chip is divided into multiple independent clusters (first cluster, second cluster, third cluster), each capable of autonomous operation with its own power supply and clock generation. This segmentation allows the system to maintain functionality even when one cluster fails, resolving the contradiction by providing redundancy through modular division rather than complete separate systems.
Solution Approach 2:
Multiple clusters are merged into a single controller chip, combining redundancy with integration. The clusters share common resources such as memory and I/O interfaces while maintaining independent processing capabilities, achieving fault tolerance without requiring completely separate system-on-chip designs.
2Reliability
If isolated power management system block is used, then reliability is improved, but functionality during fault becomes insufficient
Solution Approach 1:
Each cluster is designed with universal functionality to perform multiple roles: primary application processing, fallback processing, and monitoring functions. The third cluster can serve as both a monitoring unit and a fallback processing unit, allowing the system to adapt its configuration based on operational needs and fault conditions.
Solution Approach 2:
The monitoring cluster proactively detects faults in other clusters before they cause system failure. By continuously monitoring power supply status and clock signals, the system can initiate failover procedures in advance, maintaining functionality without waiting for complete system failure.
3Reliability
If multiple clusters with separate power supplies are implemented, then fault tolerance is improved, but manufacturing complexity increases
Solution Approach 1:
The chip is manufactured as a single monolithic substrate divided into distinct cluster segments. Each cluster contains its own power management and clock generation circuits, allowing standardized manufacturing processes while achieving post-manufacturing redundancy. This segmentation enables fault tolerance without requiring complex multi-die assembly processes.
Data Source
AI summary
A controller chip includes a first cluster including one or more first controller units, a first power supply grid, a first clock tree structure to supply one or more clock signals, and at least a first power supply input. A second cluster includes one or more second controller units, a second power supply grid, a second clock tree structure to supply one or more clock signals, and at least a second power supply input. A monitoring cluster includes a monitoring circuit configured to: monitor the power supply and the clock signal supply of each of the first cluster and second cluster, and in the event of determining at least one of a power supply failure or a clock signal supply failure in one cluster of the first cluster or the second cluster, indicate the failure to the other cluster to take one or more actions.

