Switch Data-Plane Device Main Die Failure Response
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High-capacity datacenter switches face reliability challenges due to the 'blast radius' effect, where a single switch failure can cause significant service outages, and traditional redundancy solutions are prohibitively expensive in large networks.
Innovation Solution
A switch data-plane device with a main die and multiple chiplets interconnected via primary and secondary interconnects, allowing for a synchronous graceful failure process where chiplets take over packet forwarding duties, reducing complexity and maintaining low-bandwidth connectivity for network maintenance without service outages.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If switch level redundancy is implemented to improve reliability, then the Mean Time Between Failures (MTBF) is exponentially reduced, but the cost becomes prohibitively expensive in large datacenter networks
Solution Approach 1:
The switch is divided into multiple independent chiplets (ingress chiplet, egress chiplet, forwarding chiplet) that can operate autonomously. This segmentation allows the system to maintain partial functionality even when the main die fails, providing reliability without requiring a complete backup switch.
Solution Approach 2:
The patent implements a preliminary graceful shutdown process that is triggered before complete system failure. The control processor detects main die failure and initiates a controlled transition to chiplet-level operation, preventing catastrophic failure and maintaining service during the transition period.
2Productivity
If a main die is used to apply primary forwarding process, then high throughput switching is achieved, but the system becomes vulnerable to single point of failure
Solution Approach 1:
The forwarding functionality is segmented between the main die (primary forwarding process) and multiple chiplets (secondary forwarding process). This segmentation eliminates the single point of failure by distributing forwarding capability across multiple independent units that can take over if the main die fails.
Solution Approach 2:
The control processor acts as an intermediary that manages the transition between primary and secondary forwarding processes. It monitors the main die's health and orchestrates the graceful shutdown and failover to chiplet-level forwarding, ensuring continuous operation.
3Reliability
If complete redundancy is implemented to eliminate service outages, then reliability is maximized, but the complexity and cost of the system becomes prohibitive
Solution Approach 1:
Instead of implementing complete redundancy with full backup switches, the patent applies partial redundancy at the chiplet level. The ingress chiplet, egress chiplet, and forwarding chiplet provide sufficient redundancy to maintain service during main die failure without the full complexity of complete system redundancy.
Solution Approach 2:
The patent uses individual chiplets as replaceable units that can fail independently without causing complete system failure. Each chiplet is a smaller, less expensive unit compared to full switch redundancy, providing cost-effective reliability through modular failure isolation.
Data Source
AI summary
A method for responding to a failure of a main die of a switch data-plane device, the method may include applying a secondary packet forwarding process by multiple chiplets, following the failure of the main die and during at least a part of an execution of a synchronous graceful process that follows the failure of the main die; wherein the multiple chiplets are interconnected to each other by a secondary interconnect; wherein the multiple chiplets and are coupled to the main die by a primary interconnect; wherein the applying of the secondary packet forwarding process is less complex than a primary forwarding process applied by the main die while the main die is functional.


