Fabric Controller Segmentation for Fault Tolerance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional fabric attached architectures face challenges with fault tolerance and data flow path redundancy due to the limitations of top-of-rack switches, which can lead to latencies and inefficient rerouting when a switch fails, rendering entire racks inoperable and limiting scalability.
Innovation Solution
Each resource node in the fabric attached architecture is associated with a switch and fabric controller, utilizing high-speed interconnects and a virtualized data-link layer to provide fault tolerance and redundancy, enabling efficient rerouting and integration of diverse resource nodes through a distributed fabric controller system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If top-of-rack switches are used to connect resource nodes, then device complexity is reduced and ease of operation is improved, but fault tolerance deteriorates and reliability worsens when a switch fails
Solution Approach 1:
The patent divides the fabric controller functionality into multiple distributed fabric controllers, each associated with specific resource nodes. This segmentation eliminates the single point of failure in top-of-rack switches, as each fabric controller operates independently and can be replaced without affecting the entire system. The segmentation principle directly resolves the contradiction by improving reliability through distribution while maintaining operational simplicity.
Solution Approach 2:
The patent introduces fabric controllers as intermediary devices between resource nodes and the fabric network, replacing the direct switch-based connectivity. These fabric controllers act as mediators that manage data flow and provide fault isolation, allowing the system to maintain ease of operation while significantly improving fault tolerance through the intermediary layer.
2Device complexity
If top-of-rack switches are used for fabric connectivity, then device complexity is reduced, but data flow path redundancy deteriorates and latency increases during rerouting
Solution Approach 1:
The patent implements preliminary action by pre-establishing multiple data flow paths through the distributed fabric controller network before failures occur. When a failure is detected, the system can immediately reroute traffic through alternative paths without the need for complex real-time recalculations, significantly reducing rerouting latency while maintaining manageable device complexity.
Solution Approach 2:
The patent introduces dynamic path selection capabilities in the distributed fabric controllers, allowing the system to adaptively route data flows around failures in real-time. This dynamic behavior enables the system to maintain low latency during rerouting events while keeping the overall device complexity manageable through automated control.
3Ease of manufacture
If top-of-rack switches are used, then ease of manufacture is improved, but scalability deteriorates and fault tolerance is reduced
Solution Approach 1:
The patent segments the fabric control functionality into distributed fabric controllers that can be independently added or removed from the system. This segmentation enables scalable expansion without requiring changes to the core switch infrastructure, allowing the system to grow while maintaining ease of manufacture through modular additions.
Solution Approach 2:
The fabric controllers are designed with universal functionality to handle diverse resource node types and communication patterns. This multi-functionality enables the system to scale to various configurations and workloads while maintaining a consistent manufacturing approach, resolving the contradiction between ease of manufacture and scalability.
Data Source
AI summary
A plurality of fabric controllers distributed throughout a fabric attached architecture and each associated with at least one resource node. The plurality of fabric controllers configured to control each associated resource node. Resources of the resource nodes are utilized in virtual environments responsive to respective fabric controllers issuing instructions received from the fabric attached architecture to respective resource nodes.


