Mesh Topology for Distributed Computing Reliability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional bus and star topologies in distributed computing systems face limitations in high bandwidth and reliability, particularly in environments prone to radiation effects, where single points of failure can compromise system integrity and throughput.
Innovation Solution
A distributed computing system with a mesh topology where each node has communication links to a primary I/O switch and two other nodes, providing redundant paths and utilizing primary and redundant I/O switching modules for high availability and reliability, along with serializer/deserializer modules for data transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If bus topology is used to interconnect computing components, then system simplicity is maintained, but data transfer throughput is limited to approximately 2 Gbps per backplane bus
Solution Approach 1:
The patent segments the monolithic bus topology into multiple independent point-to-point links organized in a mesh pattern. Instead of one shared bus, the system divides communication paths into multiple dedicated channels (e.g., 40 Gbps links) that can operate simultaneously, thereby increasing total throughput while maintaining manageable complexity through modular organization
Solution Approach 2:
The patent transitions from a one-dimensional bus architecture to a two-dimensional mesh topology. Nodes are arranged in a grid where each node connects to multiple neighbors in different dimensions, enabling parallel data paths and significantly increasing throughput capacity without proportionally increasing system complexity
2Productivity
If multiple backplane busses are employed to achieve higher throughput, then data transfer capacity increases, but I/O challenges and system complexity increase
Solution Approach 1:
The patent merges multiple high-speed point-to-point links into a unified mesh topology that provides both high throughput and simplified I/O access. Instead of managing multiple separate busses with their own controllers and protocols, the mesh topology integrates these connections into a single coherent architecture where nodes can access multiple resources through standardized link interfaces
Solution Approach 2:
The mesh topology provides universal connectivity where each node can directly or indirectly access any other node through multiple equivalent paths. This multi-functional capability allows the same physical infrastructure to support various communication patterns (one-to-one, one-to-many, many-to-many) without requiring separate dedicated bus structures for each scenario
3Device complexity
If traditional star topology with single switching function is used, then system simplicity is maintained, but reliability is compromised due to single point of failure
Solution Approach 1:
The patent applies local quality by giving each node in the mesh topology redundant connection capabilities. Instead of relying on a central switch, each node locally maintains multiple independent paths to other nodes, ensuring that if one path or neighboring node fails, communication can continue through alternative local routes without affecting overall system reliability
4Reliability
If dual star topology is used to provide redundancy, then reliability improves, but the topology still has choke points that restrict speed and efficiency
Solution Approach 1:
The patent segments the centralized switching function into distributed intelligence across multiple nodes in the mesh. Instead of packets funneling through central switches (creating choke points), each node can independently route packets along multiple parallel paths, eliminating speed restrictions while maintaining redundancy for high availability
Solution Approach 2:
The mesh topology adds spatial dimensions to packet routing, allowing data to flow through multiple dimensional paths simultaneously. This eliminates the single-dimension bottleneck of star topologies where all traffic must pass through central switches, enabling higher speeds while maintaining reliability through alternative routing dimensions
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Methods, systems, and devices for distributed computing are provided. Clusters of nodes are provided, each node have a communication link to a primary I/O switch as well as to two other nodes, thereby providing redundant alternative communication paths between different components of the system. Primary and redundant I/O switching modules may provide further redundancy for high availability and high reliability applications, such as applications that may be subjected to the environment as would be found in space, including radiation effects. Nodes in a cluster may provide data storage, processing, and/or input/output functions, as well as one or more alternate communications paths between system components. Multiple clusters of nodes may be coupled together to provide enhanced performance and/or reliability.