Ethernet Fabric Data Processing System for Petabyte Scale Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data processing systems are inefficient in handling petabyte-scale data sets, leading to impractical analysis due to limitations in high bandwidth access and parallel processing throughput.
Innovation Solution
A scalable data processing system architecture featuring multiple CPU subsystems, memory complexes, and an Ethernet switch fabric, with cache coherence mechanisms and software-implemented flash translation layer policies, enables efficient parallel access and management of large data sets across interconnected memory leaves and branches.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional data processing systems are used, then system simplicity is maintained, but high bandwidth access to petabyte-scale data sets is inefficient and impractical
Solution Approach 1:
The system is divided into multiple CPU subsystems, each with its own local memory, connected through an Ethernet switch fabric. This segmentation allows parallel access to distributed data sets, enabling high bandwidth access to petabyte-scale data by distributing the data across multiple independent memory subsystems rather than relying on a single centralized system.
Solution Approach 2:
The patent transitions from a traditional hierarchical memory architecture to a distributed parallel architecture using Ethernet networking. By adding the network dimension with multiple CPU subsystems connected via Ethernet switch fabric, the system achieves scalable high bandwidth access that conventional single-system architectures cannot provide.
2Productivity
If conventional parallel processing architectures are used, then processing throughput is limited, but system complexity remains manageable
Solution Approach 1:
The Ethernet switch fabric serves multiple functions: it provides interconnection between CPU subsystems, enables data transfer between distributed memory systems, and supports both storage and computation operations. This universal interconnection infrastructure simplifies the overall system architecture while enabling high parallel processing throughput through standardized Ethernet protocols.
3Productivity
If data is distributed across multiple memory targets, then access bandwidth increases, but data management complexity increases
Solution Approach 1:
The Ethernet switch fabric acts as an intermediary that manages data distribution and access across multiple memory targets. It handles the complexity of routing data requests to the appropriate CPU subsystem and memory location, while presenting a simplified interface to applications. This mediator approach enables high bandwidth access to distributed data without requiring applications to directly manage the underlying distribution complexity.
Data Source
AI summary
According to one embodiment, a data processing system includes a plurality of processing units, each processing unit having one or more processor cores. The system further includes a plurality of memory roots, each memory root being associated with one of the processing units. Each memory root includes one or more branches and a plurality of memory leaves to store data. Each of the branches is associated with one or more of the memory leaves and to provide access to the data stored therein. The system further includes a memory fabric coupled to each of the branches of each memory root to allow each branch to access data stored in any of the memory leaves associated with any one of remaining branches.


