Indexing External Memory in Reconfigurable Compute Fabric
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computer architectures face performance and capacity constraints due to the time and energy required for data movement between processors and memory, limiting advancements beyond transistor scaling.
Innovation Solution
The implementation of memory-centric compute topologies, specifically compute-near-memory (CNM) systems, which integrate processors with memory or data storage components, utilizing hybrid threading processors and reconfigurable compute fabrics to facilitate low-latency operations through custom compute fabrics and programmable atomic units.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is stored in external memory and accessed via bus, then memory capacity is improved, but access time and energy consumption increase
Solution Approach 1:
The patent segments memory into multiple banks (first bank, second bank, etc.) that can be accessed independently and simultaneously. Each compute element can access different memory banks in parallel, reducing the time penalty of large memory capacity while maintaining high data availability through spatial segmentation of the memory space.
Solution Approach 2:
The patent introduces a new dimension of parallelism by organizing memory access across multiple compute elements that can simultaneously access different memory locations. Instead of sequential access through a single bus, the system uses a mesh network topology where multiple data paths exist concurrently, transforming the access model from one-dimensional sequential to multi-dimensional parallel access.
2Productivity
If data is moved between processors and memory, then computational capacity is improved, but energy consumption increases
Solution Approach 1:
The patent merges computation and memory access functions into an integrated architecture where compute elements are directly coupled to memory banks through a mesh network. This combination eliminates the need for separate data movement operations between distinct processor and memory units, reducing energy consumption while maintaining high computational capacity through tight integration of storage and processing functions.
Solution Approach 2:
Each compute element can independently access memory banks without requiring centralized data movement control. The distributed architecture allows compute elements to self-service their data needs directly from local and remote memory banks through the mesh network, eliminating energy-intensive centralized data shuffling and reducing overall system energy consumption.
3Device complexity
If conventional Von Neumann architecture is used, then system simplicity is maintained, but performance is constrained
Solution Approach 1:
The patent implements a dynamically reconfigurable mesh network that can adapt its connectivity and data flow patterns based on computational requirements. The system transitions from the static, fixed data path of conventional Von Neumann architecture to a dynamic topology where compute elements and memory banks can communicate through multiple adaptable paths, enabling high performance while maintaining manageable complexity through systematic design.
Solution Approach 2:
The mesh network architecture provides universal connectivity where any compute element can access any memory bank through multiple possible paths. This multi-functional communication infrastructure replaces the specialized, single-purpose data bus of conventional architectures, delivering superior performance while maintaining system simplicity through a unified, scalable interconnection topology that handles diverse computational workloads.
Data Source
AI summary
Various examples are directed to systems and methods in which a flow controller of a first synchronous flow may receive an instruction to execute a first loop using the first synchronous flow. The flow controller may determine a first iteration index for a first iteration of the first loop. The flow controller may send, to a first compute element of the first synchronous flow, a first synchronous message to initiate a first synchronous flow thread for executing the first iteration of the first loop. The first synchronous message may comprise the iteration index. The first compute element may execute an input/output operation at a first location of a first compute element memory indicated by the first iteration index.


