Shared Scratchpad Memory With DMA and Load-Store Paths for Neural Cores
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing neural network hardware circuits face inefficiencies in data communication and bandwidth utilization due to the lack of shared memory resources among processor cores, leading to suboptimal performance and increased off-chip data transfer penalties.
Innovation Solution
A hardware circuit architecture that incorporates a shared memory system with DMA and load-store data paths, allowing parallel data routing between processor cores and vector registers, enhancing bandwidth and reducing the need for dedicated wires for data transfers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If dedicated memory resources are allocated to each processor core, then data access speed is improved, but device complexity and wire requirements increase
Solution Approach 1:
The patent merges memory resources into a shared memory system that is accessible by multiple processor cores simultaneously. The shared memory is coupled to multiple processor cores through a common interface, eliminating the need for dedicated memory resources and extensive wiring for each core while maintaining high data access speeds through parallel access capabilities.
Solution Approach 2:
The shared memory system serves multiple processor cores universally, acting as a common resource that can be accessed by any core in the system. This multi-functional approach allows the same memory resources to support multiple cores without requiring separate dedicated memory structures for each core.
2Productivity
If more memory resources are provided for data communication, then bandwidth utilization is improved, but device complexity increases
Solution Approach 1:
The patent combines multiple memory resources into a unified shared memory structure that serves all processor cores. This consolidation increases bandwidth utilization by allowing parallel data transfers between the shared memory and multiple cores simultaneously, while reducing overall device complexity by eliminating redundant memory structures.
Solution Approach 2:
The shared memory system introduces a new dimension of parallelism by enabling simultaneous data transfers between the shared memory and multiple processor cores through separate data paths. This dimensional approach to memory organization increases bandwidth utilization without proportionally increasing complexity.
3Productivity
If parallel data paths are implemented between shared memory and processor cores, then throughput is improved, but device complexity increases
Solution Approach 1:
The patent implements parallel data paths between the shared memory and multiple processor cores, creating additional dimensions for data flow. This allows simultaneous data transfers to occur in parallel, significantly improving throughput while the modular structure of the shared memory interface keeps data path complexity manageable.
Solution Approach 2:
The shared memory system is segmented into multiple independent data paths that can operate in parallel. Each data path provides a separate channel for data transfer between the shared memory and individual processor cores, enabling parallel operations that improve throughput while keeping each individual data path relatively simple.
Data Source
AI summary
Methods, systems, and apparatus, including computer-readable media, are described for a hardware circuit configured to implement a neural network. The circuit includes a first memory, respective first and second processor cores, and a shared memory. The first memory provides data for performing computations to generate an output for a neural network layer. Each of the first and second cores include a vector memory for storing vector values derived from the data provided by the first memory. The shared memory is disposed generally intermediate the first memory and at least one core and includes: i) a direct memory access (DMA) data path configured to route data between the shared memory and the respective vector memories of the first and second cores and ii) a load-store data path configured to route data between the shared memory and respective vector registers of the first and second cores.


