Shared Scratchpad Memory With DMA and Load-Store Paths for Neural Cores

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing neural network hardware circuits face inefficiencies in data communication and bandwidth utilization due to the lack of shared memory resources among processor cores, leading to suboptimal performance and increased off-chip data transfer penalties.

Innovation Solution

A hardware circuit architecture that incorporates a shared memory system with DMA and load-store data paths, allowing parallel data routing between processor cores and vector registers, enhancing bandwidth and reducing the need for dedicated wires for data transfers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If dedicated memory resources are allocated to each processor core, then data access speed is improved, but device complexity and wire requirements increase

Engineering Contradiction:
Improvedata access speedVSAvoidwire requirements
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent merges memory resources into a shared memory system that is accessible by multiple processor cores simultaneously. The shared memory is coupled to multiple processor cores through a common interface, eliminating the need for dedicated memory resources and extensive wiring for each core while maintaining high data access speeds through parallel access capabilities.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The shared memory system serves multiple processor cores universally, acting as a common resource that can be accessed by any core in the system. This multi-functional approach allows the same memory resources to support multiple cores without requiring separate dedicated memory structures for each core.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If more memory resources are provided for data communication, then bandwidth utilization is improved, but device complexity increases

Engineering Contradiction:
Improvebandwidth utilizationVSAvoidmemory resource complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent combines multiple memory resources into a unified shared memory structure that serves all processor cores. This consolidation increases bandwidth utilization by allowing parallel data transfers between the shared memory and multiple cores simultaneously, while reducing overall device complexity by eliminating redundant memory structures.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The shared memory system introduces a new dimension of parallelism by enabling simultaneous data transfers between the shared memory and multiple processor cores through separate data paths. This dimensional approach to memory organization increases bandwidth utilization without proportionally increasing complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If parallel data paths are implemented between shared memory and processor cores, then throughput is improved, but device complexity increases

Engineering Contradiction:
ImprovethroughputVSAvoiddata path complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements parallel data paths between the shared memory and multiple processor cores, creating additional dimensions for data flow. This allows simultaneous data transfers to occur in parallel, significantly improving throughput while the modular structure of the shared memory interface keeps data path complexity manageable.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The shared memory system is segmented into multiple independent data paths that can operate in parallel. Each data path provides a separate channel for data transfer between the shared memory and individual processor cores, enabling parallel operations that improve throughput while keeping each individual data path relatively simple.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12367383B2Shared scratchpad memory with parallel load-store
Publication Date: 2025.07.22 GOOGLE LLC
  • US12367383B2 patent drawing
  • US12367383B2 patent drawing
  • US12367383B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer-readable media, are described for a hardware circuit configured to implement a neural network. The circuit includes a first memory, respective first and second processor cores, and a shared memory. The first memory provides data for performing computations to generate an output for a neural network layer. Each of the first and second cores include a vector memory for storing vector values derived from the data provided by the first memory. The shared memory is disposed generally intermediate the first memory and at least one core and includes: i) a direct memory access (DMA) data path configured to route data between the shared memory and the respective vector memories of the first and second cores and ii) a load-store data path configured to route data between the shared memory and respective vector registers of the first and second cores.