Dynamically Programmable Memory Distribution for Neural Network Bandwidth

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current hardware platforms face challenges in achieving memory access performance and efficiency for neural network processing, which is critical for computationally and data-intensive artificial intelligence tasks.

Innovation Solution

A high bandwidth memory system is designed with multiple memory units arranged around a processing component, allowing for dynamic programmable distribution schemes and configurable access units to efficiently transfer data, thereby optimizing memory utilization and reducing latency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data is distributed across multiple memory units using a fixed distribution scheme, then memory bandwidth utilization is limited, but implementing a dynamically programmable distribution scheme increases device complexity

Engineering Contradiction:
Improvememory bandwidth utilizationVSAvoiddistribution scheme configuration
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements a dynamically programmable distribution scheme where the distribution pattern can be changed based on different data access patterns and workload requirements. The system can switch between various distribution schemes (e.g., round-robin, random, sequential) to optimize memory bandwidth utilization for different neural network operations, resolving the contradiction between achieving high productivity and managing device complexity through flexibility rather than fixed rigid structures

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes distribution parameters dynamically based on the specific memory access patterns required for different neural network layers and operations. By adjusting distribution parameters such as stride size, distribution pattern, and memory unit mapping, the system optimizes bandwidth utilization without requiring complete redesign of the memory architecture, thus improving productivity while controlling complexity through parameter adjustment rather than structural overhaul

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If multiple processing elements use the same distribution scheme, then implementation is simplified, but memory units may not be efficiently utilized leading to reduced performance

Engineering Contradiction:
Improveimplementation simplicityVSAvoidmemory unit utilization
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent applies different distribution schemes to different processing elements or memory access patterns based on their specific requirements. For example, convolutional layers might use one distribution scheme while fully connected layers use another, allowing each processing element to have optimized memory access patterns tailored to its local workload, thereby improving overall memory unit utilization while maintaining implementation simplicity through modular scheme selection

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically selects and applies appropriate distribution schemes based on the current operational context and workload characteristics. This dynamic adaptation allows the system to optimize memory utilization for each processing element's specific tasks while maintaining a unified and simple implementation framework, resolving the contradiction between ease of manufacture and productivity through context-aware dynamic configuration

Inventive Principle:
Principle #15Dynamics

3Device complexity

If data is accessed sequentially from memory, then memory access pattern is simple, but memory access latency increases for neural network processing

Engineering Contradiction:
Improvememory access patternVSAvoidmemory access latency
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The system performs preliminary data distribution and prefetching operations to optimize memory access patterns before actual computation begins. By pre-distributing data across memory units according to the required access pattern and using prediction mechanisms to anticipate future access needs, the system reduces memory access latency without introducing complex runtime decision-making, thus resolving the contradiction between simple access patterns and reduced latency through advance preparation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms that monitor memory access patterns and adjust data distribution strategies accordingly. By analyzing actual access behavior and comparing it with predicted patterns, the system can dynamically refine its data distribution approach to minimize latency while maintaining relatively simple access patterns, resolving the contradiction through iterative optimization based on observed performance feedback

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11663043B2High bandwidth memory system with dynamically programmable distribution scheme
Publication Date: 2023.05.30 META PLATFORMS INC
  • US11663043B2 patent drawing
  • US11663043B2 patent drawing
  • US11663043B2 patent drawing

AI summary

A system comprises a processor coupled to a plurality of memory units. Each of the plurality of memory units includes a request processing unit and a plurality of memory banks. The processor includes a plurality of processing elements and a communication network communicatively connecting the plurality of processing elements to the plurality of memory units. At least a first processing element of the plurality of processing elements includes a control logic unit and a matrix compute engine. The control logic unit is configured to access data from the plurality of memory units using a dynamically programmable distribution scheme.