Memory Die Local and Global Processor Architecture for Latency Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The separation of processors from memory in three-dimensional semiconductor structures leads to latency in data transmission between processors and memory, which existing technologies have not adequately addressed.

Innovation Solution

A memory die is designed with both local processors for executing local calculations on data within specific banks and a global processor to control and combine the results from multiple banks, optimizing data processing and reducing latency through a multi-die databus and bus gating circuits for efficient bandwidth utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If processors are separated from memory in three-dimensional semiconductor structures, then manufacturing and structural organization are simplified, but data transmission latency between processors and memory increases

Engineering Contradiction:
Improvestructural organizationVSAvoiddata transmission latency
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The patent divides the processing system into local processors (integrated with specific memory banks) and a global processor (separate from memory). Each memory bank is paired with a dedicated local processor that performs computations directly on stored data, while the global processor handles overall coordination. This segmentation reduces data transmission latency by enabling local processing without requiring constant communication with the separate global processor.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces local processors as intermediary components between memory banks and the global processor. These local processors act as mediators that can perform computations directly on data within their associated memory banks, reducing the need for data to be transmitted to and from the globally separated processor, thereby addressing the latency issue while maintaining the structural benefits of separation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If local processors are integrated with memory banks, then data processing speed increases through reduced latency, but device complexity increases due to additional processing components

Engineering Contradiction:
Improvedata processing speedVSAvoidprocessing component integration
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent segments processing functions into local processors (integrated with individual memory banks) and a global processor (handling overall control). Each local processor is dedicated to specific memory banks, enabling parallel processing operations that increase overall data processing speed while distributing complexity across multiple simpler components rather than one complex centralized processor.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a traditional hierarchical processor-memory architecture to a distributed parallel architecture where processing occurs at multiple levels (local processors at the bank level, global processor at the system level). This dimensional change in architectural organization enables simultaneous processing operations across multiple banks, increasing speed while managing complexity through distributed functionality.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If multiple local processors operate on different memory banks, then parallel processing capability increases, but bus bandwidth utilization becomes more challenging to manage

Engineering Contradiction:
Improveparallel processing capabilityVSAvoidbus bandwidth management
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the bus system into dedicated local buses connecting each local processor to its associated memory banks, and a separate global bus for the global processor. This segmentation allows parallel processors to operate independently on their dedicated buses without contending for shared bandwidth, enabling efficient parallel processing while simplifying bandwidth management through physical separation of communication paths.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces local processors as intermediary components that handle data processing locally before results need to be communicated globally. This intermediary layer reduces the bandwidth requirements of the global bus by filtering and processing data locally, managing bus bandwidth complexity through hierarchical data flow control where only necessary data traverses the global communication path.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11276459B2Memory die including local processor and global processor, memory device, and electronic device
Publication Date: 2022.03.15 SAMSUNG ELECTRONICS CO LTD
  • US11276459B2 patent drawing
  • US11276459B2 patent drawing
  • US11276459B2 patent drawing

AI summary

A memory die includes a first bank including first memory cells; a second bank including second memory cells; a first local processor connected with first bank local input/output lines through which first local bank data of the first bank are transmitted, and configured to execute a first local calculation on the first local bank data; a second local processor connected with second bank local input/output lines through which second local bank data of the second bank are transmitted, and configured to execute a second local calculation on the second local bank data; and a global processor configured to control the first bank, the second bank, the first local processor, and the second local processor and to execute a global calculation on a first local calculation result of the first local calculation and a second local calculation result of the second local calculation.