Chiplet LLC Hierarchy Using NUCA for Die-to-Die Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional cache management approaches for processors with LLC on multiple memory cache dies do not account for varying die-to-die interface lengths, leading to increased LLC latencies and reduced processing performance.
Innovation Solution
Implement a non-uniform cache access (NUCA) technique by grouping LLC memory cache dies based on their die-to-die interface lengths and assigning different FIFO pointer separations to create an LLC hierarchy, prioritizing critical data to shorter-latency MCDs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional cache management approaches are used for processors with LLC on multiple memory cache dies, then the system structure is simple, but LLC latencies increase and processing performance decreases
Solution Approach 1:
The patent segments the LLC into multiple Memory Cache Dies (MCDs) organized in a hierarchical structure with different levels. Each MCD is assigned to specific processor cores based on their physical proximity, creating segmented cache regions that reduce access latency for nearby cores while maintaining overall cache capacity distribution.
Solution Approach 2:
The patent implements local quality by assigning different FIFO pointer separations to different MCDs based on their die-to-die interface lengths. MCDs with shorter interface lengths receive smaller pointer separations, optimizing access speed for locally proximate memory regions while adapting to varying physical characteristics of different cache die locations.
2Quantity of substance
If memory cache dies are distributed across the processor, then cache capacity is increased, but die-to-die interface lengths vary causing increased latencies
Solution Approach 1:
The patent introduces a hierarchical dimension to the cache structure by organizing MCDs into multiple levels (LLC hierarchy). This vertical layering allows the system to accommodate increased cache capacity across multiple dies while managing access latency through level-based prioritization, where L0/L1 levels provide faster access for critical data and lower levels provide capacity for less frequently accessed data.
Solution Approach 2:
The patent dynamically adjusts the FIFO pointer separation parameter for each MCD based on its die-to-die interface length. By changing this timing parameter according to physical distance, the system optimizes access latency for each memory cache die while maintaining overall system capacity, effectively compensating for varying physical characteristics through parameter adaptation.
Data Source
AI summary
An accelerated processor includes a processor core die including a plurality of compute units, the plurality of compute units including a first level (L1) cache. The accelerated processor also includes a plurality of memory cache dies coupled to the processor core die, the plurality of memory cache dies including a last level cache (LLC) such as a level 3 (L3) cache. The accelerated processor includes an LLC controller to issue a cache access request to the LLC and, based on a latency of the cache access request, direct the cache access request to a subset of the plurality of memory cache dies.


