Spiral Cache Systolic Move-to-Front Reorganization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current cache memory systems face a trade-off between access time and the number of frequently accessed values, with larger L1 caches increasing access time and smaller caches increasing the number of misses, leading to higher latency in higher-order cache levels, and existing solutions either require complex routing circuits or impose fixed worst-case access latencies.
Innovation Solution
A spiral cache architecture that dynamically reorganizes values using a move-to-front heuristic, placing most-recently accessed values at the center and less frequently accessed values towards the periphery, allowing for multiple outstanding requests without routing or switching delays, and reducing worst-case access latency by exploiting the dimensionality of Euclidean space.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If the L1 cache size is increased to store more frequently accessed values, then the number of frequently accessed values available at shortest access time is improved, but the access time increases due to physical wiring constraints and signal propagation speed limits
Solution Approach 1:
The cache is divided into multiple independently accessible banks or modules, each with its own access path. This segmentation allows parallel access to different cache locations, effectively increasing the number of simultaneously accessible values without proportionally increasing the access time for each individual value.
Solution Approach 2:
The patent introduces a spatial dimension to the cache organization by arranging cache banks in a two-dimensional grid or array structure. This allows access to proceed in multiple directions (row-wise, column-wise, or diagonal), reducing the effective access path length and enabling more values to be accessed in parallel without linearly increasing access time.
2Loss of time
If the L1 cache size is reduced to decrease access time, then the access time is improved, but the number of frequently accessed values that are not stored in the L1 cache increases, leading to more misses in higher-order cache levels
Solution Approach 1:
By segmenting the cache into multiple smaller banks, the system achieves a larger total capacity while maintaining short access times for each bank. The segmented structure allows the L1 cache to hold more frequently accessed values across multiple banks without increasing the access time beyond what would be required for a single small cache.
3Quantity of substance
If traditional cache control algorithms are used to maintain most frequently accessed values in lower-order caches, then the frequently accessed values are kept in L1 cache, but complex routing circuits and LRU logic are required
Solution Approach 1:
The cache system employs self-organizing mechanisms where cache lines automatically migrate between banks or levels based on their access patterns without requiring complex external control logic. This self-service approach reduces the need for sophisticated routing circuits and LRU management while still achieving effective caching of frequently accessed values.
4Productivity
If multiple requests are pipelined to improve throughput, then productivity is improved, but fixed worst-case access latencies and buffering are required to control the flow of pipelined information
Solution Approach 1:
The patent utilizes a two-dimensional cache organization that enables multiple request pipelines to operate simultaneously in different spatial directions. This dimensional approach allows the system to handle multiple outstanding requests without requiring complex buffering and flow control mechanisms, as the spatial separation of request paths naturally manages the pipelined information flow.
Data Source
Figure 1A~1C
Figure 2
Figure 3
AI summary
A tiled storage array provides reduction in access latency for frequently-accessed values by re-organizing to always move a requested value to a front-most storage element of array. The previous occupant of the front-most location is moved backward according to a systolic pulse, and the new occupant is moved forward according to the systolic pulse, preserving the uniqueness of the stored values within the array, and providing for multiple in-flight access requests within the array. The placement heuristic that moves the values according to the systolic pulse can be implemented by control logic within identical tiles, so that the placement heuristic moves the values according to the position of the tiles within the array. The movement of the values can be performed via only next-neighbor connections of adjacent tiles within the array.