Compression Status Cache for Off-Chip Memory Area Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data processing systems face challenges in supporting large amounts of attached storage without incurring the die area costs associated with storing large numbers of directly mapped on-chip compression status bits, as the size of on-chip storage for compression status bits is expensive and cannot be easily scaled with the increase in attached memory.

Innovation Solution

A compression status cache is implemented to store compression information for blocks of memory in external memory, allowing for dynamic swapping of compression status for active buffers and enabling a large amount of attached memory to be allocated as compressible memory blocks without the corresponding die area cost, with compression status bits stored off-chip and cached in the compression status cache.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If on-chip compression status RAM is used to store compression status bits, then compression status information is readily accessible to memory interface circuitry, but die area cost increases linearly with the size of attached memory

Engineering Contradiction:
Improveaccess speedVSAvoiddie area
Core Design Contradiction:
SpeedVSArea of stationary object

Solution Approach 1:

The patent extracts the compression status storage function from the on-chip compression status RAM and relocates it to off-chip attached memory. Only a small cache portion remains on-chip to store actively accessed compression status information. This separation allows the majority of compression status data to be stored externally, dramatically reducing die area while maintaining fast access through the cache mechanism.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces a compression status cache as an intermediary between the on-chip memory interface circuitry and the off-chip compression status data. This cache acts as a buffer that holds recently accessed compression status information, providing fast access without requiring large on-chip storage. The cache mediates between the speed requirements of the memory interface and the area constraints of the on-chip resources.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Volume of stationary object

If attached memory size is increased to support more data, then storage capacity increases, but the size of on-chip compression status RAM must also increase proportionally

Engineering Contradiction:
Improvestorage capacityVSAvoidon-chip storage size
Core Design Contradiction:
Volume of stationary objectVSArea of stationary object

Solution Approach 1:

The patent extracts the bulk compression status storage requirement from the on-chip domain and places it in the off-chip attached memory. This allows storage capacity to scale with attached memory without proportionally increasing on-chip resources. The on-chip compression status cache maintains only a small subset of compression status data for active memory blocks, decoupling the scaling relationship between storage capacity and on-chip area.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the compression status storage into two parts: a small on-chip cache for actively accessed compression status information and a large off-chip storage area for the complete compression status data. This segmentation allows the system to support large storage capacity while maintaining small on-chip footprint, as each segment serves a different functional purpose and scales independently.

Inventive Principle:
Principle #1Segmentation

3Area of stationary object

If compression status bits are stored off-chip in attached memory, then die area cost is reduced, but access time increases due to additional memory access requirements

Engineering Contradiction:
Improvedie areaVSAvoidaccess time
Core Design Contradiction:
Area of stationary objectVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-loading compression status information into the on-chip compression status cache before it is needed by the memory interface circuitry. The cache is proactively updated with compression status data for memory blocks that are likely to be accessed soon, based on memory access patterns. This preliminary caching minimizes access time when the compression status information is actually needed, as it is already resident in the fast on-chip cache rather than requiring a slow off-chip access.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The compression status cache serves as an intermediary that resolves the time-delay issue of off-chip storage. By maintaining a copy of recently accessed compression status information in the on-chip cache, the system eliminates the need for frequent slow off-chip accesses. The cache intermediary provides fast access to compression status data while the bulk storage resides off-chip, thus maintaining speed performance while reducing die area.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8862823B1Compression status caching
Publication Date: 2014.10.14 NVIDIA CORP
  • US8862823B1 patent drawing
  • US8862823B1 patent drawing
  • US8862823B1 patent drawing

AI summary

One embodiment of the present invention sets forth a compression status cache configured to store compression information for blocks of memory stored within an external memory. A data cache unit is configured to request, in response to a cache miss, compressed data from the external memory based on compression information stored in the compression status bit cache. The compression status for active buffers is dynamically swapped into the compression status cache as needed. Different compression formats may be specified for one or more tiles within an active buffer. One advantage of the disclosed compression status cache is that a lame amount of attached memory may be allocated as compressible memory blocks, without incurring a corresponding die area cost because a portion of the compression status stored off chip in attached memory is cached in the compression status cache.