Miss Address Buffer Segmentation for Cache Coherence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data caching systems face inefficiencies and inflexibility due to their reliance on a fully associative miss address buffer (MAB), which limits the number of miss requests that can be accommodated, leading to CPU stalls and computational waste when the MAB fills up.

Innovation Solution

The method leverages data cache tags to track miss requests, allocating an entry in the MAB only temporarily and using a fill-pending flag to indicate active requests, allowing for quicker release of MAB resources and increased flexibility by moving miss request information to cacheline tags, thereby reducing the time spent in the MAB and preventing processor stalls.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a fully associative miss address buffer (MAB) is used to track data cache miss requests, then the system can maintain cache coherence and avoid duplicate requests, but the MAB fills up quickly limiting the number of miss requests that can be accommodated, leading to CPU stalls and computational waste

Engineering Contradiction:
Improvecache coherenceVSAvoidprocessor throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the MAB into multiple smaller sets (e.g., 4 sets instead of 1 fully associative buffer). Each set independently tracks miss requests for a portion of the address space. This segmentation increases the total capacity of the MAB system while maintaining the coherence tracking function, allowing more miss requests to be accommodated without CPU stalls.

Inventive Principle:
Principle #1Segmentation

2Productivity

If the MAB size is increased to accommodate more miss requests, then processor stalls are reduced, but the complexity and resource consumption of the buffer increases

Engineering Contradiction:
Improveprocessor throughputVSAvoidMAB structure complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

Instead of one large complex fully associative MAB, the system uses multiple smaller sets with simpler structures. Each set can be implemented with fewer tags and less complex associativity requirements, reducing the complexity per unit while increasing total capacity through parallelism.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a single-dimension fully associative structure to a multi-dimensional segmented structure. By organizing the MAB into multiple sets that can be indexed differently, the system achieves higher capacity without proportionally increasing the complexity of individual buffer elements.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Device complexity

If the MAB is made smaller to reduce complexity, then the buffer is more manageable, but miss requests are lost when the buffer fills up, causing CPU stalls

Engineering Contradiction:
ImproveMAB manageabilityVSAvoidCPU stall time
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

By dividing the MAB into multiple smaller sets, the system maintains manageability of individual buffer components while collectively providing sufficient capacity to handle many more miss requests. This prevents CPU stalls by ensuring the segmented structure as a whole can accommodate the workload without overflow.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12072803B1Systems and methods for tracking data cache miss requests with data cache tags
Publication Date: 2024.08.27 ADVANCED MICRO DEVICES INC
  • US12072803B1 patent drawing
  • US12072803B1 patent drawing
  • US12072803B1 patent drawing

AI summary

The disclosed computer-implemented method for tracking miss requests using data cache tags can include generating a data cache miss request associated with data requested in connection with a cacheline and allocating a miss address buffer entry for the miss request. Additionally, the method can include, setting a fill-pending flag associated with the cacheline in response to the data associated with the data cache miss request being absent from a first data cache, and de-allocating the miss address buffer entry. In the event that another load or store operation requests the same data associated with the cacheline while the fill-pending flag is set, the method can include monitoring for a fill response associated with the miss request until the fill response is received. Upon receipt of the fill response, the method can include re-setting the fill-pending flag associated with the cacheline.