Memory Rule Circuitry Configuration for Sub-NUMA Clusters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In systems with multiple cache devices, maintaining cache coherency across different cache slices is challenging, particularly in Sub-NUMA clustering (SNC) mode, where efficient access to LLC and memory devices is required while minimizing latency and unnecessary memory rule circuitry.

Innovation Solution

The implementation of a memory cluster with memory decoder circuitry that is only associated with memory devices accessible to the memory cluster, along with firmware that enables or disables programming per-cluster, allows for efficient multi-cast programming of memory access rules. This approach reduces the amount of memory rule circuitry needed and minimizes latency by ensuring that only relevant memory rules are programmed for each cluster.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If memory decoders are programmed to access all memory devices for all clusters, then cache coherency is maintained across all memory devices, but the amount of memory rule circuitry increases and area is occupied

Engineering Contradiction:
Improvecache coherencyVSAvoidCHA area
Core Design Contradiction:
ReliabilityVSArea of stationary object

Solution Approach 1:

The system divides memory access rules into cluster-specific segments. Each CHA slice is configured with memory decoders that only program rules for memory devices accessible to its specific cluster, rather than all memory devices. This segmentation reduces the amount of memory rule circuitry needed in each CHA slice while maintaining cache coherency within each cluster's memory domain.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements local quality by making each CHA slice's memory decoder programming specific to its cluster's memory devices. Instead of uniform programming across all CHAs, each cluster receives tailored memory access rules localized to its memory domain, reducing unnecessary circuitry in regions that don't need access to certain memory devices.

Inventive Principle:
Principle #3Local quality

2Productivity

If multi-cast programming is used to program all CHA slices, then memory access rules are programmed across all clusters, but unnecessary rules are programmed for clusters that don't need them

Engineering Contradiction:
Improveprogramming efficiencyVSAvoidmemory rule circuitry
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The multi-cast programming process is segmented by cluster. The system programs memory decoders in a cluster-specific manner, where each CHA slice receives programming only for its relevant memory devices. This prevents unnecessary memory rule circuitry from being programmed in CHA slices that don't need access to certain memory devices, reducing overall device complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by programming only the necessary subset of memory rules for each cluster rather than programming all possible rules. Each CHA slice receives partial programming limited to its cluster's memory devices, avoiding the excessive action of programming all memory devices for all clusters.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If CHA slices are configured to access memory devices outside their cluster, then memory access flexibility increases, but latency increases due to cross-cluster accesses

Engineering Contradiction:
Improvememory access flexibilityVSAvoidmemory access latency
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system implements local quality by configuring each CHA slice to access only memory devices within its own cluster's memory domain. This localization ensures that memory accesses remain within the same NUMA domain, minimizing latency by avoiding cross-cluster accesses. The memory decoder programming is tailored to each cluster's specific memory devices, providing optimal local access performance.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12253947B2Technologies for configuration of memory ranges
Publication Date: 2025.03.18 INTEL CORP
  • US12253947B2 patent drawing
  • US12253947B2 patent drawing
  • US12253947B2 patent drawing

AI summary

Examples described herein relate to programming a memory rule for a home agent, wherein the programming a memory rule for a home agent comprises: receiving at least one memory rule programming and based on a cluster associated with the home agent, configuring a memory rule register using a memory rule programming from among the at least one memory rule programming. In some examples, receiving at least one memory rule programming includes receiving a first memory rule programming and receiving a second memory rule programming. In some examples, a mask is applied to reject the first memory rule programming; and applying the mask to accept the second memory rule programming and program the memory rule for the home agent.