Memory Rule Circuitry Configuration for Sub-NUMA Clusters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In systems with multiple cache devices, maintaining cache coherency across different cache slices is challenging, particularly in Sub-NUMA clustering (SNC) mode, where efficient access to LLC and memory devices is required while minimizing latency and unnecessary memory rule circuitry.
Innovation Solution
The implementation of a memory cluster with memory decoder circuitry that is only associated with memory devices accessible to the memory cluster, along with firmware that enables or disables programming per-cluster, allows for efficient multi-cast programming of memory access rules. This approach reduces the amount of memory rule circuitry needed and minimizes latency by ensuring that only relevant memory rules are programmed for each cluster.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If memory decoders are programmed to access all memory devices for all clusters, then cache coherency is maintained across all memory devices, but the amount of memory rule circuitry increases and area is occupied
Solution Approach 1:
The system divides memory access rules into cluster-specific segments. Each CHA slice is configured with memory decoders that only program rules for memory devices accessible to its specific cluster, rather than all memory devices. This segmentation reduces the amount of memory rule circuitry needed in each CHA slice while maintaining cache coherency within each cluster's memory domain.
Solution Approach 2:
The patent implements local quality by making each CHA slice's memory decoder programming specific to its cluster's memory devices. Instead of uniform programming across all CHAs, each cluster receives tailored memory access rules localized to its memory domain, reducing unnecessary circuitry in regions that don't need access to certain memory devices.
2Productivity
If multi-cast programming is used to program all CHA slices, then memory access rules are programmed across all clusters, but unnecessary rules are programmed for clusters that don't need them
Solution Approach 1:
The multi-cast programming process is segmented by cluster. The system programs memory decoders in a cluster-specific manner, where each CHA slice receives programming only for its relevant memory devices. This prevents unnecessary memory rule circuitry from being programmed in CHA slices that don't need access to certain memory devices, reducing overall device complexity.
Solution Approach 2:
The patent applies partial action by programming only the necessary subset of memory rules for each cluster rather than programming all possible rules. Each CHA slice receives partial programming limited to its cluster's memory devices, avoiding the excessive action of programming all memory devices for all clusters.
3Adaptability or versatility
If CHA slices are configured to access memory devices outside their cluster, then memory access flexibility increases, but latency increases due to cross-cluster accesses
Solution Approach 1:
The system implements local quality by configuring each CHA slice to access only memory devices within its own cluster's memory domain. This localization ensures that memory accesses remain within the same NUMA domain, minimizing latency by avoiding cross-cluster accesses. The memory decoder programming is tailored to each cluster's specific memory devices, providing optimal local access performance.
Data Source
AI summary
Examples described herein relate to programming a memory rule for a home agent, wherein the programming a memory rule for a home agent comprises: receiving at least one memory rule programming and based on a cluster associated with the home agent, configuring a memory rule register using a memory rule programming from among the at least one memory rule programming. In some examples, receiving at least one memory rule programming includes receiving a first memory rule programming and receiving a second memory rule programming. In some examples, a mask is applied to reject the first memory rule programming; and applying the mask to accept the second memory rule programming and program the memory rule for the home agent.


