Final Level Cache Controller for HBM Memory Pooling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern data center servers face inefficiencies due to underutilization of DRAM and CPU resources, with a significant portion of DRAM going unused and CPU cores underutilized, leading to memory and CPU bandwidth limitations that hinder processing capabilities, especially in applications like artificial intelligence.
Innovation Solution
A data storage and access system utilizing a final level cache (FLC) architecture that includes multiple cache modules with shared memory pooling, where each CPU socket requires minimal bandwidth, utilizing low-power DDR memory and a switch-accessible memory pool to efficiently share resources across multiple CPUs, reducing cache miss rates and increasing memory access speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If DRAM main memory size is increased to meet memory demanding applications, then memory capacity is improved, but memory cost and resource utilization deteriorate
Solution Approach 1:
The patent merges DRAM main memory resources across multiple CPU sockets into a shared memory pool, allowing memory to be dynamically allocated and shared among different applications and processors. This consolidation enables higher memory utilization by allowing memory demanding applications to access pooled memory resources while reducing total memory requirements through efficient sharing.
Solution Approach 2:
The shared memory pool serves multiple CPU sockets and various applications simultaneously, making the memory resource universal rather than dedicated to a single processor or application. This multi-functional memory pool can dynamically serve different computational workloads, improving overall system efficiency.
2Productivity
If CPU socket bandwidth capacity is increased to support more cores, then processing capacity is improved, but CPU socket complexity and cost deteriorate
Solution Approach 1:
Multiple CPU sockets are merged into a unified system that shares a common memory pool through the fabric interconnect. This consolidation allows the system to achieve high processing capacity through multiple cores while reducing the bandwidth infrastructure complexity by sharing memory resources rather than providing dedicated high-bandwidth paths to each socket.
3Adaptability or versatility
If memory pooling bandwidth is increased to support more CPU sockets, then resource sharing capability is improved, but memory bandwidth requirements and cost deteriorate
Solution Approach 1:
The fabric interconnect acts as an intermediary between CPU sockets and the shared memory pool, enabling efficient memory access without requiring excessive bandwidth between each CPU socket and the memory pool. The fabric manages memory access requests from multiple CPUs, providing adaptability and versatility in resource sharing while controlling bandwidth requirements through intelligent routing and resource management.
Data Source
AI summary
A memory system, operating under the HBM standard, comprising a memory stack having layers of memory dies, on a base die. The base die is in communication with the memory stack and further comprises final level cache (FLC) controller. The FLC controller configured to receive the data request for requested data from a requesting element and process the data request to determine if the requested data is stored in the memory stack. Responsive to the requested data being stored in the memory stack, retrieve the requested data from the memory stack, transmit the requested data to the processor, and update a recently used tag associated with the requested data. Responsive to the requested data not being stored in the memory stack, the final level cache controller retrieves the requested data from an external memory, transmits the requested data to the processor, and stores the requested data in the memory stack.


