Multicore Cache Management via Ownership-Based Coherence
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Multicore processors face challenges in managing cache access and streaming data efficiently, particularly in parallel processing environments, where existing solutions like FPGAs and ASICs are costly and power-intensive, and current cache coherence techniques require complex directory management and frequent broadcasts.
Innovation Solution
Implementing a cache management system in multicore processors with a hierarchical cache structure, where each core has a cache with status-based storage locations for different types of data, and using a tiled architecture with dynamic routing to optimize data transfer and reduce overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If FPGAs are used for customized logic circuits, then reconfigurability is improved, but cost and power consumption increase significantly
Solution Approach 1:
The system divides the processing architecture into multiple independent cores, each with its own cache and control logic. This segmentation allows each core to be optimized for specific tasks while sharing common resources like L3 cache and interconnect, reducing overall power consumption compared to a monolithic FPGA implementation while maintaining reconfigurability through software.
Solution Approach 2:
The patent uses replicated cache structures (L1, L2, L3 caches) across multiple cores, where each core has its own private cache copies. This copying approach allows parallel processing with local data storage, eliminating the need for continuous FPGA reconfiguration and reducing power consumption while maintaining adaptability through software-controlled cache management.
2Adaptability or versatility
If FPGAs are used for customized logic circuits, then reconfigurability is improved, but performance deteriorates compared to ASICs
Solution Approach 1:
The multicore architecture segments processing into dedicated cores with private caches, allowing parallel execution of multiple tasks simultaneously. This segmentation enables ASIC-like performance for each core while the overall system maintains FPGA-like reconfigurability through software control of cache policies and data placement.
Solution Approach 2:
The system implements dynamic cache management where cache allocation, data placement, and access policies are adjusted in real-time based on workload characteristics. This dynamic behavior allows the system to optimize performance for different applications without physical reconfiguration, bridging the gap between static ASIC performance and reconfigurable FPGA flexibility.
3Reliability
If traditional cache coherence techniques are used, then cache coherence is maintained, but device complexity increases due to directory management
Solution Approach 1:
The patent extracts the complex directory management functionality from the cache coherence protocol and replaces it with a simplified ownership-based model. Each cache line has a single owner core that is responsible for coherence, eliminating the need for centralized directory structures and reducing system complexity while maintaining coherence reliability.
Solution Approach 2:
The cache coherence system operates through self-service mechanisms where cache lines automatically track their own ownership state and transfer ownership through simple protocol messages. This self-service approach eliminates the need for complex external directory management, reducing device complexity while ensuring cache coherence through automated ownership tracking.
4Reliability
If traditional cache coherence techniques are used, then cache coherence is maintained, but loss of time increases due to frequent broadcasts
Solution Approach 1:
The patent removes the broadcast mechanism from the cache coherence protocol and replaces it with targeted point-to-point communications between the owning core and accessing cores. This extraction of broadcasts eliminates unnecessary network traffic and reduces latency while maintaining coherence through direct ownership-based validation.
Solution Approach 2:
The owning core acts as an intermediary that handles all coherence requests for its owned cache lines. Instead of broadcasting to all cores, requests are routed directly to the owning core which then validates and responds to accessing cores. This intermediary approach dramatically reduces communication latency while maintaining coherence reliability.
Data Source
AI summary
Managing data in a computing system comprising one or more cores includes: providing a cache in each of one or more of the cores that includes multiple storage locations; storing data of a first type of multiple types of data in a selected storage location of a first cache of a first core that is selected according to status information associated with the first cache, and updating the status information; and storing data of a second type of the multiple types of data in a storage location within a subset of fewer than all of the storage locations of the first cache and managing the status information to ensure that subsequent data of the second type received by the first core for storage in the first cache is stored in the storage location within the subset.


