CXL Cache-Coherent Switch Architecture for Low-Latency Resource Sharing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges in efficiently processing large datasets due to limitations in storage and processing capabilities, with processors becoming bottlenecks in component-to-component traffic, leading to suboptimal resource utilization and increased latency.
Innovation Solution
Implementing a cache coherent switch on chip using the Compute Express Link (CXL) interconnect standard to facilitate resource sharing and caching between various components, bypassing the need for processor control, thus enabling low-latency memory access and cache coherency across accelerators, memories, and other devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If processor control is used for component-to-component traffic, then system compatibility and control are maintained, but processor burden increases and latency increases
Solution Approach 1:
The patent introduces a cache coherent switch as an intermediary device that handles component-to-component traffic independently from the processor. The switch maintains cache coherency protocols and routing functions, allowing processors to offload traffic management while the switch acts as a mediator maintaining system-wide coherence without requiring processor involvement for each transaction.
2Reliability
If processor control is used for component-to-component traffic, then system control is maintained, but latency increases due to processor bottleneck
Solution Approach 1:
The cache coherent switch serves as a dedicated intermediary that handles routing and coherency maintenance for component-to-component traffic, eliminating the need for processor involvement in each transaction. This dedicated path reduces latency by providing direct routing while the switch maintains system control through cache coherency protocols.
Solution Approach 2:
The patent segments the system into distinct functional domains: the processor handles high-level computation and memory management, while the cache coherent switch handles low-level routing and cache coherency maintenance. This segmentation allows parallel operation of control functions and data processing, reducing overall system latency.
3Reliability
If traditional memory access is used, then data availability is ensured, but access speed is limited by processor bottleneck
Solution Approach 1:
The patent implements cache memory that pre-loads and stores frequently accessed data before the processor needs it. The cache coherent switch enables background data movement and pre-positioning of data in cache locations, so when the processor needs data, it is already available in fast cache memory rather than requiring slow main memory access.
4Productivity
If resource pooling is implemented, then resource utilization improves, but system complexity increases
Solution Approach 1:
The cache coherent switch is designed as a universal intermediary that can handle multiple types of traffic (memory accesses, I/O operations, inter-processor communication) and maintain cache coherency across diverse component types. This multi-functionality allows resource pooling of accelerators, storage, and memory while the single switch architecture manages all connections, avoiding the need for separate control logic for each resource type.
Data Source
AI summary
Described herein are systems, methods, and products utilizing a cache coherent switch on chip. The cache coherent switch on chip may utilize Compute Express Link (CXL) interconnect open standard and allow for multi-host access and the sharing of resources. The cache coherent switch on chip provides for resource sharing between components while independent of a system processor, removing the system processor as a bottleneck. Cache coherent switch on chip may further allow for cache coherency between various different components. Thus, for example, memories, accelerators, and/or other components within the disclose systems may each maintain caches, and the systems and techniques described herein allow for cache coherency between the different components of the system with minimal latency.


