Coherent Link Layer Flit Multiplexing for Bandwidth Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current interconnect architectures, such as PCIe and QPI, face challenges in achieving high performance while minimizing power consumption and coherence overhead, particularly in supporting coherent memory access and efficient data transfer across accelerators and processors.
Innovation Solution
The proposed solution extends the Intel Accelerator Link (IAL) architecture by using a combination of IAL.io, IAL.cache, and IAL.mem protocols to implement a Coherence Bias Model, which facilitates high-performance accelerators with reduced coherence overhead, and employs flit-based data transfer with sub-flit multiplexing, all-data flits, and multi-data headers to optimize link efficiency and bandwidth.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional interconnect architectures (PCIe, QPI) are used for coherent memory access, then device compatibility and standardization are improved, but coherence overhead and power consumption increase
Solution Approach 1:
The patent segments the interconnect protocol into multiple specialized protocols (IAL.io, IAL.cache, IAL.mem) that handle different types of traffic separately. This segmentation allows coherent memory access to use optimized protocols with reduced overhead while maintaining compatibility through standardized interfaces.
Solution Approach 2:
The patent introduces an intermediary layer (the IAL architecture) between the accelerator and main memory that handles coherence management. This intermediary reduces the coherence overhead on the main interconnect by localizing coherence operations and using efficient flit-based transfer mechanisms.
2Adaptability or versatility
If PCIe protocol is used for data transfer, then device interoperability is improved, but transfer efficiency and bandwidth utilization deteriorate
Solution Approach 1:
The patent changes key parameters of the data transfer protocol by implementing flit-based transfer with sub-flit multiplexing, all-data flits, and multi-data headers. These parameter changes enable higher transfer efficiency and bandwidth utilization while maintaining device interoperability through standardized protocol interfaces.
Solution Approach 2:
The patent introduces dynamic protocol selection where the system can switch between different IAL protocols (IAL.io, IAL.cache, IAL.mem) based on the type of data transfer required. This dynamic approach optimizes transfer efficiency for different workloads while maintaining broad device interoperability.
3Adaptability or versatility
If coherent memory access is implemented across accelerators, then system integration is improved, but coherence overhead and complexity increase
Solution Approach 1:
The patent segments coherence management into specialized protocols (IAL.cache, IAL.mem) that handle different coherence operations. This segmentation reduces the complexity of implementing coherent memory access across accelerators by providing dedicated, optimized pathways for different types of coherence traffic.
Solution Approach 2:
The patent enables accelerators to self-manage their coherence operations through the IAL protocol stack, reducing the burden on the main system architecture. The flit-based transfer mechanism allows accelerators to autonomously handle data transfer and coherence maintenance, simplifying overall system integration.
4Productivity
If high bandwidth data transfer is achieved, then link utilization is improved, but latency may increase
Solution Approach 1:
The patent uses periodic action through its credit-based flow control mechanism that operates in regular cycles. Credits are granted and consumed periodically, allowing the system to maintain high link utilization while managing latency through predictable, periodic resource allocation and data transfer opportunities.
Data Source
AI summary
Systems, methods, and devices can include link layer logic that is to identify, by a link layer device, first data received from the memory in a first protocol format, identify, by the link layer device, second data received from the cache in a second protocol format, multiplex, by the link layer device, a portion of the first data and a portion of the second data to produce multiplexed data; and generate, by the link layer device, a flow control unit (flit) that includes the multiplexed data.


