Dual Interface Coherent Memory Bandwidth Expander
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in achieving high bandwidth and maintaining memory coherency for accelerating artificial intelligence (AI) workloads and other computationally intensive memory use applications.
Innovation Solution
A dual issue coherent computational memory bandwidth expander is introduced, which includes a coherent acceleration functional unit (CAFU) that offloads CPU functions and accelerates them on a configurable integrated circuit. This solution utilizes two interfaces between the CAFU and the device coherency agent (DCOH) to double bandwidth while maintaining memory coherency rules, employing a processor control finite state machine to control data set issuance across these interfaces.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a single interface is used between CAFU and DCOH, then device complexity is reduced, but processing bandwidth is limited
Solution Approach 1:
The patent divides the data transmission path into multiple separate interfaces (first interface and second interface) between the CAFU and DCOH. Each interface handles distinct data sets independently, allowing parallel transmission streams that collectively double the processing bandwidth compared to a single interface configuration.
Solution Approach 2:
The patent transitions from a single-dimensional data path to a multi-dimensional parallel architecture by implementing multiple interfaces operating simultaneously. This dimensional expansion enables concurrent data transmission through different pathways, effectively multiplying the bandwidth capacity without proportionally increasing interface complexity.
2Productivity
If multiple interfaces are used to increase bandwidth, then processing bandwidth is doubled, but maintaining memory coherency becomes more difficult
Solution Approach 1:
The finite state machine implements coherency checking mechanisms that monitor and validate data transmissions across multiple interfaces. By continuously checking coherency qualifications and providing feedback control, the system maintains memory coherency rules even as bandwidth increases through parallel interface operations.
Solution Approach 2:
The finite state machine acts as an intermediary controller between the multiple interfaces and the DCOH, coordinating data flow and enforcing coherency rules. This mediator ensures that parallel transmissions through multiple interfaces do not compromise memory coherency by managing the interactions between different data streams.
3Speed
If CPU functions are offloaded to CAFU, then processing speed is improved, but system complexity increases
Solution Approach 1:
The patent extracts specific computational functions from the CPU and relocates them to the CAFU (coherent acceleration functional unit). This extraction accelerates computationally intensive tasks while keeping the CPU architecture relatively simple, as only the necessary functions are offloaded rather than redesigning the entire system.
Solution Approach 2:
The CAFU is designed as a universal acceleration unit that can handle multiple types of computationally intensive workloads including AI workloads and other memory-intensive applications. This multi-functionality allows a single added component to provide broad acceleration capabilities across different application domains, reducing the need for multiple specialized units.
Data Source
AI summary
An integrated circuit includes a device coherency circuit, first and second traffic generator processor circuits, first and second interfaces, and a processor control finite state machine circuit that causes the first traffic generator processor circuit to perform first coherency data validation for first data sets and that causes the second traffic generator processor circuit to perform second coherency data validation for second data sets. The first traffic generator processor circuit transmits first traffic for the first data sets to the device coherency circuit through the first interface. The second traffic generator processor circuit transmits second traffic for the second data sets to the device coherency circuit through the second interface.


