DIMC Chiplet Architecture for Low-Latency AI Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Network Interface Cards (NICs) face challenges in scaling across distributed AI accelerators for applications like NLP and computer vision, with RDMA over Converged Ethernet and InfiniBand solutions being complex and limited by PCIe fabric topologies, hindering low-latency performance in Generative AI inferences.
Innovation Solution
A digital in-memory compute (DIMC) accelerator system using a chiplet architecture with modular chiplets, each comprising tiles and slices, optimized for high throughput computations and memory bandwidth, dynamically switching precision levels, and integrating computational functions and memory fabric to accelerate transformer models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If RDMA over Converged Ethernet combined with InfiniBand is used to facilitate multi-node GPU accelerator communication, then communication capability is improved, but system complexity increases due to complex shared address space requirements
Solution Approach 1:
The system is divided into discrete chiplet devices that can be independently configured and deployed. Each chiplet represents a modular unit that can be scaled and organized in different topologies, simplifying the overall system architecture while maintaining communication capabilities.
Solution Approach 2:
The chiplet architecture provides a universal interface that can work across different network configurations and topologies. The standardized chiplet design allows the same building blocks to serve multiple functions and be deployed in various scenarios without requiring complex specialized configurations.
2Loss of time
If RDMA over Converged Ethernet or InfiniBand fabric solutions are used to meet low-latency demands, then communication speed is improved, but deployment complexity increases
Solution Approach 1:
By segmenting the system into standardized chiplet devices, deployment becomes more manageable. Each chiplet can be independently tested, configured, and deployed, reducing the overall deployment complexity while maintaining low-latency performance through optimized inter-chiplet communication.
Solution Approach 2:
The system allows dynamic adjustment of communication parameters and precision levels to optimize performance for different workloads. This flexibility enables the system to adapt to varying latency requirements without requiring complex reconfiguration or redeployment.
3Adaptability or versatility
If PCIe fabric topologies are used for intra-switch communication, then connectivity is improved, but scalability is limited by available PCIe lanes
Solution Approach 1:
The system transitions from a centralized PCIe fabric to a distributed chiplet-based architecture. This segmentation allows each chiplet to have direct communication paths, eliminating the bottleneck of shared PCIe lanes and enabling linear scaling as more chiplets are added to the system.
Solution Approach 2:
The architecture moves from a two-dimensional PCIe switch fabric to a multi-dimensional chiplet network topology. This allows communication paths to extend in multiple directions and layers, providing scalability beyond the constraints of traditional PCIe lane limitations.
4Measurement precision
If precision is increased to maintain computational accuracy, then computation quality is improved, but power consumption increases
Solution Approach 1:
The system dynamically adjusts precision levels based on the specific computational workload and accuracy requirements. This allows the system to use higher precision only when necessary for maintaining computational accuracy, while using lower precision for operations where it is sufficient, thereby optimizing power consumption.
Solution Approach 2:
The chiplet architecture supports configurable precision parameters that can be changed based on workload characteristics. This enables the system to adapt precision levels to match computational needs, reducing power consumption by avoiding unnecessary high-precision computations while maintaining required accuracy.
Data Source
AI summary
A digital in-memory compute (DIMC) accelerator system using a chiplet architecture. The system includes a host device is configured to compile computational workload data for a target application obtained from data gathering devices into an instruction set architecture (ISA) graph to be executed by a plurality of accelerator apparatuses. Each such accelerator includes a plurality of chiplets, each of which includes a plurality of tiles, and each such tile includes a plurality of slices, a central processing unit (CPU), and a DIMC device configured to perform high throughput computations using the ISA graph to process the computational workload. The target application can include natural language processing (NLP), autonomous reasoning/decision-making, video/image processing, cybersecurity/fraud detection, manufacturing/industrial processes, agentic artificial intelligence (AI), or smart cities/Internet of Things (IoT). And the data gathering devices can include a web-scrapers, a dataset readers, a crowdsourcing devices, a sensors, a simulation devices, an IoT network, and others.


