NOC-Out Tree Interconnect for Low-Latency Many-Core Scalability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Many-core processors face performance limitations due to Network-On-Chip (NOC)-induced delays in server workloads, particularly because of the high area expense and latency associated with mesh interconnects, which hinder seamless scalability as core counts increase.
Innovation Solution
The NOC-Out interconnect fabric uses tree-based topologies for bilateral core-to-cache communication, reducing the need for direct core-to-core connectivity and leveraging the communication pattern in server workloads to minimize interconnect delays and area costs, with a flattened butterfly network connecting cache banks for low-latency and low-area overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a mesh interconnect is used to connect cores and cache, then the system can support many-core scalability, but the network area and communication latency increase significantly
Solution Approach 1:
The interconnect fabric is segmented into a request tree and a reply tree, with each tree handling one direction of communication. This segmentation allows independent optimization of each tree's routing paths and reduces overall network complexity and area requirements compared to a full mesh interconnect.
Solution Approach 2:
The patent introduces cache tiles as intermediary nodes between cores and the centralized cache. These cache tiles receive requests from multiple cores via the request tree, forward them to the cache, and return replies via the reply tree, eliminating the need for direct core-to-core connectivity and reducing network area.
2Adaptability or versatility
If a mesh interconnect is used to connect cores and cache, then the system can support many-core scalability, but the communication latency increases due to NOC-induced delays
Solution Approach 1:
The patent transitions from a two-dimensional mesh topology to a tree-based hierarchical structure with separate request and reply dimensions. This dimensional change allows optimized routing paths for each communication direction, reducing the number of hops and latency compared to mesh interconnect.
Solution Approach 2:
The request tree is designed with pre-established routing paths from cores to cache, and the reply tree has pre-established return paths. This preliminary structuring of communication paths eliminates the need for dynamic routing decisions at runtime, reducing communication latency.
3Loss of time
If low-diameter topologies are used to overcome mesh performance limitations, then communication latency is reduced, but the area expense increases significantly
Solution Approach 1:
The patent merges the functionality of multiple mesh interconnect components into a unified tree-based structure. By combining request forwarding and reply returning into hierarchical trees with shared infrastructure, the design achieves low-diameter communication paths without the area expense of fully connected low-diameter topologies.
Solution Approach 2:
The interconnect fabric is optimized with local quality by creating dedicated request trees and reply trees with optimized routing for their specific purposes. Each tree is locally optimized for its communication direction, achieving low latency without requiring global low-diameter connectivity across all nodes.
Data Source
AI summary
A NOC comprises a die having a cache and a core area, a plurality of core tiles arranged in the core area in a plurality of subsets, at least one cache memory bank arranged in the cache area, whereby the at least one cache memory bank is distinct from each of the plurality of core files. The NOC further comprises an interconnect fabric comprising a request tree to connect to a first cache memory bank of the at least one cache memory bank, each core tile of a first one of the subsets, the first subset corresponding to the first cache memory bank, such that each core tile is connected to the first cache memory bank only, and a reply tree to connect the first cache memory bank to each core tile of the first subset.


