Coherence Planes for Chip Multi-Threaded Processor Interconnects
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The multidrop bus architecture is inadequate for handling chip multithreading (CMT) due to its limited bandwidth and latency, which hampers performance and coherence maintenance among multiple processor cores.
Innovation Solution
A system comprising multiple processor cores with coherency control circuitry and coherence units that manage intranode and internode coherence, utilizing external interfaces for communication and coherence message transmission, and a coherence hub to route messages and manage source identifiers for efficient coherence activity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a multidrop bus is used to interconnect processor cores, then the system is simple to implement, but the bandwidth is limited and latency increases
Solution Approach 1:
The patent segments the interconnect system into multiple independent coherence planes (e.g., Plane 0, Plane 1, Plane 2, Plane 3) that operate in parallel. Each plane handles a subset of coherence traffic, effectively dividing the total bandwidth requirement into manageable segments. This segmentation allows the system to achieve high bandwidth without requiring a single complex high-speed bus, thus resolving the contradiction between ease of implementation and bandwidth requirements.
Solution Approach 2:
The patent introduces a new dimension to the interconnect architecture by adding multiple coherence planes that operate simultaneously. Instead of increasing the speed of a single bus (one-dimensional improvement), the system adds parallel dimensions of coherence traffic handling. This dimensional expansion allows multiple processor cores to access memory and maintain coherence without being bottlenecked by a single bus's bandwidth limitations.
2Ease of manufacture
If a multidrop bus is used to interconnect processor cores, then the system is simple to implement, but the latency is increased
Solution Approach 1:
By segmenting coherence traffic into multiple planes, the patent reduces the amount of traffic each plane must handle individually. This segmentation decreases contention and waiting time for bus access, thereby reducing latency. The simple multidrop bus topology is preserved within each plane, maintaining ease of implementation while the parallel plane structure reduces overall latency.
Solution Approach 2:
The patent enables continuous coherence operations by allowing multiple planes to operate simultaneously and independently. While one plane is handling coherence traffic, other planes can be servicing different coherence requests in parallel. This continuity eliminates idle waiting periods and reduces latency without requiring complex arbitration logic that would compromise implementation simplicity.
3Device complexity
If multiple processor cores share a multidrop bus, then the system structure is simple, but coherency cannot be maintained
Solution Approach 1:
The patent segments the coherence maintenance function into multiple independent planes, each capable of handling coherence protocols (such as MESI or MOESI) for a subset of processor cores. This segmentation allows coherency to be maintained reliably across multiple cores while keeping each plane's structure relatively simple. The segmentation isolates coherence conflicts to individual planes, preventing system-wide coherency failures.
Solution Approach 2:
The patent introduces coherence planes as intermediary layers between processor cores and the multidrop bus. These planes act as mediators that manage coherence protocols, buffer coherence traffic, and ensure proper ordering of memory operations. The intermediary planes absorb the complexity of coherency maintenance, allowing the overall system structure to remain relatively simple while achieving reliable coherency across multiple processor cores.
Data Source
AI summary
In one embodiment, a node comprises a plurality of processor cores, coherency control circuitry coupled to the plurality of processor cores, and at least one coherence unit coupled to the coherency control circuitry. Each processor core is configured to have a plurality of threads active and each processor core includes at least one first level cache. The coherency control circuitry is configured to manage intranode coherency among the plurality of processor cores. The coherency unit is configured to couple to an external interface of the node, and is configured to transmit and receive coherence messages on the external interface to maintain coherency with at least one other node having one or processor cores and a coherence unit. In another embodiment, a system comprises an interconnect and a plurality of nodes coupled to the interconnect.


