Blade Server Clock Synthesizer Architecture for Scalable QPI Interconnects

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current parallel computing architectures face scalability limitations in data access and parallelism due to bottlenecks and latencies, particularly in blade and rack server systems, where the number of processing elements is restricted by the complexity of routing common reference clocks and trace lengths, limiting the practical scalability of QPI interconnects.

Innovation Solution

Implementing a scalable common reference-clocking architecture using clock synthesizer boards to provide synchronized reference clock signals across multiple blades or servers via differential cables, reducing signal propagation delay and enabling greater scalability in QPI and PCIe interconnects by simplifying cable routing and using signal propagation delay elements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a common reference clock is distributed to multiple blades or servers via routing traces, then clock synchronization is achieved, but the complexity of routing and trace length management increases, limiting scalability

Engineering Contradiction:
Improveclock synchronizationVSAvoidrouting complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the clock distribution system by providing each blade or server with its own dedicated reference clock source instead of routing a single common clock through complex traces. Each processing element has an independent clock source that is synchronized through a simplified protocol, eliminating the need for complex routing infrastructure while maintaining synchronization reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a clock synchronization protocol that acts as an intermediary mechanism between individual clock sources and the processing elements. This protocol manages the synchronization function centrally, allowing each blade or server to have its own clock source while still achieving coordinated operation through the synchronization protocol's timing adjustments.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If QPI interconnects are used for high-speed data access, then data access bandwidth is improved, but scalability is limited by the complexity of routing common reference clocks

Engineering Contradiction:
Improvedata access bandwidthVSAvoidscalability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent applies segmentation to decouple the high-speed data transport function from the clock distribution function. Each blade or server maintains its own reference clock source, allowing QPI interconnects to operate at full speed without being constrained by a centralized clock routing architecture. This segmentation enables linear scalability with the number of processing elements while preserving high data access bandwidth.

Inventive Principle:
Principle #1Segmentation

3Productivity

If the number of processing elements is increased, then parallel processing capability is improved, but signal propagation delay increases, limiting practical scalability

Engineering Contradiction:
Improveparallel processing capabilityVSAvoidsignal propagation delay
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent extracts the clock generation function from the centralized system and distributes it to individual processing elements. Each blade or server generates its own reference clock signal locally, eliminating the need for long inter-blade clock distribution traces. This extraction of the clock generation function dramatically reduces signal propagation delay while enabling increased numbers of processing elements to operate in parallel.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS9261897B2Scalable, common reference-clocking architecture using a separate, single clock source for blade and rack servers
Publication Date: 2016.02.16 INTEL CORP
  • US9261897B2 patent drawing
  • US9261897B2 patent drawing
  • US9261897B2 patent drawing

AI summary

Scalable, common reference-clocking architecture and method for blade and rack servers. A common reference clock source is configured to provide synchronized clock input signals to a plurality of blades in a blade server or servers in a rack server. The reference clock signals are then used for clock operations related to serial interconnect links between blades and/or servers, such as QuickPath Interconnect (QPI) links or PCIe links. The serial interconnect links may be routed via electrical or optical cables between blades or servers. The common reference clock input and inter-blade or inter-server interconnect scheme is scalable, such that the plurality of blades or servers can be linked together in communication. Moreover, when QPI links are used, coherent memory transactions across blades or servers are provided, enabling fine grained parallelism to be used for parallel processing applications.