Blade Server Clock Synthesizer Architecture for Scalable QPI Interconnects
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current parallel computing architectures face scalability limitations in data access and parallelism due to bottlenecks and latencies, particularly in blade and rack server systems, where the number of processing elements is restricted by the complexity of routing common reference clocks and trace lengths, limiting the practical scalability of QPI interconnects.
Innovation Solution
Implementing a scalable common reference-clocking architecture using clock synthesizer boards to provide synchronized reference clock signals across multiple blades or servers via differential cables, reducing signal propagation delay and enabling greater scalability in QPI and PCIe interconnects by simplifying cable routing and using signal propagation delay elements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a common reference clock is distributed to multiple blades or servers via routing traces, then clock synchronization is achieved, but the complexity of routing and trace length management increases, limiting scalability
Solution Approach 1:
The patent segments the clock distribution system by providing each blade or server with its own dedicated reference clock source instead of routing a single common clock through complex traces. Each processing element has an independent clock source that is synchronized through a simplified protocol, eliminating the need for complex routing infrastructure while maintaining synchronization reliability.
Solution Approach 2:
The patent introduces a clock synchronization protocol that acts as an intermediary mechanism between individual clock sources and the processing elements. This protocol manages the synchronization function centrally, allowing each blade or server to have its own clock source while still achieving coordinated operation through the synchronization protocol's timing adjustments.
2Productivity
If QPI interconnects are used for high-speed data access, then data access bandwidth is improved, but scalability is limited by the complexity of routing common reference clocks
Solution Approach 1:
The patent applies segmentation to decouple the high-speed data transport function from the clock distribution function. Each blade or server maintains its own reference clock source, allowing QPI interconnects to operate at full speed without being constrained by a centralized clock routing architecture. This segmentation enables linear scalability with the number of processing elements while preserving high data access bandwidth.
3Productivity
If the number of processing elements is increased, then parallel processing capability is improved, but signal propagation delay increases, limiting practical scalability
Solution Approach 1:
The patent extracts the clock generation function from the centralized system and distributes it to individual processing elements. Each blade or server generates its own reference clock signal locally, eliminating the need for long inter-blade clock distribution traces. This extraction of the clock generation function dramatically reduces signal propagation delay while enabling increased numbers of processing elements to operate in parallel.
Data Source
AI summary
Scalable, common reference-clocking architecture and method for blade and rack servers. A common reference clock source is configured to provide synchronized clock input signals to a plurality of blades in a blade server or servers in a rack server. The reference clock signals are then used for clock operations related to serial interconnect links between blades and/or servers, such as QuickPath Interconnect (QPI) links or PCIe links. The serial interconnect links may be routed via electrical or optical cables between blades or servers. The common reference clock input and inter-blade or inter-server interconnect scheme is scalable, such that the plurality of blades or servers can be linked together in communication. Moreover, when QPI links are used, coherent memory transactions across blades or servers are provided, enabling fine grained parallelism to be used for parallel processing applications.


