Fault Tolerant Time Synchronization in Multi-Processor Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multiprocessor systems face challenges in maintaining time synchronization across multiple processing nodes due to varying conductor lengths and paths, leading to timing delays and the inability to combine separate processing blocks into a single symmetric multi-processor system, with a lack of reconfigurable distribution networks and redundancy to handle failures and skew in timing signals.

Innovation Solution

A method for distributing a synchronizing time-of-day signal using two oscillators, where a designated master chip generates and transmits an immediate time signal to neighboring slave chips, which internally delay it to compensate for propagation delays, ensuring redundancy and error detection and recovery mechanisms to maintain synchronization even in failure scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a centralized clock or oscillator is used for time synchronization across all processors, then time synchronization is achieved, but a single point of failure is created that can idle the entire multiprocessor system

Engineering Contradiction:
Improvetime synchronization reliabilityVSAvoidcentralized clock distribution network
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system divides the centralized clock function into distributed oscillators located at different nodes (e.g., drawer masters and slave chips). Each node has its own local oscillator, eliminating the single point of failure in a centralized clock while maintaining synchronization through inter-node timing signals.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically selects which oscillator serves as the master based on operational status and configuration. The master drawer master TOD chip can be reassigned if it fails, and the system can adapt to different topologies (e.g., changing which drawer is the master drawer), providing flexibility and fault tolerance.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If separate timing clocks are used for disparate collections of nodes to avoid centralized clock overhead, then scalability is improved, but timing skew between nodes increases making synchronization difficult

Engineering Contradiction:
Improvesystem scalabilityVSAvoidtime synchronization precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system pre-calculates and programs delay values into delay elements at each slave chip based on the expected propagation delay from the master. This preliminary adjustment compensates for varying conductor lengths and paths, ensuring that timing signals arrive synchronously across all nodes despite physical distribution differences.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system monitors the validity of timing signals from the master drawer master TOD chip and can detect when the master fails or when timing skew exceeds thresholds. This feedback mechanism allows the system to reassign master roles or adjust configurations to maintain synchronization precision.

Inventive Principle:
Principle #23Feedback

3Productivity

If the master chip transmits immediate time signals to all slave chips, then time synchronization is maintained, but propagation delays cause timing skew at different nodes

Engineering Contradiction:
Improvetime synchronization speedVSAvoidtiming accuracy at slave chips
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

Delay elements at each slave chip are pre-programmed with specific delay values that compensate for the propagation time from the master. This preliminary adjustment ensures that even though signals travel different distances, all slave chips receive and process the timing signal at the correct synchronized moment.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7487377B2Method and apparatus for fault tolerant time synchronization mechanism in a scaleable multi-processor computer
Publication Date: 2009.02.03 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US7487377B2 patent drawing
  • US7487377B2 patent drawing
  • US7487377B2 patent drawing

AI summary

Redundant time-of-day (TOD) oscillators are aligned, within a master oscillator path, to local logic oscillator and used to create independent step-sync signals. A step checker validates and provides selection signals to identify which of the TOD oscillators operates according to a criterion. Independent step-sync signals are transmitted to several sibling chips. Local step and sync signals are delayed to arrive at TOD register nearly synchronous with TOD registers in sibling chips. A slave oscillator path may be used to select time signals generated in a sibling chip, whereby the master oscillator path is deselected. A primary control register set may be used to configure which among several chips is a master chip using the master oscillator path. All remaining chips are slave chips. All segments of the topology are redundant. One of multiple possible alternate topologies is defined in a secondary control register set. Commands and TOD values are passed on the fabric at predefined time increment boundaries to establish, restore, or maintain synchronization across all chips.