Compute-Adaptive Clock Management for ML Accelerators

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large 2D array-based machine learning accelerators face significant compute-timing bottlenecks and pipeline stalls due to stringent timing requirements, which nullify the benefits of previous dynamic-timing techniques, especially in concurrent or parallel processing of long instructions.

Innovation Solution

The implementation of a compute-adaptive clock management system using a multi-phase multi-domain clocking scheme with a data detection and timing control circuit that dynamically selects clock phases across loosely synchronized clock domains, enabling elastic clock-chain synchronization and adaptive clock management within clusters of processing elements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a large 2D array-based accelerator is used for concurrent parallel processing, then processing throughput is improved, but compute-timing bottlenecks and pipeline stalls increase due to stringent timing requirements

Engineering Contradiction:
Improveprocessing throughputVSAvoidtiming reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent divides the 2D array of processing elements into multiple independent clock domains, where each clock domain can be managed separately. This segmentation allows timing adjustments to be made locally within each domain without affecting the entire array, thereby maintaining high throughput while reducing timing bottlenecks through localized control.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic clock phase selection within each clock domain, allowing the system to adapt timing characteristics in real-time based on computational requirements. This dynamic adjustment enables the accelerator to maintain reliable timing operation across varying workloads while preserving high processing throughput.

Inventive Principle:
Principle #15Dynamics

2Use of energy by moving object

If dynamic-timing techniques are applied to improve energy efficiency, then energy consumption is reduced, but compute-timing bottlenecks continuously trigger critical path adaptation or pipeline stalls in large 2D arrays

Engineering Contradiction:
Improveenergy efficiencyVSAvoidprocessing throughput
Core Design Contradiction:
Use of energy by moving objectVSProductivity

Solution Approach 1:

The patent applies different clock phases and timing characteristics to different local regions (clock domains) within the 2D array based on their specific computational needs. This local quality approach allows energy-efficient dynamic timing adjustment in regions where it benefits performance, while maintaining stable timing in regions where it would cause bottlenecks, thereby preserving overall throughput.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent incorporates timing monitoring and feedback mechanisms that detect compute-timing bottlenecks and pipeline stalls, then automatically adjust clock phase selection in affected clock domains. This feedback control enables the system to maintain energy efficiency by applying dynamic timing adjustments only when and where they improve performance without causing bottlenecks.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If multiple clock domains with loose synchronization are used, then clock-chain synchronization flexibility is improved, but system complexity increases

Engineering Contradiction:
Improveclock management flexibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal clock phase selection mechanism that can be applied across all clock domains using the same basic approach. Each clock domain uses identical phase selection logic and control structures, allowing the system to achieve high synchronization flexibility through a standardized, multi-functional framework rather than requiring complex domain-specific solutions for each clock domain.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11520371B2Compute-adaptive clock management for machine learning accelerators
Publication Date: 2022.12.06 NORTHWESTERN UNIV
  • US11520371B2 patent drawing
  • US11520371B2 patent drawing
  • US11520371B2 patent drawing

AI summary

A system for clock management in an m columns×n rows array-based accelerators. Each row of the array may include a clock domain that clocks runtime clock cycles for the m processing elements. The clock domain includes a data detection and timing control circuit which is coupled to a common clock phase bus which provides a local clock source in multiple selectable phases, wherein the data detection and timing control circuit is configured to select a clock phase to clock a next clock cycle for a next concurrent data processing by the m processing elements. Each of m processing elements is coupled to access data from a first memory and a second memory and to generate respective outputs from each of the m processing elements to a corresponding m processing element of a same column in a subsequent neighboring row for the next processing in the next clock cycle.