Scaled Compute Fabric for Deep Learning via Flow-Based Wavelet Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current advancements in deep learning technologies face challenges in improving accuracy, performance, and energy efficiency, particularly in neural network training and inference processes.

Innovation Solution

The implementation of a scaled compute fabric with processing elements that perform flow-based computations on wavelets, utilizing floating-point units with programmable exponent bias and stochastic rounding, and data structure descriptors for efficient communication and processing across a 2D mesh network.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional deep learning architectures are used, then implementation is simpler, but accuracy and performance are limited

Engineering Contradiction:
ImproveaccuracyVSAvoidarchitecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms discrete neural network operations into continuous parameter-based computations. By representing neural network states as continuous parameters and using differential equations to model transformations, the system achieves higher accuracy while maintaining manageable complexity through mathematical abstraction rather than architectural complexity.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional mechanical/neural network architecture with a field-based mathematical model. Instead of discrete node-and-edge architectures, the invention uses continuous fields and differential equations to represent and transform data, substituting structural complexity with mathematical elegance.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If more computational resources are allocated, then performance improves, but energy consumption increases

Engineering Contradiction:
ImproveperformanceVSAvoidenergy efficiency
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent changes the computational paradigm from discrete iterative operations to continuous parameter transformations. By using differential equations and field-based representations, the system achieves higher performance per unit energy because continuous mathematical operations can be executed more efficiently than traditional discrete neural network computations.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If training time is reduced, then productivity increases, but accuracy may deteriorate

Engineering Contradiction:
Improvetraining speedVSAvoidaccuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent enables continuous computation throughout the training process rather than discrete batch processing. By using continuous differential equations to model neural network transformations, the system maintains accurate gradient information continuously, allowing faster training convergence without sacrificing accuracy because the mathematical model preserves information that would otherwise be lost in discrete approximation steps.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS11328207B2Scaled compute fabric for accelerated deep learning
Publication Date: 2022.05.10 CEREBRAS SYSTEMS INC
  • US11328207B2 patent drawing
  • US11328207B2 patent drawing
  • US11328207B2 patent drawing

AI summary

Techniques in advanced deep learning provide improvements in one or more of accuracy, performance, energy efficiency, and cost. In a first embodiment, a scaled array of processing elements is implementable with varying dimensions of the processing elements to enable varying price/performance systems. In a second embodiment, an array of clusters communicates via high-speed serial channels. The array and the channels are implemented on a Printed Circuit Board (PCB). Each cluster comprises respective processing and memory elements. Each cluster is implemented via a plurality of 3D-stacked dice, 2.5D-stacked dice, or both in a Ball Grid Array (BGA). A processing portion of the cluster is implemented via one or more Processing Element (PE) dice of the stacked dice. A memory portion of the cluster is implemented via one or more High Bandwidth Memory (HBM) dice of the stacked dice.