Photonic Tensor Processor for Scalable DNN Acceleration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current deep neural network (DNN) hardware faces limitations in scalability and energy efficiency due to constraints on weight storage and reconfiguration, leading to suboptimal latencies and energy consumption during matrix-vector multiplication, which is a critical operation in DNN tasks.

Innovation Solution

A combined free-space and integrated opto-electronic DNN accelerator is developed, featuring a receiver array with static weighting devices that consume no power during regular operation, allowing for scalable and efficient matrix-vector multiplication with tens of attojoules of energy and tens of nanoseconds latency, enabling next-generation DNNs and other machine learning applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If weight stationary dataflow is used to achieve low latency, then the number of weights that can be stored on hardware is limited to approximately 1,000×1,000 due to scalability constraints, but this limits the ability to handle modern workloads requiring larger matrices

Engineering Contradiction:
ImprovelatencyVSAvoidmatrix size capacity
Core Design Contradiction:
Loss of timeVSAdaptability or versatility

Solution Approach 1:

The patent transitions from a two-dimensional weight matrix storage approach to a three-dimensional architecture by adding the time dimension through temporal multiplexing._weights are stored in a 2D array and accessed across multiple time steps to handle larger effective matrix sizes, allowing the system to process N×N matrices where N exceeds the physical hardware dimensions by utilizing time-multiplexed weight stationary dataflow

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The weight matrix is divided into multiple blocks that are processed sequentially across different time steps. Each block can be stored in the limited hardware memory, and by segmenting the overall computation into temporal segments, the system handles larger matrices than the physical weight storage capacity would otherwise allow

Inventive Principle:
Principle #1Segmentation

2Productivity

If free-space optical matrix multiplication accelerators with fixed weighting masks are used, then matrix-vector multiplication can be performed, but the system cannot be reconfigured after model updates

Engineering Contradiction:
Improvematrix-vector multiplication speedVSAvoidreconfigurability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic reconfigurability by replacing fixed weighting masks with programmable phase modulators that can be updated electronically. The weighting masks are no longer static physical structures but dynamic optical elements controlled by electronic signals, allowing the system to adapt to different neural network models while maintaining high-speed optical computation

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

An intermediary electronic control layer is introduced between the optical computation engine and the weight storage. This intermediary consists of phase modulators that translate electronic weight values into optical phase adjustments, enabling reconfiguration of the optical matrix multiplication without changing the physical optical hardware

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If RC time constraints in electronic circuits are used to store weights, then weight stationary dataflow can be implemented, but scalability is limited to approximately 1,000×1,000 weights

Engineering Contradiction:
Improveweight stationary operationVSAvoidweight matrix size
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent substitutes electronic weight storage and access mechanisms with optical weight storage using spatial light modulators. Instead of relying on electronic RC time constants that limit speed and scale, the system uses optical fields to represent and manipulate weights, eliminating the RC time bottleneck and enabling scaling to much larger matrix dimensions while maintaining weight stationary operation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Power

If optical components with control constraints are used to implement weight stationary arrays, then some level of computation can be performed, but scalability and programmability are severely limited

Engineering Contradiction:
Improvecomputation capabilityVSAvoidprogrammability and scalability
Core Design Contradiction:
PowerVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal optical computing platform where a single optical engine can perform different matrix multiplication operations by loading different weight matrices through programmable phase modulators. The system is not dedicated to a single function but can be reconfigured to handle various neural network architectures and sizes, achieving universality across different computational tasks

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution achieves orders of magnitude better energy and latency performance compared to current state-of-the-art technologies, enabling next-generation DNNs and impacting other fields like Ising machines and complex optimization tasks.

Implementation Method 1

Each photodetector in each array of photodetectors is configured to emit a photocurrent in response to detecting light representing a corresponding element of an input vector

Methodology Applied
Scientific EffectPhotoelectric effect: Photoelectric Effect

Implementation Method 2

Each static weighting device in each array of static weighting devices is operably coupled to a corresponding photodetector in the array of photodetectors and configured attenuate the photocurrent emitted by the corresponding photodetector by an amount proportional to a corresponding element of a weight matrix

Methodology Applied
Scientific EffectOptical absorption: Absorption (EM radiation)

Implementation Method 3

Each modulator in each array of modulators is operably coupled to a corresponding wire in the corresponding array of wires and configured to modulate an amplitude of a corresponding wavelength-division multiplexed (WDM) beam of light in proportion to the sum of the weight photocurrents from the corresponding wire

Methodology Applied
Scientific EffectElectro-optic effect: Electro-Optic Effects

Implementation Method 4

the broadband photodetector is in optical communication with the optical bus and configured to incoherently sum the WDM beams of light

Methodology Applied
Scientific EffectIncoherent summation: Interference

Data Source

PatentUS11546077B2Scalable, ultra-low-latency photonic tensor processor
Publication Date: 2023.01.03 MASSACHUSETTS INST OF TECH
  • US11546077B2 patent drawing
  • US11546077B2 patent drawing
  • US11546077B2 patent drawing

AI summary

Deep neural networks (DNNs) have become very popular in many areas, especially classification and prediction. However, as the number of neurons in the DNN increases to solve more complex problems, the DNN becomes limited by the latency and power consumption of existing hardware. A scalable, ultra-low latency photonic tensor processor can compute DNN layer outputs in a single shot. The processor includes free-space optics that perform passive optical copying and distribution of an input vector and integrated optoelectronics that implement passive weighting and the nonlinearity. An example of this processor classified the MNIST handwritten digit dataset (with an accuracy of 94%, which is close to the 96% ground truth accuracy). The processor can be scaled to perform near-exascale computing before hitting its fundamental throughput limit, which is set by the maximum optical bandwidth before significant loss of classification accuracy (determined experimentally).