In Situ Backpropagation for Photonic Neural Network Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for training photonic neural networks are inefficient, as they often rely on external computer simulations that depend on accurate model representations and are not scalable for large systems, or involve sequential perturbation of parameters, which is highly inefficient.

Innovation Solution

The method involves calculating a loss for an input to a photonic ANN, computing an adjoint input, measuring intensities in optical interference units, computing gradients from these measurements, and tuning phase shifters within the OIUs to optimize the network, allowing for in situ training and efficient backpropagation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If training is performed using external computer simulations, then model accuracy can be maintained, but training efficiency and scalability deteriorate

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The photonic neural network performs its own training computations directly on the hardware platform using optical signals, eliminating the need for external computer simulations. The system uses in-situ measurement of optical intensities and automatic differentiation to compute gradients and update parameters, allowing the network to train itself efficiently on the target hardware while maintaining model accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces electronic computation mechanisms with optical mechanisms for training. Instead of using electronic computers to simulate and train the neural network, the system uses photonic circuits to perform matrix multiplications, activation functions, and gradient computations directly in the optical domain, achieving both accuracy and efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If brute force in situ computation is used, then training can be performed on the photonic circuit, but computational efficiency deteriorates due to sequential parameter perturbation

Engineering Contradiction:
Improvein situ training capabilityVSAvoidcomputational efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system implements automatic differentiation by measuring optical intensities and using these measurements to compute gradients through the computational graph. The measured intensities provide feedback about the loss function, which is then used to calculate gradients with respect to all parameters simultaneously, enabling efficient in-situ training without sequential perturbation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent pre-computes adjoint states by propagating loss gradients backward through the network before performing parameter updates. This preliminary computation of the adjoint input allows all parameter gradients to be calculated in parallel from measured intensities, avoiding the need for sequential parameter perturbation during the training step.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If sequential parameter perturbation is used, then gradient computation can be performed, but training time increases significantly for large systems

Engineering Contradiction:
Improvegradient computation accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary propagation of the adjoint state through the entire network before computing any parameter gradients. This pre-computation allows all gradients to be obtained simultaneously from a single set of intensity measurements, rather than sequentially perturbing each parameter, dramatically reducing training time while maintaining gradient accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent merges the computation of all parameter gradients into a single computational pass by using the adjoint method. Instead of computing gradients for each parameter separately through sequential perturbation, the system computes all gradients simultaneously by propagating the adjoint state backward and combining results from measured intensities, reducing training time for large systems.

Inventive Principle:
Principle #5Merging (Combining)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables highly efficient in situ training of photonic neural networks, simplifies the implementation of backpropagation, and allows for parallel tuning of parameters, significantly improving training speed and energy efficiency compared to existing methods.

Implementation Method 1

a hardware implementation using optical signals has also been considered. Photonic implementations benefit from the fact that, due to the non-interacting nature of photons, linear operations—like the repeated matrix multiplications found in every neural network algorithm—can be performed in parallel

Methodology Applied
Scientific EffectOptical interference: Interference

Implementation Method 2

tuning phase shifters of the OIU based on the computed gradient

Methodology Applied
Scientific EffectPhase modulation: Phase Modulation

Data Source

PatentUS12026615B2Training of photonic neural networks through in situ backpropagation
Publication Date: 2024.07.02 THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
  • US12026615B2 patent drawing
  • US12026615B2 patent drawing
  • US12026615B2 patent drawing

AI summary

Systems and methods for training photonic neural networks in accordance with embodiments of the invention are illustrated. One embodiment includes a method for training a set of one or more optical interference units (OIUs) of a photonic artificial neural network (ANN), wherein the method includes calculating a loss for an original input to the photonic ANN, computing an adjoint input based on the calculated loss, measuring intensities for a set of one or more phase shifters in the set of OIUs when the computed adjoint input and the original input are interfered with each other within the set of OIUs, computing a gradient from the measured intensities, and tuning phase shifters of the OIU based on the computed gradient.