Optical Inference in Programmable Switches for Low-Latency AI

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI-based services face high inference latency due to the need for large Deep Neural Networks (DNNs) that are typically executed in the cloud or on edge servers, which are limited by memory, power, and computing constraints, leading to increased packet propagation and processing delays.

Innovation Solution

In-network optical inference (IOI) performs inference tasks using programmable packet switches equipped with optical computing hardware, such as the Intel Tofino 2 switch, to perform matrix multiplication and nonlinear activation functions, reducing latency and power consumption by processing data closer to the user.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If inference processing is performed on cloud or edge servers using CPU, then reliability and accuracy are maintained, but inference latency increases due to packet propagation delay and CPU processing speed limitations

Engineering Contradiction:
Improveinference accuracyVSAvoidinference latency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces the mechanical/electrical CPU processing system with an optical computing system. Optical packet switches perform matrix multiplication operations using light-based computing, substituting the traditional electrical signal processing mechanism. This enables parallel processing at optical speeds, reducing inference latency while maintaining accuracy through preserved computational functionality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transitions from sequential CPU processing to parallel optical processing. By using optical packet switches that can perform multiple matrix multiplication operations simultaneously across different spatial dimensions, the system achieves parallelism that dramatically reduces inference latency compared to sequential electrical processing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If large Deep Neural Networks are deployed in cloud servers, then inference accuracy is improved, but processing speed is limited by CPU clock frequency and packet queueing delay

Engineering Contradiction:
Improveinference accuracyVSAvoidpacket processing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces CPU-based electrical processing with optical computing for matrix multiplication operations. Optical packet switches can process packets at speeds limited only by optical signal propagation and switching capabilities, achieving billions of packets per second compared to millions per second for CPU processing, thus dramatically improving productivity while supporting large DNNs for high accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If inference processing is offloaded to cloud servers, then device computing limitations are overcome, but energy consumption increases due to data transmission and server processing

Engineering Contradiction:
Improvecomputing capabilityVSAvoidenergy consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent introduces optical packet switches as intermediary devices that perform inference processing directly at the network edge. These switches act as mediators between user devices and cloud servers, enabling computation to occur closer to the data source. This reduces the energy consumption associated with long-distance data transmission and centralized server processing while providing the computing capability needed for large DNNs.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Productivity

If programmable packet switches are used for inference processing, then processing speed increases significantly, but device complexity increases due to integration of optical computing hardware

Engineering Contradiction:
Improvepacket processing throughputVSAvoidsystem integration complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges optical computing hardware with programmable packet switch functionality into integrated devices. By combining the matrix multiplication capabilities of optical computing units with the packet switching and routing functions in a single system, the patent achieves high processing throughput while managing complexity through functional integration rather than separate components.

Inventive Principle:
Principle #5Merging (Combining)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

IOI significantly reduces end-to-end inference latency and energy consumption by leveraging optical computing in programmable switches, achieving speeds 1000 times faster than CPU cores and reducing costs while maintaining inference accuracy.

Implementation Method 1

modulating, with a first modulator, an optical pulse with a waveform proportional to the input vector element; modulating, with a second modulator in optical communication with the first modulator, the optical pulse with a waveform proportional to the weight vector

Methodology Applied
Scientific EffectOptical modulation: Electro-Optic Effects

Implementation Method 2

detecting the optical pulse with a photodetector

Methodology Applied
Scientific EffectPhotoelectric effect: Photoelectric Effect

Data Source

PatentUS12634606B2In-network optical inference
Publication Date: 2026.05.19 NTT RESEARCH INC
  • US12634606B2 patent drawing
  • US12634606B2 patent drawing
  • US12634606B2 patent drawing

AI summary

In-network Optical Inference (IOI) provides low-latency machine learning inference by leveraging programmable switches and optical matrix multiplication. IOI uses a transceiver module, called a Neuro Transceiver, with an optical processor to perform linear operations, such as matrix multiplication, in the optical domain. IOI's transceiver modules can be plugged into programmable packet switches, which are programmed to perform non-linear activations in the electronic domain and to respond to inference queries. Processing inference queries at the programmable packet switches inside the network, without sending them to cloud or edge inference servers, significantly reduces end-to-end inference latency experienced by users.