Electro-Photonic Network-on-Chip for Low-Power ML Data Movement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing demand for artificial intelligence computing, particularly in machine learning, is outpacing the available processing capacity, leading to inefficient energy consumption due to high latency and power inefficiencies in data movement between chips, especially in serializer/deserializer blocks and multiply-accumulate operations.
Innovation Solution
Implementing a hybrid electro-photonic network-on-chip (NoC) within circuit packages, utilizing bidirectional photonic channels for intra-chip and inter-chip communications, combined with a novel dot product engine and clocking scheme to reduce power consumption and increase processing speed by minimizing data movement and optimizing MAC operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If data is moved between chips using SerDes blocks over electrical interconnects or optical fibers, then communication between chips is enabled, but significant energy is expended in moving data within the chip to the SerDes and then from the SerDes into other chips
Solution Approach 1:
The patent introduces photonic channels as an intermediary communication medium between processing elements. Instead of converting data to serial bit streams through SerDes blocks for transmission, the system uses optical photons to carry data directly between processing elements via photonic channels, eliminating the energy-intensive SerDes conversion process while maintaining high-speed communication capability
Solution Approach 2:
The patent replaces the electrical/mechanical SerDes conversion system with a photonic transmission system. Data is transmitted as optical signals through photonic channels rather than being converted to electrical serial bit streams, substituting the mechanical/electrical conversion process with optical transmission that consumes less energy
2Productivity
If traditional hardware implementations are used for ML models, then processing capability is provided, but the implementations are relatively power-inefficient in performing multiply-accumulate operations
Solution Approach 1:
The patent segments the computational system into processing elements that are spatially distributed and interconnected via photonic channels. Each processing element can perform MAC operations locally, and the segmentation allows parallel execution of multiple MAC operations across different processing elements, improving overall productivity while reducing per-operation power consumption through localized computation
Solution Approach 2:
The patent adds a photonic dimension to the traditional electrical computation architecture. By introducing photonic channels for data transmission between processing elements, the system creates a hybrid electro-photonic architecture that enables efficient MAC operations through optical data movement, reducing the power consumption associated with electrical signal transmission
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach significantly reduces power consumption and increases processing speed by maximizing data locality and reducing energy losses, while enabling efficient execution of machine learning models like neural networks with minimal data movement and latency.
Implementation Method 1
utilizing bidirectional photonic channels for intra-chip and inter-chip communications
Implementation Method 2
bidirectional photonic channels for intra-chip and inter-chip communications
Data Source
AI summary
Various embodiments provide for electro-photonic networks, including a plurality of processing elements connected by bidirectional photonic channels, suited for implementing neural-network models. Weights of the model may be preloaded into memory of the processing elements based on assignments of neural nodes to processing elements implementing them, and routers of the processing elements can be configured to stream activations between the processing elements based on a predetermined flow of activations in the model.


