Dual-Neuron Optical Computing System for Large-Scale Neural Network Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing optoelectronic neural networks face challenges in optimizing complex network models and parameters, limiting their ability to handle large-scale datasets and practical applications, due to the inherent limitations of silicon-based electronic computing chips and immature optical nonlinear technologies.
Innovation Solution
A method and system for dual-neuron large-scale intelligent optical computing are introduced, which involve constructing an optoelectronic computing unit, abstracting the optical neuron as an artificial neuron, constructing an optoelectronic neural network, and performing global optimization using end-to-end gradient descent, thereby optimizing physical parameters of the optical neuron and computing the neural network based on these optimized parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If silicon-based electronic computing chips are used for neural network computation, then computing power can be achieved, but the computing power development has reached a physical limit and cannot satisfy requirements of large-scale intelligent algorithms
Solution Approach 1:
The patent replaces silicon-based electronic computing with optical computing systems that use light fields for neural network computation. The optical computing system employs spatial light modulators, Fourier optics, and photodetector arrays to perform linear operations, while electronic computing handles non-linear operations, achieving both high computing power and adaptability to large-scale algorithms.
Solution Approach 2:
The patent creates a composite optoelectronic computing system combining optical computing components (spatial light modulators, optical lenses, photodetector arrays) with electronic computing components. This hybrid system leverages the advantages of both domains: optical computing provides high-speed linear operations and parallelism, while electronic computing handles non-linear transformations, together satisfying the requirements of large-scale intelligent algorithms.
2Speed
If existing optical computing technologies are used for neural network computation, then high speed and high parallelism can be achieved, but the optical nonlinear technology is not mature and cannot implement complex non-linear operations
Solution Approach 1:
The patent uses electronic computing to perform non-linear operations that are difficult to implement in pure optical systems. The electronic computing component processes non-linear transformations of the optical signals, compensating for the immaturity of optical nonlinear technology while maintaining the high speed advantage of optical computing for linear operations.
Solution Approach 2:
The patent implements a hybrid optoelectronic architecture where optical computing handles linear operations (matrix-vector multiplications, convolutions) and electronic computing handles non-linear operations (activations, pooling). This division of labor allows the system to achieve high computing speed through optical parallelism while using mature electronic technology for non-linear transformations.
3Adaptability or versatility
If existing optoelectronic neural networks are implemented with simulation training and physical deployment, then a relationship between input-output can be modeled, but the network parameters are difficult to optimize and cannot be trained on modern large-scale datasets
Solution Approach 1:
The patent segments the optimization process into two stages: a simulation training stage where the optical computing system is modeled mathematically and parameters are optimized using standard neural network optimization tools on large-scale datasets, and a physical deployment stage where the optimized parameters are implemented in the physical optical system. This segmentation enables training on ImageNet and other large datasets while simplifying the physical implementation.
Solution Approach 2:
The patent creates a mathematical model (copy) of the optical computing system that can be simulated on conventional computers. This simulation model allows extensive training and parameter optimization using large-scale datasets before physical deployment, avoiding the need to directly optimize physical optical components while maintaining the ability to handle complex datasets.
4Device complexity
If existing optical neural networks are limited to simple datasets like MNIST and Fashion-MNIST, then the network model can be simplified, but the network performance has a large distance from practical application requirements
Solution Approach 1:
The patent separates the training process (simulation domain) from the deployment process (physical domain). During training, the system can use simple datasets to develop the optical computing architecture, then leverage the mathematical model to train on large-scale datasets like ImageNet. The simulation-physical deployment segmentation allows the system to achieve both architectural simplicity and practical performance.
Data Source
AI summary
A method and a device for dual-neuron large-scale intelligent optical computing are disclosed. The method includes: constructing an optoelectronic computing unit, and obtaining an optical neuron by modeling the optoelectronic computing unit; abstracting the optical neuron as an artificial neuron based on a mathematical abstraction manner; constructing an optoelectronic neural network by connecting a plurality of artificial neurons, and performing a global optimization on network parameters of all the artificial neurons by applying end-to-end gradient descent on the optoelectronic neural network; optimizing physical parameters of the optical neuron based on the artificial neuron after the global optimization; and computing the optoelectronic neural network based on the optimized parameters of the optical neuron.


