VCSEL-based Coherent and Scalable Efficient Deep Learning

The three-dimensional optical neural network architecture addresses the limitations of CMOS-based accelerators by utilizing laser transmitters, coherent detection, and holographic data movement, achieving high computational density and energy efficiency.

JP2025517719APending Publication Date: 2025-06-10MASSACHUSETTS INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024567603
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-13
Filing Date
2023-05-15
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Current CMOS-based neural network accelerators face challenges in achieving high computational density, low energy consumption, and efficient non-linear activation due to limitations in wire capacitance and data movement bottlenecks.

Method used

A three-dimensional optical neural network architecture utilizing coherent on-chip microscale laser transmitters, coherent detection for weighted accumulation, and holographic data movement to achieve high computational density, low energy consumption, and efficient non-linear activation.

Benefits of technology

The proposed architecture achieves the highest computing density, energy efficiency better than 1 fJ/MAC, and negligible energy consumption for non-linear activation, significantly surpassing the limitations of traditional CMOS-based systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517719000001_ABST
    Figure 2025517719000001_ABST
Patent Text Reader

Abstract

The exponential growth in deep learning models is difficult for existing computing hardware. Optical neural networks (ONNs) have the potential for ultra-high bandwidth and accelerate machine learning tasks with little loss in data movement. Scaling up ONNs involves improving scalability, energy efficiency, computational density, and in-line non-linearity. However, achieving all these criteria remains an open challenge. Here, we demonstrate a three-dimensional spatio-temporal multiplexed ONN architecture based on a high-density array of microscale vertical-cavity surface-emitting lasers (VCSELs). VCSELs injection-locked coherently to a master laser operate at gigahertz data rates with a π-phase shift voltage at the 10-millivolt level. Optical non-linearity is incorporated into the ONN using coherent detection of optical interference between VCSELs without adding an energy cost. Utilizing large-scale spatial holographic data fan-out (N = 1,000), the system enables large-scale parallel operation with a computational density reaching 10 TOPS / (mm 2 ?s), which is a 100x improvement compared to state-of-the-art CMOS microprocessors. The energy consumption reaches 1 femtojoule / OPS, which is 1,000x better than its CMOS counterpart. Neurons are encoded in time steps and are freely scalable in large numbers (up to one million). As wafer-scale integration progresses in the future, the inventors' technology opens the way to low-energy, high-throughput large-scale optical computers and incorporates non-linearity into machine learning applications ranging from data centers to edge devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. Patent Application No. 63 / 341,601, filed on May 13, 2022, the entire contents of which are incorporated herein by reference.

[0002] (Government Support) This invention was made with government support under W911NF - 17 - 1 - 0527 awarded by the Army Research Office. The government has certain rights in this invention.

Background Art

[0003] An artificial neural network is a computational system that mimics the way a biological brain processes information. Such systems are constructed to learn, combine, and summarize information from large datasets. A deep neural network (DNN) is an artificial neural network that has multiple layers between its input and output. Each layer performs matrix - vector multiplication and non - linear activation. Due to both the progress of the DNN process and the improvement of computing power, DNNs have brought about a revolution in information processing in applications including image, object, and speech recognition, game play, medicine, and physical chemistry.

[0004] Motivated by the desire to address the problem of increasing complexity, the sizes of DNNs and other machine learning models have been increasing exponentially, and some have reached over 100 billion trainable parameters. In contrast, due to practical limitations in the number of transistors and energy consumption in data movement, it is becoming increasingly difficult to expand the computing capacity using complementary metal-oxide-semiconductor (CMOS) neural network accelerators. It is necessary to develop alternative approaches that utilize qualitatively different technologies to continue scaling computing power over the next few decades.

Summary of the Invention

[0005] Generally, a neural network accelerator or other computing machine that implements a DNN or other artificial neural network should meet five criteria: high computational density (C1), low energy consumption (C2), in-line non-linearity (C3), expandable hardware (C4) that can accommodate a large number of neurons (C5). State-of-the-art microprocessors such as graphics processing units (GPUs) and application-specific integrated circuits (ASICs) are limited by the wire capacitance of electronic interconnections in terms of (C1) computational density

Number

Number

[0006] An optical neural network (ONN) is very promising for potentially orders-of-magnitude improvements in reducing these data movement bottlenecks due to its large optical bandwidth and ability to move data with low loss. Recent advances in ONN have led to photonic integrated circuits, neural connections with 3D printed phase masks, matrix multiplication under lightless conditions, and high throughput by frequency multiplexing. However, it remains difficult to achieve high data density in a high-density device with low energy consumption. Also, due to the weakness of typical optical non-linearity, it can be difficult to execute non-linear activation functions that are selected depending on specific tasks performed by the neural network in the optical domain with low light intensity.

[0007] Inspired by the axon-synapse-dendrite architecture of biological neural systems, a DNN architecture is introduced that achieves the criteria C1 to C5 in a three-dimensional architecture using (i) a coherent on-chip microscale laser transmitter as high-speed (e.g., GHz speed) "laser axons", (ii) coherent detection for weighted accumulation as low-energy "laser synapses", and (iii) holographic data movement as an optical dendrite fan-out. This coherent DNN, also called VCSEL ONN, based on an array of vertical-cavity surface emitting lasers (VCSELs) as an "axon-synapse-dendrite" microscale transceiver, (C1)

Number

Table 1

[0008] An exemplary ONN may include an input layer for receiving inputs, a plurality of fully connected layers for performing inference processing on the inputs, and an output layer for returning the output of the inference processing. Each fully connected layer may include an array of VCSELs that optically communicates with a diffractive optical element (DOE) and an array of photodetectors. During operation, the first VCSEL in the array of VCSELs emits a first beam modulated with an input vector. The other (second) VCSELs in the array of VCSELs are modulated with the weights of the optical neural network and emit a second beam coherent with the first beam. The DOE fans out the first beam, and the array of photodetectors detects the interference between each fanned-out copy of the first beam and the second beam from the DOE. The array of photodetectors can generate a photocurrent proportional to, where, and are the amplitude and phase of the first beam, and, and are, and, and are

Number

Number

Number

Number

Number

[0009] The optical neural network may have a computational density of at least [Number] and / or can operate with an energy consumption of 1 fJ / OPS. The array of VCSELs can be monolithically integrated with the DOE and the array of photodetectors. The array of VCSELs can be modulated with a half-wavelength voltage of less than about 10 mV and / or at a modulation rate of at least 1 Gb / s. The array of VCSELs can also be injection-locked to a leader laser. The optical neural network may also include a second diffractive optical element that optically communicates with at least one of the second VCSELs in the array of VCSELs and may fan out the second beam emitted by the second VCSEL.

[0010] The array of photodetectors [Number] may be configured to generate an output proportional to where [Number] and [Number] are the amplitude and phase of the first beam, and [Number] and [Number] are the amplitude and phase of the j-th second beam. The array of photodetectors may also [Number] may be configured to generate an output proportional to, where

Number

Number

[0011] Such an optical neural network can perform inference processing by adjusting the phase of the first beam emitted by the first VCSEL in an array of VCSELs having activation vectors to the optical neural network. In the array of VCSELs, the phase of the second beam emitted by the second VCSEL can be adjusted by each element of the weight matrix of the optical neural network. A diffractive optical element, or other optical component, fans out a copy of the first beam to each homodyne receiver, which detects the homodyne interference of the first beam and the copy of each second beam.

[0012] All combinations of the foregoing concepts and other concepts described in more detail below (provided such concepts are not mutually inconsistent) are part of the subject matter of the invention disclosed herein. In particular, all combinations of the subject matter recited in the claims at the end of this disclosure are part of the subject matter of the invention disclosed herein. Terms used herein that may also appear in any disclosure incorporated herein by reference should be given the meaning that most closely matches the particular concepts disclosed herein.

Brief Description of the Drawings

[0013] Those skilled in the art will understand that the drawings are presented mainly for illustrative purposes and are not intended to limit the scope of the subject matter of the invention described herein. The drawings are not necessarily to scale and, in some cases, various aspects of the subject matter of the invention disclosed herein may be shown exaggerated or enlarged in the drawings to facilitate understanding of the various features. In the drawings, like reference characters generally refer to like features (e.g., elements that are functionally and / or structurally similar).

[0014]

Fig. 1A

Fig. 1B

Fig. 1C

Fig. 1D

Fig. 2A-2B

Fig. 2C

Fig. 2D

Fig. 2E

Fig. 2F

Fig. 2G

Fig. 2H

Fig. 3A

Fig. 3B

Fig. 3C

Fig. 3D

Fig. 3E

Fig. 4A-4B

Fig. 4C

Fig. 4D

Fig. 4E

Fig. 5

Mode for Carrying Out the Invention

[0015] (Optical Neural Network (ONN) with Photonic Tensor Core) Figures 1A - 1D show the ONN architecture 100. As shown in Figure 1A, this includes an n - layer sequence (top), each of which calculates the product of an activation vector (1×k) and a weight matrix (W). Each layer can be optically implemented as a phase - encoded photonic tensor core 110 (bottom). Similar to the axon - synapse - dendrite structure in a biological neural network, in each photonic tensor core 110, a laser transmitter (axon) 112 receives an activation vector (1×k) encoded in its amplitude using a phase using amplitude - shift keying (ASK) or phase - shift keying (PSK). The beam of the axon laser 112 is fanned out to copies of parallel operations (the "dendrite") j using a diffractive optical element (DOE) 114, a phase mask, a beam splitter, or another suitable optical component. The elements of the weight matrix mapped at time step k to an array of j's PSK - encoded laser transmitters (the "synapse") 116 are applied to the activation data via homodyne coherent detection based on the photoelectric effect with an array of homodyne receivers 120, one of which is shown in detail in Figure 1D. Each weighted laser beam beats with a copy of the input laser beam on the corresponding homodyne receiver 120, generating a homodyne product between the two laser fields. The resulting photocurrent is accumulated over time step k, and the accumulated photocurrent

Number

Number

[0016] The PSK-encoded "synapse" laser transmitter 116 can be implemented as an array of injection-locked VCSELs, and the "axon" laser transmitter 112 can be implemented as one of the VCSELs in the array. The VCSEL acts as a phase modulator with negligible amplitude perturbation and accompanied by phase adjustment based on the thermo-optical effect at low data rates (e.g., within 1 MHz) and based on free-carrier injection at higher data rates (e.g., 10 MS / s to GHz). The VCSEL emission (

Number

Number

Number

Number

Number

Number

[0017] The weighting can be encoded into the drive voltage of the "synapse" VCSEL 116. This adjusts the detuning of the VCSEL from the leader laser,

Number

[0018] Figure 1B shows the optoelectronic connection between the layers of the ONN100. As described above, the (phase) modulated outputs from the axonal VCSELs 112 and one synaptic VCSEL 116 interfere at the homodyne receiver 120. The analog-to-digital converter (ADC) 140 converts the analog electronic domain output of the homodyne receiver 120 into a serialized digital signal suitable for storage in the memory 142. The memory 142 passes the stored digital values to the digital-to-analog converter (DAC) 144, and the digital-to-analog converter (DAC) 144 converts the stored digital values into an analog electronic domain signal suitable for driving the axonal VCSEL 112' in the next layer of the ONN100.

[0019] Alternatively, serialization can be potentially realized optically by adding an optical path delay to each read channel where the delay time τ is the data period. The delay is achieved with negligible loss in the optical domain. The integrated photonic voltage can drive the input VCSELs in the next layer without additional ADCs and DACs.

[0020] By fan-out of the activation vectors and the weightings, the computational density of the ONN can

Number

Number

Number

[0021] Figure 1C shows how both the activation vector and the weights can be fan-out for matrix-matrix multiplication with increased arithmetic density. In Figure 1C, an array of five VCSELs computes the product of an input matrix with input vector k = 2

Number

Number

[0022] The homodyne receiver 120 can be used for either linear operation or non-linear operation. To perform a linear multiplication, the input vector

Number

Number

Number

[0023] For linear multiplication, the homodyne receiver can be implemented as a balanced homodyne receiver 120'. This balanced homodyne receiver 120' includes a 2×2 beam splitter 122, and the 2×2 beam splitter 122 receives a beam from one of its input ports being the "axon" laser 112 and the "synapse" laser 116, and its output ports are coupled to a pair of photodetectors 124. Circuit 126 takes the difference of the analog outputs of the photodetectors 124. In this case, the output of the "axon" laser 112 is amplitude modulated by the external amplitude modulator 126 with the activation vector

Number

[0024] The output of the balanced homodyne receiver 120' is the serialized product of the input vector and the weight matrix and can undergo element-by-element non-linear activation like a general neural network, i.e., the optical domain

Number

[0025] Alternatively, ONN100 can operate according to different neural network models, and the homodyne receiver 120 performs a non-linear operation that combines linear weighting with non-linear activation.

Number

Number

[0026] Based on space-time multiplexing and fan-out data copying, the system is optimized for high-density and energy-efficient computing. This uses time step i and coherent receiver j to perform matrix-vector multiplication. The axonal input laser 112 is shared among j channels (j-time parallelism), and the number of VCSELs and photodetectors is

Number

Number

Number

Number

[0027] (VCSEL-based ONN implementation) Figures 2A - 2H show the implementation of a VCSEL-based ONN architecture 100 in the state-of-the-art silicon CMOS technology shown in FIG. 1. FIG. 2A shows an integrated photonic tensor core 210, also called a computational engine, implemented as an optoelectronic co-packaged monolithic device in a three-dimensional design. The integrated photonic tensor core 210 includes an array of individually addressable VCSELs 212 for axon-synapse data encoding. The VCSEL array 212 is electrically and physically coupled to a CMOS driver 216 on one side, and on the other (optical output) side, the VCSEL array 212 is coupled to a holographic phase mask 214 by a first layer 213 of a transparent polymer. A second layer 215 of the transparent polymer couples the holographic phase mask 214 to a CMOS detector array 220.

[0028] Figure 2B is a photograph of several 5×5 VCSEL arrays suitable for use in a photonic tensor core. The VCSEL arrays are excellent components for next-generation ONN, and have (i) a high integration device density with, for example, 80×80μm electrical wire bonding per VCSEL as shown in Figure 2B, (ii) a high modulation bandwidth with, for example, a modulation data rate of 25GS / sec as shown in Figure 2E, and (iii) high scalability that can be manufactured on a wafer scale, is compatible with state-of-the-art optical interconnections, and thus meets criteria C1 and C4 in Table 1. When utilized for dendritic fanout (to N copies) based on a holographic phase mask, the energy consumption (C2) for data encoding decreases by a multiple of N. Nonlinearity (C3) is induced by coherent detection using phase encoding. 2 When utilized for dendritic fanout (to N copies) based on a holographic phase mask, the energy consumption (C2) for data encoding decreases by a multiple of N. Nonlinearity (C3) is induced by coherent detection using phase encoding.

[0029] The 5×5 VCSEL array of Figure 2B is fabricated as a semiconductor heterostructure microresonator and has two AlGaAs / GaAs distributed Bragg reflectors as cavity mirrors and a stack of InGaAs quantum wells as the gain medium. The cavity array was patterned by ultraviolet (UV) lithography and etched by inductively coupled plasma reactive ion beam.

[0030] Each cavity has an outer diameter of 30μm and was oxidized to an aperture of 4.5μm to suppress higher-order transverse modes. To improve laser stability, the entire chip was covered with a polymer layer and the region of the VCSEL cavity was opened again. The Au-deposited p-contact of each VCSEL was wire-bonded to a signal pad that was connected to a printed circuit board (e.g., the CMOS driver 216 in Figure 2A) linked to an external driver (not shown). All VCSELs sharing a common ground (the gold bar in Figure 2B). The VCSEL has a cross-section with an ellipticity of 1%, which enables a polarized laser output with an improved extinction ratio.

[0031] During operation, a laser driver (e.g., CMOS driver 216) forward-biased a VCSEL (e.g., 2 VDC) beyond the lasing threshold of the laser driver and applied a small AC voltage (e.g., within 10 mV) signal modulation to the VCSEL. Each VCSEL emitted 100 μW of light with a wall plug efficiency of 25%. The modulation bandwidth of each VCSEL was approximately 2 GHz (3 dB), which was limited by the photon lifetime of the VCSEL cavity (e.g.,

Number

Number

Number

[0032] Figure 2C shows the fan-out of the beam from the "axonal" VCSEL 212 (lower right corner). Limited by the large dimension of the fan-out DOE 214, the output of the corner laser was separated from the weighted beam and fanned out by the DOE 214 (beam separation is not necessary with a compact DOE). The DOE 214 imprints a phase pattern on the beam profile, thereby fanning out into an array of beams / spots. For example, Figure 2F is an image of the Fourier plane of a DOE that fans out the beam into a 32×32 array of spots.

[0033] A beam splitter (BS) 236 was used for coherent homodyne detection to the corresponding photodetectors of a 5×5 photodetector (PD) array 220

Number

Number

[0034] Figure 2E shows analog data encoded by a single VCSEL beam (each spot in Figure 2F) modulated at 25 GS / s. The MIT logo in the upper plot was constructed from a time series of 2,000 samples. A 28×28 pixel image with a handwritten digit (center row) was flattened and encoded over a duration of 31.36 ns. The frequency response of the injection-locked VCSEL in the thermal region (within 10 MHz) was approximately 10 dB stronger than that in the free-carrier region. To separate from the thermal effects at low modulation frequencies, the data was modulated with a high-frequency local oscillator. This data modulation scheme is not required when the VCSEL operates at a high data rate.

[0035] Figure 2C also shows the free-space injection locking of VCSEL 212 to the leader laser 230. The VCSELs were emitted at a wavelength of approximately 974 ± 0.1 nm across the array. This excellent wavelength uniformity enabled parallel injection locking to the leader laser 230 (e.g., another VCSEL) across the array. The injection that locks VCSEL 212 to the leader laser 230 established the mutual coherence between VCSELs 212 for homodyne detection.

[0036] The first diffractive optical element (DOE) 232 in the Fourier plane of the coupling lens 242 divides the injection lock beam emitted by the leader laser 230 into a 3×3 array of injection lock beams having a grid spacing or pitch equal to the pitch of the VCSEL array 212. A polarization beam splitter (PBS) 234 reflects the array of injection lock beams into the VCSELs 212 through the coupling lens 242. Since the PBS 234 is rotated 45 degrees with respect to the polarization state of the output of the VCSEL 212, half of the power of the injection lock beam is coupled to the VCSEL 212 and locks the phase of the VCSEL 212. The front DBR of the VCSEL cavity reflects the remaining half of the power of the injection lock beam. The PBS 234 rejects this reflected light to avoid generating unwanted interference in the homodyne detector. The VCSEL 212 is tuned to the target wavelength using an electronic DC forward bias from a VCSEL driver (CMOS driver 216), facilitating simultaneous injection locking of the entire VCSEL array 212. Injection locking can be confirmed by monitoring the beat note between the leader laser 230 and each VCSEL 212.

[0037] Alternatively, the VCSEL 212 can be injection locked in a waveguide-based architecture as shown in FIG. 2D. In this waveguide-based architecture, the leader laser 230' is coupled to the waveguide 234, which guides the injection lock beam from the leader laser 230' to the back of the VCSEL 212. The gratings 236a and 236b couple portions of the injection lock beam from the waveguide 234 into the VCSELs 212a and 212b, respectively, through the rear DBR of the VCSEL. Each rear DBR couples a portion of the injection lock beam to the corresponding VCSEL cavity and returns the remaining injection lock beam back towards the corresponding grating.

[0038] The phase of the injection-locked VCSEL 212 is given by the frequency detuning between the leader laser 230 and the free-running frequency of the VCSEL, where

Number

Number

Number

Number

[0039] Figure 2H is a plot of the VCSEL detuning versus the leader laser detuning. This shows the injection-locking range. This is measured by monitoring the beat note between the leader laser and each VCSEL. An injection power of 500 nW per VCSEL results in an injection-locking range of 1.7 GHz and a phase shift voltage

Number

Number

[0040] (Nonlinearity and Computational Accuracy of Homodyne Interference) Figures 3A - 3E show detection - based optical homodyne nonlinearity suitable for use in a VCSEL - based ONN. As shown in Figure 3A, this detection - based optical homodyne nonlinearity is implemented using a photodetector 324 to detect the homodyne interference of phase - modulated beams from an “axon” VCSEL 112 and a “synapse” VCSEL 116 combined by a beam combiner 322. The plot in Figure 3B shows that the intensity of the homodyne nonlinearity is adjusted by programming the phase of the weighted VCSEL 116. Since homodyne detection relies on the photoelectric effect where electrons are raised to the conduction band by absorbed photons, the process is almost instantaneous and involves a time delay of tens of attoseconds. The resulting delay can be less than a femtosecond, as short as the optical pulse per symbol. This is in contrast to the nanosecond delays of digital non - linearities, electro - optical non - linearities, and cavity - or atom - based optical non - linearities. Its implementation using a photodetector is free of instrument complexity (e.g., ultra - short laser pulses) and is ultra - compact.

[0041] Figure 3B shows different weighting values

Number

[0042] (VCSEL-based ONN Inference Demo) Figures 4A through 4E show DNN inference by an axon-synapse-dendrite VCSEL-based ONN architecture trained with 1,000 test images of handwritten digits from the Modified National Institute of Standards and Technology (MNIST) database. For this purpose, the inventors developed a training process using PyTorch with an exclusive non-linear weighting function.

[0043] Figure 4A shows the training model 400 itself, which includes one input layer 410, two fully connected hidden layers 412a and 412b, and an output layer 414. The input layer 410 includes 784 neurons corresponding to a full-size MNIST image with a handwritten digit (e.g., the handwritten "9" shown on the left in Figure 4A). In each fully connected hidden layer 412, the matrix-vector multiplication is calculated by a custom nonlinear synaptic weighting function. The output layer 414 includes 10 neurons, each neuron represents a digit (0-9), and the prediction of the training model of the digit represented by the input is given by the number of the neuron having the maximum value.

[0044] For training, as shown on the left in Figure 4B, each 28×28 pixel test image is flattened and encoded into the phase of the input VCSEL at a driving voltage of 4mv in 784 time steps. Each weight vector has one weight vector for each synaptic VCSEL, is of the same size, is flattened, and is sent to the synaptic VCSEL. Parallel spatial multiplexing makes it possible to process all the weight vectors simultaneously. However, it is limited by the number of high-speed arbitrary waveform generator (AWG) channels available for generating data, and the data was acquired with 10 VCSELs modulated at a speed of 100 MS / second. By switching the AWG channels and translating the VCSEL chip to different arrays in the x and y directions, a total of 100 VCSELs from 5 VCSEL arrays are used to calculate the second hidden layer 412b.

[0045] Figure 4B shows the interference signal and digital calculation results between the image data and the weight vector. There is an excellent agreement between the measured and calculated interference signals for the image data and the weight vector. The interference signal is integrated with the integrated receiver over time in each channel, resulting in 100 integrated values as the input vector to the output layer 414. The signal-to-noise ratio (SNR) in the time trace is 135 and is limited by photon shot noise. The photocurrent of each channel is accumulated over time by a time integrator. The integrated values from 100 channels are serialized to form an input vector that is supplied to the next layer. The weighting of the output layer 414 is implemented with 10 weighted VCSELs, and the interference signal is integrated.

[0046] Figure 4C shows the real-time integration of the interference signal in the output layer 414. The result of the image classification was directly read from the voltage levels of the integrated values of 10 VCSEL channels, as shown in Figure 4D. When inferences were performed on a random dataset of 1,000 images with a total of 158.8 million operations, an accuracy (93.1 ± 2%) that was not statistically distinguishable from the model accuracy (95.2%) was obtained in the simulation.

[0047] (Performance of the VCSEL-based ONN) Figure 5 is a plot of energy efficiency versus computational density for state-of-the-art neural network hardware, including the VCSEL-based ONN disclosed herein. Google Tensor Processing Unit (TPU), NVIDIA Graphics Processing Unit (GPU), and Graphcore have 1 pJ / OP (Graphcore) and 0.35 tera OP / (mm 2·s) (NVIDIA A100), an application-specific integrated circuit (ASIC) optimized for deep learning tasks, having energy efficiency and computing density respectively. Regarding the ONN optical performance, the energy efficiency is the power consumption in laser generation and data encoding, while the computing density is calculated from the chip area of matrix operations. Regarding the overall system performance of ONN, the energy consumption and computing density consider laser generation, data encoding, non-linear activation, data readout, signal amplification, ADC, DAC, and memory access. In the case of VCSEL-based ONN, spatial fan-out and time-domain fan-in can reduce the energy limit by electronics.

[0048] VCSEL-based ONN enables efficient computing with a low-energy VCSEL transmitter and optical parallelism. The clock rate of an example of VCSEL-based ONN is limited by the VCSEL bandwidth (as demonstrated in Figure 3) to 1 GS / s. Due to the ultra-low operation (such as Vπ = 4 mV), the data encoded by the VCSEL modulator consumes very little power, for example, 3.7 nanowatts (3.7 attojoules per symbol at 1 GS / second), which is six orders of magnitude lower than the power consumption of ONN in a thermo-optic phase shifter, microring resonator, optical attenuator, and electro-optic modulator, each of which consumes several milliwatts of power. As a result, the main optical energy consumption in VCSEL-based ONN is usually for laser generation.

[0049] The VCSEL source is an efficient laser generator having a wall plug efficiency of 25% or higher, or a higher efficiency (e.g., over 57%). The theoretical lower limit for the laser output is given by the number of photons necessary to generate a homodyne signal with sufficient computational precision bits, which is ultimately limited by the SNR required for detection. The integration time of the receiver, in contrast to conventional amplified detectors, is read out only after accumulating over several time steps to improve the SNR. In off-the-shelf technology, the thermal noise limit of computing from integrated detection is 200 photons / OP (equivalent to 40 aJ / OP). In experimental demonstration, the VCSEL emitted 100 μW. The resulting optical energy efficiency, including the power for laser generation and data modulation, is 2.5 fJ / OP (due to the advantage of fanout). The optical energy efficiency of 2.5 fJ / OP for the VCSEL-based ONN is at least 140 times higher than that of state-of-the-art integrated ONNs.

[0050] The VCSEL-based ONN incurs energy costs from electronic digital-to-analog converters (DACs), analog-to-digital converters (ADCs), signal amplification, and memory access. The energy of the DAC per use, and memory access, are reduced by a factor of j by spatial parallel processing via laser fanout. The readout electronics, including the ADC, transimpedance amplifier, and integrator, are triggered once after time integration. The energy cost per use is wasted by a total of 2i intervening operations. Thereby, the overall energy efficiency of the system, including both electronic and optical consumption, is 7 fJ / OP, which is over 100 times higher than that of state-of-the-art electronic microprocessors. Similar to the fanout of the input laser, the weighted VCSELs can be spatially fanout (by a factor of k), which reduces the energy for weighting in the same order as the input encoding.

[0051] VCSEL-based ONNs have high computational density due to the compactness and density of VCSEL arrays in a three-dimensional architecture. VCSELs are excellent candidates for high-density computing, having a pitch of 80 μm per fabricated device. Nano- or micro-pillar lasers, which may have a diameter exceeding 1 μm and a pitch exceeding 10 μm, offer similar advantages and can be used instead of VCSEL arrays. The computational density of the VCSEL-based ONN demonstrated here reaches 25 tera-OP / (mm 2 ·s), which is approximately two orders of magnitude higher than the computational density of its electronic counterparts. In electronic circuits, improving throughput density is difficult because heat dissipation per chip area is limited. The higher energy efficiency of VCSEL-based ONNs enables higher throughput density. In other ONN configurations, high throughput density involves tiling photonic devices at high density, which often leads to severe crosstalk between adjacent channels and a decrease in computational accuracy. The channel crosstalk of VCSEL-based ONNs is reduced or eliminated using VCSEL modulators with ultra-low half-wave voltage.

[0052] VCSEL-based ONNs operate with ultra-low latency for non-linear activation due to detection-based non-linearity. In VCSEL-based ONNs, each detection event instantaneously generates a photocurrent, which is accumulated in a time integrator for i time steps before being read out. The transit time of photoelectrons to a charging capacitor, which leads to the delay of a standard photodetector from a photodiode, is negligible compared to the integration time. Thus, the delay due to non-linear activation is negligible. The processing time is dominated by data encoding and time integration and can be 30 ns for full-size MNIST images at a clock rate of 25 GS / s.

[0053] (Conclusion) Although various embodiments of the present invention have been described and illustrated herein, those skilled in the art will readily conceive of various other means and / or structures for performing the functions and / or obtaining the results and / or advantages described herein, and such variations and / or modifications are each to be regarded as being within the scope of the embodiments of the present invention described herein. Further, in general, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on the particular application in which the teachings of the present invention are used, which will be readily understood by those skilled in the art. Those skilled in the art will be able to recognize, or ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. Accordingly, the foregoing embodiments are presented by way of example only, and within the scope of the appended claims and their equivalents, the embodiments of the present invention may be practiced otherwise than as specifically described and claimed. Embodiments of the invention related to the present disclosure are directed to the individual features, systems, articles, materials, kits, and / or methods described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the invention of the present disclosure if such features, systems, articles, materials, kits, and / or methods do not mutually conflict.

[0054] Also, the various concepts of the present invention can be embodied in one or more methods, and examples thereof are provided. The acts performed as part of the method may be ordered in any suitable manner. Accordingly, embodiments may be constructed in which acts are performed in an order different from that illustrated, including performing some acts simultaneously, even though some acts are shown as sequential acts in the illustrated embodiments.

[0055] All definitions defined and used herein are to be understood as controlling the dictionary definitions, definitions of incorporated by reference documents, and / or ordinary meaning of the defined terms.

[0056] In the specification and claims, the indefinite articles "a" and "an" used herein should be understood to mean "at least one" unless clearly indicated otherwise.

[0057] In the specification and claims, the phrase "and / or" used herein should be understood to mean "either or both" of the elements so combined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. The plurality of elements listed with "and / or" should be construed in the same manner, i.e., as "one or more" of the elements so combined. Other elements other than those specifically identified by the "and / or" clause may optionally exist, whether related or unrelated to the specifically identified elements. Thus, by way of non-limiting example, a reference to "A and / or B", when used in combination with open-ended language such as "comprising", may refer in one embodiment to only A (optionally including elements other than B), in another embodiment to only B (optionally including elements other than A), in yet another embodiment to both A and B (optionally including other elements), and so on.

[0058] As used in this specification and the claims, "or" shall be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" is inclusive, i.e., it shall be interpreted as including at least one of several or enumerated elements, and optionally, additional unenumerated items, as well as including two or more of them. Conversely, terms explicitly stated, such as "only one of" or "exactly one of", or "consisting of" when used in the claims, shall refer to including exactly one of several or enumerated elements. Generally, the term "or" as used in this specification shall be interpreted only as indicating an exclusive alternative (i.e., "either one or the other but not both") when preceded by an exclusive term such as "either", "one of", "only one of", "exactly one of". "Consisting essentially of" as used in the claims shall have the ordinary meaning as used in the field of patent law.

[0059] As used in this specification and the claims, the term "at least one" in reference to a list of one or more elements means at least one element selected from one or more of the elements in the list of elements, but does not necessarily include at least one of every element specifically recited in the list of elements, and is not to be construed as excluding any combinations of elements in the list of elements. This definition also allows that elements other than those specifically identified in the list of elements to which the phrase "at least one" refers, whether related or unrelated to the specifically identified elements, may optionally be present. Thus, by way of non-limiting example, "at least one of A and B" (or equivalently, "at least one of A or B", or equivalently, "at least one of A and / or B") can, in one embodiment, refer to at least one, optionally two or more, A's with no B's present (and optionally including elements other than B), in another embodiment, refer to at least one, optionally two or more, B's with no A's present (and optionally including elements other than A), and in yet another embodiment, refer to at least one, optionally two or more, A's, and at least one, optionally two or more, B's (and optionally including other elements as needed), etc.

[0060] In the claims and in the above specification, all transitional phrases, such as "comprising," "including," "having," "has," "containing," "involving," "retaining," "consisting of," etc., are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases "consisting of" and "consisting essentially of" are to be considered closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent and Trademark Office's Manual of Patent Examining Procedure, Section 2111.03.

Claims

1. A photonic neural network comprising: a first vertical-cavity surface-emitting laser (VCSEL) that emits a first beam phase-modulated by an activation vector; and a second VCSEL that emits a second beam coherent with the first beam and phase-modulated by the weighting of a weight matrix of the photonic neural network, an array of vertical-cavity surface-emitting lasers (VCSELs); a diffractive optical element optically communicating with the first VCSEL to fan out the first beam; an array of photodetectors optically communicating with the VCSEL array and the diffractive optical element to detect interference between each fanned-out copy of the first beam from the diffractive optical element and the second beam A photonic neural network comprising the above components.

2. The photonic neural network according to claim 1, wherein the VCSEL array comprises the diffractive optical element and the photodetector array and is monolithically integrated.

3. The photonic neural network according to claim 1, wherein the VCSEL array is configured to be modulated with a half-wave voltage of less than 10 mV.

4. The photonic neural network according to claim 1, wherein the VCSEL array is configured to be modulated at a speed of at least 1 Gb / s.

5. The array of photodetectors is configured to 【Number 1】 generate an output proportional to, where 【Number 2】 , and [Number 3] are the amplitude and phase of the first beam, respectively, 【Number 4】 , and 【Number 5】 are the amplitude and phase of the j-th second beam, respectively. The photonic neural network according to claim 1.

6. The array of photodetectors is configured to 【Number 6】 generate an output proportional to, where 【Number 7】 represents the element of the i-th activation vector, 【Number 8】 represents the ij-th element of the weight matrix. The photonic neural network according to claim 1.

7. The photonic neural network has a computational density of at least 【Number 9】 . The photonic neural network according to claim 1.

8. The photonic neural network according to claim 1, configured to operate with an energy consumption of 1 fJ / OPS.

9. A leader laser optically communicating with an array of VCSELs, further comprising a leader laser for injection locking the array of VCSELs, the optical neural network according to claim 1.

10. The diffractive optical element is a first diffractive optical element, The optical neural network according to claim 1, further comprising a second diffractive optical element optically communicating with one of the second VCSELs in the array of VCSELs and fan-out the second beam emitted by one of the second VCSELs.

11. Adjusting the phase of a first beam emitted by a first vertical cavity surface emitting laser (VCSEL) within an array of VCSELs using an activation vector to the optical neural network; Adjusting the phase of a second beam emitted by a second VCSEL in the array of VCSELs with each element of the weight matrix of the optical neural network; Fan-out a copy of the first beam to each homodyne receiver; Detecting homodyne interference between the copy of the first beam and each second beam at the homodyne receiver An inference processing method comprising:

12. The method of claim 11, wherein adjusting the phase of the first beam includes applying a half-wavelength voltage of less than 10 mV to the first VCSEL.

13. The method of claim 11, wherein adjusting the phase of the first VCSEL includes applying modulation at a rate of at least 1 Gb / s.

14. The step of detecting the homodyne interference includes 【Number 10】 Generating an electrical signal proportional to, where 【Number 11】 , and 【Number 12】 Are the amplitude and phase of the first beam, respectively, 【Number 13】 , and 【Number 14】 Are the amplitude and phase of the j-th second beam, respectively, the inference processing method according to claim 11.

15. The step of detecting the homodyne interference includes 【Number 15】 Generating an electrical signal proportional to, where 【Number 16】 Represents the i-th element of the activation vector, 【Number 17】 Represents the ij-th element of the weight matrix, the inference processing method according to claim 11.

16. The inference processing occurs at a computational density of at least 【Number 18】 , the inference processing method according to claim 11.

17. The inference processing method according to claim 11, wherein the inference processing occurs with an energy consumption of 1 fJ / OPS or less.

18. The inference processing method according to claim 11, further comprising the step of injection locking the first VCSEL and the second VCSEL to a leader laser.

19. The inference processing method according to claim 11, further comprising the step of fan - out at least one copy of the second beam to at least some of the homodyne receivers.

20. An optical neural network, an input layer for receiving an input to the optical neural network, a plurality of fully - connected layers operably coupled to the input layer for performing inference processing on the input to the optical neural network, wherein each of the plurality of fully - connected layers over time 【Number 19】 emits an axonal beam phase - modulated by a serialized activation vector with an axonal VCSEL, 【Number 20】 over time 【Number 21】 an array of vertical - cavity surface - emitting lasers (VCSELs) including a synaptic VCSEL that emits a synaptic beam phase - modulated by the weighting of each of the serialized weight matrices, a diffractive optical element optically communicating with the axonal VCSEL and fan - out each copy of the axonal beam to the synaptic beam, an array of homodyne receivers optically communicating with the axonal VCSEL and the synaptic VCSEL via the diffractive optical element and detecting homodyne interference between each copy of the axonal beam and the synaptic beam, wherein the array of homodyne receivers 【Number 22】 generates an electrical signal proportional to, where 【Number 23】 represents the i - th element of the activation vector, 【24 Points】 represents the ij - th element of the weight matrix, and an array of homodyne receivers, a plurality of fully - connected layers including a leader laser optically communicating with the VCSEL and injection - locking the VCSEL, an output layer operably coupled to the plurality of fully - connected layers and returning an output of the inference processing comprising an optical neural network.