Optoelectronic Computing Systems

The optical processor system addresses inefficiencies in electronic neuromorphic computing by utilizing optical signals for matrix multiplication, achieving fast and scalable neuromorphic computing operations.

JP7744826B2Active Publication Date: 2025-09-26LIGHTELLIGENCE PTE LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2021518033
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-03-19
Filing Date
2019-06-04
Publication Date
2025-09-26
Estimated Expiration
2039-06-04

AI Technical Summary

Technical Problem

Existing neuromorphic computing systems rely heavily on electronic circuits for matrix multiplication, which are inefficient and limited in scalability and speed, while optical signals are underutilized for significant computational operations.

Method used

An optical processor system that includes a memory unit, digital-to-analog converter, optical modulators, optical matrix multiplication unit, and analog-to-digital converter to perform artificial neural network calculations using optical signals, enabling fast and efficient matrix multiplication with loop periods of 1 ns or less.

Benefits of technology

The system achieves high-speed and efficient neuromorphic computing by leveraging optical signals for matrix multiplication, enhancing computational performance and scalability beyond traditional electronic methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007744826000084
    Figure 0007744826000084
  • Figure 0007744826000085
    Figure 0007744826000085
  • Figure 0007744826000086
    Figure 0007744826000086
Patent Text Reader

Abstract

The system and method include providing input information in an electronic format, converting at least a portion of the electronic input information into an optical input vector, optically converting the optical input vector into an optical output vector based on optical matrix multiplication, converting the optical output vector into an electronic format, and electronically applying a nonlinear transformation to the electrically converted optical output vector to provide output information in the electronic format. In some examples, a set of multiple input values ​​is encoded onto respective optical signals carried by an optical waveguide. For each of at least two subsets of one or more optical signals, a corresponding set of one or more replication modules divides the subset of one or more optical signals into two or more replicas of the optical signal. For each of at least two replicas of a first subset of one or more optical signals, a corresponding multiplication module multiplies one or more optical signals of the first subset with one or more matrix element values ​​using optical amplitude modulation. For two or more results of the multiplication modules, an addition module produces an electrical signal representing the sum of the results of two or more of the multiplication modules.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 62 / 680,944, filed June 5, 2018, U.S. Provisional Application No. 62 / 744,706, filed October 12, 2018, U.S. Provisional Application No. 62 / 792,144, filed January 14, 2019, and U.S. Provisional Application No. 62 / 820,562, filed March 19, 2019. The entire disclosures of the above applications are incorporated herein by reference.

[0002] The present disclosure relates to optoelectronic computing systems. [Background technology]

[0003] Neuromorphic computing is a technique that mimics the behavior of the brain in the electronic domain. A prominent approach to neuromorphic computing is the artificial neural network (ANN), which is a collection of artificial neurons that are interconnected in a specific way to process information in a manner similar to how the brain functions. ANNs have found use in a wide range of applications, including artificial intelligence, speech recognition, text recognition, natural language processing, and various forms of pattern recognition.

[0004] An ANN has an input layer, one or more hidden layers, and an output layer. Each layer has nodes or artificial neurons, and the nodes are interconnected between layers. Each node in the hidden layer performs a weighted sum of signals received from nodes in the previous layer and performs a nonlinear transformation ("activation") of the weighted sum to generate an output. The weighted sum may be calculated by performing matrix multiplication steps. Therefore, computing an ANN typically involves multiple matrix multiplication steps, which are typically performed using electronic integrated circuits.

[0005] Calculations performed on electronic data, encoded in analog or digital form on electrical signals (e.g., voltage or current), are typically implemented using electronic computing hardware, such as analog or digital electronic devices, electronic circuit boards, or other electronic circuits implemented in integrated circuits (e.g., processors, application-specific integrated circuits (ASICs), or systems-on-chips (SoCs)). Optical signals have been used to transport data over long distances and over shorter distances (e.g., within data centers). Operations performed on such optical signals often occur in the context of optical data transport, such as in devices used to switch or filter optical signals in networks. The use of optical signals in computing platforms has been more limited. Various components and systems for all-optical computing have been proposed. Such systems may involve conversion from electrical signals at the input and back to electrical signals at the output, but may not use both types of signals (electrical and optical) for significant operations performed in the computation. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] US Patent Application Publication No. 2017 / 0351293 Summary of the Invention [Means for solving the problem]

[0007] In general, in a first aspect, a system includes a memory unit configured to store a dataset and a plurality of neural network weights; a digital-to-analog converter (DAC) unit configured to generate a plurality of modulator control signals and generate a plurality of weight control signals; an optical processor, the optical processor including a laser unit configured to generate a plurality of optical outputs; a plurality of optical modulators coupled to the laser unit and the DAC unit, the plurality of optical modulators configured to generate an optical input vector by modulating the plurality of optical outputs generated by the laser unit based on the plurality of modulator control signals; an optical matrix multiplication unit coupled to the plurality of optical modulators and the DAC unit, the optical matrix multiplication unit configured to convert the optical input vector into an optical output vector based on the plurality of weight control signals; and an optical matrix multiplication unit coupled to the optical matrix multiplication unit and configured to convert the optical output vector into an optical output vector. an analog-to-digital conversion (ADC) unit coupled to the light detection unit and configured to convert the plurality of output voltages into a plurality of digitized light outputs; and a controller including an integrated circuit configured to perform operations including receiving, from the computer, an artificial neural network calculation request including an input data set and a first plurality of neural network weights, the input data set including a first digital input vector; storing the input data set and the first plurality of neural network weights in a memory unit; and generating, through a DAC unit, a first plurality of modulator control signals based on the first digital input vector and a first plurality of weight control signals based on the first plurality of neural network weights.

[0008] Embodiments of the system may include one or more of the following features. For example, the operations may further include obtaining, from the ADC unit, a first plurality of digitized optical outputs corresponding to the optical output vector of the optical matrix multiplication unit, the first plurality of digitized optical outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector; and storing, in the memory unit, the first transformed digital output vector.

[0009] The system may have a first loop period defined as the time elapsed between storing the input data set and the first plurality of neural network weights in the memory unit and storing the first converted digital output vector in the memory unit, and the first loop period may be 1 ns or less.

[0010] In some implementations, the operations may further include outputting an artificial neural network output generated based on the first transformed digital output vector.

[0011] In some implementations, the operations may further include generating, through a DAC unit, a second plurality of modulator control signals based on the first converted digital output vector.

[0012] In some implementations, the artificial neural network computation request can further include a second plurality of neural network weights, and the operations can further include generating, through the DAC unit, a second plurality of weight control signals based on the second plurality of neural network weights based on obtaining the first plurality of digitized optical outputs. The first plurality of neural network weights and the second plurality of neural network weights can correspond to different layers of the artificial neural network.

[0013] In some implementations, the input data set may further include a second digital input vector, and the operations may further include generating, through the DAC unit, a second plurality of modulator control signals based on the second digital input vector, obtaining, from the ADC unit, a second plurality of digitized optical outputs corresponding to the optical output vector of the optical matrix multiplication unit, where the second plurality of digitized optical outputs form the second digital output vector, performing a nonlinear transformation on the second digital output vector to generate a second transformed digital output vector, storing the second transformed digital output vector in the memory unit, and outputting an artificial neural network output generated based on the first transformed digital output vector and the second transformed digital output vector. The optical output vector of the optical matrix multiplication unit results from the second optical input vector generated based on the second plurality of modulator control signals that are transformed by the optical matrix multiplication unit based on the first mentioned plurality of weight control signals.

[0014] In some implementations, the system may further include an analog nonlinearity unit disposed between the light detection unit and the ADC unit, the analog nonlinearity unit configured to receive a plurality of output voltages from the light detection unit, apply a nonlinear transfer function, and output a plurality of converted output voltages to the ADC unit, and the operations further include obtaining, from the ADC unit, a first plurality of converted digitized output voltages corresponding to the plurality of converted output voltages, the first plurality of converted digitized output voltages forming a first converted digital output vector; and storing the first converted digital output vector in the memory unit.

[0015] In some implementations, the controller integrated circuit may be configured to generate the first plurality of modulator control signals at a rate of 8 GHz or greater.

[0016] In some implementations, the system may further include an analog memory unit disposed between the DAC unit and the plurality of optical modulators, the analog memory unit configured to store analog voltages and output the stored analog voltages, and an analog nonlinearity unit disposed between the light detection unit and the ADC unit, the analog nonlinearity unit configured to receive the plurality of output voltages from the light detection unit, apply a nonlinear transfer function to the plurality of output voltages, and output a plurality of converted output voltages. The analog memory unit may include a plurality of capacitors.

[0017] In some implementations, the analog memory unit may be configured to receive and store a plurality of converted output voltages of the analog nonlinearity unit and output the stored plurality of converted output voltages to the plurality of optical modulators, and the operations may further include storing the plurality of converted output voltages of the analog nonlinearity unit in the analog memory unit based on generating the first plurality of modulator control signals and the first plurality of weight control signals; outputting the stored converted output voltages through the analog memory unit; obtaining a second plurality of converted digitized output voltages from the ADC unit, where the second plurality of converted digitized output voltages form a second converted digital output vector; and storing the second converted digital output vector in the memory unit.

[0018] In some implementations, the input dataset of the artificial neural network computation request may include multiple digital input vectors. The laser unit may be configured to generate multiple wavelengths. The multiple optical modulators may include a bank of optical modulators configured to generate multiple optical input vectors, each of the banks corresponding to one of the multiple wavelengths and generating a respective optical input vector having the respective wavelength, and an optical multiplexer configured to combine the multiple optical input vectors into a combined optical input vector including the multiple wavelengths. The optical detection unit may further be configured to demultiplex the multiple wavelengths and generate multiple demultiplexed output voltages. The operations may include obtaining multiple digitized demultiplexed optical outputs from the ADC unit, where the multiple digitized demultiplexed optical outputs form multiple first digital output vectors, each of the multiple first digital output vectors corresponding to one of the multiple wavelengths; performing a nonlinear transformation on each of the multiple first digital output vectors to generate multiple transformed first digital output vectors; and storing the multiple transformed first digital output vectors in the memory unit. Each of the plurality of digital input vectors may correspond to one of the plurality of optical input vectors.

[0019] In some implementations, the artificial neural network computational request may include multiple digital input vectors. The laser unit may be configured to generate multiple wavelengths. The multiple optical modulators may include a bank of optical modulators configured to generate multiple optical input vectors, each of the banks corresponding to one of the multiple wavelengths and generating a respective optical input vector having the respective wavelength, and an optical multiplexer configured to combine the multiple optical input vectors into a combined optical input vector including the multiple wavelengths. The operations may include obtaining, from the ADC unit, a first plurality of digitized optical outputs corresponding to an optical output vector including the multiple wavelengths, the first plurality of digitized optical outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector; and storing the first transformed digital output vector in the memory unit.

[0020] In some implementations, the DAC unit may include a 1-bit DAC subunit configured to generate multiple 1-bit modulator control signals. The resolution of the ADC unit may be 1 bit. The resolution of the first digital input vector may be N bits. The operations may include decomposing the first digital input vector into N 1-bit input vectors, each of the N 1-bit input vectors corresponding to one of the N bits of the first digital input vector; generating, via the 1-bit DAC subunit, a sequence of N 1-bit modulator control signals corresponding to the N 1-bit input vectors; obtaining, from the ADC unit, a sequence of N digitized 1-bit optical outputs corresponding to the sequence of N 1-bit modulator control signals; constructing an N-bit digital output vector from the sequence of N digitized 1-bit optical outputs; performing a nonlinear transformation on the constructed N-bit digital output vector to generate a transformed N-bit digital output vector; and storing the transformed N-bit digital output vector in a memory unit.

[0021] In some implementations, the memory unit may include a digital input vector memory configured to store a digital input vector and including at least one SRAM, and a neural network weight memory configured to store a plurality of neural network weights and including at least one DRAM.

[0022] In some implementations, the DAC unit may include a first DAC subunit configured to generate a plurality of modulator control signals and a second DAC subunit configured to generate a plurality of weight control signals, wherein the first DAC subunit and the second DAC subunit are different.

[0023] In some implementations, the laser unit may include a laser source configured to generate light and an optical power splitter configured to split the light generated by the laser source into a plurality of optical outputs, each of the plurality of optical outputs having substantially equal power.

[0024] In some implementations, the plurality of optical modulators may include one of an MZI modulator, a ring resonator modulator, or an electro-absorption modulator.

[0025] In some implementations, the light detection unit may include a plurality of light detectors and a plurality of amplifiers configured to convert photocurrents generated by the light detectors into a plurality of output voltages.

[0026] In some implementations, the integrated circuit may be an application specific integrated circuit.

[0027] In some implementations, the optical matrix multiplication unit may include an array of input waveguides for receiving an optical input vector, an optical interference unit in optical communication with the array of input waveguides for performing a linear transformation of the optical input vector into a second array of optical signals, and an array of output waveguides in optical communication with the optical interference unit for directing the second array of optical signals, wherein at least one input waveguide in the array of input waveguides is in optical communication with each output waveguide in the array of output waveguides via the optical interference unit.

[0028] In some implementations, the optical interference unit may include a plurality of interconnected Mach-Zehnder interferometers (MZIs), each MZI among the plurality of interconnected MZIs including a first phase shifter configured to change a division ratio of the MZI and a second phase shifter configured to shift a phase of an output of one of the MZIs, and the first phase shifter and the second phase shifter are coupled to a plurality of weight control signals.

[0029] In another aspect, a system includes an optical processor including: a memory unit configured to store a dataset and a plurality of neural network weights; a driver unit configured to generate a plurality of modulator control signals and generate a plurality of weight control signals; an optical processor configured to generate a plurality of optical outputs; a laser unit configured to generate a plurality of optical modulators coupled to the laser unit and the driver unit, the plurality of optical modulators configured to generate an optical input vector by modulating the plurality of optical outputs generated by the laser unit based on the plurality of modulator control signals; an optical matrix multiplication unit coupled to the plurality of optical modulators and the driver unit, the optical matrix multiplication unit configured to convert the optical input vector into an optical output vector based on the plurality of weight control signals; and an optical detection unit coupled to the optical matrix multiplication unit and configured to generate a plurality of output voltages corresponding to the optical output vector; a comparator unit coupled to the optical detection unit and configured to convert the plurality of output voltages into a plurality of digitized 1-bit optical outputs; receiving an artificial neural network computation request including an input dataset and a first plurality of neural network weights, the input dataset including a first digital input vector having an N-bit resolution; storing the input dataset and the first plurality of neural network weights in a memory unit; decomposing the first digital input vector into N 1-bit input vectors, each of the N 1-bit input vectors corresponding to one of the N bits of the first digital input vector; generating, through a driver unit, a sequence of N 1-bit modulator control signals corresponding to the N 1-bit input vectors; obtaining, from a comparator unit, a sequence of N digitized 1-bit optical outputs corresponding to the sequence of N 1-bit modulator control signals; constructing an N-bit digital output vector from the sequence of N digitized 1-bit optical outputs; and performing a nonlinear transformation on the constructed N-bit digital output vector to generate a transformed N-bit digital output vector.and storing the converted N-bit digital output vector.

[0030] In another aspect, a method for performing an artificial neural network computation in a system having an optical matrix multiplication unit configured to convert an optical input vector to an optical output vector based on a plurality of weight control signals includes the steps of receiving an artificial neural network computation request from a computer, the artificial neural network computation request including an input data set and a first plurality of neural network weights, the input data set including a first digital input vector; storing the input data set and the first plurality of neural network weights in a memory unit; and receiving, through a digital-to-analog converter (DAC) unit, the first plurality of modulator control signals and the first plurality of neural network weights based on the first digital input vector. generating a first plurality of weight control signals based on the network weights; obtaining, from an analog-to-digital conversion (ADC) unit, a first plurality of digitized optical outputs corresponding to the optical output vector of the optical matrix multiplication unit, the first plurality of digitized optical outputs forming a first digital output vector; performing, by a controller, a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector; storing, in a memory unit, the first transformed digital output vector; and outputting, by the controller, an artificial neural network output generated based on the first transformed digital output vector.

[0031] In another aspect, a method includes providing input information in an electronic format, converting at least a portion of the electronic input information to an optical input vector, optically converting the optical input vector to an optical output vector based on an optical matrix multiplication, converting the optical output vector to an electronic format, and electronically applying a nonlinear transformation to the electronically converted optical output vector to provide output information in an electronic format.

[0032] Embodiments of the method may include one or more of the following features. For example, the method may further include repeating the electrical-to-optical conversion, the optical conversion, the optical-to-electrical conversion, and the electronically applied nonlinear conversion on new electronic input information corresponding to the provided output information in electronic format.

[0033] In some implementations, the optical matrix multiplication for the first optical transformation and the optical matrix multiplication for the repeated optical transformation may be the same and may correspond to the same layer of the artificial neural network.

[0034] In some implementations, the optical matrix multiplication for the initial optical transformation and the optical matrix multiplication for the repeated optical transformations may be different and may correspond to different layers of the artificial neural network.

[0035] In some implementations, the method may further include repeating the electrical-to-optical conversion, the optical conversion, the optical-to-electrical conversion, and the electronically applied nonlinear conversion for different portions of the electronic input information, wherein the optical matrix multiplication for the first optical conversion and the optical matrix multiplication for the repeated optical conversions are the same and correspond to a first layer of the artificial neural network.

[0036] In some implementations, the method may further include providing intermediate information in electronic format based on electronic output information produced by the first layer of the artificial neural network for multiple portions of the electronic input information, and repeating electrical-to-optical conversion, optical conversion, optical-to-electrical conversion, and electronically applied nonlinear conversion for each different portion of the electronic intermediate information, wherein the optical matrix multiplication for the initial optical conversion and the optical matrix multiplication of the repeated optical conversions for the different portions of the electronic intermediate information are the same and correspond to a second layer of the artificial neural network.

[0037] In another aspect, the system includes an optical processor including a passive diffractive optical element configured to convert an optical input vector or matrix into an optical output vector or matrix representing the result of matrix processing applied to a predetermined vector defined by the optical input vector or matrix and the arrangement of the diffractive optical element.

[0038] Embodiments of the system may include one or more of the following features: For example, the matrix processing may include matrix multiplication of an optical input vector or matrix with a predetermined vector defined by the arrangement of diffractive optical elements.

[0039] In some implementations, the optical processor may include an optical matrix processing unit including an array of input waveguides for receiving an optical input vector, an optical interference unit comprising a passive diffractive optical element, the optical interference unit in optical communication with the array of input waveguides and configured to perform a linear conversion of the optical input vector into a second array of optical signals, and an array of output waveguides in optical communication with the optical interference unit for directing the second array of optical signals, wherein at least one input waveguide in the array of input waveguides is in optical communication with each output waveguide in the array of output waveguides via the optical interference unit.

[0040] In some implementations, the optical interference unit may include a substrate having at least one of holes or stripes, where the holes have dimensions in the range of 100 nm to 10 μm and the width of the stripes is in the range of 100 nm to 10 μm.

[0041] In some implementations, the optical interference unit may include a substrate having passive diffractive optical elements arranged in a two-dimensional configuration, the substrate comprising at least one of a planar substrate or a curved substrate.

[0042] In some implementations, the substrate may include a planar substrate parallel to the direction of light propagation from the array of input waveguides to the array of output waveguides.

[0043] In some implementations, the optical processor may include an optical matrix processing unit including a matrix of input waveguides for receiving an optical input matrix, an optical interference unit comprising a passive diffractive optical element, the optical interference unit in optical communication with the matrix of input waveguides and configured to perform a linear transformation of the optical input matrix into a second matrix of optical signals, and a matrix of output waveguides in optical communication with the optical interference unit for deriving the second matrix of optical signals, wherein at least one input waveguide in the matrix of input waveguides is in optical communication with each output waveguide in the matrix of output waveguides via the optical interference unit.

[0044] In some implementations, the optical interference unit may include a substrate having at least one of holes or stripes, where the holes have dimensions in the range of 100 nm to 10 μm and the width of the stripes is in the range of 100 nm to 10 μm.

[0045] In some implementations, the optical interference unit can include a substrate having passive diffractive optical elements arranged in a three-dimensional configuration.

[0046] In some implementations, the substrate may have the shape of at least one of a cube, a cylinder, a prism, or an irregular volume.

[0047] In some implementations, the optical processor may include an optical interference unit including a hologram having a passive diffractive optical element, the optical processor configured to receive modulated light representing an optical input matrix and continuously transform the light as it passes through the hologram until the light emerges from the hologram as an optical output matrix.

[0048] In some implementations, the optical interference unit may include a substrate having a passive diffractive optical element, the substrate comprising at least one of silicon, silicon oxide, silicon nitride, quartz, lithium niobate, a phase change material, or a polymer.

[0049] In some implementations, the optical interference unit may include a substrate having a passive diffractive optical element, the substrate comprising at least one of a glass substrate or an acrylic substrate.

[0050] In some implementations, the passive diffractive optical element may be formed in part by a dopant.

[0051] In some implementations, matrix processing may refer to the processing of input data represented by optical input vectors by a neural network.

[0052] In some implementations, the optical processor may include a laser unit configured to generate a plurality of optical outputs; a plurality of optical modulators coupled to the laser unit and configured to generate an optical input vector by modulating the plurality of optical outputs generated by the laser unit based on a plurality of modulator control signals; an optical matrix processing unit coupled to the plurality of optical modulators, the optical matrix processing unit comprising a passive diffractive optical element configured to convert the optical input vector into an optical output vector based on a plurality of weights defined by the passive diffractive optical element; and an optical detection unit coupled to the optical matrix processing unit and configured to generate a plurality of output electrical signals corresponding to the optical output vector.

[0053] In some implementations, the passive diffractive optical elements may be arranged in a three-dimensional configuration, the plurality of light modulators comprises a two-dimensional array of light modulators, and the light detection unit comprises a two-dimensional array of light detectors.

[0054] In some implementations, the optical matrix processing unit may include a housing module for supporting and protecting the array of input waveguides, the optical interference unit, and the array of output waveguides, and the optical processor comprises a receiving module configured to receive the optical matrix processing unit, the receiving module comprising a first interface for enabling the optical matrix processing unit to receive optical input vectors from the multiple optical modulators and a second interface for enabling the optical matrix processing unit to transmit optical output vectors to the optical detection unit.

[0055] In some implementations, the plurality of output electrical signals may include at least one of a plurality of voltage signals or a plurality of current signals.

[0056] In some implementations, the system may include a memory unit; a digital-to-analog converter (DAC) unit configured to generate a plurality of modulator control signals; an analog-to-digital conversion (ADC) unit coupled to the light detection unit and configured to convert the plurality of output electrical signals into a plurality of digitized outputs; and a controller including an integrated circuit configured to perform operations including receiving, from a computer, an artificial neural network computation request comprising an input dataset, the input dataset comprising a first digital input vector; storing the input dataset in the memory unit; and generating, through the DAC unit, a first plurality of modulator control signals based on the first digital input vector.

[0057] In another aspect, a method includes 3D printing an optical matrix processing unit comprising a passive diffractive optical element configured to convert an optical input vector or matrix into an optical output vector or matrix representing the result of matrix processing applied to a predetermined vector defined by the optical input vector or matrix and the arrangement of the diffractive optical element.

[0058] In another aspect, a method includes using one or more laser beams to generate a hologram comprising a passive diffractive optical element configured to convert an optical input vector or matrix into an optical output vector or matrix representing the result of matrix processing applied to a predetermined vector defined by the optical input vector or matrix and the arrangement of the diffractive optical element.

[0059] In another aspect, a system includes an optical processor comprising passive diffractive optical elements arranged in one dimension, the passive diffractive optical elements configured to convert an optical input into an optical output representing the result of matrix processing applied to a predetermined vector defined by the optical input and the arrangement of the diffractive optical elements.

[0060] Implementations of the system may include one or more of the following features: For example, the matrix processing may include a matrix multiplication of the optical input and a predetermined vector defined by the arrangement of the diffractive optical elements.

[0061] In some implementations, the optical processor may include an optical matrix processing unit including an input waveguide for receiving an optical input, an optical interference unit comprising a passive diffractive optical element, the optical interference unit in optical communication with the input waveguide and configured to perform a linear transformation of the optical input, and an output waveguide in optical communication with the optical interference unit for directing an optical output.

[0062] In some implementations, the optical interference unit may include a substrate having at least one of holes or gratings, and the hole or grating elements may have dimensions in the range of 100 nm to 10 μm.

[0063] In another aspect, a system includes an optical processor including a memory unit; a digital-to-analog converter (DAC) unit configured to generate a plurality of modulator control signals; an optical processor including a laser unit configured to generate a plurality of optical outputs; a plurality of optical modulators coupled to the laser unit and the DAC unit, the plurality of optical modulators configured to generate an optical input vector by modulating the plurality of optical outputs generated by the laser unit based on the plurality of modulator control signals; an optical matrix processing unit coupled to the plurality of optical modulators, the optical matrix processing unit comprising a passive diffractive optical element configured to convert the optical input vector into an optical output vector based on a plurality of weights defined by the passive diffractive optical element; and an optical detection unit coupled to the optical matrix processing unit and configured to generate a plurality of output electrical signals corresponding to the optical output vector. The system further includes an analog-to-digital conversion (ADC) unit coupled to the light detection unit and configured to convert the plurality of output electrical signals into a plurality of digitized optical outputs; and a controller including an integrated circuit configured to perform operations including receiving, from the computer, an artificial neural network computation request comprising an input dataset, the input dataset comprising a first digital input vector; storing the input dataset in a memory unit; and generating, through the DAC unit, a first plurality of modulator control signals based on the first digital input vector.

[0064] Embodiments of the system may include one or more of the following features: For example, the matrix processing unit may include a passive diffractive optical element configured to convert an optical input vector into an optical output vector representing a product of a matrix multiplication of a digital input vector and a predetermined vector defined by the passive diffractive optical element.

[0065] In some implementations, the operations further include obtaining, from the ADC unit, a first plurality of digitized optical outputs corresponding to an optical output vector of the optical matrix processing unit, wherein the first plurality of digitized optical outputs form a first digital output vector; performing a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector; and storing the first transformed digital output vector in the memory unit.

[0066] In some implementations, the system may have a first loop period defined as the time elapsed between storing the input data set in the memory unit and storing the first converted digital output vector in the memory unit, and the first loop period may be 1 ns or less.

[0067] In some implementations, the operations may further include outputting an artificial neural network output generated based on the first transformed digital output vector.

[0068] In some implementations, the operations may further include generating, through a DAC unit, a second plurality of modulator control signals based on the first converted digital output vector.

[0069] In some implementations, the input data set may further include a second digital input vector, and the operations may further include generating, through the DAC unit, a second plurality of modulator control signals based on the second digital input vector; obtaining, from the ADC unit, a second plurality of digitized optical outputs corresponding to the optical output vector of the optical matrix processing unit, wherein the second plurality of digitized optical outputs form the second digital output vector; performing a nonlinear transformation on the second digital output vector to generate a second transformed digital output vector; storing the second transformed digital output vector in a memory unit; and outputting an artificial neural network output generated based on the first transformed digital output vector and the second transformed digital output vector, wherein the optical output vector of the optical matrix processing unit results from the second optical input vector generated based on the second plurality of modulator control signals transformed by the optical matrix processing unit based on a plurality of weights defined by the passive diffractive optical element.

[0070] In some implementations, the system may further include an analog nonlinearity unit disposed between the photodetection unit and the ADC unit, the analog nonlinearity unit configured to receive the plurality of output electrical signals from the photodetection unit, apply a nonlinear transfer function, and output a plurality of converted output electrical signals to the ADC unit, and the operations may further include obtaining, from the ADC unit, a first plurality of converted digitized output electrical signals corresponding to the plurality of converted output electrical signals, the first plurality of converted digitized output electrical signals forming a first converted digital output vector; and storing the first converted digital output vector in the memory unit.

[0071] In some implementations, the controller integrated circuit may be configured to generate the first plurality of modulator control signals at a rate of 8 GHz or greater.

[0072] In some implementations, the system may further include an analog memory unit disposed between the DAC unit and the plurality of optical modulators, the analog memory unit configured to store analog voltages and output the stored analog voltages; and an analog nonlinearity unit disposed between the photodetection unit and the ADC unit, the analog nonlinearity unit configured to receive the plurality of output electrical signals from the photodetection unit, apply a nonlinear transfer function, and output a plurality of converted output electrical signals.

[0073] In some implementations, the analog memory unit may include multiple capacitors.

[0074] In some implementations, the analog memory unit may be configured to receive and store a plurality of converted output electrical signals of the analog nonlinearity unit and output the stored plurality of converted output electrical signals to a plurality of optical modulators, and the operations may further include storing the plurality of converted output electrical signals of the analog nonlinearity unit in the analog memory unit based on generating the first plurality of modulator control signals; outputting the stored converted output electrical signals through the analog memory unit; obtaining a second plurality of converted digitized output electrical signals from the ADC unit, where the second plurality of converted digitized output electrical signals form a second converted digital output vector; and storing the second converted digital output vector in the memory unit.

[0075] In some implementations, the input dataset of the artificial neural network computation request may include a plurality of digital input vectors, the laser unit may be configured to generate a plurality of wavelengths, and the plurality of optical modulators may include a bank of optical modulators configured to generate a plurality of optical input vectors, each of the bank corresponding to one of the plurality of wavelengths and generating a respective optical input vector having a respective wavelength, and an optical multiplexer configured to combine the plurality of optical input vectors into a combined optical input vector comprising the plurality of wavelengths. The optical detection unit may be further configured to demultiplex the plurality of wavelengths and generate a plurality of demultiplexed output electrical signals, and the operations may include obtaining a plurality of digitized demultiplexed optical outputs from the ADC unit, where the plurality of digitized demultiplexed optical outputs form a plurality of first digital output vectors, each of the plurality of first digital output vectors corresponding to one of the plurality of wavelengths; performing a nonlinear transformation on each of the plurality of first digital output vectors to generate a plurality of transformed first digital output vectors; and storing the plurality of transformed first digital output vectors in the memory unit, where each of the plurality of digital input vectors may correspond to one of the plurality of optical input vectors.

[0076] In some implementations, the artificial neural network computation request may include a plurality of digital input vectors, the laser unit configured to generate a plurality of wavelengths, the plurality of optical modulators may include a bank of optical modulators configured to generate a plurality of optical input vectors, each of the bank corresponding to one of the plurality of wavelengths and generating a respective optical input vector having the respective wavelength, and an optical multiplexer configured to combine the plurality of optical input vectors into a combined optical input vector comprising the plurality of wavelengths. The operations may include obtaining, from the ADC unit, a first plurality of digitized optical outputs corresponding to an optical output vector comprising the plurality of wavelengths, the first plurality of digitized optical outputs forming a first digital output vector, performing a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector, and storing the first transformed digital output vector in the memory unit.

[0077] In some implementations, the DAC unit may include a 1-bit DAC unit configured to generate multiple 1-bit modulator control signals, where the ADC unit may have a resolution of 1 bit and the first digital input vector may have a resolution of N bits. The operations may include decomposing the first digital input vector into N 1-bit input vectors, where each of the N 1-bit input vectors corresponds to one of the N bits of the first digital input vector, generating, via the 1-bit DAC unit, a sequence of N 1-bit modulator control signals corresponding to the N 1-bit input vectors, obtaining, from the ADC unit, a sequence of N digitized 1-bit optical outputs corresponding to the sequence of the N 1-bit modulator control signals, constructing an N-bit digital output vector from the sequence of N digitized 1-bit optical outputs, performing a nonlinear transformation on the constructed N-bit digital output vector to generate a transformed N-bit digital output vector, and storing the transformed N-bit digital output vector in a memory unit.

[0078] In some implementations, the memory unit may include a digital input vector memory configured to store the digital input vector and comprising at least one SRAM.

[0079] In some implementations, the laser unit may include a laser source configured to generate light and an optical power splitter configured to split the light generated by the laser source into a plurality of optical outputs, each of the plurality of optical outputs having substantially equal power.

[0080] In some implementations, the plurality of optical modulators may include one of an MZI modulator, a ring resonator modulator, or an electro-absorption modulator.

[0081] In some implementations, the light detection unit may include a plurality of light detectors and a plurality of amplifiers configured to convert photocurrents generated by the light detectors into a plurality of output electrical signals.

[0082] In some implementations, the integrated circuit may include an application specific integrated circuit.

[0083] In some implementations, the optical matrix processing unit may include an array of input waveguides for receiving an optical input vector, an optical interference unit in optical communication with the array of input waveguides for performing a linear conversion of the optical input vector into a second array of optical signals, the optical interference unit comprising a passive diffractive optical element, and an array of output waveguides in optical communication with the optical interference unit for directing the second array of optical signals, wherein at least one input waveguide in the array of input waveguides is in optical communication with each output waveguide in the array of output waveguides via the optical interference unit.

[0084] In another aspect, a system includes an optical processor including a memory unit, a driver unit configured to generate a plurality of modulator control signals, an optical processor including a laser unit configured to generate a plurality of optical outputs, a plurality of optical modulators coupled to the laser unit and the driver unit, the plurality of optical modulators configured to generate an optical input vector by modulating the plurality of optical outputs generated by the laser unit based on the plurality of modulator control signals, an optical matrix processing unit coupled to the plurality of optical modulators and the driver unit, the optical matrix processing unit comprising a passive diffractive optical element configured to convert the optical input vector into an optical output vector based on a plurality of weight control signals defined by the passive diffractive optical element, and an optical detection unit coupled to the optical matrix processing unit and configured to generate a plurality of output electrical signals corresponding to the optical output vector.The system also includes a comparator unit coupled to the light detection unit and configured to convert the plurality of output electrical signals into a plurality of digitized 1-bit optical outputs; and a controller including an integrated circuit configured to perform operations including receiving from the computer an artificial neural network computation request comprising an input dataset, the input dataset comprising a first digital input vector having N-bit resolution; storing the input dataset in the memory unit; decomposing the first digital input vector into N 1-bit input vectors, each of the N 1-bit input vectors corresponding to one of the N bits of the first digital input vector; generating through the driver unit a sequence of N 1-bit modulator control signals corresponding to the N 1-bit input vectors; obtaining from the comparator unit a sequence of N digitized 1-bit optical outputs corresponding to the sequence of N 1-bit modulator control signals; constructing an N-bit digital output vector from the sequence of N digitized 1-bit optical outputs; performing a nonlinear transformation on the constructed N-bit digital output vector to generate a converted N-bit digital output vector; and storing the converted N-bit digital output vector in the memory unit.

[0085] Embodiments of the system may include one or more of the following features: For example, the optical matrix processing unit may include an optical matrix multiplication unit configured to convert the optical input vector into an optical output vector representing a product of a matrix multiplication of an input vector represented by the optical input vector and a predetermined vector defined by the diffractive optical element.

[0086] In another aspect, a method for performing an artificial neural network computation in a system having an optical matrix processing unit includes receiving, from a computer, an artificial neural network computation request comprising an input data set comprising a first digital input vector; storing the input data set in a memory unit; generating, through a digital-to-analog converter (DAC) unit, a first plurality of modulator control signals based on the first digital input vector; and converting the optical input vector into an optical output vector by using the optical matrix processing unit comprising an arrangement of diffractive optical elements, wherein the optical output vector is defined by the optical input vector and the arrangement of diffractive optical elements. obtaining, from an analog-to-digital conversion (ADC) unit, a first plurality of digitized optical outputs corresponding to the optical output vector of the optical matrix processing unit, the first plurality of digitized optical outputs forming a first digital output vector; performing, by a controller, a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector; storing, in a memory unit, the first transformed digital output vector; and outputting, by the controller, an artificial neural network output generated based on the first transformed digital output vector.

[0087] Embodiments of the method may include one or more of the following features. For example, converting the optical input vector to an optical output vector may include converting the optical input vector to an optical output vector that represents a product of a matrix multiplication of the digital input vector and a predetermined vector defined by the arrangement of diffractive optical elements.

[0088] In another aspect, a method includes providing input information in an electronic format, converting at least a portion of the electronic input information into an optical input vector, optically converting the optical input vector into an optical output vector based on optical matrix processing by an optical processor comprising a passive diffractive optical element, converting the optical output vector into an electronic format, and electronically applying a nonlinear transformation to the electronically converted optical output vector to provide output information in an electronic format.

[0089] Embodiments of the method may include one or more of the following features. For example, optically converting the optical input vector to the optical output vector may include optically converting the optical input vector to the optical output vector based on an optical matrix multiplication of a digital input vector represented by the optical input vector and a predetermined vector defined by a passive diffractive optical element.

[0090] In some implementations, the method may further include repeating the electrical-to-optical conversion, the optical conversion, the optical-to-electrical conversion, and the electronically applied nonlinear conversion on new electronic input information corresponding to the provided output information in electronic format.

[0091] In some implementations, the optical matrix processing for the initial optical transformation and the optical matrix processing for the repeated optical transformations may be the same and may correspond to the same layer of the artificial neural network.

[0092] In some implementations, the method may further include repeating the electrical-to-optical conversion, the optical conversion, the optical-to-electrical conversion, and the electronically applied nonlinear conversion for different portions of the electronic input information, wherein the optical matrix processing for the initial optical conversion and the optical matrix processing for the repeated optical conversions are the same and may correspond to a layer of the artificial neural network.

[0093] In another aspect, the system includes an optical matrix processing unit configured to process an input vector of length N, the optical matrix processing unit comprising N+2 layers of directional couplers and N layers of phase shifters, where N is a positive integer.

[0094] Embodiments of the system may include one or more of the following features: For example, the optical matrix processing unit may include only N+2 layers of directional couplers.

[0095] In some implementations, the optical matrix processing unit may include an optical matrix multiplication unit.

[0096] In some implementations, the optical matrix processing unit may include a substrate and interconnected interferometers disposed on the substrate, each interferometer comprising an optical waveguide disposed on the substrate, and the directional coupler and phase shifter being part of the interconnected interferometers.

[0097] In some implementations, the optical matrix processing unit may include a layer of attenuators following the last layer of directional couplers.

[0098] In some implementations, a layer of attenuators may include N attenuators.

[0099] In some implementations, the system may include one or more homodyne detectors for detecting the output from the attenuator.

[0100] In some implementations, N=3, and the optical matrix processing unit may include an input terminal configured to receive an input vector, a first layer of directional couplers coupled to the input terminal, a first layer of phase shifters coupled to the first layer of directional couplers, a second layer of phase shifters coupled to the first layer of phase shifters, a second layer of phase shifters coupled to the second layer of directional couplers, a third layer of directional couplers coupled to the second layer of phase shifters, a third layer of phase shifters coupled to the third layer of directional couplers, a fourth layer of directional couplers coupled to the third layer of phase shifters, and a fifth layer of directional couplers coupled to the fourth layer of directional couplers.

[0101] In some implementations, N=4, and the optical matrix processing unit may include an input terminal configured to receive an input vector, and first, second, third, and fourth layers of directional couplers, each followed by a layer of phase shifters, where the first layer of directional couplers is coupled to the input terminal, the second to last layers of directional couplers are coupled to the fourth layer of phase shifters, and the last layer of directional couplers is coupled to the second to last layers of directional couplers.

[0102] In some implementations, N=8, and the optical matrix processing unit may include an input terminal configured to receive an input vector and eight layers of directional couplers each followed by a layer of phase shifters, with the first layer of directional couplers coupled to the input terminal, the second to last layers of directional couplers coupled to the eighth layer of phase shifters, and the final layer of directional couplers coupled to the second to last layers of directional couplers.

[0103] In some implementations, the optical matrix multiplication unit may include an input terminal configured to receive an input vector and N layers of directional couplers each followed by a layer of phase shifters, with the first layer of directional couplers coupled to the input terminal, the second to last layers of directional couplers coupled to the Nth layer of phase shifters, and the final layer of directional couplers coupled to the second to last layers of directional couplers.

[0104] In some implementations, N is an even number.

[0105] In some implementations, each i-th layer of directional couplers includes N / 2 directional couplers, where i is an odd number, and each j-th layer of directional couplers includes N / 2-1 directional couplers, where j is an even number.

[0106] In some implementations, for each ith layer of directional couplers, where i is an odd number, the kth directional coupler may be coupled to the (2k-1)th output and the 2kth output of the previous layer, where k is an integer from 1 to N / 2.

[0107] In some implementations, for each jth layer of directional couplers, where j is an even integer, the mth directional coupler may be coupled to the 2mth output and the (2m+1)th output of the previous layer, where m is an integer from 1 to N / 2-1.

[0108] In some implementations, each of the i-th layers of phase shifters may include N phase shifters, where i is an odd number, and each of the j-th layer of phase shifters may include N-2 phase shifters, where j is an even number.

[0109] In some implementations, N may be an odd number.

[0110] In some implementations, each layer of directional couplers may include (N-1) / 2 directional couplers.

[0111] In some implementations, each layer of phase shifters may include N-1 phase shifters.

[0112] In another aspect, a system includes a generator configured to generate a first dataset, the generator comprising an optical matrix processing unit; and a discriminator configured to receive a second dataset comprising data from the first dataset and data from a third dataset, wherein the data in the first dataset has similar characteristics to the characteristics of the data in the third dataset, and to classify the data in the second dataset as data from the first dataset or data from the third dataset.

[0113] Embodiments of the method may include one or more of the following features: For example, the optical matrix processing unit may include at least one of (i) the optical matrix multiplication unit described above, (ii) the passive diffractive optical element described above, or (iii) the optical matrix processing unit described above.

[0114] In some implementations, the third data set may include real data, the generator is configured to generate synthetic data that resembles the real data, and the discriminator is configured to classify the data as real data or synthetic data.

[0115] In some implementations, the generator may be configured to generate a dataset for training at least one of an autonomous vehicle, a medical diagnostic system, a fraud detection system, a weather forecasting system, a financial prediction system, a facial recognition system, a speech recognition system, or a product defect detection system.

[0116] In some implementations, the generator may be configured to generate an image that resembles an image of at least one of a real object or a real scene, and the discriminator is configured to classify the received image as (i) an image of the real object or the real scene, or (ii) a synthetic image generated by the generator.

[0117] In some implementations, the real objects may include at least one of people, animals, cells, tissues, or products, and the real scene comprises a scene encountered by the vehicle.

[0118] In some implementations, the discriminator may be configured to classify whether the received image is (i) an image of real people, real animals, real cells, real tissues, real products, or real scenes encountered by a vehicle, or (ii) a synthetic image generated by the generator.

[0119] In some implementations, the vehicle may include at least one of a motorcycle, a car, a truck, a train, a helicopter, an airplane, a submarine, a ship, or a drone.

[0120] In some implementations, the generator may be configured to generate an image of a tissue or cell associated with at least one of a human disease, an animal disease, or a plant disease.

[0121] In some implementations, the generator may be configured to generate an image of tissue or cells associated with a disease in a person, the disease comprising at least one of cancer, Parkinson's disease, sickle cell anemia, heart disease, circulatory disease, diabetes, chest disease, or skin disease.

[0122] In some implementations, the generator may be configured to generate an image of tissue or cells associated with cancer, which may include at least one of skin cancer, breast cancer, lung cancer, liver cancer, prostate cancer, or brain cancer.

[0123] In some implementations, the system may further include a random noise generator configured to generate random noise that is provided as input to the generator, the generator configured to generate the first data set based on the random noise.

[0124] In another aspect, a system includes a random noise generator configured to generate random noise and a generator configured to generate data based on the random noise, the generator comprising an optical matrix processing unit.

[0125] Embodiments of the system may include one or more of the following features: For example, the optical matrix processing unit may include at least one of (i) the optical matrix multiplication unit described above, (ii) the passive diffractive optical element described above, or (iii) the optical matrix processing unit described above.

[0126] In another aspect, a system includes a photonics circuit configured to perform a logic function on two input signals, the photonics circuit including: a first directional coupler having two input terminals and two output terminals, the two input terminals configured to receive the two input signals; a first pair of phase shifters configured to modify the phase of the signals at the two output terminals of the first directional coupler; a second directional coupler having two input terminals and two output terminals, the two input terminals configured to receive signals from the first pair of phase shifters; and a second pair of phase shifters configured to modify the phase of the signals at the two output terminals of the second directional coupler.

[0127] Embodiments of the method may include one or more of the following features: For example, the phase shifter may include a rotating phase shifter in the photonics circuit.

[0128]

number

[0129] The method may be configured to perform the following steps:

[0130] In some implementations, when input signals x1 and x2 are provided to two input terminals of the first directional coupler, the phase shifter causes the photonics circuit to

[0131]

number

[0132] The method may be configured to perform the following steps:

[0133] In some implementations, the photonics circuit includes a rotating

[0134]

number

[0135] a first photodetector configured to generate an absolute value of the signal from the second pair of phase shifters to perform

[0136] In some implementations, the photonics circuit compares the output signal of the first photodetector with a threshold to generate a binary value and output the binary value to the photonics circuit.

[0137]

number

[0138] The signal may include a comparator configured to generate:

[0139] In some implementations, the photonics circuit may include a circuit in which an output signal of the photodetector is fed back to an input terminal of the first directional coupler, passed through the first directional coupler, the first pair of phase shifters, the second directional coupler, and the second pair of phase shifters, detected by the photodetector, and used to generate an operation in the photonics circuit.

[0140]

number

[0141] which produces AND(x1, x2) and OR(x1, x2).

[0142] In some implementations, the photonics circuit includes a third directional coupler having two input terminals and two output terminals, the two input terminals configured to receive signals from the second pair of phase shifters, a third pair of phase shifters configured to modify the phase of the signal at the two output terminals of the third directional coupler, a fourth directional coupler having two input terminals and two output terminals, the two input terminals configured to receive signals from the third pair of phase shifters, a fourth pair of phase shifters configured to modify the phase of the signal at the two output terminals of the fourth directional coupler, and a fourth pair of phase shifters configured to generate an absolute value of the signal from the fourth pair of phase shifters and perform an operation on the photonics circuit.

[0143]

number

[0144] and a second photodetector configured to perform the operation AND(x1, x2) and OR(x1, x2).

[0145] In some implementations, the system may include a bitonic sorter configured such that the sorting function of the bitonic sorter is performed using photonics circuitry.

[0146] In some implementations, the system may include a device configured to perform a hash function using photonics circuitry.

[0147] In some implementations, the hash function may include Secure Hash Algorithm 2 (SHA-2).

[0148] In general, systems for performing computations produce computational results using different types of operations, each performed on a signal (e.g., an electrical or optical signal), where the underlying physical properties of the operation (e.g., in terms of energy consumption and / or speed) are best suited for that signal. For example, three such operations are replication, addition, and multiplication. As explained in more detail below, replication may be performed using optical power division, addition may be performed using current-based summation, and multiplication may be performed using optical amplitude modulation. An example of a computation that may be performed using these three types of operations is multiplying a vector by a matrix (e.g., as utilized by artificial neural network computations). Various other computations may be performed using these operations, which represent a general set of linear operations by which various computations can be performed, including, but not limited to, vector-vector dot product, vector-vector element-wise multiplication, vector-scalar element-wise multiplication, or matrix-matrix element-wise multiplication. Some of the examples described herein show techniques and configurations for vector-matrix multiplication, but corresponding techniques and configurations may be used for any of these types of computations.

[0149] Embodiments may have one or more of the following features.

[0150] Optoelectronic computing systems using both electrical and optical signals as described herein may facilitate increased flexibility and / or efficiency. In the past, there have been potential challenges associated with combining optical (or photonics) integrated devices with electrical (or electronic) integrated devices on a common platform (e.g., a common semiconductor die, or multiple semiconductor dies combined in a controlled collapsed chip connection or "flip-chip" configuration). Such potential challenges may include, for example, input / output (I / O) packaging or temperature control. In systems such as those described herein, this potential challenge may increase when used with a relatively large number of optical input ports and a relatively large number of electrical output ports (e.g., four or more optical input / output ports, 200 or more electrical input / output ports). These potential challenges may be mitigated using appropriate system design. For example, the system may use a high-density packaging configuration that controls thermal expansion between different material types (e.g., semiconductor materials such as silicon, glass materials such as silicon dioxide or "silica," ceramic materials, etc.) using an encapsulated housing that acts as a temperature control (e.g., thermoelectric cooling) and / or heat sink and provides some degree of sealing. Such temperature stabilization techniques can limit different coefficients of thermal expansion (CTE) and the resulting misalignment between the system ports and the ports of the packaged high-density fiber array.

[0151] For the replication operation, the splitting of optical power is passive, so no electrical power needs to be consumed to perform the operation. In addition, the frequency bandwidth of an electrical splitter has a limit associated with the RC time constant. In comparison, the frequency bandwidth of an optical splitter is virtually infinite. As described in more detail below, different types of optical power splitters can be used, including waveguide optical splitters or free-space beam splitters.

[0152] For a multiplication operation, one value may be encoded as an optical signal and the other value may be encoded as an amplitude scaling factor (e.g., multiplication by a value ranging from 0 to 1). After the scaling factor is set, multiplication in the optical domain has less (or no) requirement for electrical signal conditioning and is therefore less constrained by electrical noise, power consumption, and bandwidth limitations. By appropriate selection of the detection scheme, a signed result can be obtained (e.g., multiplication by a value between -1 and +1), as explained in more detail below.

[0153] For summation operations, different techniques can be used to achieve a result in which the magnitude of the current flowing through the conductors is determined based on the sum of different contributions. In situations where a current signal is incoming, when two or more conductors carrying those incoming current signals overlap at some intersection, the single conductor carrying the outgoing current signal represents the sum of those input current signals. In situations where an optical signal is incoming, two or more light waves of different wavelengths impinge on a detector, and the current signal carried on the photocurrents produced by the detector represents the sum of the power in the incoming optical signals. Both produce an electrical signal (e.g., current) as an output representing the sum, but one uses current as an input (current-input-based summation, also known as "electrical summation" performed in the "electrical domain"), while the other uses light waves as an input (optical-input-based summation, also known as "optoelectronic summation" performed in the "optoelectronic domain"). However, in some embodiments, current-input-based summation is used instead of optical-input-based summation, which allows a single optical wavelength to be used in the system, eliminating potentially complex elements of the system that may be required to provide and maintain multiple wavelengths.

[0154] Combinations of these basic operations performed by these modules can be arranged to provide devices that perform linear operations, such as vector-matrix multiplication, with arbitrary matrix element magnitudes. Other implementations of matrix multiplication using optical signals and interferometers for combining the optical signals using optical interference (without using replication or addition modules as described herein) have been limited to providing vector-matrix multiplication with some constraints (e.g., unitary or diagonal matrices). In addition, some other implementations may rely on large-scale phase alignment of multiple optical signals when they propagate through a relatively large number of optical elements (e.g., optical modulators). Alternatively, implementations described herein may be able to relax such phase alignment constraints by converting optical signals to electrical signals after propagation through fewer optical elements (e.g., after propagation through only a single optical amplitude modulator), which allows the use of optical signals with reduced coherence or even incoherent optical signals for optical modulators that do not rely on constructive / destructive interference.

[0155] For time-domain encoding of optical and electrical signals, as described in more detail below, analog electronic circuits may be optimized for operation at specific power levels, which may be useful when the circuits are operating at high speeds. Such time-domain encoding may be useful, for example, to reduce any challenges that may be associated with precisely controlling a relatively large number of distinct intensity levels for each symbol. Instead, a relatively constant amplitude may be used (for the "on" level) (at or near zero amplitude for the "off" level), while precise control of the duty cycle is applied in the time domain across multiple time slots within a single symbol duration.

[0156] Modules can be conveniently manufactured on a large scale by integrating optical and electronic devices on a common substrate (e.g., a silicon chip). Routing signals on the substrate as optical signals instead of electrical signals and consolidating photodetectors on one portion of the substrate can help eliminate long electronic wiring and its associated challenges (e.g., parasitic capacitance, inductance, and crosstalk).

[0157] For embodiments of the system using submatrix multiplication, each element of the output vector may be calculated simultaneously using a different device (e.g., a different core, a different processor, a different computer, a different server), helping to alleviate any potential constraints such as memory walls and helping the overall system scale for very large matrices. In some embodiments, each submatrix may be multiplied by a corresponding subvector using a different device. A total sum may then be calculated by collecting or accumulating the summands from the different devices. Advantageously, intermediate results in the form of optical signals can be transported between devices, even if the devices are separated by relatively long distances.

[0158] Other aspects include other combinations of the above-listed features and other features, expressed as methods, apparatus, systems, program products, and in other ways.

[0159] Particular embodiments of the subject matter described herein may be implemented to achieve one or more of the following advantages: The throughput, latency, or both of the ANN computation may be improved. The power efficiency of the ANN computation may be improved.

[0160] In another aspect, an apparatus includes: a plurality of optical waveguides, wherein a set of multiple input values ​​is encoded onto respective optical signals carried by the optical waveguides; a plurality of replication modules, wherein for each of at least two subsets of the one or more optical signals, a corresponding set of one or more of the replication modules is configured to divide the subset of the one or more optical signals into two or more replicas of the optical signals; a plurality of multiplication modules, wherein for each of at least two replicas of a first subset of the one or more optical signals, a corresponding multiplication module is configured to multiply one or more optical signals of the first subset with one or more matrix element values ​​using optical amplitude modulation, at least one of the multiplication modules includes an optical amplitude modulator including an input port and two output ports, and pairs of associated optical signals are provided from the two output ports such that a difference between the amplitudes of the associated optical signals corresponds to a result of multiplying the input values ​​by the signed matrix element values; and one or more adder modules, wherein for two or more results of the multiplication modules, a corresponding one of the adder modules is configured to produce an electrical signal representing a sum of the two or more results of the multiplication modules.

[0161] Embodiments of the apparatus may include one or more of the following features: For example, an input value in the set of multiple input values ​​encoded on each optical signal may represent an element of an input vector that has been multiplied by a matrix that includes one or more matrix element values.

[0162] In some implementations, a set of multiple output values ​​may be encoded onto each electrical signal produced by one or more summation modules, and an output value in the set of multiple output values ​​may represent an element of an output vector that results from multiplying an input vector by a matrix.

[0163] In some implementations, each of the optical signals carried by the optical waveguides may include light waves having a common wavelength that is substantially the same for all of the optical signals.

[0164] In some implementations, the replication modules may include at least one replication module including an optical splitter that sends a predetermined percentage of the power of the light wave at the input port to a first output port and sends the remaining percentage of the power of the light wave at the input port to a second output port.

[0165] In some implementations, the optical splitter may include a waveguide optical splitter that sends a predetermined percentage of the power of the light waves guided by the input optical waveguide to a first output optical waveguide and sends the remaining percentage of the power of the light waves guided by the input optical waveguide to a second output optical waveguide.

[0166] In some implementations, a guided mode of the input optical waveguide can be adiabatically coupled to a guided mode of each of the first and second output optical waveguides.

[0167] In some implementations, the optical splitter may include a beam splitter that includes at least one surface that transmits a predetermined percentage of the power of the light wave at the input port and reflects the remaining percentage of the power of the light wave at the input port.

[0168] In some implementations, at least one of the plurality of optical waveguides may include an optical fiber coupled to an optical coupler that couples a guided mode of the optical fiber to a propagation mode in free space.

[0169] In some implementations, the multiplication module may include at least one coherence-susceptible multiplication module configured to multiply one or more optical signals of the first subset with one or more matrix element values ​​using optical amplitude modulation based on interference between optical waves having a coherence length at least as long as a propagation distance through the coherence-susceptible multiplication module.

[0170] In some implementations, the coherence-susceptible multiplication module may include a Mach-Zehnder interferometer (MZI) that splits light waves guided by an input optical waveguide into a first optical waveguide arm of the MZI and a second optical waveguide arm of the MZI, the first optical waveguide arm including a phase shifter that provides a relative phase shift with respect to a phase delay of the second optical waveguide arm, and the MZI that combines the light waves from the first optical waveguide arm and the second optical waveguide arm into at least one output optical waveguide.

[0171] In some implementations, the MZI can combine lightwaves from the first and second optical waveguide arms into first and second output optical waveguides, respectively; the first photodetector can receive the lightwaves from the first output optical waveguide to generate a first photocurrent; the second photodetector can receive the lightwaves from the second output optical waveguide to generate a second photocurrent; and the result of the coherence-sensitive multiplication module can include the difference between the first photocurrent and the second photocurrent.

[0172] In some implementations, the coherence-susceptible multiplication module may include one or more ring resonators, including at least one ring resonator coupled to a first optical waveguide and at least one ring resonator coupled to a second optical waveguide.

[0173] In some implementations, the first photodetector can receive lightwaves from the first optical waveguide to generate a first photocurrent, the second photodetector can receive lightwaves from the second optical waveguide to generate a second photocurrent, and the result of the coherence-sensitive multiplication module can include the difference between the first photocurrent and the second photocurrent.

[0174] In some implementations, the multiplication modules may include at least one coherence-insensitive multiplication module configured to multiply one or more optical signals of the first subset with one or more matrix element values ​​using optical amplitude modulation based on absorption of energy in a light wave.

[0175] In some implementations, the coherence-insensitive multiplication module may include an electro-absorption modulator.

[0176] In some implementations, the one or more summing modules may include at least one summing module including: (1) two or more input conductors each carrying an electrical signal in the form of an input current, the amplitude of which represents a respective result of a respective one of the multiplication modules; and (2) at least one output conductor carrying an electrical signal in the form of an output current proportional to the sum of the input currents, the electrical signal representing the sum of the respective results.

[0177] In some implementations, the two or more input and output conductors may include wires that meet at one or more intersections between the wires, and the output current may be substantially equal to the sum of the input currents.

[0178] In some implementations, at least a first one of the input currents may be provided in the form of at least one photocurrent generated by at least one photodetector receiving an optical signal generated by a first one of the multiplication modules.

[0179] In some implementations, the first input current may be provided in the form of a difference between two photocurrents generated by different respective photodetectors that receive different optical signals both generated by the first multiplication module.

[0180] In some implementations, one of the replicas of the first subset of one or more optical signals may consist of a single optical signal on which one of the input values ​​is encoded.

[0181] In some implementations, the multiplication module corresponding to the replica of the first subset may multiply the encoded input value with a single matrix element value.

[0182] In some implementations, one of the replicas of the first subset of one or more optical signals may include more than one of the optical signals and fewer than all of the optical signals on which multiple input values ​​are encoded.

[0183] In some implementations, the multiplication modules corresponding to the replicas of the first subset may multiply the encoded input values ​​with different respective matrix element values.

[0184] In some implementations, different multiplication modules corresponding to different respective copies of the first subset of one or more optical signals may be stored in different devices that are in optical communication for transmitting one of the copies of the first subset of one or more optical signals between the different devices.

[0185] In some implementations, two or more of the plurality of optical waveguides, two or more of the plurality of replication modules, two or more of the plurality of multiplication modules, and at least one of the one or more addition modules may be disposed on a common device substrate.

[0186] In some implementations, the device may perform vector-matrix multiplication, where the input vector may be provided as a set of optical signals and the output vector may be provided as a set of electrical signals.

[0187] In some implementations, the device may further include an accumulator that integrates an input electrical signal corresponding to the output of the multiplication module or the addition module, and the input electrical signal may be encoded using time-domain coding that uses on-off amplitude modulation within each of a plurality of time slots, and the accumulator may produce an output electrical signal that is encoded with more than two amplitude levels across the plurality of time slots, corresponding to different duty cycles of the time-domain coding.

[0188] In some implementations, two or more of the multiplication modules each correspond to a different subset of one or more optical signals.

[0189] In some implementations, the apparatus may further include a multiplication module configured, for each replica of a second subset of one or more optical signals that differs from an optical signal in the first subset of one or more optical signals, to multiply the one or more optical signals of the second subset with one or more matrix element values ​​using optical amplitude modulation.

[0190] In another aspect, a method includes encoding a set of multiple input values ​​for respective optical signals; for each of at least two subsets of the one or more optical signals, splitting the subset of the one or more optical signals into two or more replicas of the optical signal using a corresponding set of one or more replica modules; for each of the at least two replicas of a first subset of the one or more optical signals, using a corresponding multiplication module to multiply one or more optical signals of the first subset with one or more matrix element values ​​using optical amplitude modulation, wherein at least one of the multiplication modules includes an optical amplitude modulator including an input port and two output ports, and pairs of associated optical signals are provided from the two output ports such that a difference between the amplitudes of the associated optical signals corresponds to a result of multiplying the input values ​​by the signed matrix element values; and, for results of two or more of the multiplication modules, using a summation module configured to produce an electrical signal representing a sum of the results of the two or more of the multiplication modules.

[0191] In another aspect, a method includes the steps of: encoding a set of input values ​​representing elements of an input vector onto respective optical signals; encoding a set of coefficients representing elements of the matrix as amplitude modulation levels of a set of optical amplitude modulators coupled to the optical signals, at least one of the optical amplitude modulators including an input port and two output ports providing a pair of associated optical signals from the two output ports such that a difference between the amplitudes of the associated optical signals corresponds to a result of multiplying the input value by a signed matrix element value; and encoding a set of output values ​​representing elements of the output vector onto respective electrical signals, at least one of the electrical signals in the form of a current whose amplitude corresponds to the sum of the respective elements of the input vector multiplied by the respective element of the row of the matrix.

[0192] Embodiments of the method may include one or more of the following features: For example, at least one of the optical signals may be provided by a first optical waveguide, and the first optical waveguide may be coupled to an optical splitter that sends a predetermined percentage of the power of the light waves guided by the first optical waveguide to a second output optical waveguide and sends the remaining percentage of the power of the light waves guided by the first optical waveguide to a third optical waveguide.

[0193] In another aspect, an apparatus includes a plurality of optical waveguides that encode a set of input values ​​representing elements of an input vector onto respective optical signals carried by the optical waveguides; a set of optical amplitude modulators coupled to the optical signals that encode a set of coefficients representing elements of a matrix as amplitude modulation levels, at least one of the optical amplitude modulators including an input port and two output ports that provide a pair of associated optical signals from the two output ports such that a difference between the amplitudes of the associated optical signals corresponds to a result of multiplying the input value by a signed matrix element value; and a plurality of summation modules that encode a set of output values ​​representing elements of the output vector onto respective electrical signals, at least one of the electrical signals being in the form of an electric current whose amplitude corresponds to the sum of the respective elements of the input vector multiplied by the respective element of a row of the matrix.

[0194] In another aspect, a method for multiplying an input vector with a given matrix includes encoding a set of input values ​​representing elements of the input vector onto each optical signal of a set of optical signals; coupling a first set of one or more devices to a first set of one or more waveguides that provide a first subset of the set of optical signals; generating a result of multiplying the values ​​encoded on the first subset of the set of optical signals with a first submatrix of the given matrix; coupling a second set of one or more devices to a second set of one or more waveguides that provide a second subset of the set of optical signals; generating a result of multiplying the values ​​encoded on the second subset of the set of optical signals with a second submatrix of the given matrix; and coupling a third set of one or more devices to provide a replica of the first subset of the set of optical signals generated by the first optical splitter. coupling a fourth set of one or more devices to the fourth set of one or more waveguides that provide a replica of the second subset of the set of optical signals produced by the second optical splitter; and generating a result of multiplying the values ​​encoded on the second subset of the set of optical signals by the fourth submatrix of the given matrix, wherein the first, second, third, and fourth submatrices concatenated together form the given matrix, and at least one output value representing an element of an output vector corresponding to the input vector multiplied with the given matrix is ​​encoded onto electrical signals produced by devices in communication with the first set of one or more devices and the second set of one or more devices.

[0195] Embodiments of the method may include one or more of the following features: For example, each pair of sets of the first set of one or more devices, the second set of one or more devices, the third set of one or more devices, and the fourth set of one or more devices may be mutually exclusive.

[0196] In another aspect, an apparatus includes a first set of one or more devices configured to receive a first set of optical signals and generate a result of multiplying values ​​encoded on the first set of optical signals with a first matrix; a second set of one or more devices configured to receive a second set of optical signals and generate a result of multiplying values ​​encoded on the second set of optical signals with a second matrix; a third set of one or more devices configured to receive a third set of optical signals and generate a result of multiplying values ​​encoded on the third set of optical signals with a third matrix; a fourth set of one or more devices configured to receive a fourth set of optical signals and generate a result of multiplying values ​​encoded on the fourth set of optical signals with a fourth matrix; and configurable connection paths between two or more of a first set of one or more devices, a second set of one or more devices, a third set of one or more devices, or a fourth set of one or more devices, wherein the first configuration of the configurable connection paths is configured to (1) provide a replica of the first set of optical signals as at least one of the second set of optical signals, the third set of optical signals, or the fourth set of optical signals, and (2) provide one or more signals from the first set of one or more devices and one or more signals from the second set of one or more devices to a summation module configured to produce an electrical signal representing a sum of values ​​encoded on the signals received by the summation module.

[0197] In another aspect, an apparatus includes a first set of one or more devices configured to receive a first set of optical signals and generate a result based on an optical amplitude modulation of one or more of the optical signals of the first set of optical signals; a second set of one or more devices configured to receive a second set of optical signals and generate a result based on an optical amplitude modulation of one or more of the optical signals of the second set of optical signals; a third set of one or more devices configured to receive a third set of optical signals and generate a result based on an optical amplitude modulation of one or more of the optical signals of the third set of optical signals; and a fourth set of optical signals and generate a result based on an optical amplitude modulation of one or more of the optical signals of the fourth set of optical signals. and a fourth set of one or more devices configured to provide one or more signals from the first set of one or more devices and one or more signals from the second set of one or more devices to a summing module configured to (1) provide a replica of the first set of optical signals as the third set of optical signals, or (2) produce an electrical signal representing a sum of values ​​encoded on the signals received by the summing module.

[0198] Embodiments of the apparatus may include one or more of the following features: For example, each pair of sets of the first set of one or more devices, the second set of one or more devices, the third set of one or more devices, and the fourth set of one or more devices may be mutually exclusive.

[0199] In some implementations, a first configuration of the configurable connection paths may be configured to provide one or more signals from the first set of one or more devices and one or more signals from the second set of one or more devices to an adder module configured to (1) provide a replica of the first set of optical signals as a third set of optical signals, and (2) produce an electrical signal representing the sum of values ​​encoded on at least two different signals received by the adder module.

[0200] In some implementations, a first configuration of configurable connection paths may be configured to provide replicas of the first set of optical signals as a third set of optical signals, and a second configuration of configurable connection paths may be configured to provide one or more signals from the first set of one or more devices and one or more signals from the second set of one or more devices to a summing module configured to produce an electrical signal representing a sum of values ​​encoded on the signals received by the summing module.

[0201] In another aspect, an apparatus includes: a plurality of optical waveguides, wherein a set of multiple input values ​​is encoded onto respective optical signals carried by the optical waveguides; a plurality of replication modules, including, for each of at least two subsets of one or more optical signals, a corresponding set of one or more replication modules configured to divide the subset of one or more optical signals into two or more replicas of the optical signal; a plurality of multiplication modules, including, for each of the at least two replicas of a first subset of the one or more optical signals, a corresponding multiplication module configured to multiply one or more optical signals of the first subset by one or more values ​​using optical amplitude modulation; and one or more summation modules, including, for two or more results of the multiplication modules, an summation module configured to produce an electrical signal representing a sum of results of two or more of the multiplication modules, the result including at least one result encoded onto the electrical signal and derived from one of the replicas of the optical signal that propagated through only a single optical amplitude modulator before being converted to an electrical signal.

[0202] In another aspect, a system includes a first unit configured to generate a plurality of modulator control signals and a processor unit, the processor unit including: a light source configured to provide a plurality of light outputs; a plurality of light modulators coupled to the light source and the first unit, the plurality of light modulators configured to generate a light input vector by modulating the plurality of light outputs provided by the light source based on the plurality of modulator control signals, the light input vector comprising a plurality of optical signals; and a matrix multiplication unit coupled to the plurality of light modulators and the first unit, the matrix multiplication unit configured to convert the light input vector into an analog output vector based on the plurality of weight control signals. The system also includes a second unit coupled to the matrix multiplication unit and configured to convert the analog output vector into a digitized output vector; and a controller including an integrated circuit, the controller configured to perform operations including receiving an artificial neural network computation request comprising an input dataset comprising a first digital input vector, receiving a first plurality of neural network weights, and generating, through the first unit, a first plurality of modulator control signals based on the first digital input vector and a first plurality of weight control signals based on the first plurality of neural network weights.

[0203] Embodiments of the system may include one or more of the following features: For example, the first unit may include a digital-to-analog converter (DAC).

[0204] In some implementations, the second unit may include an analog-to-digital converter (ADC).

[0205] In some implementations, the system may include a memory unit configured to store the dataset and the plurality of neural network weights.

[0206] In some implementations, the controller integrated circuit may be further configured to perform operations including storing the input data set and the first plurality of neural network weights in the memory unit.

[0207] In some implementations, the first unit may be configured to generate a plurality of weight control signals.

[0208] In some implementations, the controller may include an application specific integrated circuit (ASIC), and receiving the artificial neural network computation request may include receiving the artificial neural network computation request from a general purpose data processor.

[0209] In some implementations, the first unit, the processing unit, the second unit, and the controller may be disposed on at least one of a multi-chip module or an integrated circuit. Receiving the artificial neural network computation request may include receiving the artificial neural network computation request from a second data processor, which may be external to the multi-chip module or the integrated circuit, and which may be coupled to the multi-chip module or the integrated circuit through a communication channel, and the processor unit may be capable of processing data at a data rate at least an order of magnitude higher than the data rate of the communication channel.

[0210] In some implementations, the first unit, the processor unit, the second unit, and the controller may be used in an optoelectronic processing loop that is repeated for multiple iterations, the optoelectronic processing loop including: (1) at least a first optical modulation operation based on at least one of the multiple modulator control signals, and at least a second optical modulation operation based on at least one of the weight control signals, and (2) at least one of (a) an electrical summation operation or (b) an electrical storage operation.

[0211] In some implementations, the optoelectronic processing loop may include an electrical storage operation, which may be performed using a memory unit coupled to the controller, and the operations performed by the controller may further include storing the input data set and the first plurality of neural network weights in the memory unit.

[0212] In some implementations, the optoelectronic processing loop may include an electrical summation operation, which may be performed using an electrical summation module within the matrix multiplication unit, and the electrical summation module may be configured to generate currents corresponding to elements of an analog output vector representing the sum of respective elements of the optical input vector multiplied by respective neural network weights.

[0213] In some implementations, the optoelectronic processing loop may include at least one signal path in which exactly one first optical modulation operation based on at least one of the plurality of modulator control signals and exactly one second optical modulation operation based on at least one of the weight control signals are performed in a single loop iteration.

[0214] In some implementations, the first optical modulation operation may be performed by one of a plurality of optical modulators coupled to a source of optical output and to the matrix multiplication unit, and the second optical modulation operation may be performed by an optical modulator included in the matrix multiplication unit.

[0215] In some implementations, the optoelectronic processing loop may include at least one signal path in which only one electrical storage is performed in a single loop iteration.

[0216] In some implementations, the source may include a laser unit configured to generate multiple light outputs.

[0217] In some implementations, the matrix multiplication unit may include an array of input waveguides for receiving an optical input vector, where the optical input vector comprises a first array of optical signals; an optical interference unit in optical communication with the array of input waveguides for performing a linear transformation of the optical input vector into a second array of optical signals; and an array of output waveguides in optical communication with the optical interference unit for directing the second array of optical signals, where at least one input waveguide in the array of input waveguides is in optical communication with each output waveguide in the array of output waveguides via the optical interference unit.

[0218] In some implementations, the optical interference unit may include a plurality of interconnected Mach-Zehnder interferometers (MZIs), each MZI among the plurality of interconnected MZIs including a first phase shifter configured to change a division ratio of the MZI and a second phase shifter configured to shift a phase of an output of one of the MZIs, and the first phase shifter and the second phase shifter are coupled to a plurality of weight control signals.

[0219] In some implementations, the matrix multiplication unit may include a plurality of replication modules, each replication module corresponding to a subset of one or more optical signals of the optical input vector and configured to split the subset of one or more optical signals into two or more replicas of the optical signal; a plurality of multiplication modules corresponding to a subset of one or more optical signals, each multiplication module configured to multiply one or more optical signals of the subset with one or more matrix element values ​​using optical amplitude modulation; and one or more addition modules, each addition module configured to produce an electrical signal representing the sum of the results of two or more of the multiplication modules.

[0220] In some implementations, at least one of the multiplication modules includes an optical amplitude modulator including an input port and two output ports, and pairs of associated optical signals may be provided from the two output ports such that the difference between the amplitudes of the associated optical signals corresponds to the result of multiplying the input value by the signed matrix element value.

[0221] In some implementations, the matrix multiplication unit may be configured to multiply the input vector with a matrix that includes one or more matrix element values.

[0222] In some implementations, a set of multiple output values ​​may be encoded on each electrical signal produced by one or more summation modules, and an output value in the set of multiple output values ​​may represent an element of an output vector that results from multiplying an input vector with a matrix.

[0223] In some implementations, the system may include a memory unit configured to store the input data set and the neural network weights, and the second unit may include an analog-to-digital converter (ADC) unit, and the operations may further include obtaining, from the ADC unit, a first plurality of digitized outputs corresponding to an analog output vector of the matrix multiplication unit, the first plurality of digitized outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector; and storing the first transformed digital output vector in the memory unit.

[0224] In some implementations, the system may have a first loop period defined as the time elapsed between storing the input data set and the first plurality of neural network weights in the memory unit and storing the first converted digital output vector in the memory unit, wherein the first loop period is 1 ns or less.

[0225] In some implementations, the operations may further include outputting an artificial neural network output generated based on the first transformed digital output vector.

[0226] In some implementations, the first unit may include a digital-to-analog converter (DAC) unit, and the operations may further include generating, through the DAC unit, a second plurality of modulator control signals based on the first converted digital output vector.

[0227] In some implementations, the first unit may include a digital-to-analog converter (DAC) unit, the artificial neural network computation request may further include a second plurality of neural network weights, and the operation may further include generating, based on obtaining the first plurality of digitized outputs, a second plurality of weight control signals based on the second plurality of neural network weights through the DAC unit.

[0228] In some implementations, the first and second pluralities of neural network weights may correspond to different layers of an artificial neural network.

[0229] In some implementations, the first unit may include a digital-to-analog converter (DAC) unit, and the input data set may further include a second digital input vector. The operations may further include generating, through the DAC unit, a second plurality of modulator control signals based on the second digital input vector; obtaining, from the ADC unit, a second plurality of digitized outputs corresponding to the output vector of the matrix multiplication unit, where the second plurality of digitized outputs form a second digital output vector; performing a nonlinear transformation on the second digital output vector to generate a second transformed digital output vector; storing the second transformed digital output vector in a memory unit; and outputting an artificial neural network output generated based on the first transformed digital output vector and the second transformed digital output vector. The output vector of the matrix multiplication unit may result from the second optical input vector, which is generated based on the second plurality of modulator control signals transformed by the matrix multiplication unit based on the first mentioned plurality of weight control signals.

[0230] In some implementations, the system may include a memory unit configured to store the input data set and the neural network weights, and the second unit may include an analog-to-digital converter (ADC) unit. The system may further include an analog nonlinearity unit disposed between the matrix multiplication unit and the ADC unit, where the analog nonlinearity unit may be configured to receive the plurality of output voltages from the matrix multiplication unit, apply a nonlinear transfer function, and output the plurality of converted output voltages to the ADC unit. The operations performed by the controller integrated circuit may further include obtaining, from the ADC unit, a first plurality of converted digitized output voltages corresponding to the plurality of converted output voltages, where the first plurality of converted digitized output voltages form a first converted digital output vector, and storing the first converted digital output vector in the memory unit.

[0231] In some implementations, the controller integrated circuit may be configured to generate the first plurality of modulated control signals at a rate of 8 GHz or greater.

[0232] In some implementations, the first unit may include a digital-to-analog converter (DAC) unit and the second unit may include an analog-to-digital converter (ADC) unit. The matrix multiplication unit may include an optical matrix multiplication unit coupled to the plurality of optical modulators and the DAC unit, the optical matrix multiplication unit configured to convert an optical input vector into an optical output vector based on the plurality of weight control signals, and an optical detection unit coupled to the optical matrix multiplication unit and configured to generate a plurality of output voltages corresponding to the optical output vectors.

[0233] In some implementations, the system may further include an analog memory unit disposed between the DAC unit and the plurality of optical modulators, the analog memory unit configured to store analog voltages and output the stored analog voltages; and an analog nonlinearity unit disposed between the light detection unit and the ADC unit, the analog nonlinearity unit configured to receive the plurality of output voltages from the light detection unit, apply a nonlinear transfer function to the plurality of output voltages, and output a plurality of converted output voltages.

[0234] In some implementations, the analog memory unit may include multiple capacitors.

[0235] In some implementations, the analog memory unit may be configured to receive and store the plurality of converted output voltages of the analog nonlinearity unit and output the stored plurality of converted output voltages to the plurality of optical modulators. The operations may further include storing the plurality of converted output voltages of the analog nonlinearity unit in the analog memory unit based on generating the first plurality of modulator control signals and the first plurality of weight control signals, outputting the stored converted output voltages through the analog memory unit, obtaining a second plurality of converted digitized output voltages from the ADC unit, where the second plurality of converted digitized output voltages form a second converted digital output vector, and storing the second converted digital output vector in the memory unit.

[0236] In some implementations, the system may include a memory unit configured to store an input dataset and neural network weights, and the input dataset for the artificial neural network computation request may include a plurality of digital input vectors. The source may be configured to generate a plurality of wavelengths. The plurality of optical modulators may include a bank of optical modulators configured to generate a plurality of optical input vectors, each of the banks corresponding to one of the plurality of wavelengths and generating a respective optical input vector having the respective wavelength, and an optical multiplexer configured to combine the plurality of optical input vectors into a combined optical input vector comprising the plurality of wavelengths. The optical detection unit may further be configured to demultiplex the plurality of wavelengths and generate a plurality of demultiplexed output voltages. The operations may further include obtaining from the ADC unit a plurality of digitized demultiplexed optical outputs, the plurality of digitized demultiplexed optical outputs forming a plurality of first digital output vectors, each of the plurality of first digital output vectors corresponding to one of the plurality of wavelengths, performing a nonlinear transformation on each of the plurality of first digital output vectors to generate a plurality of transformed first digital output vectors, and storing the plurality of transformed first digital output vectors in a memory unit. Each of the plurality of digital input vectors may correspond to one of the plurality of optical input vectors.

[0237] In some implementations, the system may include a memory unit configured to store the input data set and the neural network weights, the second unit may include an analog-to-digital converter (ADC) unit, and the artificial neural network computation request may include a plurality of digital input vectors. The source may be configured to generate a plurality of wavelengths. The plurality of optical modulators may include a bank of optical modulators configured to generate a plurality of optical input vectors, each of the bank corresponding to one of the plurality of wavelengths and generating a respective optical input vector having the respective wavelength, and an optical multiplexer configured to combine the plurality of optical input vectors into a combined optical input vector comprising the plurality of wavelengths. The operations may include obtaining, from the ADC unit, a first plurality of digitized optical outputs corresponding to an optical output vector comprising the plurality of wavelengths, the first plurality of digitized optical outputs forming a first digital output vector, performing a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector, and storing the first transformed digital output vector in the memory unit.

[0238] In some implementations, the first unit may include a digital-to-analog converter (DAC) unit, the second unit may include an analog-to-digital converter (ADC) unit, and the DAC unit may include a 1-bit DAC subunit configured to generate a plurality of 1-bit modulator control signals. The ADC unit may have a resolution of 1 bit, and the first digital input vector may have a resolution of N bits. The operations may include decomposing a first digital input vector into N 1-bit input vectors, each of the N 1-bit input vectors corresponding to one of the N bits of the first digital input vector; generating a sequence of N 1-bit modulator control signals corresponding to the N 1-bit input vectors through a 1-bit DAC subunit; obtaining a sequence of N digitized 1-bit optical outputs corresponding to the sequence of N 1-bit modulator control signals from an ADC unit; constructing an N-bit digital output vector from the sequence of N digitized 1-bit optical outputs; performing a nonlinear transformation on the constructed N-bit digital output vector to generate a transformed N-bit digital output vector; and storing the transformed N-bit digital output vector in a memory unit.

[0239] In some implementations, the system may include a memory unit configured to store the input data set and the neural network weights. The memory unit may include a digital input vector memory configured to store the digital input vector and comprising at least one SRAM, and a neural network weight memory configured to store the neural network weights and comprising at least one DRAM.

[0240] In some implementations, the first unit may include a digital-to-analog converter (DAC) unit including a first DAC subunit configured to generate a plurality of modulator control signals and a second DAC subunit configured to generate a plurality of weight control signals, wherein the first DAC subunit and the second DAC subunit are different.

[0241] In some implementations, the light source may include a laser source configured to generate light and an optical power splitter configured to split the light generated by the laser source into a plurality of optical outputs, each of the plurality of optical outputs having substantially equal power.

[0242] In some implementations, the plurality of optical modulators may include one of an MZI modulator, a ring resonator modulator, or an electro-absorption modulator.

[0243] In some implementations, the light detection unit may include a plurality of light detectors and a plurality of amplifiers configured to convert photocurrents generated by the light detectors into a plurality of output voltages.

[0244] In some implementations, the integrated circuit may be an application specific integrated circuit.

[0245] In some implementations, the apparatus may include a plurality of optical waveguides coupled between the optical modulator and the matrix multiplication unit, and the optical input vector may include a set of a plurality of input values ​​encoded onto respective optical signals carried by the optical waveguides, and each of the optical signals carried by one of the optical waveguides may include an optical wave having a common wavelength that is substantially the same for all of the optical signals.

[0246] In some implementations, the replication modules may include at least one replication module including an optical splitter that sends a predetermined percentage of the power of the light wave at the input port to a first output port and sends the remaining percentage of the power of the light wave at the input port to a second output port.

[0247] In some implementations, the optical splitter may include a waveguide optical splitter that sends a predetermined percentage of the power of the light waves guided by the input optical waveguide to a first output optical waveguide and sends the remaining percentage of the power of the light waves guided by the input optical waveguide to a second output optical waveguide.

[0248] In some implementations, a guided mode of the input optical waveguide can be adiabatically coupled to a guided mode of a propagating mode of each of the first and second output optical waveguides.

[0249] In some implementations, the optical splitter may include a beam splitter that includes at least one surface that transmits a predetermined percentage of the power of the light wave at the input port and reflects the remaining percentage of the power of the light wave at the input port.

[0250] In some implementations, at least one of the plurality of optical waveguides may include an optical fiber coupled to an optical coupler that couples a guided mode of the optical fiber to a propagation mode in free space.

[0251] In some implementations, the multiplication module may include at least one coherence-susceptible multiplication module configured to multiply one or more optical signals of the first subset with one or more matrix element values ​​using optical amplitude modulation based on interference between optical waves having a coherence length at least as long as a propagation distance through the coherence-susceptible multiplication module.

[0252] In some implementations, the coherence-susceptible multiplication module may include a Mach-Zehnder interferometer (MZI) that splits light waves guided by an input optical waveguide into a first optical waveguide arm of the MZI and a second optical waveguide arm of the MZI, the first optical waveguide arm including a phase shifter that provides a relative phase shift with respect to a phase delay of the second optical waveguide arm, and the MZI can combine the light waves from the first optical waveguide arm and the second optical waveguide arm into at least one output optical waveguide.

[0253] In some implementations, the MZI can combine lightwaves from the first and second optical waveguide arms into first and second output optical waveguides, respectively; the first photodetector can receive the lightwaves from the first output optical waveguide to generate a first photocurrent; the second photodetector can receive the lightwaves from the second output optical waveguide to generate a second photocurrent; and the result of the coherence-sensitive multiplication module can include the difference between the first photocurrent and the second photocurrent.

[0254] In some implementations, the coherence-susceptible multiplication module may include one or more ring resonators, including at least one ring resonator coupled to a first optical waveguide and at least one ring resonator coupled to a second optical waveguide.

[0255] In some implementations, the first photodetector can receive lightwaves from the first optical waveguide to generate a first photocurrent, the second photodetector can receive lightwaves from the second optical waveguide to generate a second photocurrent, and the result of the coherence-sensitive multiplication module can include the difference between the first photocurrent and the second photocurrent.

[0256] In some implementations, the multiplication modules may include at least one coherence-insensitive multiplication module configured to multiply one or more optical signals of the first subset with one or more matrix element values ​​using optical amplitude modulation based on absorption of energy in a light wave.

[0257] In some implementations, the coherence-insensitive multiplication module may include an electro-absorption modulator.

[0258] In some implementations, the one or more summing modules may include at least one summing module including: (1) two or more input conductors each carrying an electrical signal in the form of an input current, the amplitude of which represents a respective result of a respective one of the multiplication modules; and (2) at least one output conductor carrying an electrical signal in the form of an output current proportional to the sum of the input currents, the electrical signal representing the sum of the respective results.

[0259] In some implementations, the two or more input and output conductors may include wires that meet at one or more intersections between the wires, and the output current may be substantially equal to the sum of the input currents.

[0260] In some implementations, at least a first one of the input currents may be provided in the form of at least one photocurrent generated by at least one photodetector receiving the optical signal generated by a first one of the multiplication modules.

[0261] In some implementations, the first input current may be provided in the form of a difference between two photocurrents generated by different respective photodetectors that receive different respective optical signals both generated by the first multiplication module.

[0262] In some implementations, one of the replicas of the first subset of one or more optical signals may consist of a single optical signal on which one of the input values ​​is encoded.

[0263] In some implementations, the multiplication module corresponding to the replica of the first subset may multiply the encoded input value with a single matrix element value.

[0264] In some implementations, one of the replicas of the first subset of one or more optical signals may include more than one of the optical signals and fewer than all of the optical signals on which multiple input values ​​are encoded.

[0265] In some implementations, the multiplication modules corresponding to the replicas of the first subset may multiply the encoded input values ​​with different respective matrix element values.

[0266] In some implementations, different multiplication modules corresponding to different respective copies of the first subset of one or more optical signals may be stored in different devices that are in optical communication for transmitting one of the copies of the first subset of one or more optical signals between the different devices.

[0267] In some implementations, two or more of the plurality of optical waveguides, two or more of the plurality of replication modules, two or more of the plurality of multiplication modules, and at least one of the one or more addition modules may be disposed on a common device substrate.

[0268] In some implementations, the device may perform vector-matrix multiplication, where the input vector may be provided as a set of optical signals and the output vector may be provided as a set of electrical signals.

[0269] In some implementations, the device may further include an accumulator that integrates an input electrical signal corresponding to the output of the multiplication module or the addition module, and the input electrical signal may be encoded using time-domain coding that uses on-off amplitude modulation within each of a plurality of time slots, and the accumulator may produce an output electrical signal that is encoded with more than two amplitude levels across the plurality of time slots, corresponding to different duty cycles of the time-domain coding.

[0270] In some implementations, two or more of the multiplication modules each correspond to a different subset of one or more optical signals.

[0271] In some implementations, the apparatus may further include a multiplication module configured, for each replica of a second subset of one or more optical signals that differs from an optical signal in the first subset of one or more optical signals, to multiply the one or more optical signals of the second subset with one or more matrix element values ​​using optical amplitude modulation.

[0272] In another aspect, a system includes a memory unit configured to store a data set and a plurality of neural network weights, and a driver unit configured to generate a plurality of modulator control signals. The system includes an optoelectronic processor including: a light source configured to provide a plurality of optical outputs, a plurality of optical modulators coupled to the light source and the driver unit, the plurality of optical modulators configured to generate an optical input vector by modulating the plurality of optical outputs generated by the light source based on the plurality of modulator control signals, a matrix multiplication unit coupled to the plurality of optical modulators and the driver unit, the matrix multiplication unit configured to convert the optical input vector into an analog output vector based on the plurality of weight control signals, and a comparator unit coupled to the matrix multiplication unit and configured to convert the analog output vector into a plurality of digitized 1-bit outputs. The system includes a controller including an integrated circuit, the controller configured to perform operations including receiving an artificial neural network computation request comprising an input dataset and a first plurality of neural network weights, the input dataset comprising a first digital input vector having N-bit resolution; storing the input dataset and the first plurality of neural network weights in a memory unit; decomposing the first digital input vector into N 1-bit input vectors, each of the N 1-bit input vectors corresponding to one of the N bits of the first digital input vector; generating, through a driver unit, a sequence of N 1-bit modulator control signals corresponding to the N 1-bit input vectors; obtaining, from a comparator unit, a sequence of N digitized 1-bit outputs corresponding to the sequence of N 1-bit modulator control signals; constructing an N-bit digital output vector from the sequence of N digitized 1-bit outputs; performing a nonlinear transformation on the constructed N-bit digital output vector to generate a transformed N-bit digital output vector; and storing the transformed N-bit digital output vector in the memory unit.

[0273] Embodiments of the system may include one or more of the following features: For example, receiving the artificial neural network computation request may include receiving the artificial neural network computation request from a general purpose computer.

[0274] In some implementations, the driver unit can be configured to generate multiple weight control signals.

[0275] In some implementations, the matrix multiplication unit may include an optical matrix multiplication unit coupled to the plurality of optical modulators and the driver unit, the optical matrix multiplication unit configured to convert an optical input vector into an optical output vector based on the plurality of weight control signals, and an optical detection unit coupled to the optical matrix multiplication unit and configured to generate a plurality of output voltages corresponding to the optical output vectors.

[0276] In some implementations, the matrix multiplication unit may include an array of input waveguides for receiving an optical input vector, an optical interference unit in optical communication with the array of input waveguides for performing a linear transformation of the optical input vector into a second array of optical signals, and an array of output waveguides in optical communication with the optical interference unit for directing the second array of optical signals, wherein at least one input waveguide in the array of input waveguides is in optical communication with each output waveguide in the array of output waveguides via the optical interference unit.

[0277] In some implementations, the optical interference unit may include a plurality of interconnected Mach-Zehnder interferometers (MZIs), each MZI among the plurality of interconnected MZIs including a first phase shifter configured to change a division ratio of the MZI and a second phase shifter configured to shift the phase of an output of one of the MZIs, and the first phase shifter and the second phase shifter may be coupled to a plurality of weight control signals.

[0278] In some implementations, the matrix multiplication unit may include: a plurality of replication modules, including, for each of at least two subsets of one or more optical signals of the optical input vector, a corresponding set of one or more replication modules configured to split the subset of one or more optical signals into two or more replicas of the optical signals; a plurality of multiplication modules, including, for each of the at least two replicas of a first subset of the one or more optical signals, a corresponding multiplication module configured to multiply one or more optical signals of the first subset with one or more matrix element values ​​using optical amplitude modulation; and one or more summation modules, including, for results of two or more of the multiplication modules, an summation module configured to produce an electrical signal representing a sum of results of two or more of the multiplication modules.

[0279] In some implementations, at least one of the multiplication modules may include an optical amplitude modulator including an input port and two output ports, and pairs of associated optical signals may be provided from the two output ports such that the difference between the amplitudes of the associated optical signals corresponds to the result of multiplying the input value by the signed matrix element value.

[0280] In some implementations, the matrix multiplication unit may be configured to multiply the input vector with a matrix that includes one or more matrix element values.

[0281] In some implementations, a set of multiple output values ​​may be encoded on each electrical signal produced by one or more summation modules, and an output value in the set of multiple output values ​​may represent an element of an output vector that results from the input vector being multiplied by a matrix.

[0282] In another aspect, a method is provided for performing an artificial neural network computation in a system having a matrix multiplication unit configured to convert an optical input vector into an analog output vector based on a plurality of weight control signals, the method including the steps of receiving an artificial neural network computation request comprising an input data set and a first plurality of neural network weights, the input data set comprising a first digital input vector, storing the input data set and the first plurality of neural network weights in a memory unit, generating first plurality of modulator control signals based on the first digital input vector and first plurality of weight control signals based on the first plurality of neural network weights, obtaining first plurality of digitized outputs corresponding to an output vector of the matrix multiplication unit, the first plurality of digitized outputs forming a first digital output vector, performing, by a controller, a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector, storing the first transformed digital output vector in the memory unit, and outputting, by the controller, an artificial neural network output generated based on the first transformed digital output vector.

[0283] Embodiments of the method may include one or more of the following features: For example, receiving the artificial neural network computation request may include receiving the artificial neural network computation request from a computer over a communications channel.

[0284] In some implementations, generating the first plurality of modulator control signals may include generating the first plurality of modulator control signals through a digital-to-analog converter (DAC) unit.

[0285] In some implementations, obtaining the first plurality of digitized outputs may include obtaining the first plurality of digitized outputs from an analog-to-digital conversion (ADC) unit.

[0286] In some implementations, the method may include applying a first plurality of modulator control signals to a plurality of optical modulators coupled to the light source and the DAC unit, and generating an optical input vector by modulating, using the plurality of optical modulators, a plurality of optical outputs generated by the laser unit based on the plurality of modulator control signals.

[0287] In some implementations, the matrix multiplication unit may be coupled to the plurality of optical modulators and the DAC unit, and the method may include using the matrix multiplication unit to convert an optical input vector into an analog output vector based on the plurality of weight control signals.

[0288] In some implementations, the ADC unit may be coupled to the matrix multiplication unit, and the method may include converting the analog output vector into a first plurality of digitized outputs using the ADC unit.

[0289] In some implementations, the matrix multiplication unit may include an optical matrix multiplication unit coupled to a plurality of optical modulators and a DAC unit. Converting the optical input vector into an analog output vector may include converting the optical input vector into an optical output vector based on a plurality of weight control signals using the optical matrix multiplication unit. The method may include generating a plurality of output voltages corresponding to the optical output vector using an optical detection unit coupled to the optical matrix multiplication unit.

[0290] In some implementations, the method may include receiving an optical input vector at an array of input waveguides; performing a linear conversion of the optical input vector into a second array of optical signals using an optical interference unit in optical communication with the array of input waveguides; and directing the second array of optical signals using an array of output waveguides in optical communication with the optical interference unit, wherein at least one input waveguide in the array of input waveguides is in optical communication with each output waveguide in the array of output waveguides via the optical interference unit.

[0291] In some implementations, the optical interference unit may include a plurality of interconnected Mach-Zehnder interferometers (MZIs), each of which may include a first phase shifter and a second phase shifter, and the first phase shifter and the second phase shifter may be coupled to a plurality of weight control signals. The method may include changing a division ratio of the MZIs using the first phase shifter and shifting a phase of an output of one of the MZIs using the second phase shifter.

[0292] In some implementations, the method may include: for each of at least two subsets of one or more optical signals of the optical input vector, splitting the subset of one or more optical signals into two or more replicas of the optical signals using a corresponding set of one or more replica modules; for each of the at least two replicas of a first subset of the one or more optical signals, using a corresponding multiplication module, multiplying one or more optical signals of the first subset with one or more matrix element values ​​using optical amplitude modulation; and for results of two or more of the multiplication modules, using an addition module to produce an electrical signal representing a sum of the results of two or more of the multiplication modules.

[0293] In some implementations, at least one of the multiplication modules may include an optical amplitude modulator including an input port and two output ports, and pairs of associated optical signals may be provided from the two output ports such that the difference between the amplitudes of the associated optical signals corresponds to the result of multiplying the input value by the signed matrix element value.

[0294] In some implementations, the method may include multiplying, using a matrix multiplication unit, the input vector with a matrix including one or more matrix element values.

[0295] In some implementations, the method may include encoding a set of multiple output values ​​for each electrical signal produced by the one or more summation modules, and using an output value from the set of multiple output values ​​to represent an element of an output vector resulting from multiplying the input vector by the matrix.

[0296] In another aspect, a method includes providing input information in an electronic format, converting at least a portion of the electronic input information to an optical input vector, opto-electronically converting the optical input vector to an analog output vector based on a matrix multiplication, and electronically applying a nonlinear transformation to the analog output vector to provide output information in an electronic format.

[0297] Embodiments of the method may include one or more of the following features. For example, the method may further include repeating the electrical-to-optical conversion, the optical-electronic conversion, and the electronically applied nonlinear conversion on new electronic input information corresponding to the provided output information in electronic format.

[0298] In some implementations, the matrix multiplication for the initial opto-electronic transformation and the matrix multiplication for the repeated opto-electronic transformations may be the same and may correspond to the same layer of the artificial neural network.

[0299] In some implementations, the matrix multiplication for the initial opto-electronic transformation and the matrix multiplication for the repeated opto-electronic transformations may be different and may correspond to different layers of the artificial neural network.

[0300] In some implementations, the method may further include repeating the electrical-to-optical conversion, the optoelectronic conversion, and the electronically applied nonlinear conversion for different portions of the electronic input information, wherein the matrix multiplication for the initial optoelectronic conversion and the matrix multiplication for the repeated optoelectronic conversions are the same and correspond to a first layer of the artificial neural network.

[0301] In some implementations, the method may further include providing intermediate information in electronic format based on the electronic output information produced for the plurality of portions of the electronic input information by the first layer of the artificial neural network, and repeating the electrical-to-optical conversion, the optoelectronic conversion, and the electronically applied nonlinear conversion for each of the different portions of the electronic intermediate information, wherein the matrix multiplication for the initial optoelectronic conversion and the matrix multiplication for the repeated optoelectronic conversion for the different portions of the electronic intermediate information may be the same and may correspond to a second layer of the artificial neural network.

[0302] In another aspect, a system for performing artificial neural network computations is provided, the system including a first unit configured to generate a plurality of vector control signals and generate a plurality of weight control signals, a second unit configured to provide an optical input vector based on the plurality of vector control signals, and a matrix multiplication unit coupled to the second unit and the first unit, the matrix multiplication unit configured to convert the optical input vector into an output vector based on the plurality of weight control signals. The system includes a controller including an integrated circuit, the controller configured to perform operations including receiving an artificial neural network computation request comprising an input data set and a first plurality of neural network weights, the input data set comprising a first digital input vector, and generating, through a first unit, a first plurality of vector control signals based on the first digital input vector and a first plurality of weight control signals based on the first plurality of neural network weights, the first unit, the second unit, the matrix multiplication unit, and the controller being used in an opto-electronic processing loop that is repeated for multiple iterations, the opto-electronic processing loop including (1) at least two optical modulation operations and (2) at least one of (a) an optical summation operation or (b) an optical storage operation.

[0303] In another aspect, a method for performing artificial neural network computation is provided, the method including the steps of providing input information in electronic format, converting at least a portion of the electronic input information into an optical input vector, and transforming the optical input vector into an output vector based on matrix multiplication using a set of neural network weights. The providing, transforming, and transforming steps are performed in an opto-electronic processing loop that is repeated for multiple iterations using different respective sets of neural network weights and different respective input information, the opto-electronic processing loop including (1) at least two optical modulation operations and (2) at least one of (a) an optical summation operation or (b) an optical storage operation.

[0304] The details of one or more embodiments of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the invention will become apparent from the description, drawings, and claims.

[0305] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. In case of conflict with any patent application or published patent application incorporated herein by reference, the present specification, including definitions, will control.

[0306] The present disclosure is best understood from the following detailed description when read in conjunction with the accompanying drawings, in which: It is emphasized, of course, that the various features of the drawings are not to scale, and to the contrary, the dimensions of the various features are arbitrarily expanded or reduced for clarity. [Brief explanation of the drawings]

[0307] [Figure 1A] FIG. 1 is a schematic diagram of an example artificial neural network (ANN) computing system. [Figure 1B] FIG. 1 is a schematic diagram of an example optical matrix multiplication unit. [Figure 1C] 1 is a schematic diagram of an exemplary configuration of interconnected Mach-Zehnder interferometers (MZIs). [Figure 1D] 1 is a schematic diagram of an exemplary configuration of interconnected Mach-Zehnder interferometers (MZIs). [Figure 1E] FIG. 1 is a schematic diagram of an example of an MZI. [Figure 1F] FIG. 1 is a schematic diagram of an example wavelength division multiplexing ANN computing system. [Figure 2A] 1 is a flowchart illustrating an example of a method for performing an ANN calculation. [Figure 2B] FIG. 2B illustrates an embodiment of the method of FIG. 2A. [Figure 3A] FIG. 1 is a schematic diagram of an example ANN computing system. [Figure 3B] FIG. 1 is a schematic diagram of an example ANN computing system. [Figure 4A] FIG. 1 is a schematic diagram of an example of an ANN computing system with 1-bit internal resolution. [Figure 4B] FIG. 4B is a mathematical representation of the operation of the ANN computational system of FIG. 4A. [Figure 5] FIG. 1 is a schematic diagram of an example artificial neural network (ANN) computing system. [Figure 6] FIG. 10 is a diagram of an example of an optical matrix multiplication unit. [Figure 7] FIG. 1 is a schematic diagram of an example artificial neural network (ANN) computing system. [Figure 8] FIG. 1 is a diagram of an example of an optical matrix multiplication unit. [Figure 9] FIG. 1 is a schematic diagram of an example artificial neural network (ANN) computing system. [Figure 10] FIG. 10 is a diagram of an example of an optical matrix multiplication unit. [Figure 11] FIG. 1 is a diagram of an example of a compact matrix multiplier unit. [Figure 12A] FIG. 10 is a comparison of photonics matrix multiplier units. [Figure 12B] FIG. 1 is a diagram of a miniature interconnected interferometer. [Figure 13] FIG. 1 is a diagram of a compact matrix multiplier unit. [Figure 14] FIG. 1 is a diagram of a light-generating adversarial network. [Figure 15] FIG. 1 is a diagram of a Mach-Zehnder interferometer. [Figure 16] FIG. 1 is a diagram of a photonics circuit. [Figure 17A] FIG. 1 is a diagram of a photonics circuit. [Figure 17B] FIG. 1 is a diagram of a photonics circuit. [Figure 18] FIG. 1 is a schematic diagram of an exemplary optoelectronic computing system. [Figure 19A] FIG. 1 is a schematic diagram of an exemplary system configuration. [Figure 19B] FIG. 1 is a schematic diagram of an exemplary system configuration. [Figure 20A] FIG. 1 is a schematic diagram of an example of a symmetric differential configuration. [Figure 20B] FIG. 2 is a circuit diagram of an example system module. [Figure 20C] FIG. 2 is a circuit diagram of an example system module. [Figure 21A] FIG. 1 is a schematic diagram of an example of a symmetric differential configuration. [Figure 21B] FIG. 1 is a schematic diagram of an example of a system configuration. [Figure 22A] 1 is a schematic diagram of an exemplary optical amplitude modulator. [Figure 22B] FIG. 1 is a schematic diagram of an example of an optical amplitude modulator with optical detection in a symmetric differential configuration. [Figure 22C] FIG. 1 is a schematic diagram of an example of an optical amplitude modulator with optical detection in a symmetric differential configuration. [Figure 22D] FIG. 1 is a schematic diagram of an example of an optical amplitude modulator with optical detection in a symmetric differential configuration. [Figure 23A] FIG. 1 is an optoelectronic circuit diagram of an exemplary system configuration. [Figure 23B] FIG. 1 is an optoelectronic circuit diagram of an exemplary system configuration. [Figure 23C] FIG. 1 is an optoelectronic circuit diagram of an exemplary system configuration. [Figure 24A] FIG. 1 is a schematic diagram of an exemplary computing system using multiple optoelectronic subsystems. [Figure 24B] FIG. 1 is a schematic diagram of an exemplary computing system using multiple optoelectronic subsystems. [Figure 24C] FIG. 1 is a schematic diagram of an exemplary computing system using multiple optoelectronic subsystems. [Figure 24D] FIG. 1 is a schematic diagram of an exemplary computing system using multiple optoelectronic subsystems. [Figure 24E] FIG. 1 is a schematic diagram of an exemplary computing system using multiple optoelectronic subsystems. [Figure 25] 1 is a flowchart illustrating an example of a method for performing an ANN calculation. [Figure 26] FIG. 1 is a schematic diagram of an example ANN computing system. [Figure 27]FIG. 1 is a schematic diagram of an example ANN computing system. [Figure 28] FIG. 1 is a schematic diagram of an example neural network computing system using a passive 2D optical matrix multiplication unit. [Figure 29] FIG. 1 is a schematic diagram of an example neural network computing system using a passive 3D optical matrix multiplication unit. [Figure 30] FIG. 1 is a schematic diagram of an example artificial neural network computing system using a passive 2D optical matrix multiplication unit with 1-bit internal resolution. [Figure 31] FIG. 1 is a schematic diagram of an example artificial neural network computing system using a passive 3D optical matrix multiplication unit with 1-bit internal resolution. [Figure 32A] FIG. 1 is a schematic diagram of an example artificial neural network (ANN) computing system. [Figure 32B] FIG. 1 is a schematic diagram of an example of an optoelectronic matrix multiplication unit. [Figure 33] 1 is a flow chart illustrating an example of a method for performing ANN calculations using an optoelectronic processor. [Figure 34] FIG. 34 illustrates an embodiment of the method of FIG. 33. [Figure 35A] FIG. 1 is a schematic diagram of an example wavelength division multiplexed ANN computing system using an optoelectronic processor. [Figure 35B] FIG. 1 is a schematic diagram of an example of a wavelength division multiplexing optical matrix multiplication unit. [Figure 35C] FIG. 1 is a schematic diagram of an example of a wavelength division multiplexing optical matrix multiplication unit. [Figure 36] FIG. 1 is a schematic diagram of an example of an ANN computing system using an optoelectronic matrix multiplication unit. [Figure 37] FIG. 1 is a schematic diagram of an example of an ANN computing system using an optoelectronic matrix multiplication unit. [Figure 38] FIG. 1 is a schematic diagram of an example artificial neural network computing system using an optoelectronic matrix multiplication unit with 1-bit internal resolution. [Figure 39A] FIG. 1 is a diagram of an example of a Mach-Zehnder modulator. [Figure 39B]39B is a graph showing the intensity versus voltage curve of the Mach-Zehnder modulator of FIG. 39A. [Figure 40] FIG. 1 is a schematic diagram of a homodyne detector. [Figure 41] 1 is a schematic diagram of a computing system including optical fibers each carrying signals having multiple wavelengths. DETAILED DESCRIPTION OF THE INVENTION

[0308] Like reference numbers and designations in the various drawings indicate like elements.

[0309] 1A shows a schematic diagram of an example artificial neural network (ANN) computing system 100. System 100 includes a controller 110, a memory unit 120, a digital-to-analog converter (DAC) unit 130, an optical processor 140, and an analog-to-digital converter (ADC) unit 160. Controller 110 is coupled to computer 102, memory unit 120, DAC unit 130, and ADC unit 160. Controller 110 includes an integrated circuit configured to control the operation of ANN computing system 100 to perform ANN computations.

[0310] The integrated circuit of controller 110 may be an application-specific integrated circuit that is specifically configured to perform steps of the ANN computation. For example, the integrated circuit may implement microcode or firmware specific to performing the ANN computation. Thus, controller 110 may have a reduced instruction set compared to general-purpose processors used in conventional computers, such as computer 102. In some implementations, the integrated circuit of controller 110 may include two or more circuits configured to perform different steps of the ANN computation.

[0311] In an exemplary operation of the ANN computing system 100, the computer 102 may issue an artificial neural network computation request to the ANN computing system 100. The ANN computation request may include neural network weights that define the ANN and an input data set to be processed by the provided ANN. The controller 110 receives the ANN computation request and stores the input data set and the neural network weights in the memory unit 120.

[0312] The input dataset may correspond to a variety of digital information to be processed by the ANN. Examples of input datasets include image files, audio files, LiDAR point clouds, and GPS coordinate sequences, and the operation of the ANN computation system 100 will be described based on receiving an image file as the input dataset. In general, the size of the input dataset may range significantly, from a few hundred data points to millions of data points or more. For example, a digital image file with a 1-megapixel resolution has approximately one million pixels, and each of the one million pixels may be a data point to be processed by the ANN. Due to the large number of data points in a typical input dataset, the input dataset is typically divided into multiple digital input vectors of smaller size that are individually processed by the optical processor 140. As an example, for a grayscale digital image, the elements of the digital input vector may be 8-bit values ​​representing the intensity of the image, and the digital input vector may have a length ranging from tens of elements (e.g., 32 elements, 64 elements) to hundreds of elements (e.g., 256 elements, 512 elements). In general, an input dataset of any size may be divided into digital input vectors of a size suitable for processing by the optical processor 140. If the number of elements in the input dataset is not divisible by the length of the digital input vector, zero padding may be used to pad the dataset so that it is divisible by the length of the digital input vector. The processed outputs of the individual digital input vectors may be processed to reconstruct a completed output that is the result of processing the input dataset through the ANN. In some implementations, the division of the input dataset into multiple input vectors and the subsequent vector-level processing may be performed using block matrix multiplication techniques.

[0313] Neural network weights are a set of values ​​that define the connectivity of the artificial neurons of an ANN, including the relative importance or weight of those connections. An ANN may include one or more hidden layers with respective sets of nodes. For an ANN with a single hidden layer, the ANN may be defined by two sets of neural network weights: one set corresponding to the connectivity between the input nodes and the nodes of the hidden layer, and a second set corresponding to the connectivity between the hidden layer and the output nodes. Each set of neural network weights describing connectivity corresponds to a matrix to be implemented by the optical processor 140. For ANNs with more than two hidden layers, additional sets of neural network weights are required to define the connectivity between the additional hidden layers. Thus, in general, the neural network weights included in an ANN computation request may include multiple sets of neural network weights representing the connectivity between various layers of the ANN.

[0314] Because the input data set to be processed is typically divided into multiple smaller digital input vectors for individual processing, the input data set is typically stored in digital memory. However, the speed of memory operations between the memory and the processor of the computer 102 is much slower than the rate at which the ANN computing system 100 can perform ANN calculations. For example, the ANN computing system 100 may perform tens to hundreds of ANN calculations during a typical memory read cycle of the computer 102. Thus, the rate at which ANN calculations can be performed by the ANN computing system 100 may be limited to less than the full processing rate if the ANN calculations by the ANN computing system 100 involve multiple data transfers between the system 100 and the computer 102 in the course of processing an ANN calculation request. For example, if the computer 102 were to access the input data set from its own memory and provide the digital input vectors to the controller 110 when requested, the operation of the ANN computing system 100 would be significantly slowed by the time required for the series of data transfers that would be required between the computer 102 and the controller 110. It should be noted that the memory access latency of the computer 102 is typically non-deterministic, which further complicates and reduces the rate at which digital input vectors can be provided to the ANN computing system 100. Additionally, processor cycles of the computer 102 may be wasted managing data transfer between the computer 102 and the ANN computing system 100.

[0315] Instead, in some implementations, the ANN computing system 100 stores the entire input data set in a memory unit 120, which is part of and dedicated for use by the ANN computing system 100. The dedicated memory unit 120 allows transactions between the memory unit 120 and the controller 110 to be specially adapted to allow a smooth, uninterrupted flow of data between the memory unit 120 and the controller 110. Such an uninterrupted flow of data can greatly improve the overall throughput of the ANN computing system 100 by allowing the optical processor 140 to perform matrix multiplications at its full processing rate rather than being limited by the slow memory operations of conventional computers such as the computer 102. Furthermore, because all of the data needed in performing the ANN computation is provided by the computer 102 to the ANN computing system 100 in a single transaction, the ANN computing system 100 can perform the ANN computation in a self-contained manner independent of the computer 102. This self-contained operation of the ANN computing system 100 offloads the computational burden from the computer 102 , eliminating external dependencies in the operation of the ANN computing system 100 and improving the performance of both the system 100 and the computer 102 .

[0316] The internal operation of ANN computing system 100 will now be described. Optical processor 140 includes laser unit 142, array of modulators 144, detection unit 146, and optical matrix multiplication (OMM) unit 150. Optical processor 140 operates by encoding a digital input vector of length N into an optical input vector of length N and propagating the optical input vector through OMM unit 150. OMM unit 150 receives the optical input vector of length N and performs an N×N matrix multiplication on the received optical input vector in the optical domain. The N×N matrix multiplication performed by OMM unit 150 is determined by the internal configuration of OMM unit 150. The internal configuration of OMM unit 150 may be controlled by an electrical signal, such as one generated by DAC unit 130.

[0317] The OMM unit 150 may be implemented in various ways. FIG. 1B shows a schematic diagram of an example of the OMM unit 150. The OMM unit 150 may include an array of input waveguides 152 for receiving an optical input vector, an optical interferometer unit 154 in optical communication with the array of input waveguides 152, and an array of output waveguides 156 in optical communication with the optical interferometer unit 154. The optical interferometer unit 154 performs a linear transformation of the optical input vector into a second array of optical signals. The array of output waveguides 156 directs the second array of optical signals output by the optical interferometer unit 154. At least one input waveguide in the array of input waveguides 152 is in optical communication with each output waveguide in the array of output waveguides 156 via the optical interferometer unit 154. For example, for an optical input vector of length N, the OMM unit 152 may include N input waveguides 152 and N output waveguides 156.

[0318] The optical interference unit may include multiple interconnected Mach-Zehnder interferometers (MZIs). Figures 1C and 1D show schematic diagrams of exemplary configurations 157 and 158 of interconnected MZIs. The MZIs may be interconnected in various ways, such as in configuration 157 or 158, to achieve linear transformation of an optical input vector received through the array of input waveguides 152.

[0319] 1E shows a schematic diagram of an example of an MZI 170. The MZI includes a first input waveguide 171, a second input waveguide 172, a first output waveguide 178, and a second output waveguide 179. Furthermore, each MZI 170 among the multiple interconnected MZIs includes a first phase shifter 174 configured to change the division ratio of the MZI 170 and a second phase shifter 176 configured to shift the phase of one output of the MZI, such as light exiting the MZI 170 through the second output waveguide 179. The first phase shifter 174 and the second phase shifter 176 of the MZI 170 are coupled to multiple weight control signals generated by the DAC unit 130. The first phase shifter 174 and the second phase shifter 176 are examples of reconfigurable elements of the OMM unit 150. Examples of reconfiguring elements include thermo-optic phase shifters or electro-optic phase shifters. Thermo-optic phase shifters operate by heating the waveguide to change the refractive index of the waveguide and cladding material, and the change in refractive index results in a change in phase. Electro-optic phase shifters operate by applying an electric field (e.g., LiNbO3, reverse-biased PN junction) or current (e.g., forward-biased PIN junction), which changes the refractive index of the waveguide material. By varying the weight control signal, the phase delay of the first phase shifter 174 and the second phase shifter 176 of each of the interconnected MZIs 170 can be varied, which reconfigures the optical interference unit 154 of the OMM unit 150 to perform a specific matrix multiplication determined by the phase delay set across the optical interference unit 154. Additional embodiments of the OMM unit 150 and the optical interference unit are disclosed in U.S. Patent Application Publication No. 2017 / 0351293(A1), entitled "APPARATUS AND METHODS FOR OPTICAL NEURAL NETWORK," which is incorporated herein by reference in its entirety.

[0320] An optical input vector is generated through laser unit 142 and modulator array 144. The optical input vector of length N has N independent optical signals, each having an intensity corresponding to the value of a respective element of the digital input vector of length N. As an example, laser unit 142 may generate N optical outputs. The N optical outputs have the same wavelength and are optically coherent. The optical coherence of the optical outputs allows them to optically interfere with each other, which is suitably utilized by OMM unit 150 (e.g., in the operation of an MZI). Furthermore, the optical outputs of laser unit 142 may be substantially identical to one another. For example, the N optical outputs may be substantially uniform in intensity (e.g., within 5%, 3%, 1%, 0.5%, 0.1%, or 0.01%) and substantially uniform in relative phase (e.g., within 10 degrees, 5 degrees, 3 degrees, 1 degree, or 0.1 degree). The uniformity of the optical output may increase the fidelity of the optical input vector to the digital input vector and increase the overall accuracy of the optical processor 140. In some implementations, the optical output of the laser unit 142 may have an optical power ranging from 0.1 mW to 50 mW per output, a wavelength in the near-infrared range (e.g., between 900 nm and 1600 nm), and a linewidth of less than 1 nm. The optical output of the laser unit 142 may be a single transverse mode optical output.

[0321] In some implementations, the laser unit 142 includes a single laser source and an optical power splitter. The single laser source is configured to generate laser light. The optical power splitter is configured to split the light generated by the laser source into N optical outputs having substantially equal intensities and phases. By splitting the single laser output into multiple outputs, optical coherence of the multiple optical outputs can be achieved. The single laser source can be, for example, a semiconductor laser diode, a vertical-cavity surface-emitting laser (VCSEL), a distributed feedback (DFB) laser, or a distributed Bragg reflector (DBR) laser. The optical power splitter can be, for example, a 1:N multimode interference (MMI) splitter, a multi-stage splitter including multiple 1:2 MMI splitters or directional couplers, or a star coupler. In some other implementations, a master-slave laser configuration can be used, in which the slave lasers are injection-locked by the master laser to have a stable phase relationship with the master laser.

[0322] The optical output of the laser unit 142 is coupled to a modulator array 144. The modulator array 144 is configured to receive an optical input from the laser unit 142 and modulate the intensity of the received optical input based on a modulator control signal, which is an electrical signal. Examples of modulators include a Mach-Zehnder interferometer (MZI) modulator, a ring resonator modulator, and an electro-absorption modulator. The modulator array 144 has N modulators, each receiving one of the N optical outputs of the laser unit 142. The modulators receive control signals corresponding to elements of a digital input vector and modulate the intensity of the light. The control signals may be generated by the DAC unit 130.

[0323] The DAC unit 130 is configured to generate multiple modulator control signals and generate multiple weight control signals under the control of the controller 110. For example, the DAC unit 130 receives from the controller 110 a first DAC control signal corresponding to a digital input vector to be processed by the optical processor 140. Based on the first DAC control signal, the DAC unit 130 generates a modulator control signal, which is an analog signal suitable for driving the array of modulators 144. The analog signal can be a voltage or a current, for example, depending on the technology and design of the modulators in the array 144. The voltage can have an amplitude ranging from ±0.1 V to ±10 V, for example, and the current can have an amplitude ranging from 100 μA to 100 mA, for example. In some implementations, the DAC unit 130 can include a modulator driver configured to buffer, amplify, or condition the analog signal so that the modulators in the array 144 can be appropriately driven. For example, some types of modulators can be driven using a differential control signal. In such cases, the modulator driver can be a differential driver that produces a differential electrical output based on a single-ended input signal. As another example, some types of modulators may have a bandwidth that is 3 dB less than the desired processing rate of optical processor 140. In such cases, the modulator driver may include pre-emphasis circuitry or other bandwidth-enhancing circuitry designed to widen the operating bandwidth of the modulator.

[0324] In some cases, the modulators in column 144 may have nonlinear transfer functions. For example, an MZI optical modulator may have a nonlinear relationship (e.g., a sinusoidal dependence) between the applied control voltage and its transmission. In such cases, the first DAC control signal may be adjusted or compensated based on the modulator's nonlinear transfer function so that a linear relationship may be maintained between the digital input vector and the generated optical input vector. Maintaining such linearity is typically important to ensure that the input to OMM unit 150 is an accurate representation of the digital input vector. In some implementations, compensation of the first DAC control signal may be performed by controller 110 via a lookup table that maps values ​​of the digital input vector to values ​​to be output by DAC unit 130 so that the resulting modulated optical signal is linearly proportional to the elements of the digital input vector. The lookup table may be generated by characterizing the modulator's nonlinear transfer function and calculating the inverse of the nonlinear transfer function.

[0325] In some implementations, the nonlinearity of the modulator and the resulting nonlinearity in the generated optical input vector can be compensated for by an ANN computational algorithm.

[0326] The optical input vector generated by the array of modulators 144 is input to the OMM unit 150. The optical input vector may be N spatially separated optical signals, each having an optical power corresponding to an element of the digital input vector. The optical power of the optical signals typically ranges, for example, from 1 μW to 10 mW. The OMM unit 150 receives the optical input vector and performs N×N matrix multiplication based on its internal configuration. The internal configuration is controlled by an electrical signal generated by the DAC unit 130. For example, the DAC unit 130 receives a second DAC control signal from the controller 110, which corresponds to the neural network weight to be implemented by the OMM unit 150. Based on the second DAC control signal, the DAC unit 130 generates a weight control signal, which is an analog signal suitable for controlling the reconfigurable element in the OMM unit 150. The analog signal may be, for example, a voltage or a current, depending on the type of reconfigurable element in the OMM unit 150. The voltage may have an amplitude ranging, for example, from 0.1 V to 10 V, and the current may have an amplitude ranging, for example, from 100 μA to 10 mA.

[0327] The array of modulators 144 may operate at a modulation rate different from the reconfiguration rate at which the OMM units 150 can be reconfigured. The optical input vector generated by the array of modulators 144 propagates through the OMM units at a significant fraction of the speed of light (e.g., 80%, 50%, or 25% of the speed of light), depending on the optical properties (e.g., effective refractive index) of the OMM units 150. In a typical OMM unit 150, the propagation time of the optical input vector ranges from 1 picosecond to tens of picoseconds, which corresponds to processing rates of tens to hundreds of GHz. Thus, the rate at which the optical processor 140 can perform matrix multiplication operations is limited in part by the rate at which it can generate the optical input vectors. Modulators with bandwidths of tens of GHz are readily available, and modulators with bandwidths exceeding 100 GHz are under development. Thus, the modulation rate of the array of modulators 144 may range from 5 GHz, 8 GHz, or tens to hundreds of GHz. To maintain operation of the array of modulators 144 at such modulation rates, the integrated circuit of the controller 110 may be configured to output control signals for the DAC units 130 at rates of, for example, 5 GHz, 8 GHz, 10 GHz, 20 GHz, 25 GHz, 50 GHz, or 100 GHz or higher.

[0328] The reconfiguration rate of OMM unit 150 can be much slower than the modulation rate, depending on the type of reconfigurable element implemented by OMM unit 150. For example, the reconfigurable element of OMM unit 150 may be a thermo-optical type that uses microheaters to adjust the temperature of the optical waveguides of OMM unit 150, which in turn affects the phase of the optical signals within OMM unit 150, leading to matrix multiplication. Thermal time constants associated with heating and cooling of the structure may limit the reconfiguration rate to, for example, hundreds of kHz to tens of MHz. Therefore, the modulator control signals for controlling the array of modulators 144 and the weight control signals for reconfiguring OMM unit 150 may have significantly different speed requirements. Furthermore, the electrical characteristics of the array of modulators 144 may differ significantly from the characteristics of the reconfigurable elements of OMM unit 150.

[0329] To accommodate different characteristics of the modulator control signal and the weight control signal, in some implementations, the DAC unit 130 may include a first DAC subunit 132 and a second DAC subunit 134. The first DAC subunit 132 may be specifically configured to generate the modulator control signal, and the second DAC subunit 134 may be specifically configured to generate the weight control signal. For example, the modulation rate of the modulator column 144 may be 25 GHz, and the first DAC subunit 132 may have a per-channel output update rate of 25 gigasamples per second (GSPS) and 8-bit or greater resolution. The reconstruction rate of the OMM unit 150 may be 1 MHz, and the second DAC subunit 134 may have an output update rate of 1 megasamples per second (MSPS) and 10-bit resolution. Implementing separate DAC subunits 132 and 134 allows independent optimization of the DAC subunits for each signal, which may lower the overall power consumption, complexity, cost, or a combination thereof, of the DAC unit 130. It should be noted that although DAC subunits 132 and 134 are described as subelements of DAC unit 130, in general, DAC subunits 132 and 134 may be integrated on a common chip or may be implemented as separate chips.

[0330] Based on the different characteristics of the first DAC subunit 132 and the second DAC subunit 134, in some implementations, the memory unit 120 may include a first memory subunit and a second memory subunit. The first memory subunit may be a memory dedicated to storing the input data set and the digital input vector and may have an operating speed sufficient to support the modulation rate. The second memory subunit may be a memory dedicated to storing the neural network weights and may have an operating speed sufficient to support the reconstruction rate of the OMM unit 150. In some implementations, the first memory subunit may be implemented using SRAM, and the second memory subunit may be implemented using DRAM. In some implementations, the first memory subunit may be implemented as part of the controller 110 or as a cache of the controller 110. In some implementations, the first memory subunit and the second memory subunit may be implemented by a single physical memory device as different address spaces.

[0331] The OMM unit 150 outputs an optical output vector of length N, which corresponds to the result of an N×N matrix multiplication of the optical input vector and the neural network weights. The OMM unit 150 is coupled to the detection unit 146, which is configured to generate N output voltages corresponding to the N optical signals of the optical output vector. For example, the detection unit 146 may include an array of N photodetectors configured to absorb the optical signals and generate photocurrents, and an array of N transimpedance amplifiers configured to convert the photocurrents into output voltages. The bandwidth of the photodetectors and transimpedance amplifiers may be set based on the modulation rate of the array of modulators 144. The photodetectors may be formed from various materials based on the wavelength of the optical output vector being detected. Example materials for the photodetectors include germanium, silicon-germanium alloys, and indium gallium arsenide (InGaAS).

[0332] The detection unit 146 is coupled to the ADC unit 160. The ADC unit 160 is configured to convert the N output voltages into N digitized optical outputs, where the N digitized optical outputs are quantized digital representations of the output voltages. For example, the ADC unit 160 may be an N-channel ADC. The controller 110 may obtain the N digitized optical outputs from the ADC unit 160, which correspond to the optical output vectors of the optical matrix multiplication unit 150. The controller 110 may form, from the N digitized optical outputs, a digital output vector of length N, which corresponds to the result of an N×N matrix multiplication of the input digital vector of length N.

[0333] The various electrical components of the ANN computing system 100 may be integrated in various ways. For example, the controller 110 may be an application-specific integrated circuit fabricated on a semiconductor die. Other electrical components, such as the memory unit 120, the DAC unit 130, the ADC unit 160, or a combination thereof, may be monolithically integrated on the semiconductor die on which the controller 110 is fabricated. As another example, two or more electrical components may be integrated as a system-on-chip (SoC). In an SoC implementation, the controller 110, the memory unit 120, the DAC unit 130, and the ADC unit 160 may be fabricated on respective dies, which may be integrated on a common platform (e.g., an interposer) that provides electrical connections between the integrated components. Such an SoC approach may enable faster data transfer between the electrical components of the ANN computing system 100 compared to an approach in which the components are separately placed and routed on a printed circuit board (PCB), thereby increasing the operating speed of the ANN computing system 100. Furthermore, an SoC approach can enable the use of different manufacturing technologies optimized for different electrical components, which can increase the performance of the various components and reduce the overall cost compared to monolithic integration approaches. Although integration of the controller 110, memory unit 120, DAC unit 130, and ADC unit 160 has been described, in general, a subset of components may be integrated, while other components are implemented as discrete components for various reasons, such as performance or cost. For example, in some implementations, the memory unit 120 may be integrated with the controller 110 as a functional block within the controller 110.

[0334] The various optical components of ANN computing system 100 may also be integrated in various ways. Examples of optical components of ANN computing system 100 include laser unit 142, modulator array 144, OMM unit 150, and photodetectors of detection unit 146. These optical components may be integrated in various ways to improve performance and / or reduce cost. For example, laser unit 142, modulator array 144, OMM unit 150, and photodetectors may be monolithically integrated on a common semiconductor substrate as a photonics integrated circuit (PIC). On a photonics integrated circuit formed based on a compound semiconductor material family (e.g., III-V compound semiconductors such as InP), lasers, modulators such as electroabsorption modulators, waveguides, and photodetectors may be monolithically integrated on a single die. Such a monolithic integration approach can reduce the complexity of aligning the inputs and outputs of various discrete optical components, which may require alignment accuracy ranging from less than one micron to several microns. As another example, the laser source of laser unit 142 may be fabricated on a compound semiconductor die, while the optical power divider of laser unit 142, modulator array 144, OMM unit 150, and photodetectors of detection unit 146 may be fabricated on a silicon die. PICs fabricated on silicon wafers, which may be referred to as silicon photonics technology, typically have higher integration density, higher lithographic resolution, and lower cost compared to III-V-based PICs. Such higher integration density may be beneficial in the fabrication of OMM unit 150, because OMM unit 150 typically includes tens to hundreds of optical components, such as power dividers and phase shifters. Furthermore, the higher lithographic resolution of silicon photonics technology may reduce the fabrication variability of OMM unit 150 and increase the precision of OMM unit 150.

[0335] The ANN computing system 100 can be implemented in a variety of form factors. For example, the ANN computing system 100 can be implemented as a coprocessor that plugs into a host computer. Such a system 100 can have, for example, a PCI Express form factor and communicate with the host computer via a PCIe bus. The host computer can host multiple coprocessor-type ANN computing systems 100 and connect to the computer 102 via a network. This type of implementation may be suitable for use in a cloud data center, where racks of servers may be dedicated to processing ANN computation requests received from other computers or servers. As another example, the coprocessor-type ANN computing system 100 can be plugged directly into the computer 102 that issues the ANN computation requests.

[0336] In some implementations, the ANN computing system 100 can be integrated into physical systems that require real-time ANN computing capabilities. For example, systems that rely heavily on real-time artificial intelligence tasks, such as autonomous vehicles, autonomous drones, object or face recognition security cameras, and various Internet-of-Things (IoT) devices, can benefit from having the ANN computing system 100 directly integrated with other subsystems of such systems. Having a directly integrated ANN computing system 100 can enable real-time artificial intelligence in devices with poor or no internet connectivity, increasing the reliability and availability of mission-critical artificial intelligence systems.

[0337] While the DAC unit 130 and the ADC unit 160 are shown as being coupled to the controller 110, in some implementations, the DAC unit 130, the ADC unit 160, or both may alternatively or additionally be coupled to the memory unit 120. For example, direct memory access (DMA) operations by the DAC unit 130 or the ADC unit 160 may reduce the computational load on the controller 110, reduce latency in reading from and writing to the memory unit 120, and further increase the operating speed of the ANN computation unit 100.

[0338] 2 shows a flowchart of an example method 200 for performing ANN calculations. The steps of process 200 may be performed by controller 110. In some implementations, various steps of method 200 may be performed in parallel, in combination, in a loop, or in any order.

[0339] At 210, an artificial neural network (ANN) computation request is received, the artificial neural network (ANN) computation request comprising an input dataset and a first plurality of neural network weights. The input dataset includes a first digital input vector. The first digital input vector is a subset of the input dataset. For example, it may be a subregion of an image. The ANN computation request may be generated by various entities, such as the computer 102. The computer may include one or more of various types of computing devices, such as a personal computer, a server computer, a vehicle computer, and a flight computer. The ANN computation request generally refers to an electrical signal that notifies or informs the ANN computation system 100 of an ANN computation to be performed. In some implementations, the ANN computation request may be divided into two or more signals. For example, a first signal may query the ANN computation system 100 to determine whether the system 100 is ready to receive the input dataset and the first plurality of neural network weights. In response to a positive response by the system 100, the computer may transmit a second signal including the input dataset and the first plurality of neural network weights.

[0340] At 220, the input data set and the first plurality of neural network weights are stored. The controller 110 may store the input data set and the first plurality of neural network weights in the memory unit 120. Storing the input data set and the first plurality of neural network weights in the memory unit 120 may enable flexibility in the operation of the ANN computation system 100, which may, for example, increase the overall performance of the system. For example, the input data set may be divided into digital input vectors of a set size and format by retrieving desired portions of the input data set from the memory unit 120. Different portions of the input data set may be processed in various orders or shuffled to enable various types of ANN computations to be performed. For example, shuffling may enable matrix multiplication using block matrix multiplication techniques when the input and output matrices are of different sizes. As another example, storing the input data set and the first plurality of neural network weights in memory unit 120 may enable queuing of multiple ANN computation requests by ANN computation system 100, which may allow system 100 to maintain operation at maximum speed without periods of inactivity.

[0341] In some implementations, the input data set may be stored in a first memory sub-unit and the first plurality of neural network weights may be stored in a second memory sub-unit.

[0342] At 230, a first plurality of modulator control signals are generated based on the first digital input vector, and a first plurality of weight control signals are generated based on the first plurality of neural network weights. Controller 110 may send first DAC control signals to DAC unit 130 for generating the first plurality of modulator control signals. DAC unit 130 generates the first plurality of modulator control signals based on the first DAC control signals, and modulator array 144 generates an optical input vector representing the first digital input vector.

[0343] The first DAC control signal may include a plurality of digital values ​​to be converted by the DAC unit 130 into the first plurality of modulator control signals. The plurality of digital values ​​may generally be related through various mathematical relationships or lookup tables according to the first digital input vector. For example, the plurality of digital values ​​may be linearly proportional to the values ​​of the elements of the first digital input vector. As another example, the plurality of digital values ​​may be related to the elements of the first digital input vector through a lookup table configured to maintain a linear relationship between the digital input vector and the optical input vector generated by the modulator array 144.

[0344] The controller 110 may send a second DAC control signal to the DAC unit 130 for generating the first plurality of weight control signals. The DAC unit 130 generates the first plurality of weight control signals based on the second DAC control signal, and the OMM unit 150 is reconfigured according to the first plurality of weight control signals to implement a matrix corresponding to the first plurality of neural network weights.

[0345] The second DAC control signal may include a plurality of digital values ​​to be converted by the DAC unit 130 into the first plurality of weight control signals. The plurality of digital values ​​may generally be related through various mathematical relationships or lookup tables according to the first plurality of neural network weights. For example, the plurality of digital values ​​may be linearly proportional to the first plurality of neural network weights. As another example, the plurality of digital values ​​may be calculated by performing various mathematical operations on the first plurality of neural network weights to generate weight control signals that may configure the OMM unit 150 to perform matrix multiplications corresponding to the first plurality of neural network weights.

[0346] In some implementations, the first plurality of neural network weights representing matrix M may be decomposed through singular value decomposition (SVD) into M=USV*, where U is an M×M unitary matrix, S is an M×N diagonal matrix with non-negative real numbers on the diagonal, and V* is the complex conjugate of the N×N unitary matrix V. In such a case, the first plurality of weight control signals may include a first plurality of OMM unit control signals corresponding to matrix V and a second plurality of OMM unit control signals corresponding to matrix S. Furthermore, OMM unit 150 may be configured to have a first OMM subunit configured to implement matrix V, a second OMM subunit configured to implement matrix S, and a third OMM subunit configured to implement matrix U, such that the entire OMM unit 150 implements matrix M. The SVD method is further described in U.S. Patent Application Publication No. 2017 / 0351293 A1, entitled "APPARATUS AND METHODS FOR OPTICAL NEURAL NETWORK," which is incorporated herein by reference in its entirety.

[0347] At 240, a first plurality of digitized optical outputs corresponding to the optical output vectors of the optical matrix multiplication unit are obtained. The optical input vectors generated by the modulator array 144 are processed and converted into optical output vectors by the OMM unit 150. The optical output vectors are detected by the detection unit 146 and converted into electrical signals that can be converted into digitized values ​​by the ADC unit 160. The controller 110 may, for example, send a conversion request to the ADC unit 160 to initiate the conversion of the voltages output by the detection unit 146 into digitized optical outputs. Once the conversion is complete, the ADC unit 160 may send the results of the conversion to the controller 110. Alternatively, the controller 110 may retrieve the results of the conversion from the ADC unit 160. The controller 110 may form a digital output vector from the digitized optical outputs that corresponds to the results of the matrix multiplication of the input digital vectors. For example, the digitized optical outputs may be organized or concatenated to have a vector format.

[0348] In some implementations, the ADC unit 160 may be configured or controlled to perform the ADC conversion based on a DAC control signal issued by the controller 110 to the DAC unit 130. For example, the ADC conversion may be configured to start a predetermined time after generation of the modulated control signal by the DAC unit 130. Such control of the ADC conversion may simplify the operation of the controller 110 and reduce the number of required control operations.

[0349] At 250, a nonlinear transformation is performed on the first digital output vector to generate a first transformed digital output vector. Nodes or artificial neurons of an ANN operate by first performing a weighted sum of signals received from nodes in the previous layer and then performing a nonlinear transformation ("activation") of the weighted sum to generate an output. Various types of ANNs may implement different types of distinguishable nonlinear transformations. Examples of nonlinear transfer functions include the rectified linear unit (RELU) function, the sigmoid function, the hyperbolic tangent function, the X^2 function, and the |X| function. Such a nonlinear transformation is performed on the first digital output by the controller 110 to generate the first transformed digital output vector. In some implementations, the nonlinear transformation may be performed by a specialized digital integrated circuit within the controller 110. For example, the controller 110 may include one or more modules or circuit blocks specifically adapted to accelerate the computation of one or more types of nonlinear transformations.

[0350] At 260, the first transformed digital output vector is stored. The controller 110 may store the first transformed digital output vector in the memory unit 120. If the input data set is divided into multiple digital input vectors, the first transformed digital output vector corresponds to the result of an ANN computation of a portion of the input data set, such as the first digital input vector. Thus, storing the first transformed digital output vector enables the ANN computation system 100 to perform and store additional computations on other digital input vectors of the input data set that will later be aggregated into a single ANN output.

[0351] At 270, an artificial neural network output generated based on the first transformed digital output vector is output. The controller 110 generates an ANN output, the ANN output being a result of processing the input dataset through the ANN defined by the first plurality of neural network weights. If the input dataset is split into multiple digital input vectors, the generated ANN output is an aggregate output that includes the first transformed digital output, but may also include additional transformed digital outputs corresponding to other portions of the input dataset. Once the ANN output is generated, it is transmitted to a computer, such as computer 102, that issued the ANN calculation request.

[0352] Various performance measures may be defined for the ANN computing system 100 implementing the method 200. Defining the performance measures may enable comparison of the performance of the ANN computing system 100 implementing the optical processor 140 with the performance of other systems for ANN computation that instead implement an electronic matrix multiplication unit. In one aspect, the rate at which the ANN computation may be performed may be indicated, in part, by a first loop period, defined as the time elapsed between step 220, storing the input data set and the first plurality of neural network weights in a memory unit, and step 260, storing the first converted digital output vector in a memory unit. This first loop period therefore includes the time it takes to convert electrical signals to optical signals (e.g., step 230), perform the matrix multiplication in the optical domain, and convert the result back to the electrical domain (e.g., step 240). Both steps 220 and 260 involve storing data in the memory unit 120, and are steps shared between the ANN computing system 100 and conventional ANN computing systems that do not involve the optical processor 140. Thus, the first loop period, which measures the transaction time between memories, may allow a realistic or fair comparison of ANN computational throughput to be made between the ANN computation system 100 and an ANN computation system that does not involve an optical processor 140, such as a system that implements an electronic matrix multiplication unit.

[0353] Due to the rate at which optical input vectors can be generated by the array of modulators 144 (e.g., 25 GHz) and the processing rate of the OMM unit 150 (e.g., >100 GHz), the first loop period of the ANN computation system 100 to perform a single ANN computation of a single digital input vector may approach the inverse of the speed of the array of modulators 144, e.g., 40 ps. After considering the latency associated with signal generation by the DAC unit 130 and ADC conversion by the ADC unit 160, the first loop period may be, for example, 100 ps or less, 200 ps or less, 500 ps or less, 1 ns or less, 2 ns or less, 5 ns or less, or 10 ns or less.

[0354] By comparison, the execution time of an Mx1 vector multiplication by an MxM matrix by an electronic matrix multiplication unit is typically proportional to M^2-1 processor clock cycles. For M=32, such a multiplication takes approximately 1024 cycles, which at a 3 GHz clock speed translates to an execution time of over 300 ns, several orders of magnitude slower than the first loop period of the ANN computational system 100.

[0355] In some implementations, method 200 further includes generating a second plurality of modulator control signals based on the first transformed digital output vector. In some types of ANN computations, a single digital input vector may be propagated repeatedly through or processed by the same ANN. An ANN that implements multi-pass processing may be referred to as a recurrent neural network (RNN). An RNN is a neural network in which the output of the network during the kth pass through the neural network is recirculated to the input of the neural network to be used as the input during the k+1th pass. RNNs may have various applications in pattern recognition tasks, such as speech recognition or handwriting recognition. Once the second plurality of modulator control signals is generated, method 200 may proceed from step 240 to step 260 to complete a second pass of the first digital input vector through the ANN. Generally, the recirculation of the transformed digital output as the digital input vector may be repeated for a preset number of cycles depending on the characteristics of the RNN received in the ANN computation request.

[0356] In some implementations, the method 200 further includes generating a second plurality of weight control signals based on the second plurality of neural network weights. In some cases, the artificial neural network computation request further includes the second plurality of neural network weights. Generally, an ANN has one or more hidden layers in addition to an input layer and an output layer. In an ANN with two hidden layers, the second plurality of neural network weights may correspond, for example, to the connectivity between the first layer of the ANN and the second layer of the ANN. To process the first digital input vector through the two hidden layers of the ANN, the first digital input vector may first be processed according to the method 200 up to step 260, where the results of processing the first digital input vector through the first hidden layer of the ANN are stored in the memory unit 120. The controller 110 then reconfigures the OMM unit 150 to perform a matrix multiplication corresponding to the second plurality of neural network weights associated with the second hidden layer of the ANN. Once the OMM unit 150 is reconfigured, the method 200 can generate a plurality of modulator control signals based on the first transformed digital output vector, thereby generating an updated optical input vector corresponding to the output of the first hidden layer. The updated optical input vector is then processed by the reconfigured OMM unit 150 corresponding to the second hidden layer of the ANN. In general, the described steps can be repeated until the digital input vector has been processed through all hidden layers of the ANN.

[0357] As previously described, in some implementations of OMM unit 150, the reconfiguration rate of OMM unit 150 may be much slower than the modulation rate of modulator array 144. In such cases, the throughput of ANN computation system 100 may be adversely affected by the length of time it takes to reconfigure OMM unit 150, during which ANN computations cannot be performed. To mitigate the impact of the relatively slow reconfiguration time of OMM unit 150, batch processing techniques may be utilized in which two or more digital input vectors are propagated through OMM unit 150 without a configuration change to amortize the reconfiguration time over a large number of digital input vectors.

[0358] FIG. 2B shows a diagram 290 illustrating an embodiment of the method 200 of FIG. 2A. In an ANN with two hidden layers, instead of processing a first digital input vector through the first hidden layer, reconfiguring the OMM unit 150 for the second hidden layer, processing the first digital input vector through the reconfigured OMM unit 150, and repeating the same for the remaining digital input vectors, all digital input vectors of the input data set may first be processed through the OMM unit 150 configured for the first hidden layer (configuration #1), as shown in the upper portion of diagram 290. Once all digital input vectors have been processed by the OMM unit 150 having configuration #1, the OMM unit 150 is reconfigured to configuration #2, which corresponds to the second hidden layer of the ANN. This reconfiguration may be much slower than the rate at which input vectors can be processed by the OMM unit 150. Once the OMM unit 150 is reconfigured for the second hidden layer, output vectors from the previous hidden layer may be processed by the OMM unit 150 in batches. For large input data sets having tens or hundreds of thousands of digital input vectors, the impact of the reconstruction time may be reduced by approximately the same factor, which may significantly reduce the portion of time spent by the ANN computing system 100 in reconstruction.

[0359] To perform batch processing, in some implementations, method 200 further includes generating, through a DAC unit, a second plurality of modulator control signals based on a second digital input vector; obtaining, from the ADC unit, a second plurality of digitized optical outputs corresponding to the optical output vector of the optical matrix multiplication unit, where the second plurality of digitized optical outputs form the second digital output vector; performing a nonlinear transformation on the second digital output vector to generate a second transformed digital output vector; and storing the second transformed digital output vector in a memory unit. The generation of the second plurality of modulator control signals may occur after step 260, for example. Furthermore, the ANN output of step 270 in this case is now based on both the first transformed digital output vector and the second transformed digital output vector. The obtaining, executing, and storing steps are similar to steps 240 through 260.

[0360] Batch processing techniques are one of several techniques for increasing the throughput of ANN computing system 100. Another technique for increasing the throughput of ANN computing system 100 is through parallel processing of multiple digital input vectors by utilizing wavelength division multiplexing (WDM). WDM is a technique for simultaneously propagating multiple optical signals of different wavelengths through a common propagation channel, such as the waveguides of OMM unit 150. Unlike electrical signals, optical signals of different wavelengths can propagate through a common channel without affecting other optical signals of different wavelengths on the same channel. Furthermore, optical signals can be added (multiplexed) or removed (demultiplexed) from the common propagation channel using well-known structures, such as optical multiplexers and demultiplexers.

[0361] In the context of the ANN computing system 100, multiple optical input vectors of different wavelengths may be independently generated, simultaneously propagated through the OMM unit 150, and independently detected to increase the throughput of the ANN computing system 100. Referring to FIG. 1F, a schematic diagram of an example wavelength division multiplexing (WDM) artificial neural network (ANN) computing system 104 is shown. The WDM ANN computing system 104 is similar to the ANN computing system 100 unless otherwise described. To implement WDM techniques, in some implementations of the ANN computing system 104, the laser unit 142 is configured to generate multiple wavelengths, such as λ1, λ2, and λ3. The multiple wavelengths may preferably be separated by a wavelength spacing large enough to allow for easy multiplexing and demultiplexing onto a common propagation channel. For example, wavelength spacings of 0.5 nm, 1.0 nm, 2.0 nm, or greater than 3.0 nm or 5.0 nm may allow for easy multiplexing and demultiplexing. On the other hand, the range between the shortest and longest wavelengths of the multiple wavelengths (the "WDM bandwidth") may preferably be small enough so that the characteristics or performance of the OMM unit 150 remain substantially the same across the multiple wavelengths. Optical components are typically dispersive, meaning that their optical properties vary depending on wavelength. For example, the power splitting ratio of an MZI may vary across wavelength. However, by designing the OMM unit 150 to have a sufficiently large operating wavelength interval and by restricting the wavelengths to be within that operating wavelength interval, the optical power vector output by the OMM unit 150 at each wavelength may be a sufficiently accurate result of the matrix multiplication performed by the OMM unit 150. The operating wavelength interval may be, for example, 1 nm, 2 nm, 3 nm, 4 nm, 5 nm, 10 nm, or 20 nm.

[0362] 39A shows a diagram of an example Mach-Zehnder modulator 3900 that can be used to modulate the amplitude of an optical signal. The Mach-Zehnder modulator 3900 includes two 1x2-port multimode interference couplers (MMI_1x2) 3902a and 3902b, two balanced arms 3904a and 3904b, and a phase shifter 3906 in one arm (or one phase shifter in each arm). When a voltage is applied to the phase shifter in one arm through a signal line 3908, there is a phase difference between the two arms 3904a and 3904b that is converted to amplitude modulation. The 1x2-port multimode interference couplers 3902a and 3902b and the phase shifter 3906 are configured as broadband photonic components, and the optical path lengths of the two arms 3904a and 3904b are configured to be equal. This allows the Mach-Zehnder modulator 3900 to operate over a wide wavelength range.

[0363] Figure 39B is a graph 3910 showing intensity versus voltage curves for a Mach-Zehnder modulator 3900 using the configuration shown in Figure 39A for wavelengths of 1530 nm, 1550 nm, and 1570 nm. Graph 3910 shows that the Mach-Zehnder modulator 3900 has similar intensity versus voltage characteristics for various wavelengths ranging from 1530 nm to 1570 nm.

[0364] Returning to FIG. 1F , modulator array 144 of WDM ANN computational system 104 includes banks of optical modulators configured to generate multiple optical input vectors, each of which corresponds to one of multiple wavelengths and generates a respective optical input vector having a respective wavelength. For example, in a system with optical input vectors of length 32 and three wavelengths (e.g., λ1, λ2, and λ3), modulator array 144 may have three banks of 32 modulators each. Furthermore, modulator array 144 also includes an optical multiplexer configured to combine the multiple optical input vectors into a combined optical input vector including multiple wavelengths. For example, the optical multiplexer may combine the outputs of the three banks of modulators at three different wavelengths into a single propagation channel, such as a waveguide, for each element of the optical input vector. Thus, returning to the example above, the combined optical input vector would have 32 optical signals, each signal including three wavelengths.

[0365] Additionally, the detection unit 146 of the WDM ANN computation system 104 is further configured to demultiplex the multiple wavelengths to generate multiple demultiplexed output voltages. For example, the detection unit 146 may include a demultiplexer configured to demultiplex three wavelengths included in each of 32 signals of the multi-wavelength optical output vector and route the three single-wavelength optical output vectors to three banks of photodetectors coupled to three banks of transimpedance amplifiers.

[0366] Additionally, the ADC unit 160 of the WDM ANN computing system 104 includes a bank of ADCs configured to convert the multiple demultiplexed output voltages of the detection unit 146, each corresponding to one of the multiple wavelengths and generating a respective digitized demultiplexed optical output. For example, the bank of ADCs may be coupled to a bank of transimpedance amplifiers of the detection unit 146.

[0367] Controller 110 may implement a method similar to method 200 but extended to support multi-wavelength operation. For example, the method may include obtaining a plurality of digitized demultiplexed optical outputs from ADC unit 160, where the plurality of digitized demultiplexed optical outputs form a plurality of first digital output vectors, each of the plurality of first digital output vectors corresponding to one of the plurality of wavelengths, performing a nonlinear transform on each of the plurality of first digital output vectors to generate a plurality of transformed first digital output vectors, and storing the plurality of transformed first digital output vectors in a memory unit.

[0368] In some cases, the ANN may be specially designed, or the digital input vector specially formed, so that the multi-wavelength optical output vector can be detected without demultiplexing. In such cases, detection unit 146 may be a wavelength-independent detection unit that does not demultiplex the multiple wavelengths of the multi-wavelength optical output vector. Thus, each of the photodetectors of detection unit 146 effectively sums the multiple wavelengths of the optical signal into a single photocurrent, and each of the voltages output by detection unit 146 corresponds to the element-by-element sum of the matrix multiplication results of the multiple digital input vectors.

[0369] Up until now, nonlinear transformations of weighted sums performed as part of ANN computations have been performed in the digital domain by controller 110. In some cases, the nonlinear transformations may be computationally intensive or power-intensive, significantly increasing the complexity of controller 110, or otherwise constraining the performance of ANN computing system 100 in terms of throughput or power efficiency. Therefore, in some implementations of ANN computing systems, the nonlinear transformations may be performed in the analog domain through analog electronic circuitry.

[0370] 3A shows a schematic diagram of an example ANN computing system 300. ANN computing system 300 is similar to ANN computing system 100, except for the addition of an analog nonlinearity unit 310. Analog nonlinearity unit 310 is disposed between detection unit 146 and ADC unit 160. Analog nonlinearity unit 310 is configured to receive the output voltage from detection unit 146, apply a nonlinear transfer function, and output a converted output voltage to ADC unit 160.

[0371] As ADC unit 160 receives the voltage nonlinearly converted by analog nonlinearity unit 310, controller 110 may obtain a converted digitized output voltage corresponding to the converted output voltage from ADC unit 160. Because the digitized output voltage obtained from ADC unit 160 has already been nonlinearly converted (“activated”), the nonlinear conversion step by controller 110 can be omitted, reducing the computational burden on controller 110. The first converted voltage obtained directly from ADC unit 160 may then be stored in memory unit 120 as a first converted digital output vector.

[0372] The analog nonlinearity unit 310 can be implemented in various ways, for example, a high gain amplifier in a feedback configuration, a comparator with an adjustable reference voltage, the nonlinear IV characteristic of a diode, the breakdown behavior of a diode, the nonlinear CV characteristic of a variable capacitor, or the nonlinear IV characteristic of a variable resistor can be used to implement the analog nonlinearity unit 310.

[0373] The use of the analog nonlinearity unit 310 can increase the performance of the ANN computing system 300, such as throughput or power efficiency, by reducing the steps that must be performed in the digital domain. Moving the nonlinear transformation steps out of the digital domain may allow for greater flexibility and improvement in the operation of the ANN computing system. For example, in a recurrent neural network, the output of the OMM unit 150 is activated and recycled to the input of the OMM unit 150. The activation is performed by the controller 110 in the ANN computing system 100, which requires digitizing the output voltage of the detection unit 146 on every single pass through the OMM unit 150. However, because the activation is now performed before digitization by the ADC unit 160, it may be possible to reduce the number of ADC conversions required when performing a recurrent neural network computation.

[0374] In some implementations, the analog nonlinearity unit 310 may be integrated into the ADC unit 160 as a nonlinear ADC unit. For example, the nonlinear ADC unit may be a linear ADC unit with a nonlinear lookup table that maps the linear digitized output of the linear ADC unit to a desired nonlinearly converted digitized output.

[0375] 3B shows a schematic diagram of an example ANN computing system 302. The ANN computing system 302 is similar to the system 300 of FIG. 3A, except that it further includes an analog memory unit 320. The analog memory unit 320 is coupled to the DAC unit 130 (e.g., through the first DAC sub-unit 132), the array of modulators 144, and the analog nonlinearity unit 310. The analog memory unit 320 includes a multiplexer having a first input coupled to the DAC unit 130 and a second input coupled to the analog nonlinearity unit 310. This allows the analog memory unit 320 to receive a signal from either the DAC unit 130 or the analog nonlinearity unit 310. The analog memory unit 320 is configured to store an analog voltage and output the stored analog voltage.

[0376] The analog memory unit 320 may be implemented in various ways. For example, an array of capacitors may be used as analog voltage storage elements. The capacitors of the analog memory unit 320 may be charged to the input voltage by a charging circuit. The storage of the input voltage may be controlled based on a control signal received from the controller 110. The capacitors may be electrically isolated from the surrounding environment to reduce charge leakage that would cause the capacitors to undesirably discharge. Additionally or alternatively, a feedback amplifier may be used to maintain the voltage stored in the capacitors. The stored voltage of the capacitors may be read by a buffer amplifier, which allows the charge stored by the capacitors to be preserved while outputting the stored voltage. These aspects of the analog memory unit 320 may be similar to the operation of a sample and hold circuit. The buffer amplifier may implement the function of a modulator driver for driving the modulator column 144.

[0377] The operation of the ANN computing system 302 will now be described. The first plurality of modulator control signals output by the DAC unit 130 (e.g., by the first DAC sub-unit 132) are first input to the modulator column 144 through the analog memory unit 320. In this step, the analog memory unit 320 may simply pass or buffer the first plurality of modulator control signals. The modulator column 144 generates an optical input vector based on the first plurality of modulator control signals, which propagates through the OMM unit 150 and is detected by the detection unit 146. The output voltage of the detection unit 146 is nonlinearly converted by the analog nonlinearity unit 310. At this point, instead of being digitized by the ADC unit 160, the output voltage of the detection unit 146 is stored by the analog memory unit 320, which is then converted into the next optical input vector to be output to the modulator column 144 and propagated through the OMM unit 150. This recursive process may be performed for a preset length of time or a preset number of cycles under the control of the controller 110. Once the recursive process is completed for a given digital input vector, the converted output voltage of the analog nonlinearity unit 310 is converted by the ADC unit 160.

[0378] The use of analog memory unit 320 can significantly reduce the number of ADC conversions during recurrent neural network computation, such as to a single ADC conversion per RNN computation for a given digital input vector. Each ADC conversion takes some time and consumes some amount of energy. Thus, the throughput of RNN computation by ANN computation system 302 can be higher than the throughput of RNN computation by ANN computation system 100.

[0379] The execution of the recurrent neural network computations may be controlled, for example, by controlling analog memory unit 320. For example, the controller may control analog memory unit 320 to store a voltage at one time and output the stored voltage at a different time. Thus, the circulation of signals from analog memory unit 320, through analog nonlinearity unit 310 to modulator bank 144, and back to analog memory unit 320 may be controlled by controller 110 by controlling the storage and reading of analog memory unit 320.

[0380] Thus, in some implementations, the controller 110 of the ANN computation system 302 may perform the steps of storing a plurality of converted output voltages of the analog nonlinearity unit through an analog memory unit based on generating the first plurality of modulator control signals and the first plurality of weight control signals; outputting the stored converted output voltages through the analog memory unit; obtaining a second plurality of converted digitized output voltages from the ADC unit, where the second plurality of converted digitized output voltages form a second converted digital output vector; and storing the second converted digital output vector in the memory unit.

[0381] Input data sets to be processed by ANN computing systems typically contain data with resolution greater than 1 bit. For example, a typical pixel in a grayscale digital image may have 8 bits of resolution, or 256 different levels. One way to represent and process this data in the optical domain is to encode the pixel's 256 different intensity levels as 256 different power levels of the optical signal input to OMM unit 150. Because optical signals are inherently analog, they are susceptible to noise and detection errors. Returning to FIG. 1A , to maintain the 8-bit resolution of the digital input vector throughout ANN computing system 100 and generate a true 8-bit digitized optical output at the output of ADC unit 160, each and every part of the signal chain can preferably be designed to replicate and maintain 8-bit resolution.

[0382] For example, DAC unit 130 may preferably be designed to support conversion of an 8-bit digital input vector into a modulator control signal with at least 8 bits of resolution, so that modulator column 144 can generate an optical input vector that faithfully represents the 8 bits of the digital input vector. Generally, the modulator control signal may need to have additional resolution beyond the 8 bits of the digital input vector to compensate for the nonlinear response of modulator column 144. Furthermore, the internal configuration of OMM unit 150 may preferably be sufficiently stabilized to ensure that the value of the optical output vector is not corrupted by any variations in the configuration of OMM unit 150. For example, the temperature of OMM unit 150 may need to be stabilized within, e.g., 5 degrees, 2 degrees, 1 degree, or 0.1 degrees. Furthermore, detection unit 146 may preferably be sufficiently quiet so as not to impair the 8-bit resolution of the optical output vector, and ADC unit 160 may preferably be designed to support digitization of analog voltages with at least 8 bits of resolution.

[0383] The power consumption and design complexity of various electronic components typically increase with bit resolution, operating speed, and bandwidth. For example, to a first approximation, the power consumption of the ADC unit 160 may scale linearly with the sampling rate and by a factor of 2^N, where N is the bit resolution of the conversion result. Furthermore, design considerations for the DAC unit 130 and the ADC unit 160 typically result in a trade-off between sampling rate and bit resolution. Thus, in some cases, it may be desirable for an ANN computation system to internally operate at a lower bit resolution than the resolution of the input data set, while maintaining the resolution of the ANN computation output.

[0384] 4A, there is shown a schematic diagram of an example of an artificial neural network (ANN) computing system 400 with 1-bit internal resolution. ANN computing system 400 is similar to ANN computing system 100, except that DAC unit 130 is now replaced with a driver unit 430 and ADC unit 160 is now replaced with a comparator unit 460.

[0385] The driver unit 430 is configured to generate a 1-bit modulator control signal and a multi-bit weighted control signal. For example, the driver circuit of the driver unit 430 may receive the binary digital output directly from the controller 110 and condition the binary signal into a two-level voltage or current output suitable for driving the array of modulators 144.

[0386] The comparator unit 460 is configured to convert the output voltage of the detection unit 146 into a digitized 1-bit optical output. For example, a comparator circuit in the comparator unit 460 may receive the voltage from the detection unit 146, compare the voltage with a preset threshold voltage, and output either a digital 0 when the received voltage is less than the preset threshold voltage, or a digital 1 when the received voltage is greater than the threshold voltage.

[0387] 4B, there is shown a mathematical representation of the operation of ANN computation system 400. The operation of ANN computation system 400 will now be described with reference to FIG. 4B. For a given ANN computation to be performed by ANN computation system 400, there is a corresponding digital input vector V and neural network weight matrix U. In this example, input vector V is a vector of length 4 with elements V0 through V3, and matrix U is a vector of weights U 00 From U 33 Each element of the vector V has 4 bits of resolution. Each 4-bit vector element has the 0th bit (bit0) through the 3rd bit (bit3) corresponding to positions 2^0 through 2^3. Thus, the decimal (base 10) value of a 4-bit vector element is calculated by adding 2^0*bit0+2^1*bit1+2^2*bit2+2^3*bit3. Thus, the input vector V is similarly converted by the controller 110 into V as shown. bit0 From V bit3 can be decomposed into

[0388] An ANN computation can then be performed by performing a series of matrix multiplications of 1-bit vectors followed by addition of the results of each matrix multiplication. For example, given a decomposed input vector V bit0 From V bit3 may be multiplied with matrix U through driver unit 430 by generating a sequence of four 1-bit modulator control signals corresponding to the four 1-bit input vectors. This then generates a sequence of four 1-bit optical input vectors, which propagate through OMM unit 150, which is configured through driver unit 430 to perform matrix multiplication of matrix U. Controller 110 may then obtain from comparator unit 460 a sequence of four digitized 1-bit optical outputs corresponding to the sequence of the four 1-bit modulator control signals.

[0389] In this case, where a 4-bit vector is decomposed into four 1-bit vectors, each vector should be processed by ANN computation system 400 four times faster than a single 4-bit vector could be processed by other ANN computation systems, such as system 100, to maintain the same effective ANN computation throughput. Such an increase in internal processing speed can be viewed as time-division multiplexing the four 1-bit vectors into a single time slot for processing the 4-bit vector. The necessary increase in processing speed may be achieved at least in part by increasing the operating speed of driver unit 430 and comparator unit 460 relative to DAC unit 130 and ADC unit 160, since a decrease in the resolution of the signal conversion process typically leads to an increase in the rate of signal conversion that can be achieved.

[0390] Although the signal conversion rate is improved by a factor of four for 1-bit operations, the resulting power consumption may be significantly reduced compared to 4-bit operations. As previously explained, the power consumption of signal conversion processes typically scales exponentially with bit resolution but linearly with conversion rate. Thus, a 16-fold reduction in power per conversion may result from a 4-fold reduction in bit resolution followed by a 4-fold increase in power due to the increased conversion rate. Overall, a 4-fold reduction in operating power may be achieved by ANN computing system 400 while maintaining the same effective ANN computational throughput, for example, compared to ANN computing system 100.

[0391] The controller 110 may then construct a 4-bit digital output vector from the four digitized 1-bit optical outputs by multiplying each of the digitized 1-bit optical outputs by a respective weight from 2^0 to 2^3. Once the 4-bit digital output vectors are constructed, the ANN computation may proceed by performing a nonlinear transformation on the constructed 4-bit digital output vector to generate a transformed 4-bit digital output vector and storing the transformed 4-bit digital output vector in the memory unit 120.

[0392] Alternatively, or in addition, in some implementations, each of the four digitized 1-bit optical outputs may be nonlinearly transformed. For example, a step function nonlinear function may be used for the nonlinear transformation. A transformed 4-bit digital output vector may then be constructed from the nonlinearly transformed digitized 1-bit optical outputs.

[0393] Although a separate ANN computing system 400 has been illustrated and described, in general, the ANN computing system 100 of Figure 1A may be designed to implement functionality similar to that of the ANN computing system 400. For example, the DAC unit 130 may include a 1-bit DAC subunit configured to generate a 1-bit modulator control signal, and the ADC unit 160 may be designed to have 1-bit resolution. Such a 1-bit ADC may be similar to or substantially equivalent to a comparator.

[0394] Furthermore, although the operation of the ANN computing system has been described with an internal resolution of 1 bit, in general, the internal resolution of the ANN computing system may be reduced to an intermediate level lower than the N-bit resolution of the input data set. For example, the internal resolution may be reduced to 2^Y bits, where Y is an integer greater than or equal to 0.

[0395] Embodiments of the subject matter and functional operations described herein may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. Embodiments of the subject matter described herein may be implemented using one or more modules of computer program instructions encoded on a computer-readable medium for execution by or to control the operation of a data processing apparatus. The computer-readable medium may be a manufactured product, such as a hard drive in a computer system or an optical disk sold at a retail store, or an embedded system. The computer-readable medium may be obtained separately and subsequently encode one or more modules of computer program instructions, such as by distribution of one or more modules of computer program instructions over a wired or wireless network. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, or one or more combinations thereof.

[0396] A computer program (also known as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted, declarative or procedural, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program may be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communications network.

[0397] The processes and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0398] While this specification contains many implementation details, these should not be construed as limitations on the scope of the invention or what may be claimed, but rather as descriptions of features specific to particular embodiments of the invention. Some features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in a certain combination, and may even be initially claimed as such, one or more features from a claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0399] Similarly, while operations are shown in the figures in a particular order, this should not be understood as requiring that such operations be performed in the particular or sequential order shown, or that all of the shown operations be performed, to achieve desirable results. In some situations, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.

[0400] Thus, specific embodiments of the present invention have been described. Other embodiments are within the scope of the following claims. In addition, the actions recited in the claims may be performed in a different order and still achieve desirable results. For example, the optical matrix multiplication unit 150 of FIG. 1A includes an optical interferometer unit 154 including multiple interconnected Mach-Zehnder interferometers. In some implementations, the optical interferometer unit may be implemented using one-dimensional, two-dimensional, or three-dimensional passive diffractive optical elements that consume little power. Compared to an optical interferometer unit including a Mach-Zehnder interferometer, an optical interferometer unit using passive diffractive optical elements may be smaller in size if the number of inputs / outputs remains the same, or may be able to handle a larger number of inputs / outputs for the same chip size. Passive diffractive optical elements may be manufactured at a lower cost than a Mach-Zehnder interferometer.

[0401] Referring to Figure 5, in some implementations, an artificial neural network computing system 500 includes a controller 110, a memory unit 120, a DAC unit 506, an optical processor 504, and an ADC unit 160. The memory unit 120 and the ADC unit 160 are similar to the corresponding components in the system 100 of Figure 1A. The optical processor 504 is configured to perform matrix computations using optical components. In the system 500, the weights for the optical matrix multiplication unit 502 are fixed. The DAC unit 506 is similar to the first DAC sub-unit 132 in the system 100 of Figure 1A.

[0402] In an exemplary operation of the ANN computation system 500, the computer 102 may issue an artificial neural network computation request to the ANN computation system 500. The ANN computation request may include an input data set to be processed by the provided ANN. The controller 110 receives the ANN computation request and stores the input data set in the memory unit 120.

[0403] In some implementations, a hybrid approach is used, such that one portion of the optical matrix multiplication unit 150 includes a Mach-Zehnder interferometer and another portion of the optical matrix multiplication unit 150 includes a passive diffractive element.

[0404] The internal operation of ANN computing system 500 will now be described. Optical processor 504 includes laser unit 142, array of modulators 144, detection unit 146, and optical matrix multiplication (OMM) unit 502. Laser unit 142, array of modulators 144, and detection unit 146 are similar to the corresponding components of system 100 of FIG. 1A. In this example, OMM unit 502 includes a two-dimensional diffractive optical element and may be implemented as a passively integrated silicon photonic chip. Optical matrix multiplication unit 502 may be configured to implement a diffractive neural network and can perform matrix multiplication with little power consumption.

[0405] Optical processor 504 operates by encoding a digital input vector of length N into an optical input vector of length N and propagating the optical input vector through OMM unit 502. OMM unit 502 receives the optical input vector of length N and performs an N×N matrix multiplication on the received optical input vector in the optical domain. The N×N matrix multiplication performed by OMM unit 502 is determined by the internal configuration of OMM unit 502. The internal configuration of OMM unit 502 includes, for example, the dimensions, arrangement, and geometry of the diffractive optical elements, as well as impurity doping, if any.

[0406] The OMM unit 502 can be implemented in various ways. FIG. 6 shows a schematic diagram of an example OMM unit 502 that uses a two-dimensional array of diffractive elements. The OMM unit 502 may include an array of input waveguides 602 for receiving optical input vectors, a two-dimensional optical interference unit 600 in optical communication with the array of input waveguides 602, and an array of output waveguides 604 in optical communication with the optical interference unit 600. The optical interference unit 600 includes a plurality of diffractive optical elements and performs a transformation (e.g., a linear transformation) of the optical input vectors into a second array of optical signals. The array of output waveguides 604 directs the second array of optical signals output by the optical interference unit 600. At least one input waveguide in the array of input waveguides 602 is in optical communication with each output waveguide in the array of output waveguides 604 via the optical interference unit 600. For example, for an optical input vector of length N, the OMM unit 502 may include N input waveguides 602 and N output waveguides 604.

[0407] In some implementations, the optical interference unit 600 includes a substrate having diffractive elements arranged two-dimensionally (e.g., in a 2D array). For example, a plurality of circular holes may be drilled or etched into the substrate. The holes have dimensions on the same order as the wavelength of the input light so that light is diffracted by the holes (or the structures defining the holes). For example, the dimensions of the holes may range from 100 nm to 2 μm. The holes may have the same or different sizes. The holes may also have other cross-sectional shapes, such as triangular, square, rectangular, hexagonal, or irregular shapes. The substrate may be made of a material that is transparent or semi-transparent to the input light, e.g., having a transmittance of 1% to 99% for the input light. For example, the substrate may be made of silicon, silicon oxide, silicon nitride, quartz, crystal (e.g., lithium niobate, LiNbO), a III-V material such as gallium arsenide or indium phosphide, an erbium-doped semiconductor, or a polymer.

[0408] In some implementations, holographic methods can be used to form two-dimensional diffractive optical elements in a substrate, which can be made of glass, quartz, or a photorefractive material.

[0409] When designing the OMM unit 502, one considers the size and placement of the diffractive elements in two dimensions (e.g., x and y directions) without considering the relative placement of the diffractive elements in a third dimension (e.g., z direction). Each diffractive element can be a three-dimensional structure, such as a hole, a cylinder, or a stripe with a certain depth formed in a substrate.

[0410] In FIG. 6 , diffractive optical elements are represented by circles. Diffractive optical elements may also have other shapes, such as triangles, squares, rectangles, or irregular shapes. Diffractive optical elements may have various sizes. Diffractive optical elements may not be located on lattice points, and their positions may be variable. The illustration in FIG. 6 is for illustrative purposes only. Actual diffractive optical elements may differ from those shown in the figure. Different configurations of diffractive optical elements may be used to perform different matrix operations, such as different matrix multiplication functions.

[0411] The configuration of the diffractive optical element can be determined using an optimization process. For example, a substrate can be divided into columns of pixels, and each pixel can be either filled with substrate material (no holes) or filled with air (holes). The pixel configuration can be iteratively modified, and simulations can be run for each pixel configuration by passing light through the diffractive optical element and evaluating the output. After simulations of all possible pixel configurations have been run, the configuration that produces results that most closely resemble the desired matrix processing is selected as the diffractive optical element configuration for the OMM unit 502.

[0412] As another example, a diffractive element may be initially configured as an array of holes. The placement, size, and shape of the holes may be slightly varied from their initial configuration. The parameters of each hole may be iteratively adjusted, and simulations may be performed to find the optimal configuration for the holes.

[0413] In some implementations, a machine learning process is used to design the diffractive optical element: an analytical function for how pixels affect input light to produce output light is determined, and an optimization process (e.g., gradient descent) is used to determine the optimal configuration of the pixels.

[0414] In some implementations, the OMM unit 502 may be implemented as a user-changeable component, and different OMM units 502 with different optical interferometry units 600 may be installed for different applications. For example, the system 500 may be configured as an optical character recognition system, and the optical interferometry units 600 may be configured to implement neural networks for performing optical character recognition. For example, a first OMM unit may have a first optical interferometry unit including a passive diffractive optical element configured to implement a first neural network for an optical character recognition engine for a first set of written languages ​​and fonts. A second OMM unit may have a second optical interferometry unit including a passive diffractive optical element configured to implement a second neural network for an optical character recognition engine for a second set of written languages ​​and fonts, and so on. When a user desires to use the system 500 to apply optical character recognition to the first set of written languages ​​and fonts, the user can insert the first OMM unit into the system. When a user desires to use system 500 to apply optical character recognition to a second set of written languages ​​and fonts, the user can replace the first OMM unit and insert a second OMM unit into the system.

[0415] For example, system 500 may be configured as a speech recognition system, and optical interferometry unit 600 may be configured to implement a neural network for performing speech recognition. For example, a first OMM unit may have a first optical interferometry unit including a passive diffractive optical element configured to implement a first neural network for a speech recognition engine for a first spoken language. A second OMM unit may have a second optical interferometry unit including a passive diffractive optical element configured to implement a second neural network for a speech recognition engine for a second spoken language, etc. When a user desires to use system 500 to recognize speech in a first spoken language, the user can insert the first OMM unit into the system. When a user desires to use system 500 to recognize speech in a second spoken language, the user can replace the first OMM unit and insert a second OMM unit into the system.

[0416] For example, system 500 may be part of a control unit of an autonomous vehicle, and optical interferometry unit 600 may be configured to implement a neural network for performing road condition recognition. For example, first OMM unit 502 may have a first optical interferometry unit including a passive diffractive optical element configured to implement a first neural network for recognizing road conditions including road signs in the United States. Second OMM unit 502 may have a second optical interferometry unit including a passive diffractive optical element configured to implement a second neural network for recognizing road conditions including road signs in Canada. Third OMM unit 502 may have a third optical interferometry unit including a passive diffractive optical element configured to implement a third neural network for recognizing road conditions including road signs in Mexico, etc. When the autonomous vehicle is used in the United States, the first OMM unit is inserted into the system. When the autonomous vehicle crosses the border into Canada, the first OMM unit is replaced and a second OMM unit is inserted into the system. Meanwhile, when the autonomous vehicle crosses the border into Mexico, the first OMM unit is replaced and a third OMM unit is inserted into the system.

[0417] For example, system 500 can be used for gene sequencing. DNA sequences can be classified using a convolutional neural network implemented using system 500, which includes passive diffractive optical elements. For example, system 500 can implement neural networks for distinguishing tumor types, predicting tumor grade, and predicting patient survival from gene expression patterns. For example, system 500 can implement neural networks to identify a subset of genes or signatures that are most predictive of the trait being analyzed. For example, system 500 can implement neural networks to predict or infer the expression levels of all genes from the profile of a subset of genes. For example, system 500 can implement neural networks for epigenomic analysis, such as predicting transcription factor binding sites, enhancer regions, and chromatin accessibility from gene sequences. For example, system 500 can implement neural networks to capture structure within gene sequences.

[0418] For example, system 500 may be configured as a medical diagnostic system, and OMM unit 502 may be configured to implement a neural network for analyzing physiological parameters to perform disease screening. For example, system 500 may be configured as a bacteria detection system, and OMM unit 502 may be configured to implement a multiplication function for analyzing DNA sequences to detect certain strains of bacteria.

[0419] In some implementations, the OMM unit 502 includes a housing (e.g., a cartridge) that protects a substrate having diffractive optical elements. The housing supports an input interface coupled to an input waveguide 602 and an output interface coupled to an output waveguide 604. The input interface is configured to receive output from the array of modulators 144, and the output interface is configured to transmit output of the OMM unit 502 to the detection unit 146. The OMM unit 502 may be designed as a module suitable for handling by the average consumer, allowing users to easily switch from one OMM unit 502 to another. Machine learning technology evolves over time. Users can update the system 500 by replacing an old OMM unit 502 and inserting a new, updated version.

[0420] Just as an optical compact disc can store digital information that can be retrieved by a CD player, an OMM unit can store neural network configurations that can be used in an optical processor. Just as an optical compact disc is a low-cost medium for distributing digital information (including audio, video, and software programs) to consumers, an OMM unit can be a low-cost medium for distributing pre-configured neural networks or matrix processing functions (e.g., multiplication, convolution, or any other linear operation) to consumers.

[0421] In some implementations, system 500 is an optical computing platform configured to be operable with OMM units provided by different companies. This allows different companies to develop different passive optical neural networks for various applications. The passive optical neural networks are sold to end users as standardized packages that can be installed on the optical computing platform, enabling system 500 to perform various intelligent functions.

[0422] In some implementations, the system may have a holder mechanism for supporting multiple OMM units 502 and may be provided with a mechanical handling mechanism for automatically exchanging the OMM units 502. The system determines which OMM unit 502 is needed for the current application and uses the mechanical handling mechanism to automatically remove the appropriate OMM unit from the holder mechanism and insert it into the optical processor 504.

[0423] For a given size optical chip, many more passive diffractive elements can be fitted on the substrate compared to using active interferometers such as Mach-Zehnder interferometers. For example, optical interferometer unit 154 of FIG. 1B using a Mach-Zehnder interferometer can be configured to handle 200×200 matrix multiplications, while optical interferometer unit 600 having the same overall size but using passive diffractive elements (each having dimensions of approximately 100 nm×100 nm) can be configured to handle 5000×5000 matrix multiplications.

[0424] Because passive diffractive optical elements consume little power, the OMM unit 502 can be used in low-power devices, such as battery-operated devices. The OMM unit 502 is suitable for edge computing. For example, the OMM unit 502 can be used in smart sensors where raw data from the sensor is processed using an optical processor that uses the OMM unit 502. The smart sensor can be configured to send processed data to a central computer server, thereby reducing the amount of raw data sent to the central computer server. By providing intelligent processing capabilities in smart sensors, faults and anomalies can be detected earlier and addressed more effectively. The OMM unit 502 is suitable for applications that require processing large matrix multiplications. The OMM unit 502 is suitable for applications where a neural network is already trained and the weights are already determined and do not need to be modified.

[0425] The substrate on which the diffractive optical element is formed can be either flat or curved. In the example of FIG. 6A , input light enters the optical interferometer unit 600 from the left, and output light exits the optical interferometer unit 600 from the right (the terms “left,” “right,” “top,” and “bottom” refer to the directions shown in the figure). In some examples, the passive diffractive optical element can be configured to cause a portion of the output light to exit the optical interferometer unit 600 from the top or bottom, or from any combination of the left, right, top, and bottom sides of the optical interferometer unit 600. The substrate of the optical interferometer unit 600 can have any of a variety of shapes, such as a square, rectangle, triangle, circle, or ellipse. The optical interferometer unit 600 can incorporate reflective elements or mirrors to change the direction of light propagation.

[0426] In some implementations, the artificial neural network computing system 500 may be modified by adding an analog nonlinearity unit 310 between the detection unit 146 and the ADC unit 160. The analog nonlinearity unit 310 is configured to receive the output voltage from the detection unit 146, apply a nonlinear transfer function, and output a converted output voltage to the ADC unit 160. The controller 110 may obtain a converted digitized output voltage from the ADC unit 160 corresponding to the converted output voltage. Because the digitized output voltage obtained from the ADC unit 160 has already been nonlinearly transformed (“activated”), the nonlinear transformation step by the controller 110 can be omitted, reducing the computational burden on the controller 110. The first converted voltage obtained directly from the ADC unit 160 may then be stored in the memory unit 120 as a first converted digital output vector.

[0427] The optical interference unit may be implemented using passive diffractive optical elements arranged in three dimensions. Referring to Figure 7, in some implementations, an artificial neural network computing system 700 has an optical processor 702 that includes a three-dimensional OMM unit 708. The system 700 includes a memory unit 120 and an ADC unit 160 that are similar to the corresponding components of the system 500 of Figure 5. The optical processor 702 is configured to perform matrix computations using diffractive optical elements arranged in three dimensions.

[0428] The optical processor 702 includes a laser unit 704 configured to output a two-dimensional array of light beams 714 and an array of two-dimensional modulators 706 configured to modulate the two-dimensional array of light beams 714 to generate a modulated two-dimensional array of light beams 716. The optical processor 702 includes an optical matrix multiplication (OMM) unit 708 having diffractive optical elements arranged in three dimensions and configured to process the modulated two-dimensional array of light beams 716 to generate a two-dimensional array of output light beams 718. The optical processor 702 includes a detection unit 710 having a two-dimensional array of optical sensors for detecting the two-dimensional array of output light beams 718. The output of the detection unit 710 is converted to a digital signal by the ADC unit 160.

[0429] For example, the 3D OMM unit 708 may be implemented as a passively integrated silicon photonic column or cube. The optical matrix multiplication unit 708 may be configured to implement a diffractive neuron network and can perform matrix multiplication with nearly no power consumption.

[0430] There are many ways to encode input data for use by optical processor 702. For example, a digital input vector of length N×N may be encoded into an optical input matrix of size N×N, which is propagated through OMM unit 708. OMM unit 708 performs an (N×N)×(N×N) matrix multiplication on the received optical input matrix in the optical domain. The (N×N)×(N×N) matrix multiplication performed by OMM unit 708 is determined by the internal configuration of OMM unit 708, including, for example, the dimensions, arrangement, and geometry of diffractive optical elements arranged in three dimensions, and impurity doping, if any.

[0431] The OMM unit 708 may be implemented in various ways. FIG. 8 shows a schematic diagram of an example of an OMM unit 708 using a three-dimensional arrangement of diffractive elements. The OMM unit 708 may include a matrix of input waveguides for receiving an optical input matrix 802, a three-dimensional optical interference unit 804 in optical communication with the matrix of input waveguides, and a matrix of output waveguides in optical communication with the optical interference unit 804 for providing an optical output matrix 806. The optical interference unit 804 includes multiple diffractive optical elements and performs a transformation (e.g., a linear transformation) of the optical input (e.g., an N×N vector or matrix) to an optical output (e.g., an N×N vector or matrix). The matrix of output waveguides directs the optical signals output by the optical interference unit 804. At least one input waveguide in the matrix of input waveguides is in optical communication with each output waveguide in the matrix of output waveguides via the optical interference unit 804. For example, for an optical input vector of length NxN, the OMM unit 708 may include NxN input waveguides and NxN output waveguides.

[0432] In some implementations, the optical interference unit 804 includes a block of substrates having diffractive elements arranged three-dimensionally (e.g., in a 3D matrix). For example, multiple holes may be drilled or etched into each of multiple substrates, which may be combined to form the block of substrates. The holes have dimensions on the same order of magnitude as the wavelength of the input light so that light is diffracted by the holes (or structures defining the holes). The holes may have the same or different sizes. The holes may also have other cross-sectional shapes, such as triangular, square, rectangular, hexagonal, or irregular shapes. In some implementations, holographic methods may be used to form three-dimensional diffractive optical elements across the block of substrates. The substrates may be made of a material that is transparent or semi-transparent to the input light, e.g., having a transmittance for the input light ranging from 1% to 99%.

[0433] When designing the OMM unit 708, the dimensions and placement of the diffractive optical element in the x-, y-, and z-directions are considered. The configuration of the diffractive optical element can be determined using an optimization process. For example, a block of substrate can be divided into a three-dimensional matrix of pixels, and each pixel can be either filled with substrate material (no holes) or filled with air (holes). The pixel configuration can be iteratively modified, and simulations can be run for each pixel configuration by passing light through the diffractive optical element and evaluating the output. After simulations of all possible pixel configurations have been run, the configuration that produces results that most closely resemble the desired matrix processing is selected as the diffractive optical element configuration for the OMM unit 708.

[0434] As another example, a diffractive element may be initially configured as a three-dimensional matrix of holes. The placement, size, and shape of the holes may be slightly varied from their initial configuration. The parameters of each hole may be iteratively adjusted, and simulations may be performed to find the optimal configuration for the holes.

[0435] In some implementations, machine learning processes are used to design three-dimensional diffractive optical elements: an analytical function for how pixels affect input light is determined, and gradient descent is used to determine the optimal configuration of pixels.

[0436] In some implementations, the OMM unit 708 may be implemented as a user-changeable component, and different OMM units 708 having different optical interferometry units 804 may be installed for different applications. For example, the system 700 may be configured as a medical diagnostic system, and the optical interferometry unit 804 may be configured to implement a neural network for analyzing physiological parameters to perform disease screening. For example, a first OMM unit may have a first optical interferometry unit including a 3D passive diffractive optical element configured to implement a first neural network for screening a first set of diseases. A second OMM unit may have a second optical interferometry unit including a 3D passive diffractive optical element configured to implement a second neural network for screening a second set of diseases, etc. The first OMM unit and the second OMM unit may be developed by different companies specializing in developing techniques for screening different diseases. When a user desires to use the system 700 to screen for a first set of diseases, the user can insert the first OMM unit into the system. When the user wishes to use the system 700 to screen for a second set of diseases, the user can replace the first OMM unit and insert a second OMM unit into the system.

[0437] For example, system 700 may be configured as an optical character recognition system, and optical interferometry unit 804 may be configured to implement a neural network for performing optical character recognition. For example, system 700 may be configured as a speech recognition system, and optical interferometry unit 804 may be configured to implement a neural network for performing speech recognition. For example, system 700 may be part of a control unit of an autonomous vehicle, and optical interferometry unit 804 may be configured to implement a neural network for performing road condition recognition.

[0438] For example, system 700 can be used for gene sequencing. DNA sequences can be classified using a convolutional neural network implemented using system 700, which includes passive diffractive optical elements. For example, system 700 can implement neural networks for distinguishing tumor types, predicting tumor grade, and predicting patient survival from gene expression patterns. For example, system 700 can implement neural networks to identify a subset of genes or signatures that are most predictive of the trait being analyzed. For example, system 700 can implement neural networks to predict or infer the expression levels of all genes from a profile of a subset of genes. For example, system 700 can implement neural networks for epigenomic analysis, such as predicting transcription factor binding sites, enhancer regions, and chromatin accessibility from gene sequences. For example, system 700 can implement neural networks to capture structure within gene sequences. For example, system 700 may be configured as a bacteria detection system, and optical interference unit 804 may be configured to implement a multiplication function for analyzing DNA sequences to detect certain strains of bacteria.

[0439] In some implementations, the OMM unit 700 includes a housing (e.g., a cartridge) that protects a substrate having 3D diffractive optical elements. The housing supports an input interface coupled to an input waveguide and an output interface coupled to an output waveguide. The input interface is configured to receive output from the modulator array 706, and the output interface is configured to transmit output of the OMM unit 708 to the detection unit 710. The OMM unit 708 may be designed as a module suitable for handling by an average consumer, allowing users to easily switch from one OMM unit 708 to another. Machine learning technology evolves over time. Users can update the system 700 by replacing an old OMM unit 708 and inserting a new, updated version.

[0440] In some implementations, system 700 is an optical computing platform configured to be operable with OMM units provided by different companies. This allows different companies to develop different 3D passive optical neural networks for a variety of applications. The 3D passive optical neural networks are sold to end users as standardized packages that can be installed on the optical computing platform, enabling system 700 to perform various intelligent functions.

[0441] In some implementations, the system may have a holder mechanism for supporting multiple OMM units 708, and a mechanical handling mechanism may be provided for automatically exchanging the OMM units 708. The system determines which OMM unit 708 is needed for the current application and uses the mechanical handling mechanism to automatically remove the appropriate OMM unit 708 from the holder mechanism and insert it into the optical processor 702.

[0442] In some implementations, the artificial neural network computing system 700 may be modified by adding an analog nonlinear unit between the detection unit 710 and the ADC unit 160. The analog nonlinear unit is configured to receive the output voltage from the detection unit 710, apply a nonlinear transfer function, and output a converted output voltage to the ADC unit 160. The controller 110 may obtain a converted digitized output voltage from the ADC unit 160 corresponding to the converted output voltage. Because the digitized output voltage obtained from the ADC unit 160 has already been nonlinearly transformed (“activated”), the nonlinear transformation step by the controller 110 can be omitted, reducing the computational burden on the controller 110. The first converted voltage obtained directly from the ADC unit 160 may then be stored in the memory unit 120 as a first converted digital output vector.

[0443] The optical interference unit may be implemented using passive diffractive optical elements arranged in one dimension. Referring to Figure 9, in some implementations, an artificial neural network computing system 900 has an optical processor 906 that includes a one-dimensional optical multiplication unit 916. The system 900 includes a memory unit 120 similar to the corresponding component of the system 100 of Figure 1A. The optical processor 906 is configured to perform the multiplication calculations using diffractive optical elements arranged in one dimension, along the axis of light propagation.

[0444] The optical processor 906 includes a laser unit 908 configured to output a laser beam 910 and a modulator 912 configured to modulate the beam 910 to generate a modulated beam 914. The optical processor 906 includes a one-dimensional optical multiplication unit 916 having a diffractive optical element arranged in one dimension and configured to process the modulated beam 914 to generate an output beam 918. The optical processor 906 includes a detection unit 920 having an optical sensor for detecting the output beam 916. The output of the detection unit 920 is converted into a digital signal by an ADC unit 930.

[0445] For example, the optical multiplication unit 916 may be implemented as a passive integrated silicon photonic waveguide with a diffractive optical element (e.g., a grating or holes). The optical multiplication unit 916 may be configured to perform the multiplication operation with nearly no power consumption.

[0446] There are many ways to encode input data for use by optical processor 906. For example, a digital input vector may be encoded as an optical input that is propagated through optical multiplication unit 916. Optical multiplication unit 916 performs multiplication on the received optical input in the optical domain. The multiplication performed by optical multiplication unit 916 is determined by the internal configuration of optical multiplication unit 916, including, for example, the dimensions, placement, and geometry of diffractive optical elements arranged in one dimension along the optical propagation path, and impurity doping, if any.

[0447] The optical multiplication unit 916 may be implemented in various ways. FIG. 10 shows a schematic diagram of an example of an optical multiplication unit 916 using a one-dimensional arrangement of diffractive elements. The optical multiplication unit 916 may include an input waveguide for receiving an optical input 1002, a one-dimensional optical interference unit 1004 in optical communication with the input waveguide, and an output waveguide in optical communication with the optical interference unit 1004 for providing an optical output 1006. The optical interference unit 1004 includes a plurality of diffractive optical elements and performs a conversion (e.g., a linear conversion) of the optical input to the optical output. The output waveguide guides the optical signal output by the optical interference unit 1004.

[0448] In some implementations, the optical interference unit 1004 includes an elongated substrate having diffractive elements arranged in one dimension along the light propagation path. For example, a plurality of holes may be drilled or etched into the substrate. The holes have dimensions on the same order of magnitude as the wavelength of the input light such that the light is diffracted by the holes (or the structures defining the holes). The holes may have the same or different sizes. The substrate may be made of a material that is transparent or semi-transparent to the input light, for example, having a transmittance for the input light ranging from 1% to 99%. In some implementations, holographic methods may be used to form the diffractive optical elements in the substrate.

[0449] When designing the light multiplication unit 1004, the dimensions and placement of the diffractive element along the propagation path of the light ray are taken into consideration. The configuration of the diffractive optical element can be determined using an optimization process. For example, a substrate can be divided into a series of pixels, and each pixel can be either filled with the substrate material (no holes) or filled with air (holes). The pixel configuration can be iteratively modified, and simulations can be performed for each pixel configuration by passing light through the diffractive optical element and evaluating the output. After simulations of all possible pixel configurations have been performed, the configuration that produces results that most closely resemble the desired multiplication process is selected as the diffractive optical element configuration for the light multiplication unit 1004.

[0450] As another example, a diffractive element may initially be configured as a series of holes. The placement and dimensions of the holes may be slightly varied from their initial configuration. The parameters of each hole may be iteratively adjusted, and simulations may be performed to find the optimal configuration for the holes.

[0451] In some implementations, a machine learning process is used to design a one-dimensional diffractive optical element: an analytical function for how pixels affect input light is determined, and gradient descent is used to determine the optimal configuration of the pixels.

[0452] In some implementations, the light multiplication unit 916 may be implemented as a user-changeable component, and different light multiplication units 916 with different light interferometry units 1004 may be installed for different applications. For example, the system 900 may be configured as a bacteria detection system, and the light interferometry unit 1004 may be configured to implement a multiplication function for analyzing DNA sequences to detect certain strains of bacteria. For example, the first light multiplication unit may have a first light interferometry unit including a 1D passive diffractive optical element configured to implement a first multiplication function for detecting a first group of bacteria. The second light multiplication unit may have a second light interferometry unit including a 1D passive diffractive optical element configured to implement a second multiplication function for detecting a second group of bacteria, and so on. The first light multiplication unit and the second light multiplication unit may be developed by different companies specializing in developing techniques for detecting different bacteria. When a user wants to use the system 900 to detect a first group of bacteria, the user can insert the first light multiplication unit into the system. When a user desires to use system 900 to detect a second group of bacteria, the user can replace the first light multiplication unit and insert a second light multiplication unit into the system. By using one-dimensional diffractive optical elements, laser unit 908, modulator 912, detection unit 920, and ADC unit 930 can be made at low cost.

[0453] In some implementations, the light multiplication unit 916 includes a housing (e.g., a cartridge) that protects a substrate having a 1D diffractive optical element. The housing supports an input interface coupled to an input waveguide and an output interface coupled to an output waveguide. The input interface is configured to receive output from the modulator 912, and the output interface is configured to send output of the light multiplication unit 916 to the detection unit 920. The light multiplication unit 916 may be designed as a module suitable for handling by an average consumer, allowing users to easily switch from one light multiplication unit 916 to another. Machine learning technology evolves over time. Users can update the system 900 by replacing an old light multiplication unit 916 and inserting a new, updated version.

[0454] In some implementations, system 900 is an optical computing platform configured to be operable with optical multiplication units provided by different companies. This allows different companies to develop different 1D passive optical multiplication functions for various applications. The 1D passive optical multiplication functions are sold to end users as standardized packages that can be installed on the optical computing platform, enabling system 900 to perform various intelligent functions.

[0455] In some implementations, the system may have a holder mechanism for supporting multiple light multiplication units 916, and a mechanical handling mechanism may be provided for automatically replacing the light multiplication units 916. The system determines which light multiplication unit 916 is needed for the current application and uses the mechanical handling mechanism to automatically remove the appropriate light multiplication unit 916 from the holder mechanism and insert it into the light processor 906.

[0456] In some implementations, the artificial neural network computing system 900 may be modified by adding an analog nonlinear unit between the detection unit 920 and the ADC unit 930. The analog nonlinear unit is configured to receive the output voltage from the detection unit 920, apply a nonlinear transfer function, and output a converted output voltage to the ADC unit 930. The controller 110 may obtain a converted digitized output voltage from the ADC unit 930 corresponding to the converted output voltage. Because the digitized output voltage obtained from the ADC unit 930 has already been nonlinearly transformed (“activated”), the nonlinear transformation step by the controller 110 can be omitted, reducing the computational burden on the controller 110. The first converted voltage obtained directly from the ADC unit 930 may then be stored in the memory unit 120 as a first converted digital output vector.

[0457] Passive chips with passive diffractive optical elements offer several advantages. First, because active components, which are typically the most voluminous components, are eliminated, any given size chip can accommodate a larger neural network. A typical useful neural network can contain millions of weights, which is difficult to implement on an active chip and may require multiple runs of data through the chip and chip reprogramming. In comparison, a single passive chip may be capable of supporting an entire neural network. Second, the very low power consumption of passive chips is important for "edge" applications, which may require a small footprint and low power consumption. Third, because passive chips do not contain active components, they can be manufactured at even lower cost.

[0458] Optical matrix multiplication units with passive diffractive optical elements can also be used in wavelength division multiplexed artificial neural network computing systems. For example, OMM unit 150 of system 104 of FIG. 1F can be replaced with an OMM unit that uses passive diffractive optical elements. In this example, second DAC subunit 134 can be eliminated.

[0459] In some implementations, the optical processors (e.g., 504, 702) can perform matrix processing other than matrix multiplication. The optical matrix multiplication units 502 and 708 can be replaced by optical matrix processing units that perform other types of matrix processing.

[0460] 25 shows a flowchart of an example method 2500 for performing ANN calculations using ANN calculation system 500, 700, or 900 including one or more optical matrix multiplication units or optical multiplication units with passive diffraction elements, such as 2D OMM unit 502, 3D OMM unit 708, or 1D OM unit 916. The steps of process 2500 may be performed at least in part by controller 110 or 902. In some implementations, the various steps of method 2500 may be performed in parallel, in combination, in a loop, or in any order.

[0461] At 2510, an artificial neural network (ANN) computation request is received, the artificial neural network (ANN) computation request comprising an input dataset. The input dataset includes a first digital input vector. The first digital input vector is a subset of the input dataset. For example, it may be a subregion of an image. The ANN computation request may be generated by various entities, such as the computer 102. The computer may include one or more of various types of computing devices, such as a personal computer, a server computer, a vehicle computer, and a flight computer. The ANN computation request generally refers to an electronic signal that notifies or informs the ANN computation system 500, 700, or 900 of an ANN computation to be performed. In some implementations, the ANN computation request may be split into two or more signals. For example, a first signal may query the ANN computation system 500, 700, or 900 to see if the ANN computation system 500, 700, or 900 is ready to receive the input dataset. In response to a positive response by the system 500, 700, or 900, the computer may transmit a second signal including the input dataset.

[0462] At 2520, the input dataset is stored. The controller 110 may store the input dataset in the memory unit 120. Storing the input dataset in the memory unit 120 may allow flexibility in the operation of the ANN computation system 500, 700, or 900, which may, for example, increase the overall performance of the system. For example, the input dataset may be divided into digital input vectors of a set size and format by retrieving desired portions of the input dataset from the memory unit 120. Different portions of the input dataset may be processed in various orders or shuffled to allow various types of ANN computations to be performed. For example, shuffling may enable matrix multiplication using block matrix multiplication techniques when the input and output matrix sizes are different. As another example, storing the input dataset in the memory unit 120 may allow the ANN computation system 500, 700, or 900 to queue multiple ANN computation requests, which may allow the system 500, 700, or 900 to maintain operation at maximum speed without periods of inactivity.

[0463] At 2530, a first plurality of modulator control signals are generated based on the first digital input vector. The controller 110 may send DAC control signals to the DAC unit 506, 712, or 904 for generating the first plurality of modulator control signals. The DAC unit 506, 712, or 904 generates the first plurality of modulator control signals based on the first DAC control signals, and the array of modulators 144, 706, or 912 generates an optical input vector representing the first digital input vector.

[0464] The first DAC control signal may include a plurality of digital values ​​to be converted by the DAC unit 506, 712, or 904 into the first plurality of modulator control signals. The plurality of digital values ​​generally correspond to the first digital input vector and may be related through various mathematical relationships or lookup tables. For example, the plurality of digital values ​​may be linearly proportional to the values ​​of the elements of the first digital input vector. As another example, the plurality of digital values ​​may be related to the elements of the first digital input vector through a lookup table configured to maintain a linear relationship between the digital input vector and the optical input vector generated by the modulator column 144, 706, or 912.

[0465] In some implementations, the 2D OMM unit 502, the 3D OMM unit 708, or the 1D OM unit 916 is configured to perform optical matrix processing or optical multiplication based on an optical input vector and multiple neural network weights implemented using passive diffraction elements. The multiple neural network weights representing the matrix M may be decomposed into M = U S V * through a singular value decomposition (SVD) method, where U is an M × M unitary matrix, S is an M × N diagonal matrix with nonnegative real numbers on the diagonal, and V * is the complex conjugate of the N × N unitary matrix V. In such a case, the passive diffraction elements may be configured to implement the matrix V, the matrix S, and the matrix U such that the OMM unit 502 or 708 collectively implements the matrix M.

[0466] At 2540, a first plurality of digitized optical outputs corresponding to the optical output vectors of the optical matrix multiplication unit or optical multiplication are obtained. The optical input vectors generated by the modulator column 144, 706, or 912 are processed and converted into optical output vectors by the 2D OMM unit 502, the 3D OMM unit 708, or the 1D OM unit 916. The optical output vectors are detected by the detection unit 146, 710, or 920 and converted into electrical signals that can be converted into digitized values ​​by the ADC unit 160 or 930. The controller 110 or 902 may, for example, send a conversion request to the ADC unit 160 or 930 to initiate the conversion of the voltages output by the detection unit 146, 710, or 920 into digitized optical outputs. Once the conversion is complete, the ADC unit 160 or 930 may send the conversion results to the controller 110 or 902. Alternatively, controller 110 or 902 may retrieve the conversion results from ADC unit 160 or 930. Controller 110 or 902 may form a digital output vector from the digitized optical outputs that corresponds to the result of a matrix or vector multiplication of the input digital vectors. For example, the digitized optical outputs may be organized or concatenated to have a vector format.

[0467] In some implementations, the ADC unit 160 or 930 may be configured or controlled to perform the ADC conversion based on a DAC control signal issued by the controller 110 or 902 to the DAC unit 506, 712, or 904. For example, the ADC conversion may be configured to start some preset time after generation of the modulated control signal by the DAC unit 506, 712, or 904. Such control of the ADC conversion may simplify the operation of the controller 110 or 902 and reduce the number of required control operations.

[0468] At 2550, a nonlinear transformation is performed on the first digital output vector to generate a first transformed digital output vector. Nodes, or artificial neurons, of an ANN operate by first performing a weighted sum of signals received from nodes in the previous layer and then performing a nonlinear transformation ("activation") of the weighted sum to generate an output. Various types of ANNs may implement different types of distinguishable nonlinear transformations. Examples of nonlinear transformation functions include the rectified linear unit (RELU) function, the sigmoid function, the hyperbolic tangent function, the X^2 function, and the |X| function. Such a nonlinear transformation is performed on the first digital output by the controller 110 or 902 to generate the first transformed digital output vector. In some implementations, the nonlinear transformation may be performed by a specialized digital integrated circuit within the controller 110 or 902. For example, the controller 110 or 902 may include one or more modules or circuit blocks specifically adapted to accelerate the computation of one or more types of nonlinear transformations.

[0469] At 2560, the first transformed digital output vector is stored. The controller 110 or 902 may store the first transformed digital output vector in the memory unit 120. If the input data set is divided into multiple digital input vectors, the first transformed digital output vector corresponds to the result of an ANN computation of a portion of the input data set, such as the first digital input vector. Thus, storing the first transformed digital output vector enables the ANN computation system 500, 700, or 900 to perform and store additional computations on other digital input vectors of the input data set that will later be aggregated into a single ANN output.

[0470] At 2570, an artificial neural network output generated based on the first transformed digital output vector is output. The controller 110 or 902 generates an ANN output, the ANN output being a result of processing the input dataset through the ANN defined by the first plurality of neural network weights. If the input dataset is split into multiple digital input vectors, the generated ANN output is an aggregated output that includes the first transformed digital output and may further include additional transformed digital outputs corresponding to other portions of the input dataset. Once the ANN output is generated, it is transmitted to a computer, such as computer 102, that issued the ANN calculation request.

[0471] The 2D OMM unit 502, the 3D OMM unit 708, or the 1D OM unit 916 can represent the weight coefficients of one hidden layer of the neural network. If the neural network has several hidden layers, additional 2D OMM units 502, 3D OMM units 708, or 1D OM units 916 can be serially coupled. FIG. 26 shows an example of an ANN computational system 2600 for implementing a neural network with two hidden layers. The first 2D optical matrix multiplication unit 2604 represents the weight coefficients of the first hidden layer, and the second 2D optical matrix multiplication unit 2606 represents the weight coefficients of the second hidden layer. The ANN computational system 2600 includes the controller 110, the memory unit 120, the DAC unit 506, and the optoelectronic processor 2602. The memory unit 120 and the DAC unit 506 are similar to the corresponding components of the system 500 of FIG. 5. The optoelectronic processor 2602 is configured to perform matrix calculations using optical and electronic components.

[0472] The optoelectronic processor 2602 includes a first laser unit 142, a first modulator array 144a, a first 2D optical matrix multiplication unit 2604, a first detection unit 146a, a first analog nonlinear unit 310a, an analog memory unit 320, a second laser unit 142b, a second modulator array 144b, a second 2D optical matrix multiplication unit 2606, a second detection unit 146b, a second analog nonlinear unit 310b, and an ADC unit 160. The operation of the first laser unit 142, the first modulator array 144a, the first detection unit 146a, the first analog nonlinear unit 310a, and the analog memory unit 320 is similar to the corresponding components shown in FIG. 3B. The first 2D OMM unit 2604 is similar to the 2D OMM 502 of FIG. 5. The output of the analog memory unit 320 drives the second modulator array 144b, which modulates the laser light from the second laser unit 142b to generate a light vector. The light vector from the second modulator array 144b is processed by the second 2D OMM unit 2606, which performs a matrix multiplication to generate a light output vector that is detected by the second detection unit 246b. The second detection unit 246b is configured to generate an output voltage corresponding to the optical signal of the light output vector from the 2D OMM unit 2606. The ADC unit 160 is configured to convert the output voltage into a digitized output voltage. The controller 110 may obtain a digitized output from the ADC unit 160 corresponding to the light output vector of the second 2D OMM unit 2606. The controller 110 may form a digital output vector from the digitized output corresponding to the result of a second matrix multiplication of a nonlinear transformation of the result of the first matrix multiplication of the input digital vector. The second laser unit 142b may be combined with the first laser unit 142a by using a light splitter to redirect a portion of the light from the first laser unit 142a to the second modulator row 144b.

[0473] The principles described above may be applied to implementing neural networks with more than two hidden layers, where the weight coefficients of each hidden layer are represented by a corresponding 2D OMM unit.

[0474] 27 shows an example of an ANN computational system 2700 for implementing a neural network with two hidden layers. A first 3D optical matrix multiplication unit 2704 represents the weight coefficients of the first hidden layer, and a second 3D optical matrix multiplication unit 2706 represents the weight coefficients of the second hidden layer. The ANN computational system 2700 includes a controller 110, a memory unit 120, a DAC unit 712, and an opto-electronic processor 2702. The memory unit 120 and the DAC unit 712 are similar to the corresponding components of the system 700 of FIG. 7. The opto-electronic processor 2702 is configured to perform matrix computations using optical and electronic components.

[0475] The optoelectronic processor 2702 includes a first laser unit 704a, a first modulator array 706a, a first 3D optical matrix multiplication unit 2704, a first detection unit 710a, a first analog nonlinear unit 310a, an analog memory unit 320, a second laser unit 704b, a second modulator array 706b, a second 3D optical matrix multiplication unit 2706, a second detection unit 710b, a second analog nonlinear unit 310b, and an ADC unit 160. The operation of the first laser unit 704a, the first modulator array 706a, the first detection unit 710a, the first analog nonlinear unit 310a, and the analog memory unit 320 is similar to the corresponding components shown in FIG. 3B. The first 3D OMM unit 2704 is similar to the 3D OMM 708 of FIG. 7. The output of the analog memory unit 320 drives a second modulator array 706b, which modulates laser light from a second laser unit 704b to generate a light vector. The light vector from the second modulator array 706b is processed by a second 3D OMM unit 2706, which performs a matrix multiplication to generate a light output vector that is detected by a second detection unit 710b. The second detection unit 710b is configured to generate an output voltage corresponding to the optical signal of the light output vector from the 3D OMM unit 2706. The ADC unit 160 is configured to convert the output voltage into a digitized output voltage. The controller 110 may obtain a digitized output from the ADC unit 160 corresponding to the light output vector of the second 3D OMM unit 2706. The controller 110 may form a digital output vector from the digitized output corresponding to the result of a second matrix multiplication of a nonlinear transformation of the result of the first matrix multiplication of the input digital vector. The second laser unit 704b may be combined with the first laser unit 704a by using a light splitter to redirect a portion of the light from the first laser unit 704a to the second modulator row 706b.

[0476] The principles described above may be applied to implementing neural networks with more than two hidden layers, where the weight coefficients of each hidden layer are represented by a corresponding 3D OMM unit.

[0477] The 2D OMM unit 502 and the 3D OMM unit 708 with passive diffractive optical elements are suitable for use in a recurrent neural network (RNN) where the output of the network during the kth pass through the neural network is recycled to the input of the neural network and used as the input during the k+1th pass, so that the weight coefficients of the neural network remain the same over multiple passes.

[0478] Figure 28 shows an example of a neural network computing system 2800, which may be used to implement a recurrent neural network. System 2800 includes an optical processor 2802 that operates in a manner similar to optical processor 140 of Figure 3B, except that OMM unit 150 is replaced by a 2D OMM unit 2804, which may be similar to 2D OMM unit 502 of Figure 6. Because the neural network weights of 2D OMM unit 2804 are fixed, system 2800 does not require second DAC subunit 134 used in system 302 of Figure 3B.

[0479] 29 shows an example of a neural network computing system 2900, which can be used to implement a recurrent neural network. System 2900 includes an optical processor 2902 that operates in a manner similar to optical processor 140 of FIG. 3B, except that laser unit 142, array of modulators 144, OMM unit 150, and decision unit 146 are replaced by laser unit 704, array of modulators 706, 3D OMM unit 2904, and detection unit 710, respectively, of FIG. 7. Because the neural network weights for 3D OMM unit 2904 are fixed, system 2900 does not require the second DAC subunit 134 used in system 302 of FIG. 3B.

[0480] Figure 30 shows a schematic diagram of an example artificial neural network computation system 3000 with 1-bit internal resolution. ANN computation system 3000 is similar to ANN computation system 400 of Figure 4A, except that OMM unit 150 is replaced by a 2D OMM unit 3004 (similar to 2D OMM unit 502 of Figure 5) and second driver subunit 434 is omitted. ANN computation system 3000 operates in a similar manner to ANN computation system 400, in that an input vector is decomposed into several 1-bit vectors, and an ANN computation can then be performed by performing a series of matrix multiplications of the 1-bit vectors followed by addition of the individual matrix multiplication results.

[0481] Figure 31 shows a schematic diagram of an example artificial neural network computation system 3100 with 1-bit internal resolution. ANN computation system 3100 is similar to ANN computation system 400 of Figure 4A, except that OMM unit 150 is replaced by 3D OMM unit 3104 (similar to 3D OMM unit 708 of Figure 7), and second driver subunit 434 is omitted. In the example of Figure 31, laser unit 142, modulator array 144, and detection unit 146 of Figure 4A are replaced by laser unit 704, modulator array 706, and detection unit 710, respectively, of Figure 7. ANN computation system 3100 operates in a similar manner to ANN computation system 400, in that an input vector is decomposed into several 1-bit vectors, and an ANN computation can then be performed by performing a series of matrix multiplications of the 1-bit vectors followed by addition of the results of the individual matrix multiplications.

[0482] The following describes the principles of an optical diffraction neural network. An optical diffraction neural network can be implemented as a small number of layers of a diffractive optical medium or a transmissive optical medium. Based on the Huygens-Fresnel principle, each point in the diffractive medium can be considered as a secondary light source. For each light source, the far-field diffraction can be described in the following equation:

[0483]

number

[0484] where the indices l and i refer to the ith neuron in the lth layer of the neural network, λ is the wavelength of light, and r is

[0485]

number

[0486] The output from each secondary light source can be written as the input times the phase and intensity modulation from the light source.

[0487]

number

[0488] where t is the transfer modulation, which is a complex term containing both amplitude and phase modulation,

number

[0489] The following describes a compact design of a compact photonic matrix multiplier unit capable of implementing general unitary matrix multiplication. Referring to FIG. 11 , the photonic matrix multiplier unit 1100 includes a modulator 1102, multiple interconnected interferometers 1104, and an attenuator 1106. The interconnected interferometers 1104 include a layer (or group or set) of directional couplers 1108a, 1108b, 1108c, 1108d, and 1108e (collectively 1108) and a layer (or group or set) of phase shifters 1110a, 1110b, 1110c, and 1110d (collectively 1110). Each layer (or group or set) of directional couplers may include one or more directional couplers. Each layer of phase shifters may include one or more phase shifters. In this example, interconnected interferometers 1104 include five layers of directional couplers 1108 and four layers of phase shifters. In other examples, photonic matrix multiplier unit 1100 may have different layers of directional couplers and phase shifters. Photonic matrix multiplier unit 1100 has directional couplers 1108 arranged in such a way that the number of layers of directional couplers 1108 is reduced compared to conventional matrix multiplier units that use interconnected Mach-Zehnder interferometers.

[0490] Here, the term "layer" in the phrases "layer of directional couplers" and "layer of phase shifters" refers to a group or set of directional couplers or phase shifters based on their placement in the photonic matrix multiplier unit 110 relative to the input and output ports. In the example of Figure 11, the input optical signal is processed by a first layer of directional couplers 1108a, then by a second layer of phase shifters 1110a, then by a third layer of directional couplers 1108b, then by a fourth layer of phase shifters 1110b, etc.

[0491] For example, a conventional matrix multiplier unit using interconnected Mach-Zehnder interferometers may require 2N layers of directional couplers, while the photonic matrix multiplier unit 1100 requires only N+2 layers of directional couplers. N refers to the number of input signals or the number of digits in the input vector. The mesh architecture used in the photonic matrix multiplier unit 1100 can have the most compact geometry for interconnected photonic interferometers capable of performing general matrix calculations.

[0492] 12A shows a diagram comparing the interconnected interferometer 1104 of the photonic matrix multiplier unit 1100 with conventional designs for various numbers of input signals. When there are four input signals, the interconnected Mach-Zehnder interferometer 1200 according to the conventional design requires eight layers of directional couplers, while the interconnected interferometer 1202 according to the new compact design requires only six layers of directional couplers. When there are three input signals, the interconnected Mach-Zehnder interferometer 1204 according to the conventional design requires six layers of directional couplers, while the interconnected interferometer 1206 according to the new compact design requires only five layers of directional couplers. When there are eight input signals, the interconnected Mach-Zehnder interferometer 1208 according to the conventional design requires 16 layers of directional couplers, while the interconnected interferometer 1210 according to the new compact d...

Claims

1. a first unit configured to generate a first plurality of modulator control signals; A processor unit, a light source configured to provide a plurality of light outputs; a plurality of optical modulators coupled to the light source and the first unit, the plurality of optical modulators configured to generate an optical input vector by modulating the plurality of optical outputs provided by the light source based on the first plurality of modulator control signals, the optical input vector including a plurality of optical signals; a matrix multiplication unit coupled to the plurality of optical modulators and the first unit, the matrix multiplication unit configured to convert the optical input vector into an analog output vector based on a first plurality of weight control signals; a processor unit comprising: a second unit coupled to the matrix multiplication unit and configured to convert the analog output vector into a digitized output vector; A controller comprising an integrated circuit, the integrated circuit comprising: receiving an artificial neural network computation request including an input data set including a first digital input vector; receiving a first plurality of neural network weights; generating, via the first unit, the first plurality of modulator control signals based on the first digital input vector; generating the first plurality of weight control signals based on the first plurality of neural network weights; a controller configured to perform operations including: A system comprising: the matrix multiplication unit: a plurality of replication modules, each of the plurality of replication modules corresponding to a subset of one or more optical signals of the optical input vector, the plurality of replication modules being configured to split the subset of the one or more optical signals into two or more replicas of the optical signals; a plurality of multiplication modules, each of the plurality of multiplication modules corresponding to a subset of one or more optical signals and configured to multiply the one or more optical signals of the subset with one or more matrix element values ​​using optical amplitude modulation; one or more summing modules, each summing module configured to produce an electrical signal representing a sum of results of two or more multiplication modules of the plurality of multiplication modules; A system comprising:

2. The system of claim 1 , wherein the first unit comprises a digital-to-analog converter (DAC).

3. The system of claim 1 or 2, wherein the second unit comprises an analog-to-digital converter (ADC).

4. 4. The system of claim 1, comprising a memory unit configured to store the dataset and the plurality of neural network weights.

5. 5. The system of claim 4, wherein the integrated circuit of the controller is further configured to perform operations including storing the input data set and the first plurality of neural network weights in the memory unit.

6. The system of claim 1 , wherein the first unit is configured to generate the plurality of weight control signals.

7. the controller comprises an application specific integrated circuit (ASIC); 6. The system of claim 1, wherein receiving an artificial neural network computation request comprises receiving an artificial neural network computation request from a general purpose data processor.

8. the first unit, the processor unit, the second unit, and the controller are disposed in at least one of a multi-chip module or an integrated circuit; 6. The system of claim 1, wherein receiving an artificial neural network computation request comprises receiving an artificial neural network computation request from a second data processor, the second data processor being external to the multi-chip module or the integrated circuit, the second data processor being coupled to the multi-chip module or the integrated circuit through a communications channel, and the processor unit being capable of processing data at a data rate at least one order of magnitude higher than a data rate of the communications channel.

9. The method of claim 8, wherein the first unit, the second unit, and the controller are used in an optoelectronic processing loop that is repeated over multiple iterations, the optoelectronic processing loop comprising: (1) at least a first optical modulation operation based on at least one of the plurality of modulator control signals and at least a second optical modulation operation based on at least one of the weight control signals; and (2) an electrical summation operation performed using an electrical summation module within the matrix multiplication unit, the electrical summation module configured to generate currents corresponding to elements of the analog output vector representing the sum of respective elements of the optical input vector multiplied by respective neural network weights; The system of claim 1 , comprising:

10. the optoelectronic processing loop includes an electrical storage operation, the electrical storage operation being performed using a memory unit coupled to the controller; 10. The system of claim 9, wherein the actions performed by the controller further comprise storing the input data set and the first plurality of neural network weights in the memory unit.

11. 10. The system of claim 9, wherein the optoelectronic processing loop includes at least one signal path, in which exactly one first optical modulation operation based on at least one of the plurality of modulator control signals and exactly one second optical modulation operation based on at least one of the weight control signals are performed in a single loop iteration.

12. 12. The system of claim 11, wherein the first light modulation operation is performed by one of the plurality of light modulators coupled to the light source of the light output and the matrix multiplication unit, and the second light modulation operation is performed by an light modulator included in the matrix multiplication unit.

13. 10. The system of claim 9, wherein the optoelectronic processing loop includes at least one signal path, in which only one electrical storage is performed in a single loop iteration.

14. The system of claim 1 , wherein the light source comprises a laser unit configured to generate the plurality of light outputs.

15. the matrix multiplication unit: an array of input waveguides for receiving the optical input vector, the optical input vector comprising a first array of optical signals; an optical interference unit in optical communication with the array of input waveguides for performing a linear transformation of the optical input vector into a second array of optical signals; an array of output waveguides in optical communication with the optical interference unit for directing the second array of optical signals, at least one input waveguide in the array of input waveguides in optical communication with each output waveguide in the array of output waveguides via the optical interference unit; and The system of claim 1 , comprising:

16. The optical interference unit is a plurality of interconnected Mach-Zehnder interferometers (MZIs), each MZI in the plurality of interconnected MZIs: A first phase shifter configured to change the division ratio of the MZI; a second phase shifter configured to shift the phase of one output of the MZI; 16. The system of claim 15, wherein the first phase shifter and the second phase shifter are coupled to the plurality of weight control signals.

17. 2. The system of claim 1, wherein at least one multiplication module of the plurality of multiplication modules comprises an optical amplitude modulator including an input port and two output ports, and wherein a pair of associated optical signals is provided from the two output ports such that a difference between the amplitudes of the associated optical signals corresponds to a result of multiplying an input value by a signed matrix element value.

18. 17. The system of claim 1, wherein the matrix multiplication unit is configured to multiply the optical input vector with a matrix comprising the one or more matrix element values.

19. 20. The system of claim 18, wherein a set of multiple output values ​​is encoded onto each electrical signal produced by the one or more summation modules, and an output value in the set of multiple output values ​​represents an element of an output vector that results from the optical input vector being multiplied by the matrix.

20. the system comprises a memory unit configured to store the input data set and the neural network weights, the second unit comprises an analog-to-digital converter (ADC) unit, and the operation comprises: obtaining a first plurality of digitized outputs from the ADC unit corresponding to the analog output vector of the matrix multiplication unit, the first plurality of digitized outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector; storing the first converted digital output vector in the memory unit; and 20. The system of claim 1, further comprising:

21. the system has a first loop period defined as the time elapsed between storing the input data set and the first plurality of neural network weights in the memory unit and storing the first converted digital output vector in the memory unit; 21. The system of claim 20, wherein the first loop period is 1 ns or less.

22. The operation is 22. The system of claim 20 or 21, further comprising outputting an artificial neural network output generated based on the first transformed digital output vector.

23. the first unit comprises a digital-to-analog converter (DAC) unit, and the operation comprises:

23. The system of claim 20, further comprising: generating, via the DAC unit, a second plurality of modulator control signals based on the first converted digital output vector.

24. the first unit comprises a digital-to-analog converter (DAC) unit, the artificial neural network computation request further comprises a second plurality of neural network weights, and the operation comprises:

24. The system of claim 20, further comprising: generating, based on obtaining the first plurality of digitized outputs, through the DAC unit, a second plurality of weight control signals based on the second plurality of neural network weights.

25. 25. The system of claim 24, wherein the first plurality of neural network weights and the second plurality of neural network weights correspond to different layers of an artificial neural network.

26. the first unit comprises a digital-to-analog converter (DAC) unit, and the input data set further comprises a second digital input vector; The operation is generating, via the DAC unit, a second plurality of modulator control signals based on the second digital input vector; obtaining a second plurality of digitized outputs from the ADC unit corresponding to the output vector of the matrix multiplication unit, the second plurality of digitized outputs forming a second digital output vector; performing a nonlinear transformation on the second digital output vector to generate a second transformed digital output vector; storing the second converted digital output vector in the memory unit; and outputting an artificial neural network output generated based on the first transformed digital output vector and the second transformed digital output vector; further comprising 26. The system of claim 20, wherein the output vector of the matrix multiplication unit results from a second optical input vector that is generated based on the second plurality of modulator control signals that are converted by the matrix multiplication unit based on the first plurality of weight control signals.

27. the system comprises a memory unit configured to store the input data set and the neural network weights, the second unit comprises an analog-to-digital converter (ADC) unit, and the system: an analog nonlinearity unit disposed between the matrix multiplication unit and the ADC unit, the analog nonlinearity unit providing a plurality of outputs from the matrix multiplication unit; configured to receive an input voltage, apply a nonlinear transfer function to the ADC unit, and output a plurality of converted output voltages to the ADC unit; The operations performed by the integrated circuit of the controller include: obtaining from the ADC unit a first plurality of converted digitized output voltages corresponding to the plurality of converted output voltages, the first plurality of converted digitized output voltages forming a first converted digital output vector; storing the first converted digital output vector in the memory unit; and 20. The system of claim 1, further comprising:

28. 28. The system of claim 1, wherein the integrated circuit of the controller is configured to generate the first plurality of modulated control signals at a rate of 8 GHz or greater.

29. the first unit comprises a digital-to-analog converter (DAC) unit, the second unit comprises an analog-to-digital converter (ADC) unit, and the matrix multiplication unit comprises: an optical matrix multiplication unit coupled to the plurality of optical modulators and the DAC unit, the optical matrix multiplication unit configured to convert the optical input vector into an optical output vector based on the plurality of weight control signals; an optical detection unit coupled to the optical matrix multiplication unit and configured to generate a plurality of output voltages corresponding to the optical output vector; 17. The system of claim 1, comprising:

30. an analog memory unit disposed between the DAC unit and the plurality of optical modulators, the analog memory unit configured to store an analog voltage and output the stored analog voltage; an analog nonlinearity unit disposed between the photodetection unit and the ADC unit, the analog nonlinearity unit configured to receive the plurality of output voltages from the photodetection unit, apply a nonlinear transfer function to the plurality of output voltages, and output a plurality of converted output voltages; 30. The system of claim 29, further comprising:

31. 31. The system of claim 30, wherein the analog memory unit comprises a plurality of capacitors.

32. the analog memory unit is configured to receive and store the plurality of converted output voltages of the analog nonlinearity unit, and output the stored plurality of converted output voltages to the plurality of optical modulators; The operation is storing the converted output voltages of the analog nonlinearity unit in the analog memory unit based on generating the first plurality of modulator control signals and the first plurality of weight control signals; outputting the stored converted output voltages through the analog memory unit; obtaining a second plurality of converted digitized output voltages from the ADC unit, the second plurality of converted digitized output voltages forming a second converted digital output vector; storing the second converted digital output vector in a memory unit; and 32. The system of claim 30 or 31, further comprising:

33. the system comprising a memory unit configured to store the input data set and the neural network weights, the input data set of the artificial neural network computation request comprising a plurality of digital input vectors; the light source is configured to generate a plurality of wavelengths; the plurality of optical modulators a bank of optical modulators configured to generate a plurality of optical input vectors, each of the banks generating a respective optical input vector having a respective wavelength corresponding to one of the plurality of wavelengths; an optical multiplexer configured to combine the plurality of optical input vectors into a combined optical input vector comprising the plurality of wavelengths; Equipped with the photodetector unit is further configured to demultiplex the plurality of wavelengths and generate a plurality of demultiplexed output voltages; The operation is obtaining a plurality of digitized demultiplexed optical outputs from the ADC unit, the plurality of digitized demultiplexed optical outputs forming a plurality of first digital output vectors, each of the plurality of first digital output vectors corresponding to one of the plurality of wavelengths; performing a nonlinear transformation on each of the plurality of first digital output vectors to generate a plurality of transformed first digital output vectors; storing the plurality of transformed first digital output vectors in the memory unit; and Including, 30. The system of claim 29, wherein each of the plurality of digital input vectors corresponds to one of the plurality of optical input vectors.

34. the system comprises a memory unit configured to store the input data set and the neural network weights, the second unit comprises an analog-to-digital converter (ADC) unit, and the artificial neural network computation request comprises a plurality of digital input vectors; the light source is configured to generate a plurality of wavelengths; the plurality of optical modulators a bank of optical modulators configured to generate a plurality of optical input vectors, each of the banks generating a respective optical input vector having a respective wavelength corresponding to one of the plurality of wavelengths; an optical multiplexer configured to combine the plurality of optical input vectors into a combined optical input vector comprising the plurality of wavelengths; Equipped with The operation is obtaining from the ADC unit a first plurality of digitized optical outputs corresponding to an optical output vector including the plurality of wavelengths, the first plurality of digitized optical outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector; storing the first converted digital output vector in the memory unit; and The system of claim 1 , comprising:

35. The first unit comprises a digital-to-analog converter (DAC) unit, the second unit comprises an analog-to-digital converter (ADC) unit, and the DAC unit comprises: a 1-bit DAC subunit configured to generate a plurality of 1-bit modulator control signals; the ADC unit has a resolution of 1 bit; the first digital input vector has a resolution of N bits; The operation is decomposing the first digital input vector into N one-bit input vectors, each of the N one-bit input vectors corresponding to one of the N bits of the first digital input vector; generating, via said 1-bit DAC subunits, a sequence of N 1-bit modulator control signals corresponding to said N 1-bit input vectors; obtaining from the ADC unit a sequence of N digitized 1-bit optical outputs corresponding to the sequence of the N 1-bit modulator control signals; constructing an N-bit digital output vector from said sequence of N digitized 1-bit optical outputs; performing a nonlinear transformation on the constructed N-bit digital output vector to generate a transformed N-bit digital output vector; storing the converted N-bit digital output vector in a memory unit; 35. The system of any one of claims 1 to 34, comprising:

36. The system comprises a memory unit configured to store the input data set and the neural network weights, the memory unit comprising: a digital input vector memory configured to store the digital input vector, the digital input vector memory comprising at least one SRAM; a neural network weight memory configured to store the plurality of neural network weights, the neural network weight memory comprising at least one DRAM; and 36. The system of any one of claims 1 to 35, comprising:

37. the first unit comprises a digital-to-analog converter (DAC) unit, the DAC unit comprising: a first DAC subunit configured to generate the plurality of modulator control signals; a second DAC subunit configured to generate the plurality of weight control signals; 37. The system of claim 1, wherein the first DAC subunit and the second DAC subunit are different.

38. The light source is a laser source configured to generate light; an optical power splitter configured to split the light generated by the laser source into the plurality of optical outputs; 38. The system of claim 1, wherein each of the plurality of optical outputs has substantially equal power.

39. 39. The system of claim 1, wherein the plurality of optical modulators comprise one of an MZI modulator, a ring resonator modulator, or an electro-absorption modulator.

40. The optical detection unit a plurality of photodetectors; a plurality of amplifiers configured to convert photocurrents generated by the photodetectors into the plurality of output voltages; 30. The system of claim 29, comprising:

41. 41. The system of any one of claims 1 to 40, wherein the integrated circuit is an application specific integrated circuit.

42. 20. The system of claim 1, further comprising a plurality of optical waveguides coupled between the optical modulator and the matrix multiplication unit, wherein the optical input vector comprises a set of a plurality of input values ​​encoded onto respective optical signals carried by the plurality of optical waveguides, and wherein each of the optical signals carried by one of the plurality of optical waveguides comprises an optical wave having a common wavelength that is substantially the same for all of the optical signals.

43. the duplication modules include at least one duplication module comprising an optical splitter that sends a predetermined percentage of the power of the light wave at an input port to a first output port and sends a remaining percentage of the power of the light wave at the input port to a second output port; 43. The system of claim 42.

44. 44. The system of claim 43, wherein the optical splitter comprises a waveguide optical splitter that sends a predetermined percentage of the power of a light wave guided by an input optical waveguide to a first output optical waveguide and sends a remaining percentage of the power of the light wave guided by the input optical waveguide to a second output optical waveguide.

45. 45. The system of claim 44, wherein a guided mode of the input optical waveguide is adiabatically coupled to a propagating mode of each of the first output optical waveguide and the second output optical waveguide.

46. 46. ​​The system of claim 43, wherein the optical splitter comprises a beam splitter including at least one surface that transmits the predetermined percentage of the power of the light wave at the input port and reflects the remaining percentage of the power of the light wave at the input port.

47. at least one of the plurality of optical waveguides comprises an optical fiber coupled to an optical coupler that couples a guided mode of the optical fiber to a propagation mode in free space; 47. A system according to any one of claims 42 to 46.

48. the multiplication modules include at least one coherence-sensitive multiplication module configured to multiply the one or more optical signals of the subset with one or more matrix element values ​​using optical amplitude modulation based on interference between optical waves having a coherence length at least as long as a propagation distance through the coherence-sensitive multiplication module.

48. A system according to any one of claims 1 to 19 and 42 to 47.

49. 49. The system of claim 48, wherein the coherence-sensitive multiplication module comprises a Mach-Zehnder interferometer (MZI), the MZI splitting light waves guided by an input optical waveguide into a first optical waveguide arm of the MZI and a second optical waveguide arm of the MZI, the first optical waveguide arm including a phase shifter that provides a relative phase shift with respect to a phase delay of the second optical waveguide arm, and the MZI combining light waves from the first optical waveguide arm and the second optical waveguide arm into at least one output optical waveguide.

50. 50. The system of claim 49, wherein the MZI combines lightwaves from the first and second optical waveguide arms into first and second output optical waveguides, respectively; a first photodetector receives lightwaves from the first output optical waveguide and generates a first photocurrent; a second photodetector receives lightwaves from the second output optical waveguide and generates a second photocurrent; and a result of the coherence-sensitive multiplication module comprises the difference between the first photocurrent and the second photocurrent.

51. 51. The system of any one of claims 48 to 50, wherein the coherence-sensitive multiplication module comprises one or more ring resonators, including at least one ring resonator coupled to a first optical waveguide and at least one ring resonator coupled to a second optical waveguide.

52. 52. The system of claim 51, wherein a first photodetector receives lightwaves from the first optical waveguide to generate a first photocurrent, a second photodetector receives lightwaves from the second optical waveguide to generate a second photocurrent, and a result of the coherence-sensitive multiplication module comprises the difference between the first photocurrent and the second photocurrent.

53. the multiplication modules include at least one coherence-insensitive multiplication module, the coherence-insensitive multiplication module configured to multiply the one or more optical signals of the subset with one or more matrix element values ​​using optical amplitude modulation based on absorption of energy in a light wave.

53. A system according to any one of claims 1 to 19 and 42 to 52.

54. 54. The system of claim 53, wherein the coherence-insensitive multiplication module comprises an electro-absorption modulator.

55. 55. The system of any one of claims 1 to 19 and 42 to 54, wherein the one or more summing modules include at least one summing module having: (1) two or more input conductors each carrying an electrical signal in the form of an input current, the amplitude of which represents a respective result of a respective one of the multiplication modules; and (2) at least one output conductor carrying an electrical signal in the form of an output current proportional to the sum of the input currents, the electrical signal representing the sum of the respective results.

56. 56. The system of claim 55, wherein the two or more input conductors and the output conductor comprise wiring that meet at one or more intersections of the wiring, and wherein the output current is substantially equal to the sum of the input currents.

57. 57. The system of claim 55 or 56, wherein at least a first one of the input currents is provided in the form of at least one photocurrent generated by at least one photodetector receiving an optical signal generated by a first one of the multiplication modules.

58. 58. The system of claim 57, wherein the first input current is provided in the form of a difference between two photocurrents generated by different respective photodetectors that receive different respective optical signals both generated by the first multiplication module.

59. A system described in any one of claims 1 to 19 and 42 to 58, wherein one of the replicas of the subset of the one or more optical signals consists of a single optical signal on which one of the input values ​​is encoded.

60. The system of claim 59, wherein the multiplication module corresponding to the copy of the subset multiplies the encoded input value with a single matrix element value.

61. A system described in any one of claims 1 to 19 and 42 to 60, wherein one of the copies of the subset of the one or more optical signals includes more than one of the optical signals and fewer than all of the optical signals on which multiple input values ​​are encoded.

62. The system of claim 61, wherein the multiplication modules corresponding to the copies of the subset multiply the encoded input values ​​with different respective matrix element values.

63. The system described in Claim 62, wherein different multiplication modules corresponding to each different replica of the subset of the one or more optical signals are stored in different devices that are in optical communication for transmitting one of the replicas of the subset of the one or more optical signals between the different devices.

64. two or more of the plurality of optical waveguides, two or more of the plurality of replication modules, two or more of the plurality of multiplication modules, and at least one of the one or more summation modules are disposed on a common device substrate.

64. A system according to any one of claims 42 to 63.

65. 65. The system of claim 64, wherein the device performs vector-matrix multiplication, the input vector being provided as a set of optical signals and the output vector being provided as a set of electrical signals.

66. 66. The system of claim 1, wherein the matrix multiplication unit comprises a multiplication module or an addition module, and the system further comprises an accumulator that integrates an input electrical signal corresponding to the output of the multiplication module or the addition module, wherein the input electrical signal is encoded by the accumulator using time-domain coding employing on-off amplitude modulation within each of a plurality of time slots, and the accumulator produces an output electrical signal that is encoded using more than two amplitude levels across the plurality of time slots corresponding to different duty cycles of the time-domain coding.

67. 67. The system of any one of claims 1 to 19 and 42 to 66, wherein two or more of the multiplication modules each correspond to a different subset of one or more optical signals.

68. 68. The system of claim 1, further comprising: a multiplication module configured to, for each replica of a second subset of one or more optical signals that is different from an optical signal in the first subset of one or more optical signals, multiply the one or more optical signals of the second subset by one or more matrix element values ​​using optical amplitude modulation.

69. 2. The system of claim 1, comprising a plurality of optical waveguides coupled between the optical modulator and the matrix multiplication unit, wherein the optical input vector comprises a set of a plurality of input values ​​encoded onto respective optical signals carried by the plurality of optical waveguides, and wherein each of the optical signals carried by one of the plurality of optical waveguides comprises a light wave having a common wavelength that is substantially the same for all of the optical signals.

70. the duplication modules include at least one duplication module comprising an optical splitter that sends a predetermined percentage of the power of the light wave at an input port to a first output port and sends a remaining percentage of the power of the light wave at the input port to a second output port; 20. A system according to any one of claims 1 to 19.

71. 71. The system of claim 70, wherein the optical splitter comprises a waveguide optical splitter that sends a predetermined percentage of the power of a light wave guided by an input optical waveguide to a first output optical waveguide and sends a remaining percentage of the power of the light wave guided by the input optical waveguide to a second output optical waveguide.

72. 72. The system of claim 71, wherein a guided mode of the input optical waveguide is adiabatically coupled to a propagating mode of each of the first output optical waveguide and the second output optical waveguide.

73. 73. The system of any one of claims 70 to 72, wherein the optical splitter comprises a beam splitter including at least one surface that transmits the predetermined percentage of the power of the light wave at the input port and reflects the remaining percentage of the power of the light wave at the input port.

74. A system as claimed in any one of claims 1 to 73, further comprising receiving the digitised output vector from the second unit.

Citation Information

Patent Citations

  • Optical real time computing element

    JP1992160612A

  • neural network

    JP1993501465A

  • Similarity calculator and information processor using the same

    JP1997113945A

  • Optical splitting process

    JP2003500719A

  • Computation using a network of optical parametric oscillators

    JP2016528611A