Method for performing computations in system
By using the optical matrix multiplication unit and passive diffraction optical components in the optoelectronic computing system, the problem of limited application of optical signals in computing platforms is solved, realizing fast artificial neural network computing and efficient photoelectric conversion.
Patent Information
- Application Number
- CN202512031345.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-19
- Filing Date
- 2019-06-04
- Publication Date
- 2026-05-15
AI Technical Summary
In existing technologies, the application of optical signals in computing platforms is limited, making it difficult to perform important operations in computing, especially in the conversion and computation processes between electrical and optical signals.
An optoelectronic computing system is employed, comprising a storage unit, a digital-to-analog conversion unit, an optical processor, an optical matrix multiplication unit, an optoelectronic detection unit, and a controller. The system performs neural network calculations using optical signals and utilizes the optical matrix multiplication unit and passive diffractive optical components to convert and process the optical input vector.
It achieves fast artificial neural network computation with a loop time of less than 1 ns, improving computational efficiency and supporting photoelectric conversion and nonlinear processing of electronic information.
Smart Images

Figure CN122047348A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on June 4, 2019, with application number 201980030464.8 and invention title "Optoelectronic Computing System".
[0002] Cross-reference to related applications
[0003] This application claims priority to U.S. Provisional Application No. 62 / 680,944, filed June 5, 2018; U.S. Provisional Application No. 62 / 744,706, filed October 12, 2018; U.S. Provisional Application No. 62 / 792,144, filed January 14, 2019; and U.S. Provisional Application No. 62 / 820,562, filed March 19, 2019. The entire disclosure of these applications is incorporated herein by reference. Technical Field
[0004] This disclosure relates to an optoelectronic computing system. Background Technology
[0005] Neuromorphic computing is a method in electronics that approximates brain operation. A prominent approach in neuromorphic computing is the artificial neural network (ANN), which is a collection of artificial neurons interconnected in a specific way to process information in a manner similar to brain function. ANNs have already found applications in a variety of fields, including artificial intelligence, speech recognition, text recognition, natural language processing, and various forms of pattern recognition.
[0006] An ANN has an input layer, one or more hidden layers, and an output layer. Each layer has nodes, or artificial neurons, and these nodes are interconnected between layers. Each node in a hidden layer performs a weighted sum of signals received from nodes in previous layers and performs a non-linear transformation ("activation") of the weighted sum to produce the output. The weighted sum can be computed by performing matrix multiplication steps. Therefore, computing an ANN typically involves multiple matrix multiplication steps, which are usually performed using electronic integrated circuits.
[0007] Computations performed on electronic data encoded in analog or digital form on electronic signals (such as voltage or current) are typically implemented using electronic computing hardware, such as analog or digital electronic devices implemented in integrated circuits (e.g., processors, application-specific integrated circuits (ASICs), or systems on a chip (SoCs)), electronic circuit boards, or other electronic circuits. Optical signals have been used to transmit data over both long and short distances (e.g., within data centers). Operations performed on such optical signals are typically performed in environments where optical data transmission occurs, such as within devices used to switch or filter optical signals in a network. The use of optical signals in computing platforms has become more restricted. Various components and systems have been proposed for all-optical computing. Such systems may include conversions from electrical signals to optical signals at inputs and outputs, but both types of signals (electrical and optical) cannot be used for critical operations performed in the computation. Summary of the Invention
[0008] In a typical first aspect, a system includes: a storage unit configured to store a dataset and multiple neural network weights; a digital-to-analog converter (DAC) unit configured to generate multiple modulator control signals and multiple weight control signals; an optical processor including a laser unit configured to generate multiple optical outputs; multiple optical modulators coupled to the laser unit and the DAC unit, the multiple optical modulators configured to generate an optical input vector by modulating the multiple optical outputs generated by the laser unit based on the multiple modulator control signals; an optical matrix multiplication unit coupled to the multiple optical modulators and the DAC unit, the optical matrix multiplication unit configured to convert the optical input vector into an optical output vector based on the multiple weight control signals; and a photodetector. A measurement unit, coupled to an optical matrix multiplication unit, is configured to generate multiple output voltages corresponding to optical output vectors; an analog-to-digital converter (ADC) unit, coupled to a photodetector unit, is configured to convert the multiple output voltages into multiple digital optical outputs; a controller, including an integrated circuit, is configured to perform the following operations: receiving from a computer an artificial neural network computation request including an input dataset and first multiple neural network weights, wherein the input dataset includes a first digital input vector; storing the input dataset and the first multiple neural network weights in a storage unit; and generating first multiple modulator control signals based on the first digital input vector and first multiple weight control signals based on the first multiple neural network weights via a DAC unit.
[0009] Implementations of the system may include one or more of the following features. For example, the operation may further include: obtaining a first plurality of digital optical outputs corresponding to the optical matrix multiplication unit from the ADC unit, the first plurality of digital optical outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector to generate a first converted digital output vector; and storing the first converted digital output vector in a storage unit.
[0010] The system may have a first loop period, which is defined as the time elapsed between the step of storing the input dataset and the first plurality of neural network weights in the storage unit and the step of storing the first transformed digital output vector in the storage unit. The first loop period may be less than or equal to 1 ns.
[0011] In some embodiments, the operation may further include: outputting the output of an artificial neural network based on the first transformed digital output vector.
[0012] In some embodiments, the operation may further include: generating a second plurality of modulator control signals based on a first converted digital output vector via a DAC unit.
[0013] In some embodiments, the artificial neural network computation request may further include a second plurality of neural network weights, and the operation may further include: generating a second plurality of weight control signals based on the second plurality of neural network weights via a DAC unit, based on the acquisition of a first plurality of digital light outputs. The first plurality of neural network weights and the second plurality of neural network weights may correspond to different layers of the artificial neural network.
[0014] In some embodiments, the input dataset may further include a second digital input vector, and the operation may further include: generating a second plurality of modulator control signals based on the second digital input vector via a DAC unit; obtaining a second plurality of digital optical outputs corresponding to the optical output vector of the optical matrix multiplication unit from an ADC unit, the second plurality of digital optical outputs forming a second digital output vector; performing a nonlinear transformation on the second digital output vector to generate a second converted digital output vector; storing the second converted digital output vector in a storage unit; and outputting the artificial neural network output generated based on the first and second converted digital output vectors. The optical output vector of the optical matrix multiplication unit is generated by the second optical input vector generated based on the second plurality of modulator control signals, the second optical input vector being transformed by the optical matrix multiplication unit based on the aforementioned plurality of weight control signals.
[0015] In some embodiments, the system may further include: an analog nonlinear unit disposed between the photodetector unit and the ADC unit, the analog nonlinear unit being configured to receive a plurality of output voltages from the photodetector unit, apply a nonlinear transfer function, and output a plurality of converted output voltages to the ADC unit, and the operation further including: obtaining a first plurality of converted digital output voltages corresponding to the plurality of converted output voltages from the ADC unit, the first plurality of converted digital output voltages forming a first converted digital output vector; and storing the first converted digital output vector in a storage unit.
[0016] In some embodiments, the integrated circuit of the controller may be configured to generate a first plurality of modulator control signals at a rate greater than or equal to 8 GHz.
[0017] In some embodiments, the system may further include: an analog storage unit disposed between the DAC unit and a plurality of optical modulators, the analog storage unit being configured to store analog voltages and output the stored analog voltages; and an analog nonlinear unit disposed between the photodetector unit and the ADC unit, the analog nonlinear unit being configured to receive a plurality of output voltages from the photodetector unit, apply a nonlinear transfer function, and output a plurality of converted output voltages. The analog storage unit may include a plurality of capacitors.
[0018] In some embodiments, the analog storage unit may be configured to receive and store a plurality of converted output voltages of an analog nonlinear unit, and to output the stored plurality of converted output voltages to a plurality of optical modulators, and the operation may further include: storing the plurality of converted output voltages of the analog nonlinear unit in the analog storage unit based on generating a first plurality of modulator control signals and a first plurality of weight control signals; outputting the stored converted output voltages through the analog storage unit; obtaining a second plurality of converted digital output voltages from an ADC unit, the second plurality of converted digital output voltages forming a second converted digital output vector; and storing the second converted digital output vector in the storage unit.
[0019] In some embodiments, the input dataset for the artificial neural network computation request may include multiple digital input vectors. A laser unit may be configured to generate multiple wavelengths. Multiple optical modulators may include: an optical modulator bank configured to generate multiple optical input vectors, each optical modulator bank corresponding to one of the multiple wavelengths and generating a corresponding optical input vector having the corresponding wavelength; and an optical multiplexer configured to combine the multiple optical input vectors into a combined optical input vector including the multiple wavelengths. A photodetector unit may be further configured to decompose the multiple wavelengths and generate multiple decomposed output voltages. Operation may include: obtaining multiple digitally decomposed optical outputs from an ADC unit, the multiple digitally decomposed optical outputs forming multiple first digital output vectors, wherein each of the multiple first digital output vectors corresponds to one of the multiple wavelengths; performing a nonlinear transformation on each of the multiple first digital output vectors to generate multiple transformed first digital output vectors; and storing the multiple transformed first digital output vectors in a storage unit. Each of the multiple digital input vectors may correspond to one of the multiple optical input vectors.
[0020] In some embodiments, an artificial neural network computation request may include multiple digital input vectors. A laser unit may be configured to generate multiple wavelengths. Multiple optical modulators may include: a group of optical modulators configured to generate multiple optical input vectors, each group corresponding to one of the multiple wavelengths and generating a corresponding optical input vector having the corresponding wavelength; and an optical multiplexer configured to combine the multiple optical input vectors into a combined optical input vector including the multiple wavelengths. Operation may include: obtaining a first plurality of digital optical outputs corresponding to an optical output vector from an ADC unit, the optical output vectors including the multiple wavelengths, the first plurality of digital optical outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector to generate a first converted digital output vector; and storing the first converted digital output vector in a storage unit.
[0021] In some embodiments, the DAC unit may include: a 1-bit DAC subunit configured to generate a plurality of 1-bit modulator control signals. The resolution of the ADC unit may be 1 bit. The resolution of the first digital input vector may be N bits. The operation may include: decomposing the first digital input vector into N 1-bit input vectors, each of the N 1-bit input vectors corresponding to one of the N bits of the first digital input vector; generating a sequence of N 1-bit modulator control signals corresponding to the N 1-bit input vectors through the 1-bit DAC subunit; obtaining a sequence of N digital 1-bit optical outputs corresponding to the sequence of N 1-bit modulator control signals from the ADC unit; constructing an N-bit digital output vector from the sequence of N digital 1-bit optical outputs; performing a nonlinear transformation on the constructed N-bit digital output vector to generate a converted N-bit digital output vector; and storing the converted N-bit digital output vector in a storage unit.
[0022] In some embodiments, the storage unit may include: a digital input vector memory configured to store digital input vectors and including at least one SRAM; and a neural network weight memory configured to store a plurality of neural network weights and including at least one DRAM.
[0023] In some embodiments, the DAC unit may include: a first DAC subunit configured to generate a plurality of modulator control signals; and a second DAC subunit configured to generate a plurality of weight control signals, wherein the first DAC subunit and the second DAC subunit are different.
[0024] In some embodiments, the laser unit may include: a laser source configured to generate light; and an optical power splitter configured to split the light generated by the laser source into a plurality of optical outputs, each of the plurality of optical outputs having substantially the same power.
[0025] In some embodiments, the plurality of optical modulators may include one of an MZI modulator, a ring resonator modulator, or an electro-absorption modulator.
[0026] In some embodiments, the photodetector unit may include: a plurality of photodetectors; and a plurality of amplifiers configured to convert the photocurrent generated by the photodetectors into a plurality of output voltages.
[0027] In some embodiments, the integrated circuit may be an application-specific integrated circuit (ASIC).
[0028] In some embodiments, the optical matrix multiplication unit may include: an input waveguide array for receiving an optical input vector; an optical interference unit, optically communicating with the input waveguide array, for performing a linear transformation that converts the optical input vector into a second optical signal array; and an output waveguide array, optically communicating with the optical interference unit, for guiding the second optical signal array, wherein at least one input waveguide in the input waveguide array is optically communicating with each output waveguide in the output waveguide array via the optical interference unit.
[0029] In some embodiments, the optical interference unit may include: a plurality of interconnected Mach-Zehnder interferometers (MZIs), each of the plurality of interconnected MZIs including: a first phase shifter configured to change the separation ratio of the MZI; and a second phase shifter configured to shift the phase of an output of the MZI, wherein the first phase shifter and the second phase shifter are coupled to a plurality of weight control signals.
[0030] In another aspect, a system includes: a storage unit configured to store a dataset and multiple neural network weights; a driver unit configured to generate multiple modulator control signals and multiple weight control signals; and an optical processor including: a laser unit configured to generate multiple optical outputs; multiple optical modulators coupled to the laser unit and the driver unit, the multiple optical modulators being configured to generate an optical input vector by modulating the multiple optical outputs generated by the laser unit based on the multiple modulator control signals; an optical matrix multiplication unit coupled to the multiple optical modulators and the driver unit, the optical matrix multiplication unit being configured to convert the optical input vector into an optical output vector based on the multiple weight control signals; and a photodetector unit coupled to the optical matrix multiplication unit and configured to generate multiple output voltages corresponding to the optical output vectors; a comparator unit coupled to the photodetector unit and configured to convert the multiple output voltages into multiple 1-bit digital optical outputs; and A controller, including an integrated circuit, is configured to perform the following operations: receive from a computer an artificial neural network computation request including an input dataset and first plurality of neural network weights, wherein the input dataset includes a first digital input vector having N-bit resolution; store the input dataset and the first plurality of neural network weights in a storage unit; decompose the first digital input vector into N 1-bit input vectors, each of the N 1-bit input vectors corresponding to one of the N bits of the first digital input vector; generate a sequence of N 1-bit modulator control signals corresponding to the N 1-bit input vectors via a driver unit; obtain a sequence of N digital 1-bit optical outputs corresponding to the sequence of N 1-bit modulator control signals from a comparator unit; construct an N-bit digital output vector from the sequence of N digital 1-bit optical outputs; perform a nonlinear transformation on the constructed N-bit digital output vector to produce a transformed N-bit digital output vector; and store the transformed N-bit digital output vector in a storage unit.
[0031] In another aspect, a method for performing artificial neural network computation in a system having an optical matrix multiplication unit configured to convert an optical input vector into an optical output vector based on a plurality of weight control signals, the method comprising: receiving an artificial neural network computation request from a computer including an input dataset and a first plurality of neural network weights, wherein the input dataset includes a first digital input vector; storing the input dataset and the first plurality of neural network weights in a storage unit; generating a first plurality of modulator control signals based on the first digital input vector and generating a first plurality of weight control signals based on the first plurality of neural network weights via a digital-to-analog converter (DAC) unit; obtaining a first plurality of digital optical outputs corresponding to the optical output vector of the optical matrix multiplication unit from an analog-to-digital converter (ADC) unit, the first plurality of digital optical outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector by a controller to generate a first transformed digital output vector; storing the first transformed digital output vector in a storage unit; and outputting an artificial neural network output generated based on the first transformed digital output vector by the controller.
[0032] In another aspect, one method includes: providing input information in an electronic format; converting at least a portion of the electronic input information into an optical input vector; converting the optical input vector into an optical output vector based on optical matrix multiplication; converting the optical output vector into an electronic format; and applying a nonlinear conversion circuit to the electronically converted optical output vector to provide output information in an electronic format.
[0033] Embodiments of the method may include one or more of the following features. For example, the method may further include: repeating electronic-to-optical converting, optical transforming, optical-to-electronic converting, and nonlinear conversion of electrical applications for new electronic input information corresponding to output information provided in electronic format.
[0034] In some embodiments, the optical matrix multiplication used for the initial optical conversion and the optical matrix multiplication used for the repeated optical conversion can be the same and can correspond to the same layer of an artificial neural network.
[0035] In some embodiments, the optical matrix multiplication used for the initial optical conversion and the optical matrix multiplication used for the repeated optical conversion can be different and can correspond to different layers of an artificial neural network.
[0036] In some embodiments, the method may further include: repeating electro-optical conversion, optical conversion, photoelectric conversion, and nonlinear conversion of electrical applications for different portions of electronic input information, wherein the optical matrix multiplication for the initial optical conversion and the optical matrix multiplication for the repeated optical conversion are the same and correspond to the first layer of the artificial neural network.
[0037] In some embodiments, the method may further include: providing intermediate information in an electronic format based on electronic output information for multiple portions of electronic input information generated by a first layer of an artificial neural network; and repeating electro-optical conversion, optical conversion, photoelectric conversion, and nonlinear conversion of electrical applications for each different portion of the electronic intermediate information, wherein the optical matrix multiplication for the initial optical conversion and the optical matrix multiplication for the repeated optical conversion associated with the different portions of the electronic intermediate information are the same and correspond to the second layer of the artificial neural network.
[0038] In another aspect, a system includes: an optical processor comprising a passive diffractive optical element, wherein the passive diffractive optical element is configured to convert an optical input vector or matrix into an optical output vector or matrix, representing the result of matrix processing applied to the optical input vector or matrix and a predetermined vector defined by the arrangement of the diffractive optical element.
[0039] Implementations of the system may include one or more of the following features. For example, matrix processing may include matrix multiplication between a light input vector or matrix and a predetermined vector defined by the arrangement of the diffractive optical components.
[0040] In some embodiments, the optical processor may include an optical matrix processing unit comprising: an input waveguide array for receiving an optical input vector; an optical interference unit including passive diffractive optical components, wherein the optical interference unit is in optical communication with the input waveguide array and configured to perform a linear transformation to convert the optical input vector into a second optical signal array; and an output waveguide array in optical communication with the optical interference unit for guiding the second optical signal array, wherein at least one input waveguide of the input waveguide array is in optical communication with each output waveguide of the output waveguide array via the optical interference unit.
[0041] In some embodiments, the optical interference unit may include a substrate having at least one of holes or stripes, wherein the size of the holes is in the range of 100 nm to 10 μm and the width of the stripes is in the range of 100 nm to 10 μm.
[0042] In some embodiments, the optical interference unit may include a substrate having passive diffractive optical components configured in a two-dimensional arrangement, and the substrate may include at least one of a planar substrate or a curved substrate.
[0043] In some embodiments, the substrate may include a planar substrate that is parallel to the light propagation direction from the input waveguide array to the output waveguide array.
[0044] In some embodiments, the optical processor may include an optical matrix processing unit comprising: an input waveguide matrix for receiving an optical input matrix; an optical interference unit including passive diffractive optical components, wherein the optical interference unit is in optical communication with the input waveguide matrix and configured to perform a linear transformation to convert the optical input matrix into a second optical signal matrix; and an output waveguide matrix in optical communication with the optical interference unit for guiding the second optical signal matrix, wherein at least one input waveguide of the input waveguide matrix is in optical communication with each output waveguide of the output waveguide matrix via the optical interference unit.
[0045] In some embodiments, the optical interference unit may include a substrate having at least one of holes or stripes, wherein the size of the holes is in the range of 100 nm to 10 μm and the width of the stripes is in the range of 100 nm to 10 μm.
[0046] In some embodiments, the optical interference unit may include a substrate having passive diffractive optical components configured in a three-dimensional arrangement.
[0047] In some embodiments, the substrate may have a shape that is at least one of cubic, columnar, prismatic, or irregular volume.
[0048] In some embodiments, the optical processor may include an optical interference unit comprising a hologram having passive diffractive optical components. The optical processor is configured to receive modulated light representing a light input matrix and to continuously convert the light as it passes through the hologram until the light exits the hologram as a light output matrix.
[0049] In some embodiments, the optical interference unit may include a substrate having passive diffraction optical components, and the substrate includes at least one of silicon, silicon oxide, silicon nitride, quartz, lithium niobate, phase change material, or polymer.
[0050] In some embodiments, the optical interference unit may include a substrate having passive diffraction optical components, and the substrate may include at least one of a glass substrate or an acrylic substrate.
[0051] In some embodiments, passive diffractive optical components may be partially formed by dopants.
[0052] In some embodiments, matrix processing can represent the neural network's processing of input data, which is represented by a light input vector.
[0053] In some embodiments, the optical processor may include: a laser unit configured to generate a plurality of optical outputs; a plurality of optical modulators coupled to the laser unit and configured to generate an optical input vector by modulating the plurality of optical outputs generated by the laser unit based on a plurality of modulator control signals; an optical matrix processing unit coupled to the plurality of optical modulators, the optical matrix processing unit including a passive diffractive optical component configured to convert the optical input vector into an optical output vector based on a plurality of weights defined by the passive diffractive optical component; and a photodetector unit coupled to the optical matrix processing unit and configured to generate a plurality of output electrical signals corresponding to the optical output vector.
[0054] In some embodiments, the passive diffractive optical components may be configured in a three-dimensional configuration, the plurality of optical modulators may include a two-dimensional optical modulator array, and the photodetector unit may include a two-dimensional photodetector array.
[0055] In some embodiments, the optical matrix processing unit may include a housing module to support and protect the input waveguide array, the optical interferometer, and the output waveguide array. The optical processor includes a receiving module configured to receive the optical matrix processing unit. The receiving module includes a first interface that enables the optical matrix processing unit to receive optical input vectors from a plurality of optical modulators, and a second interface that enables the optical matrix processing unit to transmit optical output vectors to the photodetector unit.
[0056] In some embodiments, the plurality of output electrical signals may include at least one of a plurality of voltage signals or a plurality of current signals.
[0057] In some embodiments, the system may include: a storage unit; a digital-to-analog converter (DAC) unit configured to generate a plurality of modulator control signals; an analog-to-digital converter (ADC) unit coupled to a photodetector unit and configured to convert a plurality of output electrical signals into a plurality of digital outputs; and a controller, including an integrated circuit, configured to perform the following operations: receiving from a computer a request for computation of an artificial neural network including an input dataset, wherein the input dataset includes a first digital input vector; storing the input dataset in the storage unit; and generating a first plurality of modulator control signals based on the first digital input vector via the DAC unit.
[0058] In another aspect, one method includes: 3D printing an optical matrix processing unit comprising a passive diffractive optical component, wherein the passive diffractive optical component is configured to convert a light input vector or matrix into a light output vector or matrix, representing the result of matrix processing applied to the light input vector or matrix and a predetermined vector defined by the arrangement of the diffractive optical component.
[0059] In another aspect, one method includes: using one or more laser beams to generate a hologram including a passive diffractive optical component, wherein the passive diffractive optical component is configured to convert an optical input vector or matrix into an optical output vector or matrix, representing the result of matrix processing applied to the optical input vector or matrix and a predetermined vector defined by the arrangement of the diffractive optical component.
[0060] In another aspect, a system includes: an optical processor comprising passive diffractive optical components arranged in a one-dimensional manner, wherein the passive diffractive optical components are configured to convert an optical input into an optical output, representing the result of matrix processing applied to the optical input and a predetermined vector defined by the arrangement of the diffractive optical components.
[0061] Implementations of the system may include one or more of the following features. For example, matrix processing may include matrix multiplication between an optical input and a predetermined vector defined by the arrangement of diffractive optical components.
[0062] In some embodiments, the optical processor may include an optical matrix processing unit comprising: an input waveguide for receiving optical input; an optical interference unit including passive diffractive optical components, wherein the optical interference unit is in optical communication with the input waveguide and is configured to perform a linear transformation of the optical input; and an output waveguide in optical communication with the optical interference unit for guiding optical output.
[0063] In some embodiments, the optical interference unit may include a substrate having at least one of a hole or a grating, the hole or grating assembly having a size in the range of 100 nm to 10 μm.
[0064] In another aspect, a system includes: a storage unit; a digital-to-analog converter (DAC) unit configured to generate a plurality of modulator control signals; and an optical processor including: a laser unit configured to generate a plurality of optical outputs; a plurality of optical modulators coupled to the laser unit and the DAC unit, the plurality of optical modulators being configured to generate an optical input vector by modulating the plurality of optical outputs generated by the laser unit based on the plurality of modulator control signals; an optical matrix processing unit coupled to the plurality of optical modulators, the optical matrix processing unit including passive diffractive optical components configured to convert the optical input vector into an optical output vector based on a plurality of weights defined by the passive diffractive optical components; and a photodetector unit coupled to the optical matrix processing unit and configured to generate a plurality of output electrical signals corresponding to the optical output vector. The system further includes: an analog-to-digital converter (ADC) unit coupled to a photodetector unit and configured to convert a plurality of output electrical signals into a plurality of digital optical outputs; and a controller, including an integrated circuit, configured to perform the following operations: receiving an artificial neural network computation request from a computer, including an input dataset, wherein the input dataset includes a first digital input vector; storing the input dataset in a storage unit; and generating a first plurality of modulator control signals based on the first digital input vector via a DAC unit.
[0065] Implementations of the system may include one or more of the following features. For example, the matrix processing unit may include a passive diffractive optics component configured to convert a light input vector into a light output vector, the light output vector representing the product of a matrix multiplication between the light input vector and a predetermined vector defined by the passive diffractive optics component.
[0066] In some embodiments, the operation further includes: obtaining a first plurality of digital optical outputs corresponding to the optical output vector of the optical matrix processing unit from the ADC unit, the first plurality of digital optical outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector; and storing the first transformed digital output vector in a storage unit.
[0067] In some embodiments, the system may have a first cycle time, which is defined as the time elapsed between the step of storing the input dataset in the storage unit and the step of storing the first transformed digital output vector in the storage unit, and wherein the first cycle time may be less than or equal to 1 ns.
[0068] In some embodiments, the operation may further include: outputting the output of an artificial neural network based on the first transformed digital output vector.
[0069] In some embodiments, the operation may further include: generating a second plurality of modulator control signals based on a first converted digital output vector via a DAC unit.
[0070] In some embodiments, the input dataset may further include a second digital input vector, and the operation may further include: generating a second plurality of modulator control signals based on the second digital input vector via a DAC unit; obtaining a second plurality of digital optical outputs corresponding to the optical output vector of the optical matrix processing unit from an ADC unit, the second plurality of digital optical outputs forming a second digital output vector; performing a nonlinear transformation on the second digital output vector to generate a second converted digital output vector; storing the second converted digital output vector in a storage unit; and outputting an artificial neural network output based on the first converted digital output vector and the second converted digital output vector, wherein the optical output vector of the optical matrix processing unit is generated by the second optical input vector generated based on the second plurality of modulator control signals, the second optical input vector being transformed by the optical matrix processing unit based on a plurality of weights defined by a passive diffractive optical component.
[0071] In some embodiments, the system may further include: an analog nonlinear unit disposed between the photodetector unit and the ADC unit, the analog nonlinear unit being configured to receive a plurality of output electrical signals from the photodetector unit, apply a nonlinear transfer function, and output a plurality of converted output electrical signals to the ADC unit, wherein the operation may further include: obtaining a first plurality of converted digital output electrical signals corresponding to the plurality of converted output electrical signals from the ADC unit, the first plurality of converted digital output electrical signals forming a first converted digital output vector; and storing the first converted digital output vector in a storage unit.
[0072] In some embodiments, the integrated circuit of the controller may be configured to generate a first plurality of modulator control signals at a rate greater than or equal to 8 GHz.
[0073] In some embodiments, the system may further include: an analog storage unit disposed between the DAC unit and a plurality of optical modulators, the analog storage unit being configured to store analog voltages and output the stored analog voltages; and an analog nonlinear unit disposed between the photodetector unit and the ADC unit, the analog nonlinear unit being configured to receive a plurality of output electrical signals from the photodetector unit, apply a nonlinear transfer function, and output a plurality of converted output electrical signals.
[0074] In some embodiments, the analog storage unit may include multiple capacitors.
[0075] In some embodiments, the analog storage unit may be configured to receive and store a plurality of converted output electrical signals of an analog nonlinear unit, and to output the stored plurality of converted output electrical signals to a plurality of optical modulators, wherein the operation may further include: storing the plurality of converted output electrical signals of the analog nonlinear unit in the analog storage unit based on generating a first plurality of modulator control signals; outputting the stored converted output electrical signals through the analog storage unit; obtaining a second plurality of converted digital output electrical signals from an ADC unit, the second plurality of converted digital output electrical signals forming a second converted digital output vector; and storing the second converted digital output vector in the storage unit.
[0076] In some embodiments, the input dataset for the artificial neural network computation request may include multiple digital input vectors, wherein the laser unit may be configured to generate multiple wavelengths, and wherein the multiple optical modulators may include: an optical modulator group configured to generate multiple optical input vectors, each optical modulator group corresponding to one of the multiple wavelengths and generating a corresponding optical input vector having the corresponding wavelength; and an optical multiplexer configured to combine the multiple optical input vectors into a combined optical input vector including the multiple wavelengths. The photodetector unit may be further configured to decompose the multiple wavelengths and generate multiple decomposed output electrical signals, and operation may include: obtaining multiple digitally decomposed optical outputs from an ADC unit, the multiple digitally decomposed optical outputs forming multiple first digital output vectors, wherein each of the multiple first digital output vectors corresponds to one of the multiple wavelengths; performing a nonlinear transformation on each of the multiple first digital output vectors to generate multiple transformed first digital output vectors; and storing the multiple transformed first digital output vectors in a storage unit, wherein each of the multiple digital input vectors corresponds to one of the multiple optical input vectors.
[0077] In some embodiments, an artificial neural network computation request may include a plurality of digital input vectors, wherein a laser unit is configured to generate a plurality of wavelengths, and wherein a plurality of optical modulators may include: an optical modulator group configured to generate a plurality of optical input vectors, each optical modulator group corresponding to one of the plurality of wavelengths and generating a corresponding optical input vector having the corresponding wavelength; and an optical multiplexer configured to combine the plurality of optical input vectors into a combined optical input vector including the plurality of wavelengths. Operation may include: obtaining a first plurality of digital optical outputs corresponding to optical output vectors from an ADC unit, the optical output vectors including the plurality of wavelengths, the first plurality of digital optical outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector; and storing the first transformed digital output vector in a storage unit.
[0078] In some embodiments, the DAC unit may include: a 1-bit DAC unit configured to generate a plurality of 1-bit modulator control signals, wherein the resolution of the ADC unit may be 1 bit, and wherein the resolution of the first digital input vector may be N bits. Operation may include: decomposing the first digital input vector into N 1-bit input vectors, each of the N 1-bit input vectors corresponding to one of the N bits of the first digital input vector; generating a sequence of N 1-bit modulator control signals corresponding to the N 1-bit input vectors through the 1-bit DAC unit; obtaining a sequence of N digital 1-bit optical outputs corresponding to the sequence of N 1-bit modulator control signals from the ADC unit; constructing an N-bit digital output vector from the sequence of N digital 1-bit optical outputs; performing a nonlinear transformation on the constructed N-bit digital output vector to generate a converted N-bit digital output vector; and storing the converted N-bit digital output vector in a storage unit.
[0079] In some embodiments, the storage unit may include: a digital input vector storage configured to store digital input vectors, and includes at least one SRAM.
[0080] In some embodiments, the laser unit may include: a laser source configured to generate light; and an optical power splitter configured to split the light generated by the laser source into a plurality of optical outputs, each of the plurality of optical outputs having substantially the same power.
[0081] In some embodiments, the plurality of optical modulators include one of an MZI modulator, a ring resonant modulator, or an electroabsorption modulator.
[0082] In some embodiments, the photodetector unit may include: a plurality of photodetectors; and a plurality of amplifiers configured to convert the photocurrent generated by the photodetectors into a plurality of output electrical signals.
[0083] In some embodiments, the integrated circuit may include an application-specific integrated circuit (ASIC).
[0084] In some embodiments, the optical matrix processing unit may include: an input waveguide array for receiving an optical input vector; an optical interference unit, optically communicating with the input waveguide array, for performing a linear conversion of the optical input vector into a second optical signal array, wherein the optical interference unit includes passive diffractive optical components; and an output waveguide array, optically communicating with the optical interference unit, for guiding the second optical signal array, wherein at least one input waveguide in the input waveguide array is optically communicating with each output waveguide in the output waveguide array via the optical interference unit.
[0085] In another aspect, a system includes: a storage unit; a driver unit configured to generate a plurality of modulator control signals; an optical processor including: a laser unit configured to generate a plurality of optical outputs; a plurality of optical modulators coupled to the laser unit and the driver unit, the plurality of optical modulators being configured to generate an optical input vector by modulating the plurality of optical outputs generated by the laser unit based on the plurality of modulator control signals; an optical matrix processing unit coupled to the plurality of optical modulators and the driver unit, the optical matrix processing unit including a passive diffractive optical component configured to convert the optical input vector into an optical output vector based on a plurality of weighted control signals defined by the passive diffractive optical component; and a photodetector unit coupled to the optical matrix processing unit and configured to generate a plurality of output electrical signals corresponding to the optical output vector. The system also includes a comparator unit coupled to a photodetector unit and configured to convert multiple output electrical signals into multiple 1-bit digital optical outputs; and a controller, including an integrated circuit, configured to perform the following operations: receiving an artificial neural network computation request from a computer, including an input dataset, wherein the input dataset includes a first digital input vector with N-bit resolution; storing the input dataset in a storage unit; decomposing the first digital input vector into N 1-bit input vectors, each of the N 1-bit input vectors corresponding to one of the N bits of the first digital input vector; generating a sequence of N 1-bit modulator control signals corresponding to the N 1-bit input vectors via a driver unit; obtaining a sequence of N 1-bit digital optical outputs corresponding to the sequence of N 1-bit modulator control signals from the comparator unit; constructing an N-bit digital output vector from the sequence of N 1-bit digital optical outputs; performing a nonlinear transformation on the constructed N-bit digital output vector to produce a converted N-bit digital output vector; and storing the converted N-bit digital output vector in a storage unit.
[0086] Implementations of the system may include one or more of the following features. For example, the optical matrix processing unit may include an optical matrix multiplication unit, a passive diffractometer configured to convert an optical input vector into an optical output vector, the optical output vector representing the product of a matrix multiplication between the input vector represented by the optical input vector and a predetermined vector defined by the passive diffracting optics component.
[0087] In another aspect, a method for performing artificial neural network computation in a system having an optical matrix processing unit includes: receiving an artificial neural network computation request from a computer, the input dataset including an input dataset comprising a first digital input vector; storing the input dataset in a storage unit; generating a first plurality of modulator control signals based on the first digital input vector via a digital-to-analog converter (DAC) unit; converting the optical input vector into an optical output vector using the optical matrix processing unit with an arrangement including passive diffractive optical components, wherein the optical output vector represents the result of matrix processing applied to the optical input vector and a predetermined vector defined by the arrangement of the diffractive optical components; obtaining a first plurality of digital optical outputs corresponding to the optical output vector of the optical matrix processing unit from an analog-to-digital converter (ADC) unit, the first plurality of digital optical outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector by a controller to generate a first transformed digital output vector; storing the first transformed digital output vector in a storage unit; and outputting an artificial neural network output generated based on the first transformed digital output vector by the controller.
[0088] Embodiments of the method may include one or more of the following features. For example, converting a light input vector into a light output vector may include converting the light input vector into a light output vector representing the product of a matrix multiplication between a digital input vector and a predetermined vector defined by the arrangement of diffractive optical components.
[0089] In another aspect, one method includes: providing input information in an electronic format; converting at least a portion of the electronic input information into an optical input vector; converting the optical input vector into an optical output vector based on optical matrix processing by an optical processor including passive diffractive optical components; converting the optical output vector into an electronic format; and applying a nonlinear conversion circuit to the electronically converted optical output vector to provide output information in an electronic format.
[0090] Embodiments of the method may include one or more of the following features. For example, converting an optical input vector into an optical output vector may include converting an optical input vector into an optical output vector based on optical matrix multiplication between a digital input vector represented by the optical input vector and a predetermined vector defined by a passive diffractive optics component.
[0091] In some embodiments, the method may further include: repeating electro-optical conversion, optical conversion, photoelectric conversion, and nonlinear conversion of electrical applications for new electronic input information corresponding to output information provided in electronic format.
[0092] In some embodiments, the optical matrix processing for the initial optical conversion and the optical matrix processing for the repeated optical conversion may be the same and may correspond to the same layer of the artificial neural network.
[0093] In some embodiments, the method may further include: repeating electro-optical conversion, optical conversion, photoelectric conversion, and nonlinear conversion of electrical applications for different portions of electronic input information, wherein the optical matrix processing for the initial optical conversion and the optical matrix processing for the repeated optical conversion may be the same and correspond to a layer of an artificial neural network.
[0094] In another aspect, a system includes an optical matrix processing unit configured to process an input vector of length N, wherein the optical matrix processing unit includes N+2 layers of directional couplers and N layers of phase shifters, and N is a positive integer.
[0095] Implementations of the system may include one or more of the following features. For example, the optical matrix processing unit may include directional couplers of no more than N+2 layers.
[0096] In some embodiments, the optical matrix processing unit may include an optical matrix multiplication unit.
[0097] In some embodiments, the optical matrix processing unit may include a substrate and an interconnecting interferometer disposed on the substrate, wherein each interferometer includes an optical waveguide disposed on the substrate, and a directional coupler and a phase shifter are part of the interconnecting interferometer.
[0098] In some embodiments, the optical matrix processing unit may include an attenuator layer following the last directional coupler layer.
[0099] In some embodiments, a single attenuator layer may include N attenuators.
[0100] In some embodiments, the system may include one or more homodyne detectors for detecting the output from the attenuator.
[0101] In some embodiments, N=3, and the optical matrix processing unit may include: an input terminal configured to receive an input vector; a first-layer directional coupler coupled to the input terminal; a first-layer phase shifter coupled to the first-layer directional coupler; a second-layer directional coupler coupled to the first-layer phase shifter; a second-layer phase shifter coupled to the second-layer directional coupler; a third-layer directional coupler coupled to the second-layer phase shifter; a third-layer phase shifter coupled to the third-layer directional coupler; a fourth-layer directional coupler coupled to the third-layer phase shifter; and a fifth-layer directional coupler coupled to the fourth-layer directional coupler.
[0102] In some embodiments, N=4, and the optical matrix processing unit may include: an input terminal configured to receive an input vector; first, second, third, and fourth directional couplers, each followed by a phase shifter, wherein the first directional coupler is coupled to the input terminal; a second-to-last directional coupler coupled to the fourth phase shifter; and a final directional coupler coupled to the second-to-last directional coupler.
[0103] In some embodiments, N=8, and the optical matrix processing unit may include: an input terminal configured to receive an input vector; eight directional couplers, each directional coupler being followed by a phase shifter, wherein the first directional coupler is coupled to the input terminal; a penultimate directional coupler coupled to the eighth phase shifter; and a final directional coupler coupled to the penultimate directional coupler.
[0104] In some embodiments, the optical matrix multiplication unit may include: an input terminal configured to receive an input vector; N layers of directional couplers, each layer of directional couplers being followed by a phase shifter, wherein the first layer of directional couplers is coupled to the input terminal; the penultimate layer of directional couplers coupled to the Nth layer of directional couplers; and the final layer of directional couplers coupled to the penultimate layer of directional couplers.
[0105] In some embodiments, N is an even number.
[0106] In some embodiments, each i-th layer directional coupler includes N / 2 directional couplers, where i is an odd number, and each j-th layer directional coupler includes N / 2-1 directional couplers, where j is an even number.
[0107] In some embodiments, for each i-th layer directional coupler where i is odd, the k-th directional coupler can be coupled to the (2k-1)-th and 2k-th outputs of the previous layer, where k is an integer from 1 to N / 2.
[0108] In some embodiments, for each j-th layer directional coupler where j is an even number, the m-th directional coupler can be coupled to the (2m)-th and (2m+1)-th outputs of the previous layer, where m is an integer from 1 to N / 2-1.
[0109] In some embodiments, each i-th layer phase shifter may include N phase shifters, where i is an odd number, and each j-th layer phase shifter may include N-2 phase shifters, where j is an even number.
[0110] In some embodiments, N may be an odd number.
[0111] In some embodiments, each layer of directional couplers may include (N-1) / 2 directional couplers.
[0112] In some embodiments, each phase shifter layer may include N-1 phase shifters.
[0113] In another aspect, a system includes: a generator configured to generate a first dataset, wherein the generator includes a light matrix processing unit; and a discriminator configured to receive a second dataset comprising data from the first dataset and data from a third dataset, wherein the data in the first dataset has similar characteristics to the data in the third dataset, and to classify the data in the second dataset as either data from the first dataset or data from the third dataset.
[0114] Embodiments of the method may include one or more of the following features. For example, the optical matrix processing unit may include at least one of the following: (i) the optical matrix multiplication unit described above, (ii) the passive diffraction optical component described above, or (iii) the optical matrix processing unit described above.
[0115] In some embodiments, the third dataset may include real data, the generator is configured to produce synthesized data that resembles real data, and the discriminator is configured to classify the data as real data or synthesized data.
[0116] In some embodiments, the generator may be configured to generate a dataset for training at least one of an autonomous vehicle, a medical diagnostic system, a fraud detection system, a weather forecasting system, a financial forecasting system, a facial recognition system, a voice recognition system, or a product defect detection system.
[0117] In some embodiments, the generator may be configured to generate an image that resembles at least one of a real object or a real scene, and the discriminator may be configured to classify the received image as (i) an image of a real object or a real scene, or (ii) a synthetic image generated by the generator.
[0118] In some embodiments, the real-world object may include at least one of a person, animal, cell, tissue, or product, and the real-world scenario may include a scenario encountered by a vehicle.
[0119] In some embodiments, the discriminator may be configured to classify received images as either (i) images of real people, real animals, real cells, real tissues, real products, or real scenes encountered by vehicles, or (ii) synthetic images generated by the generator.
[0120] In some embodiments, the means of transport may include at least one of a motorcycle, a car, a truck, a train, a helicopter, an airplane, a submarine, a ship, or a drone.
[0121] In some embodiments, the generator may be configured to generate images of tissues or cells associated with at least one of human, animal, or plant diseases.
[0122] In some embodiments, the generator may be configured to generate images of tissues or cells associated with human diseases, including at least one of cancer, Parkinson's disease, sickle cell anemia, heart disease, cardiovascular disease, diabetes, chest disease, or skin disease.
[0123] In some embodiments, the generator may be configured to generate images of cancer-related tissues or cells, and the cancer may include at least one of skin cancer, breast cancer, lung cancer, liver cancer, prostate cancer, or brain cancer.
[0124] In some embodiments, the system may further include a random noise generator configured to generate random noise input to the generator, and the generator is configured to generate a first dataset based on the random noise.
[0125] In another aspect, a system includes: a random noise generator configured to generate random noise; and a generator configured to generate data based on the random noise, wherein the generator includes a light matrix processing unit.
[0126] Embodiments of the system may include one or more of the following features. For example, the optical matrix processing unit may include (i) the optical matrix multiplication unit described above, (ii) the passive diffraction optical component described above, or (iii) at least one of the optical matrix processing units described above.
[0127] In another aspect, a system includes: an optical circuit configured to execute a logic function on two input signals, the optical circuit including: a first directional coupler having two inputs and two outputs, the two inputs being configured to receive the two input signals; a first pair of phase shifters configured to modify the phase of the signals at the two outputs of the first directional coupler; a second directional coupler having two inputs and two outputs, the two inputs being configured to receive signals from the first pair of phase shifters; and a second pair of phase shifters configured to modify the phase of the signals at the two outputs of the second directional coupler.
[0128] Embodiments of the method may include one or more of the following features. For example, a phase shifter may be configured to cause the optical circuitry to perform rotation:
[0129] .
[0130] In some embodiments, when input signals x1 and x2 are provided to the two input terminals of the first directional coupler, the phase shifter can be configured to cause the optical circuitry to perform operation:
[0131] .
[0132] In some embodiments, the optical circuit may include a first photodetector configured to generate the absolute value of a signal from a second pair of phase shifters, causing the optical circuit to perform an operation:
[0133] .
[0134] In some embodiments, the optical circuit may include a comparator configured to compare the output signal of the first photodetector with a threshold to generate a binary value that causes the optical circuit to produce an output.
[0135] .
[0136] In some embodiments, the optical circuit may include a feedback mechanism configured to feed back the output signal of the photodetector to the input of a first directional coupler, and through the first directional coupler, a first pair of phase shifters, a second directional coupler, and a second pair of phase shifters, and detected by the photodetector to cause the optical circuit to perform operation.
[0137] ,
[0138] It produces outputs AND(x1,x2) and OR(x1,x2).
[0139] In some embodiments, the optical circuit may include: a third directional coupler having two inputs and two outputs, the two inputs being configured to receive signals from a second pair of phase shifters; a third pair of phase shifters configured to modify the phase of the signals at the two outputs of the third directional coupler; a fourth directional coupler having two inputs and two outputs, the two inputs being configured to receive signals from the third pair of phase shifters; a fourth pair of phase shifters configured to modify the phase of the signals at the two outputs of the fourth directional coupler; and a second photodetector configured to generate the absolute value of the signal from the fourth pair of phase shifters, causing the optical circuit to perform operation:
[0140] ,
[0141] It produces outputs AND(x1,x2) and OR(x1,x2).
[0142] In some embodiments, the system may include a bitonic sorter configured to use optical circuitry to execute the sorting function of the bitonic sorter.
[0143] In some embodiments, the system may include means configured to use optical circuitry to perform a hash function.
[0144] In some embodiments, the hash function may include secure hash algorithm 2 (SHA-2).
[0145] Generally, systems used to perform computations employ different types of operations to produce computational results. Each operation is performed on a signal (e.g., an electrical or optical signal) that best suits the fundamental physical characteristics of the operation (e.g., in terms of energy consumption and / or speed). Three such operations are, for example, copying, summation, and multiplication. Copying can be performed using optical power splitting, summation can be performed using electrical current-based summation, and multiplication can be performed using optical amplitude modulation, as described in more detail below. An example of a computation that can be performed using these three types of operations is multiplying a vector by a matrix (e.g., as used in artificial neural network computations). These operations can be used to perform a variety of other computations. These operations represent a general set of linear operations that can perform various computations, including but not limited to: vector-vector dot product, vector-vector element-wise multiplication, vector-scalar element-wise multiplication, or matrix-matrix element-wise multiplication. Some examples described herein show techniques and configurations for vector-matrix multiplication, but the corresponding techniques and configurations can be used for any of these types of computations.
[0146] It may have one or more of the following advantages.
[0147] The optoelectronic computing systems using electrical and optical signals described herein can facilitate increased flexibility and / or efficiency. In the past, there may have been potential challenges associated with combining optical (or photonic) integrated devices with electrical (or electronic) integrated devices on a common platform (e.g., a common semiconductor die, or multiple semiconductor dies combined in a controlled collapsed chip connection or flip-chip arrangement). These potential challenges, for example, might include input / output (I / O) packaging or temperature control. For the systems described herein, these potential challenges may be amplified when used with a relatively large number of optical input ports and a relatively large number of electrical output ports (e.g., four or more optical input / output ports, 200 or more electrical input / output ports). These potential challenges can be mitigated using appropriate system design. For example, the system can use a high-density packaging arrangement that employs temperature control (e.g., thermoelectric cooling) to manage the thermal expansion between different material types (e.g., semiconductor materials (e.g., silicon), glass materials (silica), ceramic materials, etc.), and / or uses an enclosing housing as a heat sink and provides a degree of sealing. This temperature stabilization technique limits the different coefficients of thermal expansion (CTE) and the resulting misalignment between the system ports and the ports of the packaged high-density fiber array.
[0148] For replication operations, since optical power splitting is passive, no power is required to perform the operation. Furthermore, the frequency bandwidth of an electrical splitter is limited by its RC time constant. In contrast, the frequency bandwidth of an optical splitter is virtually unlimited. Different types of optical power splitters can be used, including waveguide optical splitters or free-space beamsplitters, as described in more detail below.
[0149] For multiplication operations, one value can be encoded as an optical signal, and the other value can be encoded as an amplitude scaling coefficient (e.g., multiplying by a value in the range of 0 to 1). After setting the scaling coefficient, the requirement for electrical signal conditioning in multiplication operations in the optical domain is reduced (or eliminated), thus reducing constraints due to electrical noise, power consumption, and bandwidth limitations. By appropriately selecting the detection scheme, signed results (e.g., multiplying by a value between -1 and +1) can be obtained, as described in more detail below.
[0150] For summation operations, different techniques can be used to determine the magnitude of the current in a conductor based on the sum of its different contributions. In the case of input current signals, when two or more conductors carrying those input current signals are combined at a junction, the single conductor carrying the output current signal represents the sum of those input current signals. In the case of input optical signals, when two or more light waves of different wavelengths illuminate a detector, the current signal carried on the photocurrent generated by the detector represents the sum of the power in the input optical signals. Both produce an electrical signal (e.g., current) as the output representing the sum, but one uses current as input (current-input-based summation, also known as "electrical summation" performed in the "electrical domain"), while the other uses light waves as input (optical-input-based summation, also known as "optoelectronic summation" performed in the "optoelectronic domain"). However, in some embodiments, summation based on current input is used instead of summation based on optical input, which allows a single optical wavelength to be used in the system, avoiding potentially complex components that might be required in the system and maintaining multiple wavelengths.
[0151] Combinations of these basic operations performed by these modules can be configured to provide means for performing linear operations, such as vector-matrix multiplication with arbitrary matrix element magnitudes. Other implementations of matrix multiplication using optical signals and interferometers for combining signals using optical interference have been limited to providing vector-matrix multiplication with certain constraints, such as unitary or diagonal matrices. Additionally, some other embodiments may rely on large-scale phase alignment of multiple optical signals as they propagate through a relatively large number of optical components (e.g., optical modulators). Alternatively, the embodiments described herein may relax this phase alignment constraint by converting the optical signals into electrical signals after propagation through fewer optical components (e.g., after propagation through no more than a single optical amplitude modulator), which allows the use of optical signals with reduced coherence, or even incoherent optical signals that do not rely on constructive / destructive interference.
[0152] The time-domain coding of optical and electrical signals, described in more detail below, allows analog electronic circuits to be optimized for operation at specific power levels, which can be helpful if the circuit operates at high speeds. This time-domain coding is useful in reducing any challenges that may be associated with precisely controlling the relatively large number of clearly distinguishable intensity levels for each symbol. Conversely, when applying precise control of the duty cycle in the time domain over multiple time slots within a single symbol duration, relatively constant amplitudes can be used (for the "on" level, an amplitude of zero or near zero is present at the "off" level).
[0153] By integrating photonic and electronic devices onto a common substrate (e.g., a silicon chip), modules can be easily manufactured on a large scale. The routing signal on the substrate is an optical signal rather than an electrical signal, and grouping photodetectors within a portion of the substrate helps avoid long electronic wiring and its associated challenges (e.g., parasitic capacitance, inductance, and crosstalk).
[0154] In embodiments of systems using submatrix multiplication, different devices (e.g., different cores, different processors, different computers, different servers) can be used to simultaneously compute each element of the output vector, helping to alleviate certain potential limitations (e.g., memory wall) and helping the entire system scale to very large matrices. In some embodiments, different devices can be used to multiply each submatrix by its corresponding subvector. The sum can then be calculated by collecting or accumulating the addends from the different devices. Intermediate results in the form of optical signals can be conveniently transmitted between devices, even if the devices are separated by relatively large distances.
[0155] Other aspects include other combinations of the above features and other features expressed as methods, devices, systems, program products, and other means.
[0156] Specific embodiments of the subject matter described in this specification may be implemented to achieve one or more of the following advantages: Improved throughput, latency, or both in ANN computations; Improved power efficiency in ANN computations.
[0157] In another aspect, an apparatus includes: a plurality of optical waveguides, wherein a plurality of input values are encoded on corresponding optical signals carried by the optical waveguides; a plurality of replication modules, wherein for each of at least two subsets of one or more optical signals, a corresponding set of the one or more replication modules is configured to divide the subset of one or more optical signals into two or more copies of the optical signals; a plurality of multiplication modules, wherein for each of at least two copies of a first subset of one or more optical signals, a corresponding multiplication module is configured to multiply the one or more optical signals of the first subset by one or more matrix element values using optical amplitude modulation, wherein at least one of the multiplication modules includes an optical amplitude modulator, the optical amplitude modulator including an input port and two output ports, and providing a pair of correlated optical signals from the two output ports such that the difference between the amplitudes of the correlated optical signals corresponds to the result of multiplying the input values by the signed matrix element values; and one or more summing modules, wherein for the results of two or more multiplication modules, a corresponding summing module is configured to generate an electrical signal representing the sum of the results of the two or more multiplication modules.
[0158] Embodiments of the device may include one or more of the following features. For example, an input value in a set of multiple input values encoded on a corresponding optical signal may represent an element of an input vector multiplied by a matrix comprising one or more matrix element values.
[0159] In some embodiments, a set of multiple output values may be encoded on corresponding electrical signals generated by one or more summing modules, and the output values in the set of multiple output values may represent elements of an output vector generated by multiplying an input vector by a matrix.
[0160] In some embodiments, each optical signal carried by an optical waveguide may include an optical wave having a common wavelength, which is substantially the same for all optical signals.
[0161] In some embodiments, the replication module may include at least one replication module having an optical splitter that sends a predetermined proportion of the power of the light wave to a first output port at an input port, and sends the remaining proportion of the power of the light wave to a second output port at an input port.
[0162] In some embodiments, the optical splitter may include a waveguide optical splitter that transmits a predetermined proportion of the power of the light wave guided by the input optical waveguide to a first output optical waveguide, and transmits the remaining proportion of the power of the light wave guided by the input optical waveguide to a second output optical waveguide.
[0163] In some embodiments, the guiding mode of the input optical waveguide can be adiabatically coupled to the guiding mode of each of the first and second output optical waveguides.
[0164] In some embodiments, the optical splitter may include a beam splitter comprising at least one surface that transmits a predetermined proportion of the power of a light wave at an input port and reflects the remaining proportion of the power of a light wave at an input port.
[0165] In some embodiments, at least one of the plurality of optical waveguides may include an optical fiber coupled to an optical coupler, the optical coupler coupling the fiber’s guiding mode to a free-space propagation mode.
[0166] In some embodiments, the multiplication module may include at least one coherence-sensitive multiplication module configured to multiply one or more optical signals of a first subset by one or more matrix element values based on the interference between light waves using optical amplitude modulation, the light waves having a coherence length at least as long as the propagation distance through the coherence-sensitive multiplication module.
[0167] In some embodiments, the coherent sensitive multiplication module may include a Mach-Zehnder interferometer (MZI) that separates the light waves guided by the input optical waveguide into a first optical waveguide arm and a second optical waveguide arm of the MZI. The first optical waveguide arm includes a phase shifter, the phase shifter generating a relative phase shift due to the phase delay of the phase shifter relative to the second optical waveguide arm, and the MZI combines the light waves from the first and second optical waveguide arms into at least one output optical waveguide.
[0168] In some embodiments, the MZI can combine light waves from the first optical waveguide arm and the second optical waveguide arm into each of the first output optical waveguide and the second output optical waveguide. The first photodetector can receive light waves from the first output optical waveguide to generate a first photocurrent, and the second photodetector can receive light waves from the second output optical waveguide to generate a second photocurrent. The result of the coherent sensitive multiplication module can include the difference between the first photocurrent and the second photocurrent.
[0169] In some embodiments, the coherent sensitive multiplication module may include one or more ring resonators, including at least one ring resonator coupled to a first optical waveguide and at least one ring resonator coupled to a second optical waveguide.
[0170] In some embodiments, a first photodetector may receive light waves from a first optical waveguide to generate a first photocurrent, a second photodetector may receive light waves from a second optical waveguide to generate a second photocurrent, and the result of the coherent sensitive multiplication module may include the difference between the first photocurrent and the second photocurrent.
[0171] In some embodiments, the multiplication module may include at least one coherence-insensitive multiplication module configured to multiply one or more optical signals of a first subset by one or more matrix element values based on energy absorption within the optical wave using optical amplitude modulation.
[0172] In some embodiments, the coherent insensitive multiplication module may include an electro-absorption modulator.
[0173] In some embodiments, one or more summing modules may include at least one summing module having the following components: (1) two or more input conductors, each input conductor carrying an electrical signal in the form of an input current, the magnitude of which represents the corresponding result of a corresponding multiplication module, and (2) at least one output conductor carrying an electrical signal representing the sum of the corresponding results in the form of an output current, the output current being proportional to the sum of the input currents.
[0174] In some embodiments, the two or more input conductors and output conductors may include multiple wires that meet at one or more nodes between the wires, and the output current is substantially equal to the sum of the input currents.
[0175] In some embodiments, at least a first input current may be provided in the form of at least one photocurrent generated by at least one photodetector, which receives an optical signal generated by a first multiplication module of the multiplication module.
[0176] In some embodiments, the first input current may be provided as the difference between two photocurrents generated by different corresponding photodetectors that receive different corresponding optical signals generated by the first multiplication module.
[0177] In some embodiments, one of the copies of a first subset of one or more optical signals may consist of a single optical signal, wherein one of the input values is encoded on the single optical signal.
[0178] In some embodiments, the multiplication module corresponding to a copy of the first subset can multiply the encoded input value by a single matrix element value.
[0179] In some embodiments, one of the copies of a first subset of one or more optical signals may include more than one optical signal, but less than the total number of all optical signals, on which multiple input values are encoded.
[0180] In some embodiments, the multiplication module corresponding to a copy of the first subset can multiply the encoded input value by different corresponding matrix element values.
[0181] In some embodiments, different multiplication modules corresponding to different copies of a first subset of one or more optical signals may be included in different devices that perform optical communication to transmit one copy of the first subset of one or more optical signals between the different devices.
[0182] In some embodiments, at least two or more of a plurality of optical waveguides, two or more of a plurality of replication modules, two or more of a plurality of multiplication modules, and at least one of one or more summation modules may be disposed on the substrate of the common device.
[0183] In some embodiments, the apparatus performs vector-matrix multiplication, wherein an input vector can be provided as a set of optical signals and an output vector can be provided as a set of electrical signals.
[0184] In some embodiments, the apparatus may further include an accumulator that combines the input electrical signals corresponding to the outputs of a multiplication module or a summing module, wherein time domain encoding may be used to encode the input electrical signals, the time domain encoding using on-off amplitude modulation in each of a plurality of time slots, and the accumulator may generate an output electrical signal encoded with more than two amplitude levels, the amplitude levels corresponding to different duty cycles of the time domain encoding on the plurality of time slots.
[0185] In some embodiments, each of two or more multiplication modules corresponds to a different subset of one or more optical signals.
[0186] In some embodiments, for each copy of a second subset of one or more optical signals that is different from the optical signals in a first subset of one or more optical signals, the apparatus may further include a multiplication module configured to multiply one or more optical signals of the second subset by one or more matrix element values using optical amplitude modulation.
[0187] In another aspect, a method includes: encoding a set of multiple input values on corresponding optical signals; for each of at least two subsets of one or more optical signals, using a corresponding set of one or more replication modules to divide the subset of one or more optical signals into two or more copies of the optical signals; for each of at least two copies of a first subset of one or more optical signals, using a corresponding multiplication module to multiply the one or more optical signals of the first subset by one or more matrix element values using optical amplitude modulation, wherein at least one multiplication module includes an optical amplitude modulator, the optical amplitude modulator includes an input port and two output ports, and provides a pair of correlated optical signals from the two output ports such that the difference between the amplitudes of the correlated optical signals corresponds to the result of multiplying the input values by the signed matrix element values; and for the results of two or more multiplication modules, using a summation module to generate an electrical signal representing the sum of the results of the two or more multiplication modules.
[0188] In another aspect, one method includes: encoding a set of input values representing elements of an input vector on a corresponding optical signal; encoding a set of coefficients representing matrix elements as amplitude modulation levels of a set of optical amplitude modulators coupled to the optical signal, including at least one optical amplitude modulator with one input port and two output ports providing a pair of correlated optical signals from the two output ports such that the difference between the amplitudes of the correlated optical signals corresponds to the result of multiplying the input values by the values of the signed matrix elements; and encoding a set of output values representing elements of an output vector on a corresponding electrical signal, wherein at least one electrical signal is in the form of a current whose amplitude corresponds to the sum of the corresponding elements of the input vector multiplied by the corresponding elements of a row of the matrix.
[0189] Embodiments of the method may include one or more of the following features. For example, at least one optical signal may be provided by a first optical waveguide, and the first optical waveguide may be coupled to an optical splitter that transmits a predetermined proportion of the power of the light wave guided by the first optical waveguide to a second output optical waveguide, and transmits the remaining predetermined proportion of the power of the light wave guided by the first optical waveguide to a third optical waveguide.
[0190] In another aspect, an apparatus includes: a plurality of optical waveguides encoded with a set of input values representing elements of an input vector on a corresponding optical signal carried by the optical waveguides; a set of optical amplitude modulators coupled to the optical signal, encoding a set of coefficients representing matrix elements as amplitude modulation levels, including at least one optical amplitude modulator with one input port and two output ports providing a pair of correlated optical signals from the two output ports such that the difference between the amplitudes of the correlated optical signals corresponds to the result of multiplying the input values by the values of signed matrix elements; and a plurality of summing modules encoded with a set of output values representing elements of an output vector on a corresponding electrical signal, wherein at least one electrical signal is in the form of a current and its amplitude corresponds to the sum of the corresponding elements of the input vector multiplied by the corresponding elements of a row of the matrix.
[0191] In another aspect, a method for multiplying an input vector by a given matrix includes: encoding a set of input values representing elements of an input vector on a corresponding optical signal of a set of optical signals; coupling a first set of one or more means to a first set of one or more waveguides providing a first subset of the set of optical signals, and producing a result of multiplying a first submatrix of the given matrix by the value encoded on the first subset of the set of optical signals; coupling a second set of one or more means to a second set of one or more waveguides providing a second subset of the set of optical signals, and producing a result of multiplying a second submatrix of the given matrix by the value encoded on the second subset of the set of optical signals; and coupling a third set of one or more means to provide a copy of the first subset of the set of optical signals generated by a first optical splitter. The third group of one or more waveguides, and the result of multiplying a third submatrix of a given matrix by a value encoded on a first subset of the optical signals; the fourth group of one or more devices is coupled to the fourth group of one or more waveguides that provide a copy of a second subset of the optical signals generated by the second optical splitter, and the result of multiplying a fourth submatrix of a given matrix by a value encoded on a second subset of the optical signals; wherein the first, second, third and fourth submatrixes connected together form a given matrix; and wherein at least one output value representing an element of an output vector is encoded on an electrical signal corresponding to the input vector multiplied by the given matrix, the electrical signal being generated by a device communicating with the first group of one or more devices and the second group of one or more devices.
[0192] Embodiments of the method may include one or more of the following features. For example, each pair of devices in the first group of one or more devices, the second group of one or more devices, the third group of one or more devices, and the fourth group of one or more devices may be mutually exclusive.
[0193] In another aspect, an apparatus includes: a first group of one or more devices configured to receive a first group of optical signals and generate a result of multiplying a first matrix by a value encoded on the first group of optical signals; a second group of one or more devices configured to receive a second group of optical signals and generate a result of multiplying a second matrix by a value encoded on the second group of optical signals; a third group of one or more devices configured to receive a third group of optical signals and generate a result of multiplying a third matrix by a value encoded on the third group of optical signals; a fourth group of one or more devices configured to receive a fourth group of optical signals and generate a result of multiplying a fourth matrix by a value encoded on the fourth group of optical signals; and a [missing information - likely a device name or function]. Configure a connection path between two or more of the following: a first group of one or more devices, a second group of one or more devices, a third group of one or more devices, or a fourth group of one or more devices, wherein a first configuration of the configurable connection path is configured to (1) provide a copy of the first group of optical signals as at least one of the second group of optical signals, the third group of optical signals, or the fourth group of optical signals, and (2) provide one or more signals from the first group of one or more devices and one or more signals from the second group of one or more devices to a summing module, the summing module being configured to generate an electrical signal representing the sum of values encoded on the signals received by the summing module.
[0194] In another aspect, an apparatus includes: a first group of one or more devices configured to receive a first group of optical signals and generate a result based on optical amplitude modulation of one or more of the first group of optical signals; a second group of one or more devices configured to receive a second group of optical signals and generate a result based on optical amplitude modulation of one or more of the second group of optical signals; a third group of one or more devices configured to receive a third group of optical signals and generate a result based on optical amplitude modulation of one or more of the third group of optical signals; a fourth group of one or more devices configured to receive a fourth group of optical signals and generate a result based on optical amplitude modulation of one or more of the fourth group of optical signals; and a configurable connection path between two or more of the first group of one or more devices, the second group of one or more devices, the third group of one or more devices, or the fourth group of one or more devices, wherein a first configuration of the configurable connection path is configured to (1) provide a copy of the first group of optical signals as the third group of optical signals, or (2) provide one or more signals from the first group of one or more devices and one or more signals from the second group of one or more devices to a summing module, the summing module being configured to generate an electrical signal representing the sum of values encoded on the signals received by the summing module.
[0195] Embodiments of the device may include one or more of the following features. For example, each pair of devices in a first group of one or more devices, a second group of one or more devices, a third group of one or more devices, and a fourth group of one or more devices may be mutually exclusive.
[0196] In some embodiments, a first configuration of the configurable connection path is configured to (1) provide a copy of the first set of optical signals as a third set of optical signals, and (2) provide one or more signals from one or more devices in the first set and one or more signals from one or more devices in the second set to a summing module, the summing module being configured to generate an electrical signal representing the sum of values encoded on at least two different signals received by the summing module.
[0197] In some embodiments, a first configuration of the configurable connection path may be configured to provide a copy of the first set of optical signals as a third set of optical signals, and a second configuration of the configurable connection path may be configured to provide one or more signals from one or more devices in the first set and one or more signals from one or more devices in the second set to a summing module, the summing module being configured to generate an electrical signal representing the sum of values encoded on the signals received by the summing module.
[0198] In another aspect, an apparatus includes: a plurality of optical waveguides, wherein a plurality of input values are encoded on respective optical signals carried by the optical waveguides; a plurality of replication modules, for each of at least two subsets of one or more optical signals, the plurality of replication modules including a corresponding set of one or more replication modules configured to divide a subset of one or more optical signals into two or more copies of the optical signals; a plurality of multiplication modules, for each of at least two copies of a first subset of one or more optical signals, the plurality of multiplication modules including a corresponding multiplication module configured to multiply one or more optical signals of the first subset by one or more values using optical amplitude modulation; and one or more summing modules, for the results of two or more multiplication modules, the one or more summing modules including a summing module configured to generate an electrical signal representing the sum of the results of the two or more multiplication modules, wherein the result includes at least one result encoded on the electrical signal, and the result is derived from a copy of the optical signal, which propagates through no more than a single optical amplitude modulator before being converted into an electrical signal.
[0199] In another aspect, a system includes: a first unit configured to generate a plurality of modulator control signals; and a processor including: a light source configured to provide a plurality of optical outputs; a plurality of optical modulators coupled to the light source and the first unit, the plurality of optical modulators configured to generate an optical input vector by modulating the plurality of optical outputs provided by the light source based on the plurality of modulator control signals, the optical input vector including a plurality of optical signals; and a matrix multiplication unit coupled to the plurality of optical modulators and the first unit, the matrix multiplication unit configured to convert the optical input vector into an analog output vector based on a plurality of weight control signals. The computing system further includes a second unit coupled to the matrix multiplication unit, and the second unit configured to convert the analog output vector into a digital output vector; and a controller including an integrated circuit configured to perform the following operations: receiving an artificial neural network computation request, the artificial neural network computation request including an input dataset including a first digital input vector; receiving a first plurality of neural network weights; and generating the first plurality of modulator control signals based on the first digital input vector and generating the first plurality of weight control signals based on the first plurality of neural network weights through the first unit.
[0200] Implementations of the system may include one or more of the following features. For example, the first unit may include a digital-to-analog converter (DAC).
[0201] In some embodiments, the second unit may include an analog-to-digital converter (ADC).
[0202] In some embodiments, the system may include a storage unit configured to store a dataset and multiple neural network weights.
[0203] In some embodiments, the integrated circuit of the controller may be further configured to perform operations including storing the input dataset and a first plurality of neural network weights in the storage unit.
[0204] In some embodiments, the first unit may be configured to generate a plurality of weight control signals.
[0205] In some embodiments, the controller may include an application-specific integrated circuit (ASIC), and receiving an artificial neural network computation request may include receiving an artificial neural network computation request from a general-purpose data processor.
[0206] In some embodiments, the first unit, processing unit, second unit, and controller may be disposed on at least one of the multi-chip module or integrated circuit. Receiving an artificial neural network computation request may include receiving the artificial neural network computation request from a second data processor, wherein the second data processor may be external to the multi-chip module or integrated circuit, the second data processor may be coupled to the multi-chip module or integrated circuit via a communication channel, and the processing unit may process data at a data rate at least an order of magnitude greater than the data rate of the communication channel.
[0207] In some embodiments, the first unit, the processing unit, the second unit, and the controller may be used in a photoelectric processing loop repeated in multiple iterations, and the photoelectric processing loop includes: (1) at least a first optical modulation operation based on at least one of a plurality of modulator control signals, and at least a second optical modulation operation based on at least one weight control signal, and (2) at least one of (a) an electrical summation operation or (b) an electrical storage operation.
[0208] In some embodiments, the photoelectric processing loop may include an electrical storage operation, and the electrical storage operation is performed using a storage unit coupled to a controller, wherein the operation performed by the controller may further include storing an input dataset and a first plurality of neural network weights in the storage unit.
[0209] In some embodiments, the photoelectric processing loop may include an electrical summation operation, which may be performed using an electrical summation module within a matrix multiplication unit, wherein the electrical summation module may be configured to generate currents corresponding to elements of an analog output vector, the elements of which represent the sum of the corresponding elements of the light input vector multiplied by the corresponding neural network weights.
[0210] In some embodiments, the photoelectric processing loop may include at least one signal path on which no more than one first optical modulation operation is performed in a single loop iteration based on at least one of a plurality of modulator control signals, and no more than one second optical modulation operation is performed in a single loop iteration based on at least one weight control signal.
[0211] In some embodiments, a first optical modulation operation may be performed by one of a plurality of optical modulators coupled to a light source and a matrix multiplication unit for optical output, and a second optical modulation operation may be performed by an optical modulator included in the matrix multiplication unit.
[0212] In some embodiments, the photoelectric processing loop may include at least one signal path on which no more than one electrical storage operation is performed in a single loop iteration.
[0213] In some embodiments, the light source may include a laser unit configured to generate multiple light outputs.
[0214] In some embodiments, the matrix multiplication unit may include: an input waveguide array for receiving an optical input vector, the optical input vector including a first optical signal array; an optical interference unit, optically in communication with the input waveguide array, for performing a linear transformation to convert the optical input vector into a second optical signal array; and an output waveguide array, optically in communication with the optical interference unit, for guiding the second optical signal array, wherein at least one input waveguide in the input waveguide array is optically in communication with each output waveguide in the output waveguide array via the optical interference unit.
[0215] In some embodiments, the optical interference unit may include: a plurality of interconnected Mach-Zehnder interferometers (MZIs), each of the plurality of interconnected MZIs including: a first phase shifter configured to change the separation ratio of the MZI; and a second phase shifter configured to shift the phase of an output of the MZI, wherein the first phase shifter and the second phase shifter are coupled to a plurality of weight control signals.
[0216] In some embodiments, the matrix multiplication unit may include: a plurality of copying modules, each copying module corresponding to a subset of one or more optical signals of an optical input vector, and configured to divide the subset of one or more optical signals into two or more copies of the optical signals; a plurality of multiplication modules, each multiplication module corresponding to a subset of one or more optical signals, and configured to multiply the subset of one or more optical signals by one or more matrix element values using optical amplitude modulation; and one or more summing modules, each summing module being configured to generate an electrical signal representing the sum of the results of two or more of the multiplication modules.
[0217] In some embodiments, at least one multiplication module includes an optical amplitude modulator, which includes an input port and two output ports, and can provide a pair of correlated optical signals from the two output ports such that the difference between the amplitudes of the correlated optical signals corresponds to the result of multiplying the input value by the value of a signed matrix element.
[0218] In some embodiments, the matrix multiplication unit may be configured to multiply an input vector by a matrix that includes one or more matrix element values.
[0219] In some embodiments, a set of multiple output values may be encoded on corresponding electrical signals generated by one or more summing modules, and the output values in the set of multiple output values may represent elements of an output vector generated by multiplying an input vector by a matrix.
[0220] In some embodiments, the computing system may include a storage unit configured to store an input dataset and neural network weights, the second unit may include an analog-to-digital converter (ADC) unit, and the operation may further include: obtaining a first plurality of digital outputs of analog output vectors corresponding to matrix multiplication units from the ADC unit, the first plurality of digital outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector to produce a first transformed digital output vector; and storing the first transformed digital output vector in the storage unit.
[0221] In some embodiments, the system has a first loop period, which is defined as the time elapsed between the step of storing the input dataset and the first plurality of neural network weights in the storage unit and the step of storing the first transformed digital output vector in the storage unit, and wherein the first loop period is less than or equal to 1 ns.
[0222] In some embodiments, the operation may further include: outputting the output of an artificial neural network based on the first transformed digital output vector.
[0223] In some embodiments, the first unit may include a digital-to-analog converter (DAC) unit, and the operation may further include: generating a second plurality of modulator control signals based on a first converted digital output vector via the DAC unit.
[0224] In some embodiments, the first unit may include a digital-to-analog conversion (DAC) unit, the artificial neural network computation request may further include a second plurality of neural network weights, and the operation may further include: generating a second plurality of weight control signals based on the second plurality of neural network weights through the DAC unit based on the acquisition of the first plurality of digital outputs.
[0225] In some embodiments, the first plurality of neural network weights and the second plurality of neural network weights may correspond to different layers of an artificial neural network.
[0226] In some embodiments, the first unit may include a digital-to-analog converter (DAC) unit, and the input dataset may further include a second digital input vector. The operation may further include: generating a second plurality of modulator control signals based on the second digital input vector via the DAC unit; obtaining a second plurality of digital outputs corresponding to the output vectors of the matrix multiplication unit from the ADC unit, the second plurality of digital outputs forming a second digital output vector; performing a nonlinear transformation on the second digital output vector to generate a second converted digital output vector; storing the second converted digital output vector in a storage unit; and outputting the artificial neural network output generated based on the first and second converted digital output vectors. The output vector of the matrix multiplication unit may be generated by a second optical input vector based on the second plurality of modulator control signals, the second optical input vector being transformed by the matrix multiplication unit based on the aforementioned plurality of weight control signals.
[0227] In some embodiments, the system may include a storage unit configured to store an input dataset and neural network weights, and a second unit may include an analog-to-digital converter (ADC) unit. The system may further include an analog nonlinear unit disposed between the matrix multiplication unit and the ADC unit, the analog nonlinear unit being configured to receive a plurality of output voltages from the matrix multiplication unit, apply a nonlinear transfer function, and output a plurality of converted output voltages to the ADC unit. Operations performed by the integrated circuit of the controller may further include: obtaining a first plurality of converted digital output voltages corresponding to the plurality of converted output voltages from the ADC unit, the first plurality of converted digital output voltages forming a first converted digital output vector; and storing the first converted digital output vector in the storage unit.
[0228] In some embodiments, the integrated circuit of the controller may be configured to generate a first plurality of modulator control signals at a rate greater than or equal to 8 GHz.
[0229] In some embodiments, the first unit may include a digital-to-analog converter (DAC) unit, and the second unit may include an analog-to-digital converter (ADC) unit. The matrix multiplication unit may include: an optical matrix multiplication unit coupled to a plurality of optical modulators and the DAC unit, the optical matrix multiplication unit being configured to convert an optical input vector into an optical output vector based on a plurality of weighted control signals; and a photodetector unit coupled to the optical matrix multiplication unit and configured to generate a plurality of output voltages corresponding to the optical output vector.
[0230] In some embodiments, the system may further include: an analog storage unit disposed between the DAC unit and a plurality of optical modulators, the analog storage unit being configured to store analog voltages and output the stored analog voltages; and an analog nonlinear unit disposed between the photodetector unit and the ADC unit, the analog nonlinear unit being configured to receive a plurality of output voltages from the photodetector unit, apply a nonlinear transfer function, and output a plurality of converted output voltages.
[0231] In some embodiments, the analog storage unit may include multiple capacitors.
[0232] In some embodiments, the analog storage unit may be configured to receive and store a plurality of converted output voltages of an analog nonlinear unit, and to output the stored plurality of converted output voltages to a plurality of optical modulators. The operation may further include: storing the plurality of converted output voltages of the analog nonlinear unit in the analog storage unit based on generating a first plurality of modulator control signals and a first plurality of weight control signals; outputting the stored converted output voltages through the analog storage unit; obtaining a second plurality of converted digital output voltages from an ADC unit, the second plurality of converted digital output voltages forming a second converted digital output vector; and storing the second converted digital output vector in the storage unit.
[0233] In some embodiments, the system may include a storage unit configured to store an input dataset and neural network weights, wherein the input dataset requested by the artificial neural network computation may include multiple digital input vectors. A light source may be configured to generate multiple wavelengths. Multiple optical modulators may include: a group of optical modulators configured to generate multiple optical input vectors, each group corresponding to one of the multiple wavelengths and generating a corresponding optical input vector having the corresponding wavelength; and an optical multiplexer configured to combine the multiple optical input vectors into a combined optical input vector including the multiple wavelengths. A photodetector unit may be further configured to decompose the multiple wavelengths and generate multiple decomposed output voltages. Operation may include: obtaining multiple digitally decomposed optical outputs from an ADC unit, the multiple digitally decomposed optical outputs forming multiple first digital output vectors, wherein each of the multiple first digital output vectors corresponds to one of the multiple wavelengths; performing a nonlinear transformation on each of the multiple first digital output vectors to generate multiple transformed first digital output vectors; and storing the multiple transformed first digital output vectors in a storage unit. Each of the multiple digital input vectors corresponds to one of the multiple optical input vectors.
[0234] In some embodiments, the system may include a storage unit configured to store an input dataset and neural network weights, a second unit may include an analog-to-digital converter (ADC) unit, and an artificial neural network computation request may include multiple digital input vectors. A light source may be configured to generate multiple wavelengths. Multiple optical modulators may include: a group of optical modulators configured to generate multiple optical input vectors, each group corresponding to one of the multiple wavelengths, and generating a corresponding optical input vector having the corresponding wavelength; and an optical multiplexer configured to combine the multiple optical input vectors into a combined optical input vector including the multiple wavelengths. Operation may include: obtaining a first plurality of digital optical outputs corresponding to optical output vectors from the ADC unit, the optical output vectors including the multiple wavelengths, the first plurality of digital optical outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector to generate a first converted digital output vector; and storing the first converted digital output vector in the storage unit.
[0235] In some embodiments, the first unit may include a digital-to-analog converter (DAC) unit, the second unit may include an analog-to-digital converter (ADC) unit, and the DAC unit may include a 1-bit DAC subunit configured to generate a plurality of 1-bit modulator control signals. The resolution of the ADC unit may be 1 bit, and the resolution of the first digital input vector may be N bits. Operation may include: decomposing the first digital input vector into N 1-bit input vectors, each of the N 1-bit input vectors corresponding to one of the N bits of the first digital input vector; generating a sequence of N 1-bit modulator control signals corresponding to the N 1-bit input vectors through the 1-bit DAC subunit; obtaining a sequence of N digital 1-bit optical outputs corresponding to the sequence of N 1-bit modulator control signals from the ADC unit; constructing an N-bit digital output vector from the sequence of N digital 1-bit optical outputs; performing a nonlinear transformation on the constructed N-bit digital output vector to generate a converted N-bit digital output vector; and storing the converted N-bit digital output vector in a storage unit.
[0236] In some embodiments, the system may include a storage unit configured to store an input dataset and neural network weights. The storage unit may include: a digital input vector storage configured to store digital input vectors and including at least one SRAM; and a neural network weight storage configured to store multiple neural network weights and including at least one DRAM.
[0237] In some embodiments, the first unit may include a digital-to-analog converter (DAC) unit, which includes: a first DAC subunit configured to generate a plurality of modulator control signals; and a second DAC subunit configured to generate a plurality of weight control signals, wherein the first DAC subunit and the second DAC subunit are different.
[0238] In some embodiments, the light source may include: a laser source configured to generate light; and an optical power splitter configured to split the light generated by the laser source into a plurality of optical outputs, each of the plurality of optical outputs having substantially the same power.
[0239] In some embodiments, the plurality of optical modulators include one of an MZI modulator, a ring resonant modulator, or an electroabsorption modulator.
[0240] In some embodiments, the photodetector unit may include: a plurality of photodetectors; and a plurality of amplifiers configured to convert the photocurrent generated by the photodetectors into a plurality of output voltages.
[0241] In some embodiments, the integrated circuit may be an application-specific integrated circuit (ASIC).
[0242] In some embodiments, the system may include a plurality of optical waveguides coupled between an optical modulator and a matrix multiplication unit, wherein the optical input vector may include a set of multiple input values encoded on corresponding optical signals carried by the optical waveguides, and each optical signal carried by one of the optical waveguides may include a light wave having a common wavelength that is substantially the same for all optical signals.
[0243] In some embodiments, the replication module may include at least one replication module having an optical splitter that sends a predetermined proportion of the power of the light wave to a first output port at an input port, and sends the remaining proportion of the power of the light wave to a second output port at an input port.
[0244] In some embodiments, the optical splitter may include a waveguide optical splitter that transmits a predetermined proportion of the power of the light wave guided by the input optical waveguide to a first output optical waveguide, and transmits the remaining proportion of the power of the light wave guided by the input optical waveguide to a second output optical waveguide.
[0245] In some embodiments, the guiding mode of the input optical waveguide can be adiabatically coupled to the guiding mode of each of the first and second output optical waveguides.
[0246] In some embodiments, the optical splitter may include a beam splitter comprising at least one surface that transmits a predetermined proportion of the power of a light wave at an input port and reflects the remaining proportion of the power of a light wave at an input port.
[0247] In some embodiments, at least one of the plurality of optical waveguides may include an optical fiber coupled to an optical coupler, the optical coupler coupling the fiber’s guiding mode to a free-space propagation mode.
[0248] In some embodiments, the multiplication module may include at least one coherent-sensitive multiplication module configured to multiply one or more optical signals of a first subset by one or more matrix element values based on the interference between light waves using optical amplitude modulation, the light waves having a coherent length at least as long as the propagation distance through the coherent-sensitive multiplication module.
[0249] In some embodiments, the coherent sensitive multiplication module may include a Mach-Zehnder interferometer (MZI) that separates light waves guided by an input waveguide into a first waveguide arm and a second waveguide arm of the MZI. The first waveguide arm includes a phase shifter that generates a relative phase shift due to a phase delay of the phase shifter relative to the second waveguide arm. The MZI may also combine light waves from the first and second waveguide arms into at least one output waveguide.
[0250] In some embodiments, the MZI can combine light waves from the first optical waveguide arm and the second optical waveguide arm into each of the first output optical waveguide and the second output optical waveguide. The first photodetector can receive light waves from the first output optical waveguide to generate a first photocurrent, and the second photodetector can receive light waves from the second output optical waveguide to generate a second photocurrent. The result of the coherent sensitive multiplication module can include the difference between the first photocurrent and the second photocurrent.
[0251] In some embodiments, the coherent sensitive multiplication module may include one or more ring resonators, the ring resonators including at least one ring resonator coupled to a first optical waveguide and at least one ring resonator coupled to a second optical waveguide.
[0252] In some embodiments, a first photodetector may receive light waves from a first optical waveguide to generate a first photocurrent, a second photodetector may receive light waves from a second optical waveguide to generate a second photocurrent, and the result of the coherent sensitive multiplication module may include the difference between the first photocurrent and the second photocurrent.
[0253] In some embodiments, the multiplication module may include at least one coherent insensitive multiplication module configured to multiply one or more optical signals of a first subset by one or more matrix element values based on energy absorption within the optical wave using optical amplitude modulation.
[0254] In some embodiments, the coherent insensitive multiplication module may include an electroabsorption modulator.
[0255] In some embodiments, one or more summing modules may include at least one summing module having the following components: (1) two or more input conductors, each input conductor carrying an electrical signal in the form of an input current, the magnitude of which represents the corresponding result of a corresponding multiplication module, and (2) at least one output conductor carrying an electrical signal representing the sum of the corresponding results in the form of an output current, the output current being proportional to the sum of the input currents.
[0256] In some embodiments, two or more input conductors and output conductors may include wires that meet at one or more nodes between the wires, and the output current is substantially equal to the sum of the input currents.
[0257] In some embodiments, at least a first input current may be provided in the form of at least one photocurrent generated by at least one photodetector, which receives an optical signal generated by a first multiplication module of the multiplication module.
[0258] In some embodiments, the first input current may be provided as the difference between two photocurrents generated by different corresponding photodetectors that receive different corresponding optical signals generated by the first multiplication module.
[0259] In some embodiments, one of the copies of a first subset of one or more optical signals may consist of a single optical signal, wherein one of the input values is encoded on the single optical signal.
[0260] In some embodiments, the multiplication module corresponding to a copy of the first subset can multiply the encoded input value by a single matrix element value.
[0261] In some embodiments, one of the copies of a first subset of one or more optical signals may include more than one optical signal, but less than the total number of all optical signals, on which multiple input values are encoded.
[0262] In some embodiments, the multiplication module corresponding to a copy of the first subset can multiply the encoded input value by different corresponding matrix element values.
[0263] In some embodiments, different multiplication modules corresponding to different copies of a first subset of one or more optical signals may be included in different devices that perform optical communication to transmit one copy of the first subset of one or more optical signals between the different devices.
[0264] In some embodiments, at least two or more of a plurality of optical waveguides, two or more of a plurality of replication modules, two or more of a plurality of multiplication modules, and at least one of one or more summation modules may be disposed on the substrate of the common device.
[0265] In some embodiments, the apparatus performs vector-matrix multiplication, wherein an input vector can be provided as a set of optical signals and an output vector can be provided as a set of electrical signals.
[0266] In some embodiments, the apparatus may further include an accumulator that combines the input electrical signals corresponding to the outputs of a multiplication module or a summing module, wherein time-domain coding may be used to encode the input electrical signals, the time-domain coding using switching amplitude modulation in each of a plurality of time slots, and the accumulator may generate an output electrical signal that is encoded with more than two amplitude levels, the amplitude levels corresponding to different duty cycles of the time-domain coding on the plurality of time slots.
[0267] In some embodiments, each of two or more multiplication modules corresponds to a different subset of one or more optical signals.
[0268] In some embodiments, for each copy of a second subset of one or more optical signals that is different from the optical signals in a first subset of one or more optical signals, the apparatus may further include a multiplication module configured to multiply one or more optical signals of the second subset by one or more matrix element values using optical amplitude modulation.
[0269] In another aspect, a system includes: a storage unit configured to store a dataset and multiple neural network weights; and a driver unit configured to generate multiple modulator control signals. The system also includes an optoelectronic processor comprising: a light source configured to provide multiple light outputs; multiple light modulators coupled to the light source and the driver unit, the multiple light modulators configured to generate a light input vector by modulating the multiple light outputs generated by the light source based on the multiple modulator control signals; a matrix multiplication unit coupled to the multiple light modulators and the driver unit, the matrix multiplication unit configured to convert the light input vector into an analog output vector based on the multiple weight control signals; and a comparator unit coupled to the matrix multiplication unit and configured to convert the analog output vector into multiple 1-bit digital outputs. The system includes a controller, which includes an integrated circuit, configured to perform the following operations: receiving an artificial neural network computation request including an input dataset and first plurality of neural network weights, wherein the input dataset includes a first digital input vector having N-bit resolution; storing the input dataset and the first plurality of neural network weights in a storage unit; decomposing the first digital input vector into N 1-bit input vectors, each of the N 1-bit input vectors corresponding to one of the N bits of the first digital input vector; generating a sequence of N 1-bit modulator control signals corresponding to the N 1-bit input vectors via a driver unit; obtaining a sequence of N 1-bit digital outputs corresponding to the sequence of N 1-bit modulator control signals from a comparator unit; constructing an N-bit digital output vector from the sequence of N 1-bit digital outputs; performing a nonlinear transformation on the constructed N-bit digital output vector to produce a transformed N-bit digital output vector; and storing the transformed N-bit digital output vector in a storage unit.
[0270] Implementations of the system may include one or more of the following features. For example, receiving an artificial neural network computation request may include receiving an artificial neural network computation request from a general purpose computer.
[0271] In some embodiments, the driver unit may be configured to generate multiple weight control signals.
[0272] In some embodiments, the matrix multiplication unit may include: an optical matrix multiplication unit coupled to a plurality of optical modulator and driver units, the optical matrix multiplication unit being configured to convert an optical input vector into an optical output vector based on a plurality of weight control signals; and a photodetector unit coupled to the optical matrix multiplication unit and configured to generate a plurality of output voltages corresponding to the optical output vector.
[0273] In some embodiments, the matrix multiplication unit may include: an input waveguide array for receiving an optical input vector; an optical interference unit, optically communicating with the input waveguide array, for performing a linear transformation that converts the optical input vector into a second optical signal array; and an output waveguide array, optically communicating with the optical interference unit, for guiding the second optical signal array, wherein at least one input waveguide in the input waveguide array is optically communicating with each output waveguide in the output waveguide array via the optical interference unit.
[0274] In some embodiments, the optical interference unit may include: a plurality of interconnected Mach-Zehnder interferometers (MZIs), each of the plurality of interconnected MZIs including: a first phase shifter configured to change the separation ratio of the MZI; and a second phase shifter configured to shift the phase of an output of the MZI, wherein the first phase shifter and the second phase shifter may be coupled to a plurality of weight control signals.
[0275] In some embodiments, the matrix multiplication unit may include: a plurality of copying modules, for each of at least two subsets of one or more optical signals of an optical input vector, the plurality of copying modules including a corresponding set of one or more copying modules configured to divide a subset of one or more optical signals into two or more copies of optical signals; a plurality of multiplication modules, for each of at least two copies of a first subset of one or more optical signals, the plurality of multiplication modules including a corresponding multiplication module configured to multiply one or more optical signals of the first subset by one or more matrix element values using optical amplitude modulation; and one or more summing modules, for the results of two or more multiplication modules, the one or more summing modules including a summing module configured to generate an electrical signal representing the sum of the results of two or more multiplication modules.
[0276] In some embodiments, at least one multiplication module may include an optical amplitude modulator, which includes an input port and two output ports, and can provide a pair of correlated optical signals from the two output ports such that the difference between the amplitudes of the correlated optical signals corresponds to the result of multiplying the input value by the value of the signed matrix element.
[0277] In some embodiments, the matrix multiplication unit may be configured to multiply an input vector by a matrix that includes one or more matrix element values.
[0278] In some embodiments, a set of multiple output values may be encoded on corresponding electrical signals generated by one or more summing modules, and the output values in the set of multiple output values may represent elements of an output vector generated by multiplying an input vector by a matrix.
[0279] In another aspect, a method is provided for performing artificial neural network computation in a system having a matrix multiplication unit configured to convert an optical input vector into an analog output vector based on a plurality of weight control signals. The method includes: receiving an artificial neural network computation request comprising an input dataset and a first plurality of neural network weights, wherein the input dataset includes a first digital input vector; storing the input dataset and the first plurality of neural network weights in a storage unit; generating a first plurality of modulator control signals based on the first digital input vector and generating a first plurality of weight control signals based on the first plurality of neural network weights; obtaining a first plurality of digital outputs corresponding to the output vector of the matrix multiplication unit, the first plurality of digital outputs forming a first digital output vector; performing a nonlinear transformation on the first digital output vector by a controller to generate a first transformed digital output vector; storing the first transformed digital output vector in a storage unit; and outputting the artificial neural network output generated based on the first transformed digital output vector by the controller.
[0280] Embodiments of the method may include one or more of the following features. For example, receiving an artificial neural network computation request may include receiving the artificial neural network computation request from a computer via a communication channel.
[0281] In some embodiments, generating the first plurality of modulator control signals may include generating the first plurality of modulator control signals via a digital-to-analog converter (DAC) unit.
[0282] In some embodiments, obtaining the first plurality of digital outputs may include obtaining the first plurality of digital outputs from an analog-to-digital converter (ADC) unit.
[0283] In some embodiments, the method may include: applying a first plurality of modulator control signals to a plurality of optical modulators coupled to a light source and a DAC unit; and using the plurality of optical modulators to generate an optical input vector by modulating a plurality of optical outputs generated by a laser unit based on the plurality of modulator control signals.
[0284] In some embodiments, the matrix multiplication unit may be coupled to multiple optical modulators and DAC units, and the method may include: using the matrix multiplication unit to convert an optical input vector into an analog output vector based on multiple weight control signals.
[0285] In some embodiments, the ADC unit may be coupled to a matrix multiplication unit, and the method may include: using the ADC unit to convert an analog output vector into a first plurality of digital outputs.
[0286] In some embodiments, the matrix multiplication unit may include an optical matrix multiplication unit coupled to multiple optical modulators and DAC units. Converting an optical input vector into an analog output vector may include using the optical matrix multiplication unit to convert the optical input vector into an optical output vector based on multiple weighted control signals. The method may include using a photodetector unit coupled to the optical matrix multiplication unit to generate multiple output voltages corresponding to the optical output vector.
[0287] In some embodiments, the method may include: receiving an optical input vector in an input waveguide array; performing a linear transformation to convert the optical input vector into a second optical signal array using an optical interference unit that is optically in communication with the input waveguide array; and guiding the second optical signal array using an output waveguide array that is optically in communication with the optical interference unit, wherein at least one input waveguide in the input waveguide array is optically in communication with each output waveguide in the output waveguide array via the optical interference unit.
[0288] In some embodiments, the optical interference unit may include: a plurality of interconnected Mach-Zehnder interferometers (MZIs), each of the plurality of interconnected MZIs including a first phase shifter and a second phase shifter, and the first and second phase shifters being coupled to a plurality of weight control signals. The method may include: changing the separation ratio of the MZI using the first phase shifter, and shifting the phase of an output of the MZI using the second phase shifter.
[0289] In some embodiments, the method may include: for each of at least two subsets of one or more optical signals of an optical input vector, dividing the subset of one or more optical signals into two or more copies of the optical signals using a corresponding set of one or more replication modules; for each of at least two copies of a first subset of one or more optical signals, multiplying the one or more optical signals of the first subset by one or more matrix element values using a corresponding multiplication module with optical amplitude modulation; and for the results of two or more multiplication modules, using a summation module to generate an electrical signal representing the sum of the results of the two or more multiplication modules.
[0290] In some embodiments, at least one multiplication module may include an optical amplitude modulator, which includes an input port and two output ports, and can provide a pair of correlated optical signals from the two output ports such that the difference between the amplitudes of the correlated optical signals corresponds to the result of multiplying the input value by the value of the signed matrix element.
[0291] In some embodiments, the method may include multiplying an input vector by a matrix that includes one or more matrix element values using a matrix multiplication unit.
[0292] In some embodiments, the method may include encoding a set of multiple output values on corresponding electrical signals generated by one or more summing modules, and using the output values from the set of multiple output values to represent elements of an output vector, which is generated by multiplying an input vector by a matrix.
[0293] In another aspect, one method includes: providing input information in an electronic format; converting at least a portion of the electronic input information into an optical input vector; photoelectrically converting the optical input vector into an analog output vector based on matrix multiplication; and applying a nonlinear conversion to the analog output vector to provide output information in an electronic format.
[0294] Embodiments of the method may include one or more of the following features. For example, the method may further include: repeating electro-optical conversion, photoelectric conversion, and nonlinear conversion of electrical applications for new electronic input information corresponding to output information provided in electronic format.
[0295] In some embodiments, the matrix multiplication for the initial photoelectric conversion and the matrix multiplication for the repeated photoelectric conversion can be the same and can correspond to the same layer of an artificial neural network.
[0296] In some embodiments, the matrix multiplication used for the initial photoelectric conversion and the matrix multiplication used for the repeated photoelectric conversion can be different and can correspond to different layers of an artificial neural network.
[0297] In some embodiments, the method may further include: repeating electro-optical conversion, photoelectric conversion, and nonlinear conversion of electrical applications for different portions of electronic input information, wherein the matrix multiplication for the initial photoelectric conversion and the matrix multiplication for the repeated photoelectric conversion are the same and correspond to the first layer of the artificial neural network.
[0298] In some embodiments, the method may further include: providing intermediate information in an electronic format based on electronic output information for multiple portions of electronic input information generated by a first layer of an artificial neural network; and repeating electro-optical conversion, photoelectric conversion, and nonlinear conversion of electrical applications for each different portion of the electronic intermediate information, wherein the matrix multiplication for the initial photoelectric conversion and the matrix multiplication for the repeated photoelectric conversion associated with the different portions of the electronic intermediate information are the same and correspond to the second layer of the artificial neural network.
[0299] In another aspect, a system is provided for performing artificial neural network computation. The system includes: a first unit configured to generate a plurality of vector control signals and a plurality of weight control signals; a second unit configured to provide an optical input vector based on the plurality of vector control signals; and a matrix multiplication unit coupled to the second unit and the first unit, the matrix multiplication unit being configured to convert the optical input vector into an output vector based on the plurality of weight control signals. The system includes a controller comprising an integrated circuit configured to perform the following operations: receiving an artificial neural network computation request including an input dataset and a first plurality of neural network weights, wherein the input dataset includes a first digital input vector; and generating the first plurality of vector control signals based on the first digital input vector and generating the first plurality of weight control signals based on the first plurality of neural network weights through the first unit; wherein the first unit, the second unit, the matrix multiplication unit, and the controller are used in a photoelectric processing loop repeated in a plurality of iterations, and the photoelectric processing loop includes: (1) at least two optical modulation operations, and (2) at least one of (a) an electrical summation operation or (b) an electrical storage operation.
[0300] In another aspect, a method is provided for performing artificial neural network computation. The method includes: providing input information in an electronic format; converting at least a portion of the electronic input information into an optical input vector; and converting the optical input vector into an output vector based on matrix multiplication using a set of neural network weights. The provisioning and conversion are performed in a photoelectric processing loop, repeated in multiple iterations using different corresponding sets of neural network weights and different corresponding input information, and the photoelectric processing loop includes: (1) at least two optical modulation operations, and (2) at least one of (a) an electrical summation operation or (b) an electrical storage operation.
[0301] Details of one or more embodiments of the subject matter described in this disclosure are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of this disclosure will become apparent from the specification and drawings.
[0302] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. In the event of any conflict with a patent application or patent application publication incorporated herein, this disclosure (including definitions) shall prevail. Attached Figure Description
[0303] This disclosure is best understood from the following detailed description when read in conjunction with the accompanying drawings. It is to be emphasized that, by convention, the various features in the drawings are not to scale. Rather, for clarity, the dimensions of the various features have been arbitrarily enlarged or reduced.
[0304] Figure 1AThis is a schematic diagram of an example of an artificial neural network (ANN) system.
[0305] Figure 1B This is a schematic diagram of an example of a light matrix multiplication unit.
[0306] Figure 1C and Figure 1D This is a schematic diagram of an example configuration of an interconnected Mach-Zehnder interferometer (MZI).
[0307] Figure 1E This is a schematic diagram of an example of MZI.
[0308] Figure 1F This is a schematic diagram of an example of a wavelength division multiplexed ANN (ANN) computing system.
[0309] Figure 2A This is a flowchart showing an example of a method used to perform ANN computation.
[0310] Figure 2B It is a display Figure 2A A diagram of one aspect of the method.
[0311] Figure 3A and Figure 3B This is a schematic diagram of an example of an ANN computing system.
[0312] Figure 4A This is a schematic diagram of an example of an ANN computing system with 1-bit internal resolution.
[0313] Figure 4B yes Figure 4A The mathematical representation of the operation of the ANN computing system.
[0314] Figure 5 This is a schematic diagram of an example of an artificial neural network (ANN) computing system.
[0315] Figure 6 This is a schematic diagram of an example of a light matrix multiplication unit.
[0316] Figure 7 This is a schematic diagram of an example of an artificial neural network (ANN) computing system.
[0317] Figure 8 This is a diagram of an example of a light matrix multiplication unit.
[0318] Figure 9 This is a schematic diagram of an example of an artificial neural network (ANN) computing system.
[0319] Figure 10This is a diagram of an example of a light matrix multiplication unit.
[0320] Figure 11 This is a diagram of an example of a compact matrix multiplier unit.
[0321] Figure 12A A diagram showing the comparison of photon matrix multiplier units is displayed.
[0322] Figure 12B This is a diagram of a compact interconnected interferometer.
[0323] Figure 13 This is a diagram of a compact matrix multiplier unit.
[0324] Figure 14 This is a diagram of an optical generative adversarial network.
[0325] Figure 15 This is a diagram of the Mach-Zehnder interferometer.
[0326] Figure 16 , Figure 17A as well as Figure 17B This is a diagram of a photonic circuit.
[0327] Figure 18 This is a schematic diagram of an example of an optoelectronic computing system.
[0328] Figure 19A and Figure 19B This is a schematic diagram of the example system configuration.
[0329] Figure 20A This is a schematic diagram illustrating an example of a symmetric differential configuration.
[0330] Figure 20B and Figure 20C This is a circuit diagram of an example system module.
[0331] Figure 21A This is a schematic diagram of an example of a symmetric differential configuration.
[0332] Figure 21B This is a schematic diagram illustrating an example of system configuration.
[0333] Figure 22A This is a schematic diagram of an example optical amplitude modulator.
[0334] Figures 22B to 22D This is a schematic diagram of an example of an optical amplitude modulator using optical detection in a symmetrical differential configuration.
[0335] Figures 23A to 23C This is a sample system configuration optoelectronic circuit diagram.
[0336] Figures 24A to 24E This is a schematic diagram of an example computing system that uses multiple optoelectronic systems.
[0337] Figure 25 This is a flowchart showing an example of a method used to perform ANN computation.
[0338] Figure 26 and Figure 27 This is a schematic diagram of an example of an ANN computing system.
[0339] Figure 28 This is a schematic diagram of an example of a neural network computing system using a passive 2D optical matrix multiplication unit.
[0340] Figure 29 This is a schematic diagram of an example of a neural network computing system using a passive 3D light matrix multiplication unit.
[0341] Figure 30 This is a schematic diagram of an example of an artificial neural network computing system with 1-bit internal resolution, where the system uses passive 2D optical matrix multiplication units.
[0342] Figure 31 This is a schematic diagram of an example of an artificial neural network computing system with 1-bit internal resolution, where the system uses a passive 3D light matrix multiplication unit.
[0343] Figure 32A This is a schematic diagram of an example of an artificial neural network (ANN) computing system.
[0344] Figure 32B This is a schematic diagram of an example of a photoelectric matrix multiplication unit.
[0345] Figure 33 This is a flowchart showing an example of a method for performing ANN calculations using a photoelectric processor.
[0346] Figure 34 It is a display Figure 33 A diagram of one aspect of the method.
[0347] Figure 35A This is a schematic diagram of an example of a wavelength division multiplexing (WDM) ANN computing system using an optoelectronic processor.
[0348] Figure 35B and Figure 35C This is a schematic diagram of an example of a wavelength division multiplexing optoelectronic matrix multiplication unit.
[0349] Figure 36 and Figure 37 This is a schematic diagram of an example of an ANN computing system using photoelectric matrix multiplication units.
[0350] Figure 38 This is a schematic diagram of an example of an artificial neural network computing system with 1-bit internal resolution, where the system uses photoelectric matrix multiplication units.
[0351] Figure 39A This is a schematic diagram of an example of a Mach-Zehnder modulator.
[0352] Figure 39B It is a display Figure 39A The intensity-voltage curve of the Mach-Zehnder modulator.
[0353] Figure 40 This is a schematic diagram of a zero-difference detector.
[0354] Figure 41 It is a schematic diagram of a computing system that includes optical fibers, each of which carries signals of multiple wavelengths.
[0355] The same reference numerals and names in the figures indicate the same components. Detailed Implementation
[0356] Figure 1A A schematic diagram of an example of an artificial neural network (ANN) computing system 100 is shown. System 100 includes a controller 110, a storage unit 120, a digital-to-analog converter (DAC) unit 130, an optical processor 140, and an analog-to-digital converter (ADC) unit 160. The controller 110 is coupled to a computer 102, the storage unit 120, the DAC unit 130, and the ADC unit 160. The controller 110 includes an integrated circuit configured to control the operation of the ANN computing system 100 to perform ANN computations.
[0357] The integrated circuit of controller 110 may be a special-purpose integrated circuit specifically configured to perform steps of ANN computation processing. For example, the integrated circuit may implement microcode or firmware specific to performing ANN computation processing. In this way, controller 110 may have a reduced instruction set relative to a general-purpose processor used in a conventional computer (e.g., computer 102). In some embodiments, the integrated circuit of controller 110 may include two or more circuits configured to perform different steps of ANN computation processing.
[0358] In an example operation of the ANN computing system 100, the computer 102 may issue an artificial neural network computing request to the ANN computing system 100. The ANN computing request may include defining the neural network weights of the ANN and the input dataset to be processed by the provided ANN. The controller 110 receives the ANN computing request and stores the input dataset and neural network weights in the storage unit 120.
[0359] The input dataset can correspond to various digital information that the ANN will process. Examples of input datasets include image files, audio files, LiDAR point clouds, and GPS coordinate sequences, and the operation of the ANN computing system 100 will be described based on the received image file as the input dataset. Typically, the size of the input dataset can vary greatly, from hundreds of data points to millions or more. For example, a digital image file with a megapixel resolution has approximately one million pixels, and each of those one million pixels can be a data point processed by the ANN. Due to the large number of data points in a typical input dataset, the input dataset is usually divided into multiple smaller-sized digital input vectors for separate processing by the optical processor 140. As an example, for a grayscale digital image, the elements of the digital input vector can be 8-bit values representing image intensity, and the digital input vector can have a length ranging from tens of elements (e.g., 32 elements, 64 elements) to hundreds of elements (e.g., 256 elements, 512 elements). Generally, an input dataset of any size can be divided into numeric input vectors of a size suitable for processing by the optical processor 140. If the number of elements in the input dataset is not divisible by the length of the numeric input vectors, zero padding can be used to pad the dataset so that it is divisible by the length of the numeric input vectors. The processing outputs of the individual numeric input vectors can be processed to reconstruct the complete output, which is the result of processing the input dataset through an ANN. In some embodiments, block matrix multiplication can be used to implement the division of the input dataset into multiple input vectors and subsequent vector-level processing.
[0360] Neural network weights are a set of values that define the connectivity of the artificial neurons in an ANN, including the relative importance or weight of those connections. An ANN may include one or more hidden layers with corresponding sets of nodes. In the case of an ANN with a single hidden layer, the ANN can be defined by two sets of neural network weights: one set corresponding to the connectivity between the input nodes and the nodes in the hidden layer, and the second set corresponding to the connectivity between the hidden layer and the output nodes. Each set of neural network weights describing connectivity corresponds to a matrix implemented by the optical processor 140. For ANNs with two or more hidden layers, additional sets of neural network weights are needed to define the connectivity between the additional hidden layers. Thus, in a typical ANN computation request, the neural network weights included may include multiple sets of neural network weights representing the connectivity between the various layers of the ANN.
[0361] Because the input dataset to be processed is typically divided into multiple smaller numeric input vectors for separate processing, the input dataset is usually stored in digital storage. However, the speed of storage operations between the storage and processor of computer 102 is significantly slower than the rate at which ANN computing system 100 can perform ANN computations. For example, ANN computing system 100 can perform dozens to hundreds of ANN computations during a typical memory read period of computer 102. Thus, if the ANN computation of ANN computing system 100 involves multiple data transfers between system 100 and computer 102 during the processing of ANN computation requests, the rate at which ANN computations can be performed by ANN computing system 100 can be limited to its full processing speed. For example, if computer 102 needs to access the input dataset from its own storage and provide the numeric input vectors to controller 110 upon request, the operation of ANN computing system 100 may be significantly slowed down by the time required for the series of data transfers between computer 102 and controller 110. It is worth noting that the memory access latency of computer 102 is typically non-deterministic, which further complicates and reduces the speed at which digital input vectors can be provided to the ANN computing system 100. Furthermore, processor cycles of computer 102 may be wasted in managing data transfers between computer 102 and the ANN computing system 100.
[0362] Conversely, in some embodiments, the ANN computing system 100 stores the entire input dataset in a storage unit 120, which is part of and dedicated to the ANN computing system 100. The dedicated storage unit 120 allows for transactions between the storage unit 120 and the controller 110 that are particularly well-suited to allow for a smooth and uninterrupted data flow between them. This uninterrupted data flow significantly improves the overall throughput of the ANN computing system 100 by allowing the optical processor 140 to perform matrix multiplications at its full processing speed, without being limited by the slow storage operations of a conventional computer (e.g., computer 102). Furthermore, because all the data required to perform ANN computations is provided to the ANN computing system 100 by computer 102 in a single transaction, the ANN computing system 100 can perform its ANN computations in a manner independent of computer 102. The unique operation of this ANN computing system 100 reduces the computational burden on the computer 102 and eliminates external dependencies in the operation of the ANN computing system 100, thereby improving the performance of both the system 100 and the computer 102.
[0363] The internal operation of the ANN computing system 100 will now be described. The optical processor 140 includes a laser unit 142, a modulator array 144, a detection unit 146, and an optical matrix multiplication (OMM) unit 150. The optical processor 140 operates by encoding a digital input vector of length N onto an optical input vector of length N and propagating the optical input vector through the OMM unit 150. The OMM unit 150 receives the optical input vector of length N and performs an N×N matrix multiplication on the received optical input vector in the optical domain. The N×N matrix multiplication performed by the OMM unit 150 is determined by the internal configuration of the OMM unit 150. The internal configuration of the OMM unit 150 can be controlled by electrical signals, such as those generated by the DAC unit 130.
[0364] OMM unit 150 can be implemented in various ways. Figure 1BA schematic diagram of an example OMM unit 150 is shown. The OMM unit 150 may include an array of input waveguides 152 to receive an optical input vector; an optical interference unit 154 in optical communication with the array of input waveguides 152; and an array of output waveguides 156 in optical communication with the optical interference unit 154. The optical interference unit 154 linearly converts the optical input vector into a second optical signal array. The array of output waveguides 156 guides the second optical signal array output by the optical interference unit 154. At least one input waveguide in the array of input waveguides 152 is in optical communication with each output waveguide in the array of output waveguides 156 via the optical interference unit 154. For example, for an optical input vector of length N, the OMM unit 150 may include N input waveguides 152 and N output waveguides 156.
[0365] The optical interference unit may include multiple interconnected Mach-Zehnder interferometers (MZIs). Figure 1C and Figure 1D Schematic diagrams of configurations 157 and 158 showing examples of interconnected MZIs are shown. MZIs can be interconnected in various ways (e.g., in configurations 157 or 158) to achieve linear transformation of the optical input vector received by the array of input waveguides 152.
[0366] Figure 1EA schematic diagram of an example MZI 170 is shown. The MZI 170 includes a first input waveguide 171, a second input waveguide 172, a first output waveguide 178, and a second output waveguide 179. Furthermore, each of the plurality of interconnected MZIs 170 includes a first phase shifter 174 configured to change the splitting ratio of the MZI 170; and a second phase shifter 176 configured to shift the phase of one output of the MZI 170, for example, light exiting the MZI 170 via the second output waveguide 179. The first phase shifter 174 and the second phase shifter 176 of the MZI 170 are coupled to a plurality of weighted control signals generated by the DAC unit 130. The first phase shifter 174 and the second phase shifter 176 are examples of reconfigurable components of the OMM unit 150. Examples of reconfigurable components include thermo-optic phase shifters or electro-optic phase shifters. Thermo-optic phase shifters operate by heating the waveguide to change the refractive index of the waveguide and cladding material, which translates into a phase change. Electro-optic phase shifters operate by applying an electric field (e.g., lithium niobate (LiNbO3), reverse-biased PN junction) or current (e.g., forward-biased PIN junction), which changes the refractive index of the waveguide material. By changing the weighting control signal, the phase delay of the first phase shifter 174 and the second phase shifter 176 of each interconnected MZI 170 can be altered. This reconfigures the optical interferometer 154 of the OMM unit 150 to achieve a specific matrix multiplication determined by the phase delay set across the entire optical interferometer 154. Additional embodiments of the OMM unit 150 and the optical interference unit are disclosed in U.S. Patent Publication No. US 2017 / 0351293A1 entitled “APPARATUS AND METHODS FOR OPTICAL NEURAL NETWORK”, which is incorporated herein by reference in its entirety.
[0367] A light input vector is generated by laser unit 142 and modulator array 144. The light input vector of length N has N independent light signals, the intensity of which corresponds to the value of the corresponding element of the digital input vector of length N. As an example, laser unit 142 can generate N light outputs. The N light outputs have the same wavelength and are optically coherent. The optical coherence of the light outputs allows them to interfere with each other, a characteristic utilized by OMM unit 150 (e.g., in MZI operation). Furthermore, the light outputs of laser unit 142 can be substantially identical to each other. For example, the N light outputs can be substantially uniform in their intensity (e.g., within 5%, 3%, 1%, 0.5%, 0.1%, or 0.01%) and their relative phase (e.g., within 10 degrees, 5 degrees, 3 degrees, 1 degree, or 0.1 degrees). Uniformity of the optical output can improve the faithfulness of the optical input vector to the digital input vector, thereby improving the overall accuracy of the optical processor 140. In some embodiments, the optical output of the laser unit 142 may have an optical power of 0.1mW to 50mW per output, a wavelength in the near-infrared range (e.g., between 900nm and 1600nm), and a linewidth of less than 1nm. The optical output of the laser unit 142 may be a single transverse-mode optical output.
[0368] In some embodiments, the laser unit 142 includes a single laser source and an optical power splitter. The single laser source is configured to generate laser light. The optical power splitter is configured to split the light generated by the laser source into N optical outputs having substantially the same intensity and phase. By splitting the single laser output into multiple outputs, optical coherence of the multiple optical outputs can be achieved. For example, the single laser source may be a semiconductor laser diode, a vertical-cavity surface-emitting laser (VCSEL), a distributed feedback (DFB) laser, or a distributed Bragg reflector (DBR) laser. For example, the optical power splitter may be a 1:N multimode interference (MMI) splitter, a multi-stage splitter including multiple 1:2 MMI splitters or directional couplers, or a star coupler. In some other embodiments, a master-slave laser configuration can be used, in which the slave laser is injected and locked by the master laser to have a stable phase relationship with the master laser.
[0369] The optical output of laser unit 142 is coupled to modulator array 144. Modulator array 144 is configured to receive optical input from laser unit 142 and modulate the intensity of the received optical input based on modulator control signals (which are electrical signals). Examples of modulators include Mach-Zehnder interferometer (MZI) modulators, ring resonator modulators, and electro-absorption modulators. Modulator array 144 has N modulators, each modulator receiving one of the N optical outputs of laser unit 142. The modulator receives control signals corresponding to elements of a digital input vector and modulates the intensity of the light. The control signals may be generated by DAC unit 130.
[0370] DAC unit 130 is configured to generate multiple modulator control signals and multiple weight control signals under the control of controller 110. For example, DAC unit 130 receives a first DAC control signal from controller 110, the first DAC control signal corresponding to a digital input vector to be processed by optical processor 140. DAC unit 130 generates modulator control signals based on the first DAC control signal, the modulator control signals being analog signals suitable for driving modulator array 144. For example, the analog signal can be a voltage or a current, depending on the technology and design of the modulator of array 144. Voltages can have amplitudes ranging from ±0.1V to ±10V, and currents can have amplitudes ranging from 100μA to 100mA. In some embodiments, DAC unit 130 may include a modulator driver configured to buffer, amplify, or modulate the analog signal such that the modulator of array 144 can be sufficiently driven. For example, certain types of modulators can be driven with differential control signals. In this case, the modulator driver can be a differential driver that generates a differential electrical output based on a single-ended input signal. As another example, some types of modulators may have a 3dB bandwidth, which is less than the desired processing rate of the optical processor 140. In this case, the modulator driver may include a pre-emphasis circuit or other bandwidth enhancement circuitry designed to extend the modulator's operating bandwidth.
[0371] In some cases, the modulator of array 144 may have a nonlinear transfer function. For example, an MZI optical modulator may have a nonlinear relationship between the applied control voltage and its transmission (e.g., sinusoidal dependence). In this case, the first DAC control signal can be adjusted or compensated based on the modulator's nonlinear transfer function so that the linear relationship between the digital input vector and the resulting optical input vector is maintained. Maintaining this linearity is generally important to ensure that the input to OMM unit 150 is an accurate representation of the digital input vector. In some embodiments, compensation of the first DAC control signal can be performed by controller 110 using a lookup table that maps the values of the digital input vector to the values to be output by DAC unit 130, such that the resulting modulated optical signal is linearly proportional to the elements of the digital input vector. The lookup table can be generated by characterizing the nonlinear transfer function of the modulator and calculating the inverse function of the nonlinear transfer function.
[0372] In some embodiments, the nonlinearity of the modulator and the nonlinearity obtained in the resulting optical input vector can be compensated by an ANN calculation algorithm.
[0373] The optical input vector generated by the modulator array 144 is input to the OMM unit 150. The optical input vector can be N spatially separated optical signals, each with an optical power corresponding to an element of the digital input vector. For example, the optical power of the optical signals is typically in the range of 1 μW to 10 mW. The OMM unit 150 receives the optical input vector and performs an N×N matrix multiplication based on its internal configuration. The internal configuration is controlled by an electrical signal generated by the DAC unit 130. For example, the DAC unit 130 receives a second DAC control signal from the controller 110, the second DAC control signal corresponding to the neural network weights to be implemented by the OMM unit 150. The DAC unit 130 generates a weight control signal based on the second DAC control signal, the weight control signal being an analog signal suitable for controlling the reconfigurable components within the OMM unit 150. For example, the analog signal can be voltage or current, depending on the type of reconfigurable components in the OMM unit 150. Voltages can have amplitudes ranging from 0.1V to 10V, and currents can have amplitudes ranging from 100 μA to 10 mA.
[0374] Modulator array 144 can operate at a modulation rate different from the reconfiguration rate of reconfigurable OMM unit 150. The optical input vector generated by modulator array 144 propagates through the OMM unit at approximately a proportion of the speed of light (e.g., 80%, 50%, or 25% of the speed of light), depending on the optical characteristics of OMM unit 150 (e.g., effective refractive index). For a typical OMM unit 150, the propagation time of the optical input vector is in the range of 1 to tens of picoseconds, corresponding to processing rates of tens to hundreds of GHz. Thus, the rate at which optical processor 140 can perform matrix multiplication operations is partially limited by the rate at which the optical input vector can be generated. Modulators with bandwidths of tens of GHz are readily available, and modulators with bandwidths exceeding 100 GHz are under development. Thus, for example, the modulation rate of modulator array 144 can be in the range of 5 GHz, 8 GHz, or tens to hundreds of GHz. In order to maintain the operation of the modulator array 144 at such a modulation rate, the integrated circuit of the controller 110 can be configured to output control signals for the DAC unit 130 at a rate greater than or equal to, for example, 5 GHz, 8 GHz, 10 GHz, 20 GHz, 25 GHz, 50 GHz or 100 GHz.
[0375] Depending on the type of reconfigurable component implemented by OMM unit 150, the reconfiguration rate of OMM unit 150 can be significantly slower than the modulation rate. For example, the reconfigurable component of OMM unit 150 can be thermo-optical, using microheaters to adjust the temperature of the optical waveguide of OMM unit 150, which in turn affects the phase of the optical signal within OMM unit 150 and results in matrix multiplication. Due to the thermal time constant associated with the heating and cooling of the structure, the reconfiguration rate can be limited to several hundred kHz to several tens of MHz. Consequently, the modulator control signal used to control modulator array 144 and the weight control signal used to reconfigure OMM unit 150 may have significantly different speed requirements. Furthermore, the electrical characteristics of modulator array 144 can be significantly different from the electrical characteristics of the reconfigurable component of OMM unit 150.
[0376] To accommodate the different characteristics of the modulator control signal and the weight control signal, in some embodiments, the DAC unit 130 may include a first DAC subunit 132 and a second DAC subunit 134. The first DAC subunit 132 may be specifically configured to generate the modulator control signal, and the second DAC subunit 134 may be specifically configured to generate the weight control signal. For example, the modulation rate of the modulator array 144 may be 25 GHz, and the first DAC subunit 132 may have a per-channel output update rate of 25 giga-samples per second (GSPS) and a resolution of 8 bits or higher. The reconfiguration rate of the OMM unit 150 may be 1 MHz, and the second DAC subunit 134 may have an output update rate of 1 mega-samples per second (MSPS) and a resolution of 10 bits. Implementing separate first DAC subunits 132 and second DAC subunits 134 allows for independent optimization of the DAC subunits for the respective signals, which can reduce the overall power consumption, complexity, cost, or a combination thereof of the DAC unit 130. It is worth noting that although the first DAC subunit 132 and the second DAC subunit 134 are described as sub-components of the DAC unit 130, in general, the first DAC subunit 132 and the second DAC subunit 134 can be integrated on a common chip or implemented as separate chips.
[0377] Based on the different characteristics of the first DAC subunit 132 and the second DAC subunit 134, in some embodiments, the storage unit 120 may include a first storage subunit and a second storage subunit. The first storage subunit may be a memory dedicated to storing the input dataset and digital input vectors, and may have an operating speed sufficient to support the modulation rate. The second storage subunit may be a memory dedicated to storing neural network weights, and may have an operating speed sufficient to support the reconfiguration rate of the OMM unit 150. In some embodiments, the first storage subunit may be implemented using SRAM, and the second storage subunit may be implemented using DRAM. In some embodiments, both the first and second storage subunits may be implemented using DRAM. In some embodiments, the first storage unit may be implemented as part of the controller 110 or as a cache of the controller 110. In some embodiments, the first and second storage subunits may be implemented by a single physical storage device in different address spaces.
[0378] OMM unit 150 outputs a light output vector of length N, corresponding to the result of an N×N matrix multiplication of the light input vector and the neural network weights. OMM unit 150 is coupled to detection unit 146, which is configured to generate N output voltages corresponding to the N light output vector. For example, detection unit 146 may include an array of N photodetectors (configured to absorb light signals and generate photocurrent) and an array of N transimpedance amplifiers (configured to convert photocurrent into output voltage). The bandwidths of the photodetectors and transimpedance amplifiers can be set based on the modulation rate of modulator array 144. The photodetectors can be formed from various materials based on the wavelength of the detected light output vector. Examples of materials used for photodetectors include germanium, silicon-germanium alloys, and indium gallium arsenide (InGaAs).
[0379] Detection unit 146 is coupled to ADC unit 160. ADC unit 160 is configured to convert N output voltages into N digital optical outputs, where the N digital optical outputs are quantized digital representations of the output voltages. For example, ADC unit 160 may be an N-channel ADC. Controller 110 can obtain the N digital optical outputs corresponding to the optical output vector of optical matrix multiplication unit 150 from ADC unit 160. Controller 110 can then form a digital output vector of length N from the N digital optical outputs, which corresponds to the result of an N×N matrix multiplication of an input digital vector of length N.
[0380] Various electronic components of the ANN computing system 100 can be integrated in various ways. For example, the controller 110 can be an application-specific integrated circuit (ASIC) fabricated on a semiconductor die. Other electronic components (e.g., memory cell 120, DAC cell 130, ADC cell 160, or combinations thereof) can be monolithically integrated on the semiconductor die on which the controller 110 is fabricated. As another example, two or more electronic components can be integrated into a system-on-a-chip (SoC). In an SoC embodiment, the controller 110, memory cell 120, DAC cell 130, and ADC cell 160 can be fabricated on their respective dies, and the respective dies can be integrated on a common platform (e.g., an interposer) that provides electrical connections between the integrated components. Compared to methods that separately arrange and route components on a printed circuit board (PCB), this SoC approach allows for faster data transfer between the electronic components of the ANN computing system 100, thereby improving the operating speed of the ANN computing system 100. Furthermore, the SoC approach allows for the use of different manufacturing techniques optimized for different electronic components, which can improve the performance of different components and reduce the overall cost of the monolithic integration approach. While the integration of controller 110, memory unit 120, DAC unit 130, and ADC unit 160 has been described, typically a subset of components can be integrated, while other components are implemented as discrete components for various reasons (e.g., performance or cost). For example, in some embodiments, memory unit 120 may be integrated with controller 110 as a functional block within controller 110.
[0381] Various optical components of the ANN computing system 100 can also be integrated in various ways. Examples of optical components in the ANN computing system 100 include a laser unit 142, a modulator array 144, an OMM unit 150, and a photodetector for the detection unit 146. These optical components can be integrated in various ways to improve performance and / or reduce cost. For example, the laser unit 142, modulator array 144, OMM unit 150, and photodetector can be monolithically integrated on a common semiconductor substrate serving as a photonic integrated circuit (PIC). On a photonic integrated circuit formed based on a compound semiconductor material system (e.g., a III-V compound semiconductor such as indium phosphide (InP)), the laser, modulator (e.g., an electroabsorption modulator), waveguide, and photodetector can be monolithically integrated on a single die. This monolithic integration method reduces the complexity of aligning the inputs and outputs of various discrete optical components, which may require alignment accuracy ranging from sub-micron to several micrometers. As another example, the laser source of laser unit 142 can be fabricated on a compound semiconductor die, while the optical power splitter, modulator array 144, OMM unit 150, and photodetector of detection unit 146 of laser unit 142 can be fabricated on a silicon die. PICs fabricated on silicon wafers (often referred to as silicon photonics technology) typically offer greater integration density, higher lithographic resolution, and lower cost compared to III-V based PICs. This greater integration density can be beneficial in the fabrication of OMM unit 150, as OMM unit 150 typically includes dozens to hundreds of optical components, such as power splitters and phase shifters. Furthermore, the higher lithographic resolution of silicon photonics technology can reduce fabrication variations in OMM unit 150, thereby improving the accuracy of OMM unit 150.
[0382] The ANN computing system 100 can be implemented in various forms. For example, the ANN computing system 100 can be implemented as a co-processor inserted into a host computer. This ANN computing system 100 may have, for example, a high-speed PCI (PCI Express) card form factor and communicate with the host computer via a PCIe bus. The host computer can host multiple co-processor-type ANN computing systems 100 and connect them to computer 102 via a network. This type of embodiment is suitable for cloud data centers, where server racks can be dedicated to processing ANN computing requests received from other computers or servers. As another embodiment, the co-processor-type ANN computing system 100 can be directly inserted into the computer 102 that issued the ANN computing request.
[0383] In some embodiments, the ANN computing system 100 can be integrated into physical systems that require real-time ANN computing capabilities. For example, systems heavily reliant on real-time artificial intelligence tasks (such as autonomous vehicles, autonomous drones, object or facial recognition security cameras, and various Internet of Things (IoT) devices) can benefit from directly integrating the ANN computing system 100 with other subsystems of such systems. Direct integration of the ANN computing system 100 enables real-time artificial intelligence in devices with poor or no network connectivity and enhances the reliability and availability of mission-critical AI systems.
[0384] Although DAC unit 130 and ADC unit 160 are shown coupled to controller 110, in some embodiments, DAC unit 130, ADC unit 160, or both may alternatively or additionally be coupled to memory unit 120. For example, direct memory access (DMA) operations of DAC unit 130 or ADC unit 160 can reduce the computational burden on controller 110 and reduce latency in reading from and writing to memory unit 120, thereby further improving the operating speed of ANN computing unit 100.
[0385] Figure 2A A flowchart illustrating an example of a method 200 for performing ANN computation is shown. The steps of method 200 can be executed by controller 110. In some embodiments, the steps of method 200 can be run in parallel, combined, cyclically, or in any order.
[0386] In step 210, an artificial neural network (ANN) computation request is received, including an input dataset and a first plurality of neural network weights. The input dataset includes a first digital input vector. The first digital input vector is a subset of the input dataset. For example, it can be a sub-region of an image. The ANN computation request can be generated by various entities (e.g., computer 102). The computer can include one or more of various types of computing devices, such as a personal computer, a server computer, a vehicle computer, and a flight computer. The ANN computation request generally refers to an electrical signal notifying or informing the ANN computation system 100 that an ANN computation is about to be performed. In some embodiments, the ANN computation request can be divided into two or more signals. For example, the first signal may query the ANN computation system 100 to check whether the system 100 is ready to receive the input dataset and the first plurality of neural network weights. In response to an affirmative response from the system 100, the computer may send a second signal including the input dataset and the first plurality of neural network weights.
[0387] In step 220, the input dataset and the first plurality of neural network weights are stored. The controller 110 can store the input dataset and the first plurality of neural network weights in storage unit 120. Storing the input dataset and the first plurality of neural network weights in storage unit 120 allows for flexibility in the operation of the ANN computing system 100, which can, for example, improve the overall performance of the system. For example, by retrieving the desired portion of the input dataset from storage unit 120, the input dataset can be divided into numeric input vectors of a defined size and format. The different portions of the input dataset can be processed in various orders or shuffled to allow for the execution of various types of ANN computations. For example, shuffling can allow matrix multiplication to be performed using block matrix multiplication techniques when the input and output matrix sizes are different. As another example, storing the input dataset and the first plurality of neural network weights in storage unit 120 allows the ANN computing system 100 to queue multiple ANN computation requests, which allows the ANN computing system 100 to maintain operation at full speed without inactive periods.
[0388] In some embodiments, the input dataset may be stored in a first storage sub-unit, and the first plurality of neural network weights may be stored in a second storage sub-unit.
[0389] In step 230, a first plurality of modulator control signals are generated based on the first digital input vector, and a first plurality of weight control signals are generated based on the first plurality of neural network weights. Controller 110 may send the first DAC control signal to DAC unit 130 to generate the first plurality of modulator control signals. DAC unit 130 generates the first plurality of modulator control signals based on the first DAC control signals, and modulator array 144 generates an optical input vector representing the first digital input vector.
[0390] The first DAC control signal may include multiple digital values to be converted by the DAC unit 130 into a first plurality of modulator control signals. These multiple digital values typically correspond to a first digital input vector and can be associated through various mathematical relationships or lookup tables. For example, the multiple digital values may be linearly proportional to the values of the elements of the first digital input vector. As another example, the multiple digital values can be associated with the elements of the first digital input vector through a lookup table configured to maintain a linear relationship between the digital input vector and the optical input vector generated by the modulator array 144.
[0391] The controller 110 can send the second DAC control signal to the DAC unit 130 to generate a first plurality of weight control signals. The DAC unit 130 generates the first plurality of weight control signals based on the second DAC control signal, and reconfigures the OMM unit 150 according to the first plurality of weight control signals to realize a matrix corresponding to the first plurality of neural network weights.
[0392] The second DAC control signal may include multiple digital values that will be converted by DAC unit 130 into a first plurality of weight control signals. These multiple digital values typically correspond to the first plurality of neural network weights and can be associated through various mathematical relationships or lookup tables. For example, the multiple digital values may be linearly proportional to the first plurality of neural network weights. As another example, the multiple digital values may be calculated by performing various mathematical operations on the first plurality of neural network weights to generate weight control signals, which can configure OMM unit 150 to perform matrix multiplication corresponding to the first plurality of neural network weights.
[0393] In some embodiments, the first plurality of neural network weights representing matrix M can be decomposed into M = USV using the singular value decomposition (SVD) method. Where U is an M×M unitary matrix, S is an M×N diagonal matrix with non-negative real numbers on its diagonal, and V It is the complex conjugate of an N×N unitary matrix V. In this case, the first plurality of weight control signals may include a first plurality of OMM unit control signals corresponding to matrix V, and a second plurality of OMM unit control signals corresponding to matrix S. Further, OMM unit 150 may be configured to have a first OMM sub-unit configured to implement matrix V, a second OMM sub-unit configured to implement matrix S, and a third OMM sub-unit configured to implement matrix U, such that OMM unit 150 as a whole implements matrix M. The SVD method is further described in U.S. Patent Publication No. US2017 / 0351293A1 entitled “APPARATUS AND METHODS FOR OPTICAL NEURAL NETWORK”, which is incorporated herein by reference in its entirety.
[0394] In step 240, a first plurality of digital optical outputs corresponding to the optical matrix multiplication unit's optical output vector are obtained. The optical input vector generated by the modulator array 144 is processed by the OMM unit 150 and converted into an optical output vector. The optical output vector is detected by the detection unit 146 and converted into an electrical signal, which can be converted into a digital value by the ADC unit 160. The controller 110 can, for example, send a conversion request to the ADC unit 160 to begin converting the voltage output by the detection unit 146 into a digital optical output. Once the conversion is complete, the ADC unit 160 can send the conversion result to the controller 110. Alternatively, the controller 110 can obtain the conversion result from the ADC unit 160. The controller 110 can form a digital output vector from the digital optical output, which corresponds to the result of matrix multiplication of the input digital vector. For example, the digital optical outputs can be organized or concatenated to have a vector format.
[0395] In some embodiments, the ADC unit 160 can be configured or controlled to perform ADC conversion based on a DAC control signal being sent from the controller 110 to the DAC unit 130. For example, the ADC conversion can be configured to begin at a preset time after the DAC unit 130 generates the modulation control signal. This control of the ADC conversion simplifies the operation of the controller 110 and reduces the number of necessary control operations.
[0396] In step 250, a nonlinear transformation is performed on the first digital output vector to produce a first transformed digital output vector. A node or artificial neuron in an ANN operates by first performing a weighted summation of signals received from nodes in previous layers, and then performing a nonlinear transformation (“activation”) of the weighted sum to produce an output. Various types of ANNs can implement various types of differentiable nonlinear transformations. Examples of nonlinear transformation functions include the rectified linear unit (RELU) function, the sigmoid function, the hyperbolic tangent function, the X^2 function, and the |X| function. This nonlinear transformation is performed on the first digital output by controller 110 to produce the first transformed digital output vector. In some embodiments, the nonlinear transformation may be performed by a dedicated digital integrated circuit within controller 110. For example, controller 110 may include one or more modules or circuit blocks particularly suited to accelerate the computation of one or more types of nonlinear transformations.
[0397] In step 260, the first transformed digital output vector is stored. The controller 110 may store the first transformed digital output vector in storage unit 120. When the input dataset is divided into multiple digital input vectors, the first transformed digital output vector corresponds to the ANN computation result of a portion of the input dataset, for example, the first digital input vector. In this way, storing the first transformed digital output vector allows the ANN computation system 100 to perform and store additional computations on other digital input vectors of the input dataset, for later aggregation into a single ANN output.
[0398] In step 270, the output is generated based on the first transformed digital output vector, forming the artificial neural network output. The controller 110 generates the ANN output, which is the result of processing the input dataset by an ANN defined by a first plurality of neural network weights. When the input dataset is divided into multiple digital input vectors, the generated ANN output is an aggregated output including the first transformed digital output, but may further include additional transformed digital outputs corresponding to other parts of the input dataset. Once the ANN output is generated, it is sent to the computer (e.g., computer 102) that initiated the ANN computation request.
[0399] Various performance metrics can be defined for the ANN computing system 100 implementing method 200. Defining performance metrics allows for comparison of the performance of the ANN computing system 100 implementing the optical processor 140 with the performance of other systems used to replace the implementation of the electronic matrix multiplication unit (EMU) in ANN computing. In one aspect, the rate at which ANN computing can be performed can be indicated in part by a first cycle period, defined as the time elapsed between step 220, which stores the input dataset and the first plurality of neural network weights in the storage unit, and step 260, which stores the first transformed digital output vector in the storage unit. Thus, the first cycle period includes the time spent converting the electrical signal into an optical signal (e.g., step 230), performing matrix multiplication in the optical domain, and converting the result back to the electrical domain (e.g., step 240). Steps 220 and 260 both involve storing data in the storage unit 120, a step shared between the ANN computing system 100 and conventional ANN computing systems without the optical processor 140. In this way, the first cycle of measuring memory-to-memory transaction time can allow for a real or fair comparison of ANN computational throughput between ANN computing system 100 and ANN computing systems without optical processor 140 (e.g., systems implementing electrical matrix multiplication units).
[0400] Because the modulator array 144 can generate optical input vectors at a rate (e.g., 25 GHz) and the OMM unit 150 can process at a rate (e.g., >100 GHz), the first cycle time of the ANN computing system 100 for performing a single ANN computation on a single digital input vector can be close to the reciprocal of the speed of the modulator array 144 (e.g., 40 ps). Taking into account the delays associated with signal generation by the DAC unit 130 and ADC conversion by the ADC unit 160, the first cycle time can be, for example, less than or equal to 100 ps, less than or equal to 200 ps, less than or equal to 500 ps, less than or equal to 1 ns, less than or equal to 2 ns, less than or equal to 5 ns, or less than or equal to 10 ns.
[0401] For comparison, the multiplication time of an M×1 vector and an M×M matrix in an electrical matrix multiplication unit is typically proportional to M^2-1 processor clock cycles. For M=32, such a multiplication would take approximately 1024 cycles, resulting in a runtime of over 300 ns at a 3 GHz clock speed, which is several orders of magnitude slower than the first cycle of an ANN computing system 100.
[0402] In some embodiments, method 200 further includes the step of generating a second plurality of modulator control signals based on the first transformed digital output vector. In some types of ANN computation, a single digital input vector can be repeatedly propagated through or processed by the same ANN. An ANN that implements multi-pass processing can be called a recurrent neural network (RNN). An RNN is a neural network in which the output of the network is recirculated back to the input of the neural network during the (k)th pass through the neural network and used as input during the (k+1)th pass. RNNs can have various applications in pattern recognition tasks, such as speech or handwriting recognition. Once the second plurality of modulator control signals have been generated, method 200 can proceed from step 240 to step 260 to complete the second pass of the first digital input vector through the ANN. Generally, depending on the characteristics of the RNN received in the ANN computation request, the transformed digital output can be repeatedly recirculated back to the digital input vector for a predetermined number of cycles.
[0403] In some embodiments, method 200 further includes the step of generating a second plurality of weight control signals based on a second plurality of neural network weights. In some cases, the artificial neural network computation request further includes a second plurality of neural network weights. Generally, in addition to the input layer and the output layer, an ANN also has one or more hidden layers. For an ANN with two hidden layers, the second plurality of neural network weights may correspond to the connectivity between the first layer and the second layer of the ANN. To process the first digital input vector through the two hidden layers of the ANN, the first digital input vector may first be processed according to method 200 up to step 260, wherein the result of processing the first digital input vector through the first hidden layer of the ANN in step 260 is stored in storage unit 120. Then, controller 110 reconfigures OMM unit 150 to perform matrix multiplication corresponding to the second plurality of neural network weights associated with the second hidden layer of the ANN. Once OMM unit 150 is reconfigured, method 200 may generate a plurality of modulator control signals based on the first converted digital output vector, which generate an updated optical input vector corresponding to the output of the first hidden layer. The updated optical input vector is then processed by the reconfigured OMM unit 150, which corresponds to the second hidden layer of the ANN. Typically, the steps described can be repeated until the digital input vector has been processed through all the hidden layers of the ANN.
[0404] As described above, in some embodiments of the OMM unit 150, the reconfiguration rate of the OMM unit 150 can be significantly slower than the modulation rate of the modulator array 144. In this case, the throughput of the ANN computing system 100 may be adversely affected by the amount of time spent reconfiguring the OMM unit 150 during periods when ANN computations cannot be performed. To mitigate the impact of the relatively slow reconfiguration time of the OMM unit 150, batch processing techniques can be utilized, in which two or more digital input vectors are propagated through the OMM unit 150 without configuration changes, to amortize the reconfiguration time over a larger number of digital input vectors.
[0405] Figure 2B Description shown Figure 2A Schematic diagram 290 illustrates aspects of method 200. For an ANN with two hidden layers, instead of processing the first digit input vector through the first hidden layer, reconfiguring the OMM unit 150 for the second hidden layer, processing the first digit input vector through the reconfigured OMM unit 150, and repeating the same operation for the remaining digit input vectors, all digit input vectors of the input dataset can first be processed by the OMM unit 150 configured for the first hidden layer (configuration #1), as shown in the upper part of schematic diagram 290. Once all digit input vectors have been processed by the OMM unit 150 with configuration #1, the OMM unit 150 is reconfigured to configuration #2, which corresponds to the second hidden layer of the ANN. This reconfiguration can be significantly slower than the rate at which the OMM unit 150 can process the input vectors. Once the OMM unit 150 is reconfigured for the second hidden layer, the output vectors from the previous hidden layer can be batch processed by the OMM unit 150. For large input datasets with tens or hundreds of thousands of digit input vectors, the impact of reconfiguration time can be reduced by roughly the same factors, which can significantly reduce the portion of time that the ANN computing system 100 spends on reconfiguration.
[0406] To achieve batch processing, in some embodiments, method 200 further includes the steps of generating a second plurality of modulator control signals based on a second digital input vector via a DAC unit; obtaining a second plurality of digital optical outputs corresponding to the optical output vectors of optical matrix multiplication units from an ADC unit, the second plurality of digital optical outputs forming a second digital output vector; performing a nonlinear transformation on the second digital output vector to generate a second converted digital output vector; and storing the second converted digital output vector in a storage unit. For example, generating the second plurality of modulator control signals may be performed after step 260. Furthermore, in this case, the ANN output of step 270 is now based on the first and second converted digital output vectors. The acquisition, execution, and storage steps are similar to steps 240 to 260.
[0407] Batch processing is one of several techniques used to increase the throughput of the ANN computing system 100. Another technique for increasing the throughput of the ANN computing system 100 is to process multiple digital input vectors in parallel using wavelength division multiplexing (WDM). WDM is a technique that simultaneously propagates multiple optical signals of different wavelengths through a common propagation channel (e.g., the waveguide of the OMM unit 150). Unlike electrical signals, optical signals of different wavelengths can propagate through a common channel without affecting other optical signals of different wavelengths on the same channel. Furthermore, known structures such as optical multiplexers and demultiplexers can be used to add (multiplex) or discard (demultiplex) optical signals from the common propagation channel.
[0408] Within the context of the ANN computing system 100, multiple optical input vectors of different wavelengths can be independently generated, simultaneously propagated through the OMM unit 150, and independently detected to enhance the throughput of the ANN computing system 100. (Refer to...) Figure 1F This diagram illustrates an example of a wavelength division multiplexing (WDM) artificial neural network (ANN) computing system 104. Unless otherwise described, the WDM ANN computing system 104 is similar to the ANN computing system 100. To implement WDM technology, in some embodiments of the ANN computing system 104, the laser unit 142 is configured to generate multiple wavelengths, for example... , as well as Multiple wavelengths can preferably be separated by a sufficiently large wavelength spacing to allow for easy multiplexing and demultiplexing onto a common propagation channel. For example, wavelength spacing greater than 0.5 nm, 1.0 nm, 2.0 nm, 3.0 nm, or 5.0 nm allows for simple multiplexing and demultiplexing. On the other hand, the range between the shortest and longest wavelengths of the multiple wavelengths (“WDM bandwidth”) can preferably be small enough that the characteristics or performance of the OMM unit 150 remain substantially the same across the multiple wavelengths. Optical components are typically dispersive, meaning their optical characteristics vary with wavelength. For example, the power separation ratio of an MZI can vary with wavelength. However, by designing the OMM unit 150 to have a sufficiently large operating wavelength window, and by limiting the wavelengths within the operating wavelength window, the optical output vector output by the OMM unit 150 at each wavelength can be a sufficiently accurate result of matrix multiplication implemented by the OMM unit 150. The operating wavelength window can be, for example, 1 nm, 2 nm, 3 nm, 4 nm, 5 nm, 10 nm or 20 nm.
[0409] Figure 39A A diagram showing an example of a Mach-Zehnder modulator 3900 that can be used to modulate the amplitude of an optical signal is presented. The Mach-Zehnder modulator 3900 includes two 1×2-port multimode interference couplers (MMI_1x2) 3902a and 3902b, two balanced arms 3904a and 3904b, and a phase shifter 3906 in one arm (or one phase shifter in each arm). When a voltage is applied to the phase shifter in one arm via signal line 3908, a phase difference that will be converted into amplitude modulation will exist between the two arms 3904a and 3904b. The 1x2-port multimode interference couplers 3902a and 3902b and the phase shifter 3906 are configured as broadband photonic components, and the optical path lengths of the two arms 3904a and 3904b are configured to be equal. This allows the Mach-Zehnder modulator 3900 to operate over a wide wavelength range.
[0410] Figure 39B Graph 3910 shows the effects of using wavelengths of 1530nm, 1550nm, and 1570nm at... Figure 39A The intensity-voltage curves of the Mach-Zehnder modulator 3900 in the configuration shown are illustrated. Graph 3910 shows that the Mach-Zehnder modulator 3900 exhibits similar intensity-voltage characteristics for different wavelengths in the range of 1530 nm to 1570 nm.
[0411] Return to reference Figure 1FThe modulator array 144 of the WDM ANN computing system 104 includes banks of optical modulators configured to generate multiple optical input vectors, each of which corresponds to one of multiple wavelengths and generates a corresponding optical input vector with that wavelength. For example, for a length of 32 and 3 wavelengths (e.g.: , as well as The system of optical input vectors includes a modulator array 144 that can have three groups of 32 modulators each. Furthermore, the modulator array 144 includes an optical multiplexer configured to combine multiple optical input vectors into a combined optical input vector comprising multiple wavelengths. For example, the optical multiplexer can combine the outputs of three modulator groups of three different wavelengths into a single propagation channel (e.g., a waveguide) for each element of the optical input vector. Thus, returning to the example above, the combined optical input vector would have 32 optical signals, each comprising three wavelengths.
[0412] Furthermore, the detection unit 146 of the WDM ANN computing system 104 is further configured to multiplex multiple wavelengths and generate multiple multiplexed output voltages. For example, the detection unit 146 may include a multiplexer configured to multiplex three wavelengths in each of 32 signals comprising a multi-wavelength optical output vector, and route the three single-wavelength optical output vectors to three sets of photodetectors coupled to three sets of transimpedance amplifiers.
[0413] Furthermore, the ADC unit 160 of the WDM ANN computing system 104 includes an ADC group configured to convert the multiple multiplexed output voltages of the detection unit 146. Each ADC group corresponds to one of multiple wavelengths and produces a corresponding digitally multiplexed optical output. For example, the ADC group may be coupled to a transimpedance amplifier group of the detection unit 146.
[0414] The controller 110 can implement a method similar to method 200, but extended to support multi-wavelength operation. For example, the method may include the steps of obtaining a plurality of digitally multiplexed optical outputs from the ADC unit 160, the plurality of digitally multiplexed optical outputs forming a plurality of first digital output vectors, each of the plurality of first digital output vectors corresponding to one of a plurality of wavelengths; performing a nonlinear transformation on each of the plurality of first digital output vectors to generate a plurality of transformed first digital output vectors; and storing the plurality of transformed first digital output vectors in a storage unit.
[0415] In some cases, ANNs can be specifically designed and digital input vectors can be specifically formed, enabling the detection of multi-wavelength optical output vectors without multiplexing. In this case, detection unit 146 can be a wavelength-insensitive detection unit that does not multiplex the multiple wavelengths of the multi-wavelength optical output vector. In this way, each photodetector of detection unit 146 effectively adds multiple wavelengths of the optical signal to a single photocurrent, and each voltage output by detection unit 146 corresponds to the element-by-element sum of the matrix multiplication results of multiple digital input vectors.
[0416] Up to this point, the weighted summation nonlinear transformation performed as part of ANN computation has been executed in the digital domain by controller 110. In some cases, the nonlinear transformation may be computationally intensive or power-intensive, significantly increasing the complexity of controller 110, or limiting the performance of the ANN computation system 100 in terms of flux or power efficiency. Therefore, in some embodiments of the ANN computation system, the nonlinear transformation can be performed in the analog domain using analog electronics.
[0417] Figure 3A A schematic diagram of an example ANN computing system 300 is shown. The ANN computing system 300 is similar to the ANN computing system 100, except that an analog nonlinear unit 310 is added. The analog nonlinear unit 310 is disposed between the detection unit 146 and the ADC unit 160. The analog nonlinear unit 310 is configured to receive the output voltage from the detection unit 146, apply a nonlinear transfer function, and output the converted output voltage to the ADC unit 160.
[0418] When the ADC unit 160 receives a voltage that has already been nonlinearly converted by the analog nonlinear unit 310, the controller 110 can obtain the corresponding converted digital output voltage from the ADC unit 160. Because the digital output voltage obtained from the ADC unit 160 has already been nonlinearly converted (“activated”), the nonlinear conversion step of the controller 110 can be omitted, thereby reducing the computational burden on the controller 110. Then, the first converted voltage obtained directly from the ADC unit 160 can be stored as a first converted digital output vector in the storage unit 120.
[0419] The analog nonlinear unit 310 can be implemented in various ways. For example, a high-gain amplifier in a feedback configuration, a comparator with an adjustable reference voltage, a nonlinear IV characteristic of a diode, a diode breakdown behavior, a nonlinear CV characteristic of a variable capacitor, or a nonlinear IV characteristic of a variable resistor can be used to implement the analog nonlinear unit 310.
[0420] Using analog nonlinear unit 310 can improve the performance of ANN computing system 300, such as flux or power efficiency, by reducing the number of steps performed in the digital domain. Moving the nonlinear transformation step out of the digital domain allows for additional flexibility and improvements in the operation of the ANN computing system. For example, in a recurrent neural network, the output of OMM unit 150 is activated and then loops back to the input of OMM unit 150. The activation step is performed by controller 110 in ANN computing system 100, which requires digitizing the output voltage of detection unit 146 each time it passes through OMM unit 150. However, because the activation step is now performed before the digitization of ADC unit 160, the number of ADC conversions required in performing recurrent neural network computations can be reduced.
[0421] In some embodiments, the analog nonlinear unit 310 may be integrated into the ADC unit 160 as a nonlinear ADC unit. For example, the nonlinear ADC unit may be a linear ADC unit with a nonlinear lookup table that maps the linear digital output of the linear ADC unit to the desired nonlinear converted digital output.
[0422] Figure 3B A schematic diagram of an example of an ANN computing system 302 is shown. The ANN computing system 302 is similar to... Figure 3A The system 300 differs in that it further includes an analog storage unit 320. The analog storage unit 320 is coupled to a DAC unit 130 (e.g., via a first DAC subunit 132), a modulator array 144, and an analog nonlinear unit 310. The analog storage unit 320 includes a multiplexer having a first input coupled to the DAC unit 130 and a second input coupled to the analog nonlinear unit 310. This allows the analog storage unit 320 to receive signals from either the DAC unit 130 or the analog nonlinear unit 310. The analog storage unit 320 is configured to store analog voltages and output the stored analog voltages.
[0423] The analog storage cell 320 can be implemented in various ways. For example, a capacitor array can be used as an analog voltage storage component. The capacitors of the analog storage cell 320 can be charged to the input voltage via a charging circuit. The storage of the input voltage can be controlled based on a control signal received from the controller 110. The capacitors can be electrically isolated from the surrounding environment to reduce charge leakage that could lead to unwanted capacitor discharge. Alternatively (or alternatively), a feedback amplifier can be used to maintain the voltage stored on the capacitors. The stored voltage of the capacitors can be read out via a buffer amplifier, which allows the charge stored by the capacitors to be retained while the stored voltage is output. These aspects of the analog storage cell 320 can be analogous to the operation of a sample and hold circuit. The buffer amplifier can implement the function of a modulator driver for driving the modulator array 144.
[0424] The operation of the ANN computing system 302 will now be described. A first plurality of modulator control signals output from DAC unit 130 (e.g., from the first DAC subunit 132) are first input to modulator array 144 via analog storage unit 320. In this step, analog storage unit 320 can simply pass or buffer the first plurality of modulator control signals. Modulator array 144 generates an optical input vector based on the first plurality of modulator control signals, which is propagated through OMM unit 150 and detected by detection unit 146. The output voltage of detection unit 146 is nonlinearly converted by analog nonlinear unit 310. Instead of being digitized by ADC unit 160, the output voltage of detection unit 146 is stored in analog storage unit 320, which is then output to modulator array 144 to be converted into the next optical input vector to be propagated through OMM unit 150. Under the control of controller 110, this recurrent processing can be performed in a preset time period or a preset number of cycles. Once the recursive processing for a given digital input vector is complete, the converted output voltage of the analog nonlinear unit 310 is converted by the ADC unit 160.
[0425] The use of analog memory unit 320 can significantly reduce the number of ADC conversions during recurrent neural network computation, for example, reducing it to a single ADC conversion for each given digital input vector RNN computation. Each ADC conversion takes time and consumes energy. In this way, the throughput of RNN computation in ANN computation system 302 can be higher than that of RNN computation in ANN computation system 100.
[0426] The execution of recurrent neural network computations can be controlled by controlling the analog storage unit 320. For example, the controller can control the analog storage unit 320 to store voltage at specific times and output the stored voltage at different times. In this way, the signal loop from the analog storage unit 320 to the modulator array 144, through the analog nonlinear unit 310, and back to the analog storage unit 320 can be controlled by the controller 110 to control the storage and readout of the analog storage unit 320.
[0427] In some embodiments, the controller 110 of the ANN computing system 302 may perform the following steps: based on generating a first plurality of modulator control signals and a first plurality of weight control signals, storing a plurality of converted output voltages of an analog nonlinear unit in an analog storage unit; outputting the stored converted output voltages in the analog storage unit; obtaining a second plurality of converted digital output voltages from an ADC unit, the second plurality of converted digital output voltages forming a second converted digital output vector; and storing the second converted digital output vector in a storage unit.
[0428] The input datasets processed by ANN computing systems typically include data with a resolution greater than 1 bit. For example, a typical pixel in a grayscale digital image may have an 8-bit resolution, or 256 different levels. One way to represent and process this data in the optical domain is to encode the 256 different intensity levels of the pixels as 256 different power levels of the optical signal input to the OMM unit 150. The optical signal is inherently an analog signal and is therefore susceptible to noise and detection errors. (Return to reference) Figure 1A In order to maintain 8-bit resolution of the digital input vector throughout the ANN computing system 100 and to produce a true 8-bit digital optical output at the output of the ADC unit 160, each part of the signal chain can preferably be designed to reproduce and maintain 8-bit resolution.
[0429] For example, DAC unit 130 may preferably be designed to support the conversion of an 8-bit digital input vector into a modulator control signal with at least 8-bit resolution, such that modulator array 144 can generate an 8-bit optical input vector that faithfully represents the digital input vector. Typically, the modulator control signal may need to have an additional resolution beyond the 8 bits of the digital input vector to compensate for the nonlinear response of modulator array 144. Furthermore, the internal configuration of OMM unit 150 may preferably be sufficiently stable to ensure that the value of the optical output vector is not corrupted by any fluctuations in the configuration of OMM unit 150. For example, the temperature of OMM unit 150 may need to be stabilized within 5 degrees, 2 degrees, 1 degree, or 0.1 degrees. Additionally, detection unit 146 may preferably have sufficiently low noise to not corrupt the 8-bit resolution of the optical output vector, and ADC unit 160 may preferably be designed to support the digitization of analog voltages with at least 8-bit resolution.
[0430] The power consumption and design complexity of various electronic components typically increase with bit resolution, operating speed, and bandwidth. For example, as a first-order approximation, the power consumption of ADC unit 160 can scale linearly with the sampling rate, and the scaling factor is 2^N, where N is the bit resolution of the conversion result. Furthermore, design considerations for DAC unit 130 and ADC unit 160 often result in a trade-off between sampling rate and bit resolution. Thus, in some cases, it may be desirable for the ANN computation system to operate internally at a lower bit resolution than the input dataset while maintaining the resolution of the ANN computation output.
[0431] Reference Figure 4A This diagram shows an example of an artificial neural network (ANN) computing system 400 with 1-bit internal resolution. The ANN computing system 400 is similar to the ANN computing system 100, except that the DAC unit 130 is now replaced by a driver unit 430 and the ADC unit 160 is now replaced by a comparator unit 460.
[0432] Driver unit 430 is configured to generate a 1-bit modulator control signal and a multi-bit weighted control signal. For example, the drive circuit of driver unit 430 can directly receive binary digital output from controller 110 and modulate the binary signal to a form suitable for driving the two-level voltage or current output of modulator array 144.
[0433] Comparator unit 460 is configured to convert the output voltage of detection unit 146 into a 1-bit digital optical output. For example, the comparison circuit of comparator unit 460 can receive a voltage from detection unit 146, compare the voltage with a preset threshold voltage, and output a digital 0 or 1 when the received voltage is less than or greater than the preset threshold voltage.
[0434] Reference Figure 4B This shows the mathematical representation of the operation of the ANN computing system 400. Now refer to... Figure 4B Describe the operation of the ANN computation system 400. For a given ANN computation to be performed by the ANN computation system 400, there exists a corresponding numerical input vector V and a neural network weight matrix U. In this embodiment, the input vector V is a vector of length 4 with elements V0 to V3, and the matrix U is a matrix with weights U... 00 To U 33 A 4×4 matrix. Each element of vector V has a resolution of 4 bits. Each 4-bit vector element has bits 0 through 3 corresponding to positions 2^0 through 2^3. Thus, through 2^0... bit0+2^1 bit1+2^2 bit2+2^3 The sum of bits 3 is used to calculate the decimal (base 10) value of the 4-bit vector element. Therefore, as shown in the figure, the input vector V can be similarly decomposed into V by the controller 110. bit0 To V bit3 .
[0435] Next, a specific ANN calculation can be performed by performing a series of matrix multiplications on a 1-bit vector and then summing the results of each matrix multiplication. For example, by generating a sequence of four 1-bit modulator control signals corresponding to four 1-bit input vectors through the driver unit 430, the decomposed input vector V can be... bit0 To V bit3 Each of the input signals is multiplied by matrix U. This, in turn, produces a sequence of four 1-bit optical input vectors, which are propagated through OMM unit 150, configured to perform matrix multiplication of matrix U via driver unit 430. Next, controller 110 can obtain a sequence of four digital 1-bit optical outputs corresponding to the sequences of the four 1-bit modulator control signals from comparator unit 460.
[0436] When a 4-bit vector is decomposed into four 1-bit vectors, each vector should be processed by the ANN computing system 400 at a speed four times faster than that of other ANN computing systems (such as system 100) that can process a single 4-bit vector, while maintaining the same effective ANN computation throughput. This increased internal processing speed can be viewed as time-division multiplexing the four 1-bit vectors into a single timeslot for processing the 4-bit vector. The required increase in processing speed can be achieved, at least in part, by the increased operating speed of the driver unit 430 and comparator unit 460 relative to the DAC unit 130 and ADC unit 160, since the reduction in resolution of signal conversion processing typically results in an increase in the achievable signal conversion rate.
[0437] Although the signal conversion rate increases fourfold in 1-bit operations, the resulting power consumption can be significantly reduced compared to 4-bit operations. As mentioned above, the power consumption of signal conversion processing typically scales exponentially with bit resolution and linearly with conversion rate. Thus, a 16-fold reduction in power per conversion could be due to a 4-fold reduction in bit resolution, followed by a 4-fold increase in power due to the increased conversion rate. In summary, ANN computing system 400 can achieve a 4-fold reduction in operating power compared to, for example, ANN computing system 100, while maintaining the same effective ANN computational throughput.
[0438] Next, controller 110 can construct a 4-bit digital output vector from the four 1-bit digital optical outputs by multiplying each digital 1-bit optical output by a corresponding weight from 2^0 to 2^3. Once the 4-bit digital output vector is constructed, ANN calculations can be performed on the constructed 4-bit digital output vector to generate a transformed 4-bit digital output vector; and the transformed 4-bit digital output vector can be stored in storage unit 120.
[0439] Alternatively (or additionally), in some embodiments, each of the four 1-bit digital optical outputs can be nonlinearly converted. For example, a step-function nonlinear function can be used for the nonlinear conversion. A converted 4-bit digital output vector can then be constructed from the nonlinearly converted 1-bit digital optical outputs.
[0440] Although a standalone ANN computing system 400 has been shown and described, generally speaking, Figure 1AThe ANN computing system 100 can be designed to perform functions similar to those of the ANN computing system 400. For example, the DAC unit 130 may include a 1-bit DAC subunit configured to generate a 1-bit modulator control signal, and the ADC unit 160 may be designed to have a 1-bit resolution. Such a 1-bit ADC can be similar to or effectively equivalent to a comparator.
[0441] Furthermore, while the operation of an ANN computation system with an internal resolution of 1 bit has been described, in general, the internal resolution of an ANN computation system can be reduced to an intermediate level below the N-bit resolution of the input dataset. For example, the internal resolution can be reduced to 2^Y bits, where Y is an integer greater than or equal to 0.
[0442] Embodiments of the subject matter and functional operation described in this disclosure can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this disclosure and their structural equivalents, or one or more combinations thereof. Embodiments of the subject matter described in this disclosure can be implemented using one or more modules of computer program instructions encoded on a computer-readable medium to be executed by or control the operation of a data processing device. The computer-readable medium can be an manufactured product (e.g., a hard disk drive in a computer system or an optical disc sold through retail channels) or an embedded system. The computer-readable medium can be separately acquired and subsequently encoded with one or more modules of computer program instructions, for example, by transmitting one or more modules of computer program instructions via a wired or wireless network. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, or a combination of one or more of these.
[0443] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suited to a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that stores other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code segments). A computer program can be deployed to execute on a single computer or on multiple computers located at a site or distributed across multiple sites interconnected by a communication network.
[0444] The processing and logic flows described in this disclosure can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and producing outputs. The processing and logic flows can also be executed by special-purpose logic circuitry, and the device can be implemented as special-purpose logic circuitry, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC).
[0445] While this disclosure contains numerous implementation details, these should not be construed as limiting the scope of this disclosure or the claims, but rather as descriptions of specific features of particular embodiments of this disclosure. Certain features described in this disclosure within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although the features described above act on certain combinations and are even initially claimed in this way, in some cases one or more features may be removed from the claimed combination, and the claimed combination may refer to a sub-combination or a variation of the sub-combination.
[0446] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or to perform all shown operations to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system components in the described embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated into a single software product or packaged into multiple software products.
[0447] Therefore, specific embodiments of this disclosure have been described. Other embodiments are within the scope of the following claims. Furthermore, the actions recited in the claims can be performed in a different order and still achieve the desired result. For example, Figure 1A The optical matrix multiplication unit 150 includes an optical interference unit 154 comprising a plurality of interconnected Mach-Zehnder interferometers. In some embodiments, the optical interference unit can be implemented using a one-dimensional, two-dimensional, or three-dimensional passive diffractive optical element that consumes virtually no power. Compared to an optical interference unit including Mach-Zehnder interferometers, an optical interference unit using a passive diffractive optical element can have a smaller size, or can handle a larger number of inputs / outputs for the same chip size, while keeping the number of inputs / outputs constant. Passive diffractive optical elements can be manufactured at a lower cost compared to Mach-Zehnder interferometers.
[0448] Reference Figure 5 In some embodiments, the artificial neural network computing system 500 includes a controller 110, a storage unit 120, a DAC unit 506, an optical processor 504, and an ADC unit 160. The storage unit 120 and the ADC unit 160 are similar to... Figure 1A The corresponding component in system 100. Optical processor 504 is configured to perform matrix calculations using optical components. In system 500, the weights of the two-dimensional optical matrix multiplication unit 502 are fixed. DAC unit 506 is similar to... Figure 1A The first DAC subunit 132 of the system 100.
[0449] In an example operation of the ANN computing system 500, the computer 102 may issue an artificial neural network computing request to the ANN computing system 500. The ANN computing request may include an input dataset to be processed by the provided ANN. The controller 110 receives the ANN computing request and stores the input dataset in the storage unit 120.
[0450] In some embodiments, a hybrid approach is used, wherein a portion of the optical matrix multiplication unit 150 includes a Mach-Zehnder interferometer, and another portion of the optical matrix multiplication unit 150 includes a passive diffraction component.
[0451] The internal operation of the ANN computing system 500 will now be described. The optical processor 504 includes a laser unit 142, a modulator array 144, a detection unit 146, and a two-dimensional optical matrix multiplication (OMM) unit 502. The laser unit 142, modulator array 144, and detection unit 146 are similar to... Figure 1A The corresponding component of system 100. In this example, the two-dimensional OMM unit 502 includes two-dimensional diffractive optical components and can be implemented as a passive integrated silicon photonic chip. The two-dimensional optical matrix multiplication unit 502 can be configured to implement a diffractive neural network and can perform matrix multiplication with almost zero power consumption.
[0452] The optical processor 504 operates by encoding a digital input vector of length N onto an optical input vector of length N, and propagating the optical input vector through a two-dimensional OMM unit 502. The two-dimensional OMM unit 502 receives the optical input vector of length N and performs an N×N matrix multiplication on the received optical input vector in the optical domain. The N×N matrix multiplication performed by the two-dimensional OMM unit 502 is determined by the internal configuration of the two-dimensional OMM unit 502. The internal configuration of the two-dimensional OMM unit 502 includes the size, position, and geometry of the diffractive optical components, as well as any doping of impurities.
[0453] Two-dimensional OMM units 502 can be implemented in various ways. Figure 6 A schematic diagram of an example of a two-dimensional OMM unit 502 using a two-dimensional diffractive component array is shown. The two-dimensional OMM unit 502 may include an array of input waveguides 602 to receive an optical input vector, a two-dimensional optical interference unit 600 in optical communication with the array of input waveguides 602, and an array of output waveguides 604 in optical communication with the optical interference unit 600. The optical interference unit 600 includes a plurality of diffractive optical components and performs a conversion (e.g., linear conversion) of the optical input vector to a second optical signal array. The array of output waveguides 604 guides the second optical signal array output by the optical interference unit 600. At least one input waveguide in the array of input waveguides 602 is in optical communication with each output waveguide in the array of output waveguides 604 via the optical interference unit 600. For example, for an optical input vector of length N, the two-dimensional OMM unit 502 may include N input waveguides 602 and N output waveguides 604.
[0454] In some embodiments, the optical interference unit 600 includes a substrate with diffraction components arranged in a two-dimensional (e.g., 2D array) configuration. For example, multiple circular holes can be drilled or etched into the substrate. The size of these holes can be on the order of magnitude of the wavelength of the input light, such that the light is diffracted by the holes (or the structure defining the holes). For example, the size of the holes can range from 100 nm to 2 μm. The holes can have the same or different sizes. The holes can also have other cross-sectional shapes, such as triangular, square, rectangular, hexagonal, or irregular shapes. The substrate can be made of a material that is transparent or translucent to the input light, for example, having a transmittance of 1% to 99% relative to the input light. For example, the substrate can be made of silicon, silicon oxide, silicon nitride, quartz, crystals (e.g., lithium niobate (LiNbO3)), III-V group materials (e.g., gallium arsenide or indium phosphide), erbium-modified semiconductors, or polymers.
[0455] In some embodiments, the holographic method can be used to form two-dimensional diffractive optical components in a substrate. The substrate can be made of glass, crystal, or photorefractive material.
[0456] When designing a two-dimensional OMM unit 502, the size and position of the diffraction components are considered in two dimensions (e.g., the X and Y directions), without considering their relative positions in the third dimension (e.g., the Z direction). Each diffraction component can be a three-dimensional structure formed in the substrate, such as a hole, column, or strip with a certain depth.
[0457] exist Figure 6 In this context, diffractive optical components are represented by circles. Diffractive optical components can also have other shapes, such as triangles, squares, rectangles, or irregular shapes. They can also have various sizes. Diffractive optical components do not necessarily reside at grid points; their positions can be varied. Figure 6 The diagrams in this paper are for illustrative purposes only. Actual diffractive optical components may differ from those shown in the diagrams. Different arrangements of diffractive optical components can be used to achieve different matrix calculations, such as different matrix multiplication functions.
[0458] Optimization processes can be used to determine the configuration of the diffractive optical components. For example, the substrate can be divided into a pixel array, and each pixel can be filled with substrate material (without holes) or filled with air (with holes). The pixel configuration can be modified iteratively, and for each pixel configuration, simulation can be performed by passing light through the diffractive optical components and evaluating the output. After performing simulations of all possible pixel configurations, the configuration that provides the closest result to the desired matrix processing is selected as the diffractive optical component configuration of the two-dimensional OMM unit 502.
[0459] As another example, the diffraction components are initially configured as an array of holes. The location, size, and shape of the holes can differ slightly from their initial configuration. The parameters of each hole can be iteratively adjusted, and simulations can be performed to find the optimal configuration of the holes.
[0460] In some embodiments, machine learning processing is used to design diffractive optical components. An analytical function is used to determine how pixels influence input light to produce output light, and optimization processes (e.g., gradient descent methods) are used to determine the optimal configuration of the pixels.
[0461] In some embodiments, the two-dimensional OMM unit 502 can be implemented as a user-changeable component, and different two-dimensional OMM units 502 with different optical interference units 600 can be installed for different applications. For example, the system 500 can be configured as an optical character recognition system, and the optical interference unit 600 can be configured to implement a neural network for performing optical character recognition. For example, a first OMM unit may have a first optical interference unit including passive diffractive optics components configured to implement a first neural network for an optical character recognition engine for a first set of written languages and fonts. A second OMM unit may have a second optical interference unit including passive diffractive optics components configured to implement a second neural network for an optical character recognition engine for a second set of written languages and fonts. When a user wants to use the system 500 to apply optical character recognition to the first set of written languages and fonts, the user can insert the first OMM unit into the system. When a user wants to use the system 500 to apply optical character recognition to the second set of written languages and fonts, the user can remove the first OMM unit and insert the second OMM unit into the system.
[0462] For example, system 500 can be configured as a speech recognition system, and optical interference unit 600 can be configured to implement a neural network for performing speech recognition. For example, a first OMM unit may have a first optical interference unit including passive diffractive optics, configured to implement a first neural network for a speech recognition engine for a first spoken language. A second OMM unit may have a second optical interference unit including passive diffractive optics, configured to implement a second neural network for a speech recognition engine for a second spoken language, etc. When a user wants to use system 500 to recognize speech in a first spoken language, the user can insert the first OMM unit into the system. When a user wants to use system 500 to recognize speech in a second spoken language, the user can remove the first OMM unit and insert the second OMM unit into the system.
[0463] For example, system 500 may be part of a control unit for an autonomous vehicle, and optical interferometry unit 600 may be configured to implement a neural network for performing road condition recognition. For example, a first OMM unit may have a first optical interferometry unit including passive diffractive optics configured to implement a first neural network for recognizing road conditions (including road signs) in the United States. A second OMM unit may have a second optical interferometry unit including passive diffractive optics configured to implement a second neural network for recognizing road conditions (including road signs) in Canada. A third OMM unit may have a third optical interferometry unit including passive diffractive optics configured to implement a third neural network for recognizing road conditions (including road signs) in Mexico, and so on. When the autonomous vehicle is used in the United States, the first OMM unit is inserted into the system. When the autonomous vehicle crosses the border and enters Canada, the first OMM unit is removed, and the second OMM unit is inserted into the system. On the other hand, when the autonomous vehicle crosses the border and enters Mexico, the first OMM unit is replaced and the third OMM unit is inserted into the system.
[0464] For example, System 500 can be used for genetic sequencing. DNA sequences can be classified using convolutional neural networks, implemented using System 500 which includes passive diffraction optics components. For example, System 500 can implement neural networks for distinguishing tumor types, predicting tumor grades, and predicting patient survival from gene expression patterns. For example, System 500 can implement neural networks for identifying subsets of genes or traits that are most predictive of the analyzed characteristics. For example, System 500 can implement neural networks for predicting or inferring the expression levels of all genes from a profile of gene subsets. For example, System 500 can implement neural networks for epigenomic analysis, such as predicting transcription factor binding sites, enhancer regions, and chromatin accessibility from gene sequences. For example, System 500 can implement neural networks for capturing structures within gene sequences.
[0465] For example, system 500 can be configured as a medical diagnostic system, and the two-dimensional OMM unit 502 can be configured to implement a neural network for analyzing physiological parameters to perform disease screening. For example, system 500 can be configured as a bacterial detection system, and the two-dimensional OMM unit 502 can be configured to implement a multiplicative function for analyzing DNA sequences to detect certain bacterial strains.
[0466] In some embodiments, the two-dimensional OMM unit 502 includes a housing (e.g., a cartridge) protecting a substrate with diffractive optical components. The housing supports an input interface coupled to an input waveguide 602 and an output interface coupled to an output waveguide 604. The input interface is configured to receive output from a modulator array 144, and the output interface is configured to send the output of the two-dimensional OMM unit 502 to a detection unit 146. The two-dimensional OMM unit 502 can be designed as a module suitable for handling by ordinary consumers, allowing users to easily switch from one two-dimensional OMM unit 502 to another. Machine learning techniques improve over time. Users can upgrade the system 500 by replacing older two-dimensional OMM units 502 and inserting newer, upgraded versions.
[0467] Similar to how optical compact discs store digital information accessible to CD players, OMM units can store neural network configurations that can be used in optical processors. Just as optical compact discs are low-cost media for distributing digital information (including audio, video, and software programs) to consumers, OMM units can be low-cost media for distributing pre-configured neural network or matrix processing functions (e.g., multiplication, convolution, or any other linear operations) to consumers.
[0468] In some embodiments, system 500 is an optical computing platform configured to operate using OMM units from different companies. This allows different companies to develop different passive optical neural networks for various applications. Passive optical neural networks are sold to end users in standardized packages, which can be installed in the optical computing platform to allow system 500 to perform various intelligent functions.
[0469] In some embodiments, the system may have a retainer mechanism for supporting a plurality of two-dimensional OMM units 502, and may provide a mechanical processing mechanism for automatically replacing the two-dimensional OMM units 502. The system determines which two-dimensional OMM unit 502 is needed for the current application, and uses the mechanical processing mechanism to automatically retrieve the appropriate OMM unit from the retainer mechanism and insert it into the optical processor 504.
[0470] For optical chips of a specific size, more passive diffraction components can be mounted on the substrate compared to using an active interferometer (such as a Mach-Zehnder interferometer). For example, using a Mach-Zehnder interferometer... Figure 1B The optical interference unit 154 can be configured to handle 200×200 matrix multiplications, while the optical interference unit 600, which has the same overall size and uses passive diffraction components (each with a size of about 100 nm×100 nm), can be configured to handle 5000×5000 matrix multiplications.
[0471] Passive diffractive optical components consume almost no power, therefore the two-dimensional OMM unit 502 can be used in low-power devices, such as battery-operated devices. The two-dimensional OMM unit 502 is suitable for edge computing. For example, the two-dimensional OMM unit 502 can be used in smart sensors, where raw data from the sensor is processed using an optical processor employing the two-dimensional OMM unit 502. The smart sensor can be configured to send the processed data to a central computer server, thereby reducing the amount of raw data sent to the central computer server. By placing intelligent processing capabilities on the smart sensor, faults and anomalies can be detected earlier and handled more effectively. The two-dimensional OMM unit 502 is suitable for applications requiring the processing of large matrix multiplications. The two-dimensional OMM unit 502 is suitable for applications where neural networks have already been trained and weights have been determined and do not require modification.
[0472] The substrate forming the diffractive optical component can be planar or curved. Figure 6 In the example, input light enters the optical interference unit 600 from the left, and output light exits the optical interference unit 600 from the right (the terms "left," "right," "upper," and "lower" refer to the directions shown in the figures). In some embodiments, passive diffractive optics can be configured such that some output light exits the optical interference unit from the upper or lower portion, or any combination of the left, right, upper, and lower sides of the optical interference unit 600. The substrate for the optical interference unit 600 can have various shapes, such as square, rectangular, triangular, circular, or elliptical. The optical interference unit 600 may include reflective components or mirrors to redirect the direction of light propagation.
[0473] In some embodiments, the artificial neural network computing system 500 can be modified by adding an analog nonlinear unit 310 between the detection unit 146 and the ADC unit 160. The analog nonlinear unit 310 is configured to receive the output voltage from the detection unit 146, apply a nonlinear transfer function, and output the converted output voltage to the ADC unit 160. The controller 110 can obtain the converted digital output voltage corresponding to the converted output voltage from the ADC unit 160. Because the digital output voltage obtained from the ADC unit 160 has already been nonlinearly converted (“activated”), the nonlinear conversion step of the controller 110 can be omitted, thereby reducing the computational burden on the controller 110. Then, the first converted voltage obtained directly from the ADC unit 160 can be stored as a first converted digital output vector in the storage unit 120.
[0474] Optical interference units can be implemented using passive diffractive optical components arranged in three dimensions. (See reference...) Figure 7In some embodiments, the artificial neural network computing system 700 has an optical processor 702, which includes a three-dimensional OMM unit 708. The system 700 includes a storage unit 120 and an ADC unit 160, which are similar to... Figure 5 The corresponding component of system 500. The optical processor 702 is configured to perform matrix calculations using diffractive optical components arranged in three dimensions.
[0475] The optical processor 702 includes a laser unit 704 configured to output a two-dimensional beam array 714, and a two-dimensional modulator array 706 configured to modulate the two-dimensional beam array 714 to generate a modulated two-dimensional beam array 716. The optical processor 702 includes a three-dimensional optical matrix multiplication (OMM) unit 708 having diffractive optical components arranged in three dimensions, and configured to process the modulated two-dimensional beam array 716 and generate a two-dimensional output beam array 718. The optical processor 702 includes a detection unit 710 with a two-dimensional optical sensor array to detect the two-dimensional output beam array 718. An ADC unit 160 converts the output of the detection unit 710 into a digital signal.
[0476] For example, the 3D OMM unit 708 can be implemented as a passive integrated silicon photonic pillar or cube. The three-dimensional optical matrix multiplication unit 708 can be configured to implement a diffractive neural network and can perform matrix multiplication with almost zero power consumption.
[0477] There are many methods for encoding input data for use by the optical processor 702. For example, a digital input vector of length N×N can be encoded onto an optical input matrix of size N×N, which is propagated through a three-dimensional OMM unit 708. The three-dimensional OMM unit 708 performs (N×N)×(N×N) matrix multiplication on the received optical input matrix in the optical domain. The (N×N)×(N×N) matrix multiplication performed by the three-dimensional OMM unit 708 is determined by the internal configuration of the three-dimensional OMM unit 708, which includes the size, position, and geometry of the diffractive optical components arranged in three dimensions, as well as the doping of impurities (if any).
[0478] The 3D OMM element 708 can be implemented in various ways. Figure 8A schematic diagram of an example of a three-dimensional OMM unit 708 using a three-dimensional arrangement of diffractive components is shown. The three-dimensional OMM unit 708 may include an input waveguide matrix for receiving an optical input matrix 802, a three-dimensional optical interference unit 804 in optical communication with the input waveguide matrix, and an output waveguide matrix in optical communication with the optical interference unit 804 for providing an optical output matrix 806. The optical interference unit 804 includes multiple diffractive optical components and performs a transformation (e.g., a linear transformation) from optical input (e.g., an N×N vector or matrix) to optical output (e.g., an N×N vector or matrix). The output waveguide matrix guides the optical signal output by the optical interference unit 804. At least one input waveguide in the input waveguide matrix is in optical communication with each output waveguide in the output waveguide matrix via the optical interference unit 804. For example, for an N×N length optical input vector, the three-dimensional OMM unit 708 may include N×N input waveguides and N×N output waveguides.
[0479] In some embodiments, the optical interference unit 804 includes a substrate block with diffractive components arranged in a three-dimensional (e.g., 3D matrix) configuration. For example, multiple holes can be drilled or etched in each of a plurality of substrate slices, and the multiple substrate slices can be combined to form the substrate block. The size of these holes can be on the order of magnitude of the wavelength of the input light, such that the light is diffracted by the holes (or the structure defining the holes). The holes can have the same or different sizes. The holes can also have other cross-sectional shapes, such as triangular, square, rectangular, hexagonal, or irregular shapes. In some embodiments, a holographic method can be used to form a three-dimensional diffractive optical component throughout the substrate block. The substrate can be made of a material that is transparent or translucent to the input light, for example, having a transmittance of 1% to 99% relative to the input light.
[0480] When designing the three-dimensional OMM unit 708, the dimensions and positions of the diffractive components in the x, y, and z directions are considered. Optimization processes can be used to determine the configuration of the diffractive optical components. For example, the substrate block can be divided into a three-dimensional pixel matrix, and each pixel can be filled with substrate material (without holes) or filled with air (with holes). The pixel configuration can be modified iteratively, and for each pixel configuration, simulations can be performed by passing light through the diffractive optical components and evaluating the output. After performing simulations of all possible pixel configurations, the configuration that provides the closest result to the desired matrix processing is selected as the diffractive optical component configuration for the three-dimensional OMM unit 708.
[0481] As another example, the diffraction components are initially configured as a three-dimensional aperture matrix. The position, size, and shape of the apertures can differ slightly from their initial configuration. The parameters of each aperture can be iteratively adjusted, and simulations can be performed to find the optimal configuration of the apertures.
[0482] In some embodiments, machine learning processing is used to design three-dimensional diffractive optical components. An analysis function is used to determine how pixels affect the input light, and a gradient descent method is used to determine the optimal pixel configuration.
[0483] In some embodiments, the three-dimensional OMM unit 708 can be implemented as a user-variable component, and different three-dimensional OMM units 708 with different optical interference units 804 can be installed for different applications. For example, the system 700 can be configured as a medical diagnostic system, and the optical interference unit 804 can be configured to implement a neural network for analyzing physiological parameters to perform disease screening. For example, the first OMM unit may have a first optical interference unit including 3D passive diffractive optics components, configured to implement a first neural network for screening a first group of diseases. The second OMM unit may have a second optical interference unit including 3D passive diffractive optics components, configured to implement a second neural network for screening a second group of diseases, and so on. The first and second OMM units can be developed by different companies that specialize in developing technologies for screening different diseases. When a user wants to use the system 700 to screen for the first group of diseases, the user can insert the first OMM unit into the system. When a user wants to use the system 700 to screen for the second group of diseases, the user can remove the first OMM unit and insert the second OMM unit into the system.
[0484] For example, system 700 can be configured as an optical character recognition system, and optical interference unit 804 can be configured to implement a neural network for performing optical character recognition. For example, system 700 can be configured as a speech recognition system, and optical interference unit 804 can be configured to implement a neural network for performing speech recognition. For example, system 700 can be part of a control unit of an autonomous vehicle, and optical interference unit 804 can be configured to implement a neural network for performing road condition recognition.
[0485] For example, System 700 can be used for gene sequencing. DNA sequences can be classified using convolutional neural networks, implemented using System 700 which includes passive diffraction optics components. For example, System 700 can implement neural networks for distinguishing tumor types, predicting tumor grades, and predicting patient survival from gene expression patterns. For example, System 700 can implement neural networks for identifying subsets of genes or traits most predictive of the analyzed characteristics. For example, System 700 can implement neural networks for predicting or inferring the expression levels of all genes from a data map of a subset of genes. For example, System 700 can implement neural networks for epigenetic analysis, such as predicting transcription factor binding sites, enhancer regions, and chromatin accessibility from gene sequences. For example, System 700 can implement neural networks for capturing structures within gene sequences. For example, System 700 can be configured as a bacterial detection system, and the optical interference unit 804 can be configured to implement a multiplicative function for analyzing DNA sequences to detect certain bacterial strains.
[0486] In some embodiments, the three-dimensional OMM unit 708 includes a housing (e.g., a case) protecting the substrate having 3D diffractive optical components. The housing supports an input interface coupled to an input waveguide and an output interface coupled to an output waveguide. The input interface is configured to receive output from the modulator array 706, and the output interface is configured to send the output of the three-dimensional OMM unit 708 to the detection unit 710. The three-dimensional OMM unit 708 can be designed as a module suitable for handling by ordinary consumers, allowing users to easily switch from one three-dimensional OMM unit 708 to another. Machine learning techniques improve over time. Users can upgrade the system 700 by replacing older three-dimensional OMM units 708 and inserting newer, upgraded versions.
[0487] In some embodiments, system 700 is an optical computing platform configured to operate using OMM units from different companies. This allows different companies to develop different 3D passive optical neural networks for various applications. The 3D passive optical neural networks are sold to end users in standardized packages, which can be installed in the optical computing platform to allow system 700 to perform various intelligent functions.
[0488] In some embodiments, the system may have a retainer mechanism for supporting a plurality of three-dimensional OMM units 708, and may provide a mechanical processing mechanism for automatically replacing the three-dimensional OMM units 708. The system determines which three-dimensional OMM unit 708 is needed for the current application, and uses the mechanical processing mechanism to automatically obtain the appropriate three-dimensional OMM unit 708 from the retainer mechanism and insert it into the optical processor 702.
[0489] In some embodiments, the artificial neural network computing system 700 can be modified by adding an analog nonlinear unit between the detection unit 710 and the ADC unit 160. The analog nonlinear unit is configured to receive the output voltage from the detection unit 710, apply a nonlinear transfer function, and output the converted output voltage to the ADC unit 160. The controller 110 can obtain the converted digital output voltage corresponding to the converted output voltage from the ADC unit 160. Because the digital output voltage obtained from the ADC unit 160 has already been nonlinearly converted (“activated”), the nonlinear conversion step of the controller 110 can be omitted, thereby reducing the computational burden on the controller 110. Then, the first converted voltage obtained directly from the ADC unit 160 can be stored as a first converted digital output vector in the storage unit 120.
[0490] Optical interference units can be implemented using passive diffractive optical components arranged in one dimension. (See reference...) Figure 9 In some embodiments, the artificial neural network computing system 900 has an optical processor 906, which includes a one-dimensional optical multiplication unit 916. The system 900 includes a storage unit 120, which is similar to... Figure 1A The corresponding component of system 100. The optical processor 906 is configured to perform multiplication calculations using diffractive optical components arranged in one dimension (along the optical propagation axis).
[0491] The optical processor 906 includes: a laser unit 908 configured to output a laser beam 910; and a modulator 912 configured to modulate the laser beam 910 to generate a modulated beam 914. The optical processor 906 includes a one-dimensional optical multiplication unit 916 having one-dimensionally arranged diffractive optical components and configured to process the modulated beam 914 and generate an output beam 918. The optical processor 906 includes a detection unit 920 having a photosensor for detecting the output beam 918. The output of the detection unit 920 is converted into a digital signal by an ADC unit 930.
[0492] For example, the one-dimensional optical multiplication unit 916 can be implemented as a passive integrated silicon photonic waveguide with diffractive optical components (e.g., gratings or holes). The one-dimensional optical multiplication unit 916 can be configured to perform multiplication operations with almost zero power consumption.
[0493] There are many methods for encoding input data for use by the optical processor 906. For example, a digital input vector can be encoded as an optical input propagating through a one-dimensional optical multiplication unit 916. The one-dimensional optical multiplication unit 916 performs multiplication on the received optical input in the optical domain. The multiplication performed by the one-dimensional optical multiplication unit 916 is determined by the internal configuration of the one-dimensional optical multiplication unit 916, including, for example, the size, position, and geometry of the diffractive optical components arranged in one dimension along the light propagation path, and the doping of impurities (if any).
[0494] One-dimensional optical multiplication unit 916 can be implemented in various ways. Figure 10 A schematic diagram illustrating an example of a one-dimensional optical multiplication unit 916 using a one-dimensional arrangement of diffractive components is shown. The one-dimensional optical multiplication unit 916 may include an input waveguide for receiving an optical input 1002, a one-dimensional optical interference unit 1004 in optical communication with the input waveguide, and an output waveguide in optical communication with the optical interference unit 1004 for providing an optical output 1006. The optical interference unit 1004 includes multiple diffractive optical components and performs a conversion (e.g., linear conversion) from optical input to optical output. The output waveguide guides the optical signal output from the optical interference unit 1004.
[0495] In some embodiments, the optical interference unit 1004 includes an elongated substrate having diffractive components arranged in one dimension along the light propagation path. For example, multiple holes can be drilled or etched in the substrate. The size of these holes can be on the order of magnitude of the wavelength of the input light, such that the light is diffracted by the holes (or the structure defining the holes). The holes can have the same or different sizes. The substrate can be made of a material that is transparent or translucent to the input light, for example, having a transmittance of 1% to 99% relative to the input light. In some embodiments, holographic methods can also be used to form the diffractive optical components in the substrate.
[0496] When designing the optical interference unit 1004, the size and position of the diffractive components along the propagation path of the light beam are considered. Optimization processes can be used to determine the configuration of the diffractive optical components. For example, the substrate can be divided into a series of pixels, and each pixel can be filled with substrate material (without holes) or filled with air (with holes). The pixel configuration can be modified iteratively, and for each pixel configuration, simulations can be performed by passing light through the diffractive optical components and evaluating the output. After performing simulations of all possible pixel configurations, the configuration that provides the closest result to the desired multiplication process is selected as the diffractive optical component configuration of the optical interference unit 1004.
[0497] In another embodiment, the diffraction component is initially configured as a series of apertures. The positions and sizes of the apertures may differ slightly from their initial configuration. The parameters of each aperture can be iteratively adjusted, and simulations can be performed to find the optimal configuration of the apertures.
[0498] In some embodiments, machine learning processing is used to design one-dimensional diffractive optical components. An analytical function is used to determine how pixels affect the input light, and a gradient descent method is used to determine the optimal pixel configuration.
[0499] In some embodiments, the one-dimensional optical multiplication unit 916 can be implemented as a user-variable component, and different one-dimensional optical multiplication units 916 with different optical interference units 1004 can be installed for different applications. For example, the system 900 can be configured as a bacterial detection system, and the optical interference unit 1004 can be configured to implement a multiplication function for analyzing DNA sequences to detect certain bacterial strains. For example, a first optical multiplication unit may have a first optical interference unit including a 1D passive diffraction optics component configured to implement a first multiplication function for detecting a first group of bacteria. A second optical multiplication unit may have a second optical interference unit including a 1D passive diffraction optics component configured to implement a second multiplication function for detecting a second group of bacteria, and so on. The first and second optical multiplication units can be developed by different companies that specialize in developing technologies for detecting different bacteria. When a user wants to use the system 900 to detect a first group of bacteria, the user can insert the first optical multiplication unit into the system. When a user wants to use system 900 to detect a second group of bacteria, the user can replace the first optical multiplier unit and insert the second optical multiplier unit into the system. By using one-dimensional diffractive optical components, the laser unit 908, modulator 912, detection unit 920, and ADC unit 930 can be manufactured at low cost.
[0500] In some embodiments, the one-dimensional optical multiplier unit 916 includes a housing (e.g., a case) protecting a substrate having 1D diffractive optical components. The housing supports an input interface coupled to an input waveguide and an output interface coupled to an output waveguide. The input interface is configured to receive output from a modulator 912, and the output interface is configured to send the output of the one-dimensional optical multiplier unit 916 to a detection unit 920. The one-dimensional optical multiplier unit 916 can be designed as a module suitable for handling by a typical consumer, allowing the user to easily switch from one one-dimensional optical multiplier unit 916 to another. Machine learning techniques improve over time. Users can upgrade the system 900 by replacing older one-dimensional optical multiplier units 916 and inserting newer, upgraded versions.
[0501] In some embodiments, system 900 is an optical computing platform configured to operate using optical multiplication units from different companies. This allows different companies to develop different 1D passive optical multiplication functions for various applications. The 1D passive optical multiplication functions are sold to end users in standardized packages that can be installed in the optical computing platform to allow system 900 to perform various intelligent functions.
[0502] In some embodiments, the system may have a holder mechanism for supporting a plurality of one-dimensional optical multiplication units 916, and may provide a mechanical processing mechanism for automatically removing the one-dimensional optical multiplication units 916. The system determines which one-dimensional optical multiplication unit 916 is needed for the current application, and uses the mechanical processing mechanism to automatically retrieve the appropriate one-dimensional optical multiplication unit 916 from the holder mechanism and insert it into the optical processor 906.
[0503] In some embodiments, the artificial neural network computing system 900 can be modified by adding an analog nonlinear unit between the detection unit 920 and the ADC unit 930. The analog nonlinear unit is configured to receive the output voltage from the detection unit 920, apply a nonlinear transfer function, and output the converted output voltage to the ADC unit 930. The controller 902 can obtain the converted digital output voltage corresponding to the converted output voltage from the ADC unit 930. Because the digital output voltage obtained from the ADC unit 930 has already been nonlinearly converted (“activated”), the nonlinear conversion step of the controller 902 can be omitted, thereby reducing the computational burden on the controller 902. Then, the first converted voltage directly obtained from the ADC unit 930 can be stored as a first converted digital output vector in the storage unit 120.
[0504] Passive chips with passive diffractive optics offer several advantages. First, because active components (often the largest parts) have been phased out, chips of any given size can contain much larger neural networks. Typically useful neural networks can include millions of weights, which is challenging to implement on active chips and may require multiple data runs and reprogramming of the chip. In contrast, a single passive chip can potentially support an entire neural network. Second, the very low power consumption of passive chips is important for "edge" applications, as such applications often require a small footprint and low power consumption. Third, passive chips can be manufactured at very low cost because they do not contain active components.
[0505] Optical matrix multiplication units with passive diffraction optics components can also be used in wavelength division multiplexing (WDM) artificial neural network computing systems. For example, Figure 1FThe OMM unit 150 of system 104 can be replaced by an OMM unit using passive diffractive optical components. In this example, the second DAC subunit 134 can be removed.
[0506] In some embodiments, the optical processors (e.g., 504, 702) may perform matrix processing other than matrix multiplication. Optical matrix multiplication units 502 and 708 may be replaced by optical matrix processing units that perform other types of matrix processing.
[0507] Figure 25 A flowchart illustrating an example of method 2500 for performing ANN computations using an ANN computation system 500, 700, or 900, which includes one or more optical matrix multiplication units or optical multiplication units with passive diffraction components, such as 2D OMM unit 502, 3D OMM unit 708, or 1D OMM unit 916. The steps of processing 2500 may be performed at least in part by controller 110 or 902. In some embodiments, the various steps of method 2500 may be run in parallel, combined, cyclically, or in any order.
[0508] In step 2510, an artificial neural network (ANN) computation request, including an input dataset, is received. The input dataset includes a first digital input vector. The first digital input vector is a subset of the input dataset. For example, it can be a sub-region of an image. The ANN computation request can be generated by various entities (e.g., computer 102). The computer can include one or more of various types of computing devices, such as personal computers, server computers, vehicle computers, and flight computers. The ANN computation request typically refers to an electrical signal notifying or informing the ANN computation system 500, 700, or 900 that an ANN computation is about to be performed. In some embodiments, the ANN computation request can be divided into two or more signals. For example, the first signal can query the ANN computation system 500, 700, or 900 to check whether the system 500, 700, or 900 is ready to receive the input dataset. In response to an affirmative response from the system 500, 700, or 900, the computer can send a second signal including the input dataset.
[0509] In step 2520, the input dataset is stored. Controller 110 can store the input dataset in storage unit 120. Storing the input dataset in storage unit 120 allows for flexibility in the operation of the ANN computing system 500, 700, or 900, for example, improving overall system performance. For example, by retrieving the desired portion of the input dataset from storage unit 120, the input dataset can be divided into numeric input vectors of a defined size and format. Different portions of the input dataset can be processed in various orders or shuffled to allow for the execution of various types of ANN computations. For example, shuffling can allow matrix multiplication using block matrix multiplication techniques when the input and output matrix sizes are different. As another example, storing the input dataset in storage unit 120 allows the ANN computing system 500, 700, or 900 to queue multiple ANN computation requests, allowing the system 500, 700, or 900 to maintain full-speed operation without inactive periods.
[0510] In step 2530, a first plurality of modulator control signals are generated based on the first digital input vector. Controller 110 may send the first DAC control signal to DAC units 506, 712, or 904 to generate the first plurality of modulator control signals. DAC units 506, 712, or 904 generate the first plurality of modulator control signals based on the first DAC control signal, and modulator array 144, 706, or modulator 912 generates an optical input vector representing the first digital input vector.
[0511] The first DAC control signal may include multiple digital values to be converted by DAC units 506, 712, or 904 into a first plurality of modulator control signals. These multiple digital values typically correspond to a first digital input vector and can be associated through various mathematical relationships or lookup tables. For example, the multiple digital values may be linearly proportional to the values of the elements of the first digital input vector. As another example, the multiple digital values may be associated with the elements of the first digital input vector through a lookup table configured to maintain a linear relationship between the digital input vector and the optical input vector generated by modulator arrays 144, 706, or modulator 912.
[0512] In some embodiments, the 2D OMM unit 502, 3D OMM unit 708, or 1D OMM unit 916 is configured to perform optical matrix processing or optical multiplication based on the optical input vector and multiple neural network weights implemented using passive diffraction components. The multiple neural network weights representing matrix M can be decomposed into M = USV using the singular value decomposition (SVD) method. Where U is an M×M unitary matrix, S is an M×N diagonal matrix with non-negative real numbers on its diagonal, and V It is the complex conjugate of the N×N unitary matrix V. In this case, the passive diffraction component can be configured to realize matrix V, matrix S, and matrix U, such that the OMM element 502 or 708 realizes matrix M as a whole.
[0513] In step 2540, a first plurality of digital optical outputs corresponding to the optical matrix multiplication unit or optical multiplication are obtained. The optical input vector generated by the modulator array 144, 706, or modulator 912 is processed by the 2D OMM unit 502, 3D OMM unit 708, or 1D OMM unit 916 and converted into an optical output vector. The optical output vector is detected by the detection unit 146, 710, or 920 and converted into an electrical signal, which can be converted into a digital value by the ADC unit 160 or 930. The controller 110 or 902 can, for example, send a conversion request to the ADC unit 160 or 930 to begin converting the voltage output by the detection unit 146, 710, or 920 into a digital optical output. Once the conversion is complete, the ADC unit 160 or 930 can send the conversion result to the controller 110 or 902. Alternatively, the controller 110 or 902 can obtain the conversion result from the ADC unit 160 or 930. Controller 110 or 902 can generate a digital output vector from the digital optical output, which corresponds to the result of matrix multiplication or vector multiplication of the input digital vector. For example, the digital optical output can be organized or connected to have a vector format.
[0514] In some embodiments, the ADC unit 160 or 930 can be set or controlled to perform ADC conversion based on a DAC control signal issued by the controller 110 or 902 to the DAC unit 506, 712, or 904. For example, the ADC conversion can be set to begin at a preset time after the DAC unit 506, 712, or 904 generates a modulation control signal. This control of the ADC conversion simplifies the operation of the controller 110 or 902 and reduces the number of necessary control operations.
[0515] In step 2550, a nonlinear transformation is performed on the first digital output vector to produce a first transformed digital output vector. A node or artificial neuron of an ANN operates by first performing a weighted summation of signals received from nodes in previous layers, and then performing a nonlinear transformation (“activation”) of the weighted sum to produce an output. Various types of ANNs can implement various types of differentiable nonlinear transformations. Examples of nonlinear transformation functions include the Modified Linear Unit (RELU) function, the sigmoid function, the hyperbolic tangent function, the X^2 function, and the |X| function. This nonlinear transformation is performed on the first digital output by controller 110 or 902 to produce the first transformed digital output vector. In some embodiments, the nonlinear transformation may be performed by a dedicated digital integrated circuit within controller 110 or 902. For example, controller 110 or 902 may include one or more modules or circuit blocks particularly suited to accelerate the computation of one or more types of nonlinear transformations.
[0516] In step 2560, the first transformed digital output vector is stored. The controller 110 or 902 may store the first transformed digital output vector in storage unit 120. When the input dataset is divided into multiple digital input vectors, the first transformed digital output vector corresponds to the ANN computation result of a portion of the input dataset, such as the first digital input vector. In this way, storing the first transformed digital output vector allows the ANN computation system 500, 700, or 900 to perform and store additional computations on other digital input vectors of the input dataset for later aggregation into a single ANN output.
[0517] In step 2570, the output is generated based on the artificial neural network output vector produced by the first transformed digital output vector. Controller 110 or 902 generates the ANN output, which is the result of processing the input dataset by an ANN defined by a first plurality of neural network weights. When the input dataset is divided into multiple digital input vectors, the generated ANN output is an aggregated output including the first transformed digital output, but may further include additional transformed digital outputs corresponding to other parts of the input dataset. Once the ANN output is generated, it is sent to the computer (e.g., computer 102) that initiated the ANN computation request.
[0518] A 2D OMM unit 502, a 3D OMM unit 708, or a 1D OMM unit 916 can represent the weight coefficients of a hidden layer in a neural network. If the neural network has multiple hidden layers, additional 2D OMM units 502, 3D OMM units 708, or 1D OMM units 916 can be coupled in series. Figure 26An example of an ANN computing system 2600 for implementing a neural network with two hidden layers is shown. A first 2D optical matrix multiplication unit 2604 represents the weight coefficients of the first hidden layer, and a second 2D optical matrix multiplication unit 2606 represents the weight coefficients of the second hidden layer. The ANN computing system 2600 includes a controller 110, a storage unit 120, a DAC unit 506, and an optoelectronic processor 2602. The storage unit 120 and the DAC unit 506 are similar to... Figure 5 The corresponding component of system 500. The optoelectronic processor 2602 is configured to perform matrix calculations using optical and electronic components.
[0519] The optoelectronic processor 2602 includes a first laser unit 142a, a first modulator array 144a, a first 2D optical matrix multiplication unit 2604, a first detection unit 146a, a first analog nonlinear unit 310a, an analog storage unit 320, a second laser unit 142b, a second modulator array 144b, a second 2D optical matrix multiplication unit 2606, a second detection unit 146b, a second analog nonlinear unit 310b, and an ADC unit 160. The operation of the first laser unit 142a, the first modulator array 144a, the first detection unit 146a, the first analog nonlinear unit 310a, and the analog storage unit 320 is similar to... Figure 3B The corresponding component is shown. The first 2D OMM unit 2604 is similar to... Figure 5 The 2D OMM 502. The output of the analog storage unit 320 drives the second modulator array 144b, which modulates the laser light from the second laser unit 142b to generate a light vector. The light vector from the second modulator array 144b is processed by the second 2D OMM unit 2606, which performs matrix multiplication and generates a light output vector, which is detected by the second detection unit 246b. The second detection unit 246b is configured to generate an output voltage corresponding to the light signal from the light output vector of the second 2D OMM unit 2606. The ADC unit 160 is configured to convert the output voltage into a digital output voltage. The controller 110 can obtain a digital output corresponding to the light output vector of the second 2D OMM unit 2606 from the ADC unit 160. The controller 110 can form a digital output vector from the digital output, which corresponds to the result of a second matrix multiplication of the result of a first matrix multiplication of the input digital vector. By using a light splitter to combine the second laser unit 142b with the first laser unit 142a, some of the light from the first laser unit 142a is redirected to the second modulator array 144b.
[0520] The above principle can be applied to implement neural networks with three or more hidden layers, where the weight coefficients of each hidden layer are represented by the corresponding 2D OMM unit.
[0521] Figure 27 An example of an ANN computing system 2700 for implementing a neural network with two hidden layers is shown. A first 3D optical matrix multiplication unit 2704 represents the weight coefficients of the first hidden layer, and a second 3D optical matrix multiplication unit 2706 represents the weight coefficients of the second hidden layer. The ANN computing system 2700 includes a controller 110, a storage unit 120, a DAC unit 712, and an optoelectronic processor 2702. The storage unit 120 and the DAC unit 712 are similar to... Figure 7 The corresponding component of system 700. The optoelectronic processor 2702 is configured to perform matrix calculations using optical and electronic components.
[0522] The optoelectronic processor 2702 includes a first laser unit 704a, a first modulator array 706a, a first 3D optical matrix multiplication unit 2704, a first detection unit 710a, a first analog nonlinear unit 310a, an analog storage unit 320, a second laser unit 704b, a second modulator array 706b, a second 3D optical matrix multiplication unit 2706, a second detection unit 710b, a second analog nonlinear unit 310b, and an ADC unit 160. The operation of the first laser unit 704a, the first modulator array 706a, the first detection unit 710a, the first analog nonlinear unit 310a, and the analog storage unit 320 is similar to... Figure 3B The corresponding component is shown. The first 3D OMM unit 2704 is similar to... Figure 7The 3D OMM 708. The output of the analog storage unit 320 drives the second modulator array 706b, which modulates the laser light from the second laser unit 704b to generate a light vector. The light vector from the second modulator array 706b is processed by the second 3D OMM unit 2706, which performs matrix multiplication and generates a light output vector, which is detected by the second detection unit 710b. The second detection unit 710b is configured to generate an output voltage corresponding to the light signal from the light output vector of the second 3D OMM unit 2706. The ADC unit 160 is configured to convert the output voltage into a digital output voltage. The controller 110 can obtain a digital output corresponding to the light output vector of the second 3D OMM unit 2406 from the ADC unit 160. The controller 110 can form a digital output vector from the digital output, which corresponds to the result of a second matrix multiplication of the result of a first matrix multiplication of the input digital vector. The second laser unit 704b can be combined with the first laser unit 704a by using an optical splitter to redirect some of the light from the first laser unit 704a to the second modulator array 706b.
[0523] The above principle can be applied to implement neural networks with three or more hidden layers, where the weight coefficients of each hidden layer are represented by the corresponding 3D OMM unit.
[0524] The 2D OMM unit 502 and 3D OMM unit 708 with passive diffractive optical components are suitable for recurrent neural networks (RNNs), wherein the output of the network during the (k)th pass through the neural network is recycled back to the input of the neural network and used as input during the (k+1)th pass, so that the weight coefficients of the neural network remain the same during multiple passes.
[0525] Figure 28 An example of a neural network computing system 2800 is shown, which can be used to implement a recurrent neural network. System 2800 includes an optical processor 2802, which operates in a manner similar to... Figure 3B The optical processor 140 operates in a manner similar to that of the optical processor 140, except that the OMM unit 150 is replaced by a 2D OMM unit 2804. The 2D OMM unit 2804 can be similar to... Figure 6 The 2D OMM unit 502. The neural network weights of the 2D OMM unit 2804 are fixed, therefore the system 2800 does not need to... Figure 3B The second DAC subunit 134 used in system 302.
[0526] Figure 29 An example of a neural network computing system 2900 is shown, which can be used to implement a recurrent neural network. System 2900 includes an optical processor 2902, which operates in a manner similar to... Figure 3B The optical processor 140 operates in a manner where, apart from the laser unit 142, modulator array 144, OMM unit 150, and detection unit 146, each is respectively... Figure 7 The laser unit 704, modulator array 706, 3D OMM unit 2904, and detection unit 710 are replaced. The neural network weights of the 3D OMM unit 2904 are fixed, therefore the system 2900 does not need to... Figure 3B The second DAC subunit 134 used in system 302.
[0527] Figure 30 A schematic diagram shows an example of an artificial neural network computing system 3000 with 1-bit internal resolution. The ANN computing system 3000 is similar to... Figure 4A The ANN computing system 400, except for the OMM unit 150, consists of 2D OMM units 3004 (which are similar to...). Figure 5 The 2D OMM unit 502 is replaced, and the second driver subunit 434 is omitted. The ANN computing system 3000 operates in a manner similar to that of the ANN computing system 400, wherein the input vector is decomposed into multiple 1-bit vectors, and certain ANN computations can then be performed by summing the results of the individual matrix multiplications after performing a series of matrix multiplications on the 1-bit vectors.
[0528] Figure 31 A schematic diagram shows an example of an artificial neural network computing system 3100 with 1-bit internal resolution. The ANN computing system 3100 is similar to... Figure 4A The ANN computing system 400, in addition to the OMM unit 150, consists of 3D OMM units 3104 (which are similar to...). Figure 7 The 3D OMM unit 708 is replaced, and the second driver subunit 434 is omitted. Figure 31 In the example, Figure 4A The laser unit 142, modulator array 144, and detection unit 146 are respectively composed of Figure 7 The laser unit 704, modulator array 706, and detection unit 710 are replaced. The ANN computing system 3100 operates in a manner similar to that of the ANN computing system 400, wherein the input vector is decomposed into multiple 1-bit vectors, and then certain ANN calculations can be performed by summing the results of each matrix multiplication after performing a series of matrix multiplications on the 1-bit vectors.
[0529] The following describes the principle of optical diffraction neural networks. Optical diffraction neural networks can be implemented using several layers of diffraction or transmission optical media. Based on the Huygens–Fresnel principle, each point in the diffraction medium can be considered a secondary light source. For each light source, far-field diffraction can be described by the following equation:
[0530] .
[0531] Here, the exponents l and i represent the i-th neuron in the l-th layer of the neural network, λ is the wavelength of light, and r is the distance, where:
[0532] .
[0533] The output from each secondary light source can be written as the input multiplied by the phase and intensity modulation of the light source:
[0534]
[0535] Here, t represents transmission modulation, which is a complex term encompassing both amplitude and phase modulation. It is the sum of inputs from all previous light sources. Overall, the output can be combined into the far-field diffraction time w and the amplitude |A| plus an additional phase term. Therefore, each point in each layer can be considered as a neuron that takes input from multiple neurons in the previous layer and adds additional phase and intensity modulation before outputting to the next layer.
[0536] The following describes a compact design for a compact photonic matrix multiplier unit that can implement general unitary matrix multiplication. (Refer to...) Figure 11The photonic matrix multiplier unit 1100 includes a modulator 1102, multiple interconnecting interferometers 1104, and an attenuator 1106. The interconnecting interferometers 1104 include directional coupler layers (or groups or sets of directional couplers) 1108a, 1108b, 1108c, 1108d, and 1108e (collectively referred to as 1108) and phase shifter layers (or groups or sets of phase shifters) 1110a, 1110b, 1110c, and 1110d (collectively referred to as 1110). Each directional coupler layer (or group or set of directional couplers) may include one or more directional couplers. Each phase shifter layer may include one or more phase shifters. In this example, the interconnecting interferometer 1104 includes five layers of directional couplers 1108 and four layers of phase shifters. In other embodiments, the photonic matrix multiplier unit 1100 may have different directional coupler and phase shifter layers. Compared to conventional matrix multiplier units that use interconnected Mach-Zehnder interferometers, the photonic matrix multiplier unit 1100 has a directional coupler 1108, which is positioned in a way that reduces the number of layers of the directional coupler 1108.
[0537] Here, the term "layer" in the phrases "directional coupler layer" and "phase shifter layer" refers to a group or set of directional couplers or phase shifters based on their positions relative to the input and output ports within the photon matrix multiplier unit 1100. Figure 11 In the example, the input optical signal is processed by the first-layer directional coupler 1108a, then by the second-layer phase shifter 1110a, then by the third-layer directional coupler 1108b, then by the fourth-layer phase shifter 1110b, and so on.
[0538] For example, a conventional matrix multiplier unit using an interconnected Mach-Zehnder interferometer might require 2N layers of directional couplers, while the photonic matrix multiplier unit 1100 requires only N+2 layers of directional couplers. N represents the number of input signals, or the number of digits in the input vector. The mesh architecture used in the photonic matrix multiplier unit 1100 likely has the most compact geometry for performing general matrix computations in an interconnected photonic interferometer.
[0539] Figure 12AA graph comparing the interconnected interferometer 1104 of the photon matrix multiplier unit 1100 with a conventionally designed interconnected interferometer for various numbers of input signals is shown. When there are four input signals, the conventionally designed interconnected Mach-Zehnder interferometer 1200 requires eight layers of directional couplers, while the new compact interconnected interferometer 1202 requires only six layers. When there are three input signals, the conventionally designed interconnected Mach-Zehnder interferometer 1204 requires six layers of directional couplers, while the new compact interconnected interferometer 1206 requires only five layers. When there are eight input signals, the conventionally designed interconnected Mach-Zehnder interferometer 1208 requires 16 layers of directional couplers, while the new compact interconnected interferometer 1210 requires only 10 layers.
[0540] Typically, when there are n input signals, a conventionally designed interconnected Mach-Zehnder interferometer requires 2n layers of directional couplers, while a newly designed compact interconnected interferometer requires only n+2 layers of directional couplers.
[0541] In a traditional design, for n input signals, there are n layers of Mach-Zehnder interferometers, and each Mach-Zehnder interferometer includes a directional coupler, followed by a pair of phase shifters, and then another directional coupler. Therefore, an n-layer Mach-Zehnder interferometer has 2n layers of directional couplers. Consequently, in a traditional design, for n input signals, n layers of phase shifters and 2n layers of directional couplers are required.
[0542] In contrast, in the new compact design, a directional coupler layer is followed by a first phase shifter layer, then another directional coupler layer, then a second phase shifter layer, then another directional coupler layer, then a third phase shifter layer, and so on. After the last phase shifter layer, there are two more directional couplers. As a result, for n input signals, there are n phase shifters and n+2 directional couplers.
[0543] Because directional couplers occupy a lot of space, reducing the number of directional couplers from 2·n to n+2 compared to traditional designs can significantly reduce the size of the photon matrix multiplier unit 1100.
[0544] Figure 12B A diagram is shown based on the newly designed compact interconnected interferometer 1212, where the number of input signals is 5.
[0545] The following describes the compact design decomposition using gradient descent. The compact design of the photon matrix multiplier described above can take any unitary matrix U, and an analytic decomposition algorithm is used to determine which phases need to be implemented using phase shifters, and thus implement matrix U. For example, phases can be extracted from a given matrix U using gradient descent. The gradient descent process is as follows: Start with a fixed matrix U and initialize random weights θ for the phase shifters of the compact design. Construct matrix U' using the compact design, i.e., U' = CompactDesign(θ). Next, consider the loss function L = |U-U'|^2 (this is the Frobenius norm of the matrix) and minimize this function using gradient descent (i.e., update θ using gradient updates).
[0546] Reference Figure 13 Homodyne detection is used (e.g., extracting the real part from the output), thus requiring an additional attenuator layer 1302 before detection to simulate the orthogonal matrix. This means that, along with θ, the diagonal weights x of the attenuator also need to be learned. In this way, the required phase and diagonal weights of U can be determined, and the decomposition can be obtained numerically.
[0547] The following describes an optical generative adversarial network (OGAN), which includes a generator configured to efficiently produce faithful data. Figure 14 An example of a light generative adversarial network 1400 is shown, wherein...
Claims
1. A method for performing computation in a system having a matrix multiplication unit configured to convert an optical input vector into an analog output vector based on a plurality of weight control signals, the method comprising: Receive a computation request, the computation request including at least one digital input vector and a first plurality of weights; The digital input vector and the first plurality of weights are stored in a storage unit; A first plurality of modulator control signals are generated based on the digital input vector, and a first plurality of weight control signals are generated based on the first plurality of weights; The optical input vector is generated based on the control signals of the first plurality of modulators; Obtain a first plurality of digital outputs corresponding to the analog output vector of the matrix multiplication unit, the first plurality of digital outputs forming a first digital output vector; The controller performs a nonlinear transformation on the first digital output vector to generate a first transformed digital output vector; The first converted digital output vector is stored in the storage unit; as well as The controller outputs the output generated based on the first converted digital output vector.
2. The method of claim 1, wherein receiving a computation request includes receiving the computation request from a computer via a communication channel, and the matrix multiplication unit is configured to process data at a data rate at least an order of magnitude greater than the data rate of the communication channel.
3. The method according to claim 1, wherein the matrix multiplication unit and the controller are disposed on at least one of a multi-chip module or an integrated circuit; Receiving the computation request includes receiving the computation request from a second data processor, wherein the second data processor is located external to the multi-chip module or the integrated circuit, and the second data processor is coupled to the multi-chip module or the integrated circuit via a communication channel. The matrix multiplication unit is capable of processing data at a data rate at least one order of magnitude greater than the data rate of the communication channel.
4. The method of claim 1, wherein the controller comprises an application-specific integrated circuit (ASIC), and Receiving the computation request includes receiving the computation request from a general-purpose data processor.
5. The method of claim 1, wherein generating the first plurality of modulator control signals comprises generating the first plurality of modulator control signals via a digital-to-analog converter (DAC) unit.
6. The method of claim 1, wherein obtaining the first plurality of digital outputs comprises obtaining the first plurality of digital outputs from an analog-to-digital converter (ADC) unit.
7. The method according to claim 5, comprising: The first plurality of modulator control signals are applied to a plurality of optical modulators, the plurality of optical modulators being coupled to a light source and a digital-to-analog converter (DAC) unit configured to generate the first plurality of modulator control signals; as well as Using the plurality of optical modulators, an optical input vector is generated by modulating the plurality of optical outputs generated by the laser unit based on the plurality of modulator control signals.
8. The method of claim 7, wherein the matrix multiplication unit is coupled to the plurality of optical modulators and the DAC unit, and the method comprises: Using the matrix multiplication unit, the optical input vector is converted into the analog output vector based on the multiple weight control signals.
9. The method of claim 6, wherein the ADC unit is coupled to the matrix multiplication unit, and the method comprises: The analog output vector is converted into the first plurality of digital outputs using the ADC unit.
10. The method of claim 7, wherein the matrix multiplication unit comprises an optical matrix multiplication unit coupled to the plurality of optical modulators and the DAC unit. Converting the optical input vector into an analog output vector includes using the optical matrix multiplication unit to convert the optical input vector into an optical output vector based on the plurality of weight control signals, and The method includes: A photoelectric detection unit coupled to the optical matrix multiplication unit is used to generate multiple output voltages corresponding to the optical output vector.
11. The method according to claim 1, comprising: The optical input vector is received by the input waveguide array; Using an optical interferometer unit that is optically in communication with the input waveguide array, a linear transformation is performed to convert the optical input vector into a second optical signal array; and The second optical signal array is guided using an output waveguide array that is optically in communication with the optical interference unit, wherein at least one input waveguide in the input waveguide array is optically in communication with each output waveguide in the output waveguide array via the optical interference unit.
12. The method of claim 11, wherein the optical interference unit comprises a plurality of interconnected Mach-Zehnder interferometers (MZIs), each of the plurality of interconnected MZIs comprising a first phase shifter and a second phase shifter, and the first phase shifter and the second phase shifter are coupled to the plurality of weight control signals. in, The method includes: The separation ratio of the MZI is changed using the first phase shifter, and The second phase shifter is used to shift the phase of one output of the MZI.
13. The method of claim 6, comprising: For each of at least two subsets of one or more optical signals of the optical input vector, the subset of one or more optical signals is divided into two or more copies of the optical signals using a corresponding set of one or more replication modules; For each of at least two copies of a first subset of one or more optical signals, the corresponding multiplication module is used to multiply one or more optical signals of the first subset by one or more matrix element values using optical amplitude modulation. as well as For the results of two or more multiplication modules, a summation module is used to generate an electrical signal representing the sum of the two or more results of the multiplication modules.
14. The method of claim 13, wherein at least one of the multiplication modules includes an optical amplitude modulator, the optical amplitude modulator including an input port and two output ports, and providing a pair of correlated optical signals from the two output ports such that the difference between the amplitudes of the correlated optical signals corresponds to the result of multiplying the input value by the value of a signed matrix element.
15. The method of claim 13, further comprising using the matrix multiplication unit to multiply the input vector by a matrix including the values of the one or more matrix elements.
16. The method of claim 15, further comprising encoding a plurality of output values on corresponding electrical signals generated by the one or more summing modules, and The output vector is generated by multiplying the input vector by the matrix, using the output values from the set of multiple output values to represent the elements of the output vector.
17. The method of claim 1, wherein receiving the computation request comprises receiving a computation request comprising an input dataset and a first plurality of neural network weights, and the input dataset comprises a first numeric input vector; Generating the first plurality of modulator control signals includes generating the first plurality of modulator control signals based on the first digital input vector; Generating the first plurality of weight control signals includes generating the first plurality of weight control signals based on the first plurality of neural network weights; The output includes the output generated by the controller based on the first converted digital output vector.
18. The method of claim 1, wherein the system has a first cycle time period, the first cycle time period being defined as the time elapsed between storing the input vector and the first plurality of weights in the storage unit and storing the first transformed digital output vector in the storage unit, and The first cycle time period is less than or equal to 1 ns.
19. The method of claim 5, further comprising generating a second plurality of modulator control signals by means of the DAC unit based on the first converted digital output vector; The computation request also includes a second set of weights; as well as The DAC unit generates a second plurality of weight control signals based on the second plurality of weights; The first plurality of weights and the second plurality of weights correspond to different layers of the artificial neural network.
20. The method of claim 1, wherein receiving the computation request comprises receiving a computation request having a plurality of digital input vectors and the first plurality of weights; Generating the first plurality of modulator control signals includes generating the first plurality of modulator control signals based on the plurality of digital input vectors; Generating the first plurality of weight control signals includes generating the first plurality of weight control signals based on the first plurality of weights; The first plurality of modulator control signals are applied to a plurality of optical modulator groups to generate a plurality of optical input vectors, wherein the plurality of optical modulator groups are coupled to a light source configured to generate light having a plurality of wavelengths, each of the optical modulator groups corresponding to one of the plurality of wavelengths, and each optical input vector corresponding to one of the plurality of wavelengths; The plurality of optical input vectors are combined into a combined optical input vector having the plurality of wavelengths; The matrix multiplication unit is used to convert the combined light input vector into a combined light output vector having the multiple wavelengths. The combined optical output vector is decomposed into multiple optical output vectors, each optical output vector corresponding to a corresponding wavelength. The plurality of optical output vectors are detected to generate a multi-path decomposed output voltage, which forms a plurality of digital output vectors, wherein each digital output vector corresponds to a corresponding wavelength. The controller performs a nonlinear transformation on each of the plurality of digital output vectors to generate a plurality of transformed digital output vectors; as well as The plurality of converted digital output vectors are stored in the storage unit; Each of the plurality of digital input vectors corresponds to one of the plurality of optical input vectors.