Method and systems for permuted diagonal computing with ultrashort pulses
Offset diagonal computing with optical delay lines and modulators addresses scaling challenges in optical neural networks, achieving high-speed, efficient processing of large-scale models by leveraging temporal and spatial multiplexing.
Patent Information
- Application Number
- PCT/US2025/039702
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-01-16
- Filing Date
- 2025-07-29
- Publication Date
- 2026-02-05
AI Technical Summary
Current optical neural networks face challenges in scaling to the size of the human brain due to complex engineering requirements, thermal management, and inefficient memory access, making it impractical to process large-scale models effectively.
A system of offset diagonal computing using optical delay lines, modulators, and detectors for matrix-matrix multiplication, enabling efficient processing of large-scale models through temporal and spatial multiplexing, with in-memory accumulation and fan-in operations.
Enables real-time processing of billion-scale model sizes with reduced component count and energy efficiency, surpassing traditional computing architectures in speed and scalability, particularly for sparse and complex-valued problems.
Smart Images

Figure US2025039702_05022026_PF_FP_ABST
Abstract
Description
METHOD AND SYSTEMS FOR PERMUTED DIAGONAL COMPUTING WITH ULTRASHORT PULSESCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the priority and benefit of U.S. Provisional Application No. 63 / 676,611; filed on July 29, 2024, titled “METHOD AND SYSTEMS FOR PERMUTED DIAGONAL COMPUTING WITH ULTRASHORT PULSES” and U.S. Provisional Application No. 63 / 745,949; filed on January 16, 2025, titled “METHOD AND SYSTEMS FOR OFFSET DIAGONAL COMPUTING WITH SPATIAL, TEMPORAL, AND WAVELENGTH MULTIPLEXING.” The entire disclosure of each of these applications is hereby incorporated by reference.FIELD
[0002] This application relates to the field of high-performance computing, in particular advanced computer algorithms and optical computing.BACKGROUND
[0003] Artificial neural networks (ANNs) inspired by the massive processing of the brain can now be found across modern society in web searching, smartphones, and games. While incredibly effective, large neural networks are incompatible with traditional computers for which Moore’s law has hit physical limits, and memory accessing adds significant lag. Specialized dense-data-optimized processors such as GPUs and TPUs now greatly surpass these bottlenecks but are power hungry and face lingering challenges in memory access speeds. Optics, which has already demonstrated superiority for interconnection tasks, has been recognized as a great fit for today’s incredible computing demands due to its low-loss propagation, high bandwidth, and lack of interference between neighboring channels. While recent spatially extended optical neuralnetwork approaches have already demonstrated useful fast and efficient calculations for small neural networks, scaling size by more than six orders of magnitude to meet current demand appears to be an insurmountable technical challenge.
[0004] While optical systems to date have demonstrated the promise of optics for highspeed ANNs, scaling for processing state-of-the-art models is challenging, if not impractical; scaling to the number of connections estimated in the human brain (-100 trillion synapses) appears implausible. For instance, demonstrated MZI and MRR systems are complex, highly calibrated nano-photonic devices and only have as few as four input nodes for input processing, whereas current ANN models trained with GPUs / TPUs require at least tens of thousands of input nodes with billions of overall trainable parameters. Diffractive ANNs have more effective nodes, corresponding to each pixel of a diffractive element, for example, but the nodes are only indirectly connected to each other which severely limits the models that can be run on this hardware. Diffractive systems used for directly connecting each node have accomplished multiplication with only 56 nodes. In addition, nonlinear activation is more difficult to achieve in free space diffractive ANNs. Overall, scaling up current spatial optical ANNs requires engineering complexity in thermal management, calibration, and control design that scales with the number of nodes squared.
[0005] This scaling challenge is in part addressed by leveraging the time domain by encoding symbols with high speeds (10s of GS / s). However, temporal systems demonstrated to date still require many well-calibrated integrated electrical components such as balanced detectors, ADCs, and capacitive integrators before their energy efficiencies can significantly surpass their digital counterparts. New approaches are required to realize the efficiencies possible with optical computing within practical design constraints as well as to leverage efficient sparse models.SUMMARY
[0006] An aspect of the application is a system of offset diagonal computing comprising: an input vector per wavelength, wherein the input vector consists of a sequence of elements of an input optical stream; one or more input modulators per wavelength, wherein the one or more input modulators receive the input vector; one or more fan-outs, wherein the one or more fanouts receive the input vector from the one or more input modulators, and wherein the one or more fan-outs divide the elements of the input vector into N total paths;optical delay lines, wherein the optical delay lines receive the input vector from thepaths;weight modulators, wherein the one or more input modulators are connected to theweight modulators through the N total paths and optical delay lines; a multiplication unit, wherein a matrix-matrix multiplication (MMM) computation is performed that is a sum of offset diagonals of a weight matrix multiplied by the input matrix; one or more fan-ins, wherein the one or more fan-ins receive one or more elements of the input vector output from one or more of the N weight modulators; and a detector per wavelength, wherein products are summed through time- synchronized incidence on the detector through one or more spatially distinct pathways from one or more of theweight modulators, and / or wherein products are summed by binning one or more neighboring pulses, and wherein the temporal binning is achieved through electronic integration.
[0007] In certain embodiments, the weight matrix is subdivided into a plurality of sets of columns of elements of the weight matrix, and the sets of columns of products of the weight matrix are summed simultaneously.
[0008] In certain embodiments, the input vector is received by one or more input modulators less than the number of diagonals in the matrix, and further wherein a number of neighboring pulses that are binned is greater than one.
[0009] In certain embodiments, the weight matrix is subdivided into a plurality of blocks of columns and rows of elements of the weight matrix, and the elements of the columns of the weight matrix are summed simultaneously.
[0010] In certain embodiments, the weight matrix is sub-divided into sets of columns and summation is performed simultaneously on the sets of columns of the weight matrix, and wherein for each given set of columns of the weight matrix that set is temporally binned by a detector dedicated to that given set.
[0011] In certain embodiments, the input vector is received by one or more input modulators less than the number of columns in the matrix, and optical fan-out is not required.
[0012] In certain embodiments, the elements of any one or more modulators can be efficiently made to not contribute to the weight matrix.
[0013] In certain embodiments, the rows or columns of the weight matrix are rearranged through arbitrary permutations of the input vector.
[0014] In certain embodiments, the modulators are selected from the group comprising electro-optical modulators such as thin-film lithium niobate, Spatial light modulators (SLMs), Absorptive modulators, Micro-ring resonator modulators, spectral waveshapers, and optical switches (for 0 / 1).
[0015] In certain embodiments, the delay lines are selected from the group comprising on-chip waveguides such as SiN-based or Si-photonic delay lines, free-space delays, and optical fiber delay lines.
[0016] In certain embodiments, wavelength dependent fan-ins (which are a type of dispersive optical elements) are selected from the group comprising free space components (optical prisms, optical diffraction-gratings), metamaterial surfaces, and on-chip components(Arrayed waveguide gratings (AWG) or on-chip Bragg-gratings).
[0017] In certain embodiments, further comprising wavelength dependent fan-outs, wherein wavelength dependent fan-outs are selected from the group comprising free space components (optical prisms, optical diffraction-gratings), metamaterial surfaces, and on-chip components (Arrayed waveguide gratings (AWG) or on-chip Bragg-gratings).
[0018] In certain embodiments, detectors are selected from the group comprising free- space or integrated detectors with optical beams incident at different angles, on-chip multi-port photodetectors (detectors which accept inputs from multiple optical waveguides), single-mode detectors followed by capacitive integrators, and multi-mode photodiodes (accept inputs from multiple spatial modes of an optical waveguide).
[0019] Another aspect of the application is a method of designing and operating an offset diagonal computing system comprising: determining a number of input modulators for use in the system, wherein each input modulator has an input and an output; determining a number of weight modulators for use in the system, wherein the weight modulators are organized in one or more sets; determining a number of detectors for use in the system, N ; dividing the output of each input modulator in N paths exiting the input modulators, wherein each of the N paths is received by a distinct weight modulator and delay line from 0 to the max delay, Md, wherein any path exiting any given input modulator becomes divided into Nd (Md+1) paths; multiplying an input vector element with weights, wherein a product is produced once multiplication is complete; routing optical paths exiting the weight modulators after multiplication is complete onto an assigned detector; spatially and / or temporally accumulating products before reading out a requisite detector signal.
[0020] In certain embodiments, calculations are batched by dividing the weight matrix with vertical cuts, and wherein the number of vertical cuts required to complete a computationdepends on the number of offset diagonals that can be processed simultaneously which is determined by the number of available weight modulators; and further comprising the steps of: forming weight submatrices by the vertical cuts; computing each submatrix; summing together the results of computing each submatrix.
[0021] In certain embodiments, any of the systems described herein implementing any of the methods described herein.
[0022] A further aspect of the application is a modular architecture comprising any one of the systems described herein, wherein the modular architecture comprises: an input matrix encoder; a plurality of optical delay lines; a plurality of compute cores; and one or more optical combining modules.
[0023] In certain embodiments, the input matrix encoder is selected from the group comprising an integrated frequency -comb system, an N frequency channel laser bank plus wavelength mux plus micro-ring resonators (MRRs) system, an N frequency channel laser bank plus electro-optic modulators system, and an electro-optic RF comb plus MRRs system.
[0024] In certain embodiments, the plurality of compute cores is selected from the group comprising an integrated AWG output system, a bulk output system, an individual AWGs plus two-layer routing plus multi-port PDs system, and an individual AWGs plus multi-layer routing plus multi-port PDs system.
[0025] In certain embodiments, the integrated AWG output system comprises input matrix elements from the encoder fanned out into a plurality of optical paths, wherein each optical path comprises on-chip delay lines, wherein the on-chip delay lines perform delays to cyclically shift the input vectors; and a stack of broadband electro-optic modulators (EOMs), wherein the EOMs simultaneously perform multiplication operations across all wavelengthchannels, wherein products of the multiplication are then incident into an arrayed waveguide grating-component, wherein the grating component both spatially separates the wavelength channels and focusses them; a linear detector array, wherein the array performs spatial summations of the products that are focused onto the array.
[0026] In certain embodiments, the arrayed waveguide grating-component comprises: an input field from an electro-optic modulator containing several wavelength channels, wherein the input field is incident into a star coupler; wherein the star coupler divides an input signal irrespective of the wavelength into several output channels, and wherein each output channel of the star coupler has a different optical path length, but the neighboring channels have a fixed path length difference; each wavelength channel accumulates a different phase in each spatial channel of the star coupler, wherein these channels are then incident onto a curved structure that functions as an optical lens, and wherein each wavelength channel by the virtue of the phase difference accumulated and through optical interference exits at a different angle and as a consequence are focused onto a different detector.
[0027] In certain embodiments, the bulk output system comprises: input matrix elements from the encoder fanned out into a plurality of optical paths, wherein each optical path comprises on-chip delay lines, outputs are then coupled into free-space with the use of on-chip grating couplers or fiber-optic collimators; a diffraction grating is employed as a wavelength selective element which splits each wavelength channel into a corresponding angled beam; and a focusing lens then focuses the beams belonging to a single wavelength channel onto an appropriate detector.
[0028] In certain embodiments, the individual AWGs plus two-layer routing plus multiport PDs system comprises: input matrix elements from the encoder fanned out into a plurality of optical paths, wherein each optical path comprises on-chip delay lines; low loss waveguidecrossings or two optical waveguide layers, where each wavelength channel is successively routed to a second layer with a separate detector, while remaining wavelength channels are allowed to continue on a first layer till all wavelength channels are routed to the second layer.
[0029] In certain embodiments, the individual AWGs plus multi-layer routing plus multiport PDs system comprises: input matrix elements from the encoder fanned out into a plurality of optical paths, wherein each optical path comprises on-chip delay lines; and cascaded AWGs, where in each electro-optic modulator is followed by several AWGs each with a finer frequency selectivity.
[0030] In certain embodiments, the optical combining modules are selected from the group comprising an electrical combining module system, an optical combining module system, a vertical stacking module system, and a vertical plus electrical plus optical combining hierarchy module system.
[0031] In certain embodiments, the electrical combining module system comprises photocurrents produced by each processing core are electrically combined by summing the currents corresponding to each wavelength channel photodetectors.
[0032] In certain embodiments, the optical combining module system comprises at least one compute core followed by an input-encoder-like system driven by the analog signal from the wavelength detectors which optically broadcasts the results of the compute core.
[0033] In certain embodiments, the vertical stacking module system comprising compute cores are optically combined by vertically stacking compute cores; the products generated on each core are summed simultaneously using a free-space grating, an optical lens and a linear detector array.
[0034] In certain embodiments, the vertical plus electrical plus optical combininghierarchy module system comprises in hierarchical order: compute cores are optically combined by vertically stacking compute cores; the products generated on each core are summed simultaneously using a free-space grating, an optical lens and a linear detector array; photocurrents produced by each processing core are electrically combined by summing the currents corresponding wavelength channel photodetectors; and every compute core is followed with an input-encoder-like system which optically broadcasts the results of the compute core;
[0035] In certain embodiments, the system performs computations with positive and negative weights comprising two output ports are spatially separated by ensuring they occupy two separate polarization modes of the optical waveguides, wherein each wavelength channel is assigned two detectors which separately accumulate the computed products from the upper and lower ports of the weight modulator, respectively; and the difference of the two photocurrents produces the requisite summed signal.
[0036] In certain embodiments, the system performs computations with positive and negative weights comprising a first star coupler and a second star coupler, wherein the first star coupler has an optical path delay from an output of an input modulator, and wherein the second star coupler has a different optical path delay from a different output of the input modulator, and wherein changing the difference in delay results in tuning of the angle at which different wavelengths exit the arrayed grating.
[0037] Another aspect of the application is a system of offset diagonal computing comprising: an input vector per wavelength, wherein the input vector consists of a sequence of elements of an input optical stream; one or more input modulators per wavelength, wherein the one or more input modulators receive the input vector; one or more fan-outs, wherein the one or more fan-outs receive the input vector from the one or more input modulators, and wherein the one or more fan-outs divide the elements of the input vector into N total paths; N weightmodulators, wherein the one or more input modulators are connected to theweight modulators through the N total paths; a multiplication unit, wherein a matrix-matrix multiplication (MMM) computation is performed from a weight matrix decomposed into columns multiplied by the input matrix; one or more fan-ins, wherein the one or more fan-ins receive one or more elements of the input vector output from one or more of theweight modulators; and a detector per wavelength, wherein products are summed through time- synchronized incidence on the detector through one or more spatially distinct pathways from one or more of the N weight modulators, and / or wherein products are summed by binning one or more neighboring pulses, and wherein the temporal binning is achieved through capacitive summation.
[0038] In certain embodiments, the weight matrix is sub-divided into sets of columns and summation is performed simultaneously on the sets of columns of the weight matrix.
[0039] In certain embodiments, the weight matrix is subdivided into a plurality of sets of rows of elements of the weight matrix, and the sets of rows of elements of the weight matrix are summed simultaneously.
[0040] A further aspect of the application is a system of offset diagonal computing comprising: an input vector per wavelength, wherein the input vector consists of a sequence of elements of an input optical stream; one or more input modulators per wavelength, wherein the one or more input modulators receive the input vector; one or more fan-outs, wherein the one or more fan-outs receive the input vector from the one or more input modulators, and wherein the one or more fan-outs divide the elements of the input vector into N total paths; N weight modulators, wherein the one or more input modulators are connected to theweight modulators through the N total paths; a multiplication unit, wherein a matrix-matrix multiplication (MMM) computation is performed from a weight matrix decomposed into rowsmultiplied by the input matrix; one or more fan-ins, wherein the one or more fan-ins receive one or more elements of the input vector output from one or more of the N™ weight modulators; and a detector per wavelength, wherein products are summed through time-synchronized incidence on the detector through one or more spatially distinct pathways from one or more of the N weight modulators, and / or wherein products are summed by binning one or more neighboring pulses, and wherein the temporal binning is achieved through capacitive summation.
[0041] In certain embodiments, the weight matrix is sub-divided into sets of columns and summation is performed simultaneously on the sets of columns of the weight matrix.
[0042] In certain embodiments, the weight matrix is subdivided into a plurality of sets of rows of elements of the weight matrix, and the sets of rows of elements of the weight matrix are summed simultaneously.
[0043] In certain embodiments herein, there may be no binning, which means only one pulse is binned and no capacitor is needed.
[0044] In a further embodiment, an optical matrix multiplier includes an optical source configured to produce a sequence of optical pulses based on an input vector. The multiplier further includes a fanout module configured to receive a sequence of optical pulses and produce multiple optical pulse sequences. The multiplier further includes multiple delay lines, each delay line configured to apply a delay to an associated one of the optical pulse sequences. The multiplier further includes multiple modulators, each modulator configured to modulate one of the delayed optical pulse sequences. The multiplier further includes at least one accumulator configured to sum at least one modulated optical pulse sequence.
[0045] Implementations of the disclosure may include one or more of the following optional features. In some examples, the at least one accumulator includes a photodetector.configured to receive the modulated optical pulse sequence and produce an output. Said at least one accumulator may include a fan-in accumulator configured to receive each optical pulse sequence of the modulated optical pulse sequences and produce an output vector that includes sums of each individual optical pulse sequence. The accumulator may include an optical cavity accumulator. In some examples, each of the modulators are configured to modulate pulses of one of the delayed optical pulse sequences based on values of a weight matrix. In some examples, the optical source is configured to simultaneously produce multiple sequences of optical pulses, each of the sequences of optical pulses at a different wavelength, each of the sequences of optical pulses based on a different input vector. Each of the modulators may be configured to modulate each of the different wavelengths of the delayed optical pulse sequences. The at least one accumulator may include a wavelength-dependent fan-out configured to receive each of the different wavelengths of the modulated optical pulse sequence and direct each different wavelength to a different photodetector. In some examples, the optical matrix multiplier further includes an optical activation unit configured to receive an output of said at least one accumulator and produce an activation output that is a nonlinear function of the received output. The optical matrix multiplier may further include an optical memory coupled to the at least one accumulator and configured to store optical pulses. The optical memory may include an optical resonator.
[0046] In an embodiment, a modular computing system includes multiple optical matrix multipliers, each of the optical matrix multipliers configured to produce an output. The modular computing system further includes a combining module configured to receive each output of the optical matrix multipliers, and produce a combined output.
[0047] Implementations of the disclosure may include one or more of the following optional features. In some examples, the modular computing system further includes multiplecoarse delay lines configured to delay to the output of each optical matrix multiplier by a different coarse delay. The combining module may include an optical accumulator configured to sum the outputs of the optical matrix multipliers.
[0048] In an embodiment, an optical matrix multiplier includes an optical source configured to produce a sequence of optical pulses based on an input vector. The optical matrix multiplier further includes a fanout module configured to receive a sequence of optical pulses and produce multiple optical pulse sequences. The optical matrix multiplier further includes multiple modulators, each modulator configured to modulate one of the optical pulse sequences. The optical matrix multiplier further includes at least one accumulator configured to sum at least one modulated optical pulse sequence.
[0049] In an embodiment, a method of designing and operating a modular computing system includes determining a number of input modulators for use in the system, wherein each input modulator includes an input and an output. The method further includes determining a number of weight modulators for use in the system, wherein the weight modulators are organized in one or more sets. The method further includes determining a number of detectors (Nd) for use in the system. The method further includes configuring the system to divide the output of each input modulator in Nd paths exiting the input modulators, wherein each of the Nd paths is received by a distinct and separate set of weight modulators. The method further includes determining a maximum delay (Md) for the system, and further configure the system to divide each path into each set of weight modulators into Md+1 additional paths followed by optical delay lines from 0 to Md, wherein any path exiting any given input modulator becomes divided into Nd * (Md+1) paths, and wherein each path includes a unique delay line. The method further includes assigning a weight modulator to each path, routing optical paths exiting the weight modulators onto an assigned detector, and operating the system to produce an output.
[0050] In an embodiment, a method of designing and operating a modular computing system includes receiving a weight matrix that includes modulation values for an optical matrix multiplier. The method further includes identifying a permuted the weight matrix and corresponding permuted input vector associated with zero modulation values for at least one modulator. The method further includes omitting the at least one modulator associated with the zero modulation values and operating the system to produce an output.
[0051] In an embodiment, a method of optical matrix multiplication includes producing a sequence of optical pulses based on an input vector, receiving, by a fanout module, a sequence of optical pulses and producing, by the fanout module, multiple optical pulse sequences, modulating each sequence of optical pulses based on a weight matrix, and summing, by at least one accumulator, at least one modulated optical pulse sequence.
[0052] Implementations of the disclosure may include one or more of the following optional features. In some examples, the method further includes delaying each sequence of optical pulses by a different delay prior to modulating each sequence of optical pulses. Producing the sequence of optical pulses may include simultaneously producing multiple sequences of optical pulses, each of the sequences of optical pulses at a different wavelength, each of the sequences of optical pulses based on a different input vector. Modulating each sequence of optical pulses may include modulating each of the different wavelengths of the optical pulse sequences. In some examples, the method further includes receiving, by an optical activation unit, an output of said at least one accumulator, and producing, by the optical activation unit, an activation output that is a nonlinear function of the received output.BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The present application can be better understood by reference to the following drawings, wherein like references numerals represent like elements. The drawings are merelyexemplary to illustrate certain features that may be used singularly or in combination with other features and the present application should not be limited to the embodiments shown.
[0054] Fig. 1 shows a traditional spatial neural network.
[0055] Fig. 2 shows (Panel A) traditional matrix vector multiplication (MVM) computers multiply each row of the weight matrix with the input vector and combine the results for the output vector. (Panel B) In this application’s permuted diagonal approach, MVM computations can be understood as the sum of offset diagonals of the weight matrix multiplied by the input vector, which is processed physically by the equivalent multiplication of a diagonal matrix with the same elements and a cyclically shifted input vector. In a proposed scheme, computations are performed by accumulating products of diagonals with cyclically shifted copies of the input vector. 14 (Panel C) Schematic of a MMM processor within this scheme. Compatible structured matrices including 14 (Panel D) diagonally sparse, 14 (Panel E) offset sparse, and 14 (Panel F) block diagonal matrices.
[0056] Fig. 3 shows a time domain all-optical processor with an in-memory cavity accumulator.
[0057] Fig. 4 shows parallel processing combined with an in-memory accumulator for processing large models at high speeds.
[0058] Fig. 5 shows highly sparse matrix small-world networks are efficiently computed with this in-memory time-domain approach.
[0059] Fig. 6 shows faster processing with parallel modulators with accumulation through optical fan-in.
[0060] Fig. 7 shows the time-domain all-optical neural network processor (Panel A; overview) consists of (Panel B) a modulator and delay stack, (Panel C) an in-memory accumulator, (Panel D) an all-optical activation unit, and (Panel E) a Kerr-resonator regeneratorunit.
[0061] Fig. 8 shows: Full Schrodinger equation simulations of an all-optical fiber-cavity accumulator. (Panel A) A stream of 10-ps Gaussian pulses separated by 10 ns is incident into a fiber-loop accumulator. Summing of N= 1000 pulses with (Panel B) real -valued amplitude (relative error is inset) and (Panel C) complex -valued amplitude (phase of summed pulses is inset).
[0062] Fig. 9 shows: (Panel A) Picosecond pulses are generated by EOM laser modulation followed by dispersion compensating fiber. (Panel B) Optical spectrum (detector limited temporal response inset) of a 12-GHz, 25-ps EO pulsed source developed in Pi’s lab.
[0063] Fig. 10 shows: (Panel A) All-optical activation function implemented through a phase-preserving nonlinear optical loop mirror consisting of two optical loop mirrors with an intermediate phase conjugation stage. (Panel B) Numerically simulated amplitude and phase (inset) transfer function. (Panel C) Pulses at the input and output of the activation unit.
[0064] Fig. 11 shows: novel fiber Kerr resonators established by the PI including stretched pulse solitons with measured (Panel A) temporal autocorrelations and (Panel B) spectrum of the shortest pulses from a fiber Kerr resonator, and chirped pulse solitons with measured (Panel C) de-chirping and (Panel D) spectrum of the highest pulse energy from a Kerr resonator.
[0065] Fig. 12 shows: Soliton buffering and regeneration. (Panel A) Stable soliton solutions as a function of drive power and frequency detuning; (Panel B): Cavity-soliton pulse duration as a function of drive power; (Panels C-D): Soliton (Panel C) temporal envelope and (Panel D) spectrum for 100 mW drive power; (Panels E-F): Pulse cleaning from (Panel E) noisy pulses with a range of durations to pulses cleaned to within 1% of the stable soliton (Panel F) with only 20 round trips of propagation for a total of 0.18ms time cost.
[0066] Fig. 13 shows: (Panel A) Single layer all-optical classifier to validate operation with combined sub-systems. (Panel B) Exceptional phase locking for over 15 hours demonstrated to stabilize Kerr-resonator cavities. (Panel C) Repetition rate locking to Ken- resonator cavities demonstrated.
[0067] Fig. 14 shows (Panel A) a fan-in accumulation design; and (Panel C) an alternative fan-in accumulation design for matrix-matrix multiplication for which each input vector is encoded on a distinct wavelength channel.
[0068] FIG. 15 shows a temporal computing paradigm.
[0069] FIG. 16 shows a spatial computing paradigm.
[0070] FIG. 17 shows an embodiment of an optical computing design process.
[0071] FIG. 18 shows another embodiment of an optical computing design process.
[0072] FIG. 19 shows three embodiments of the optical computing design process highlighting design limitations.
[0073] FIG. 20 shows an embodiment of re-ordering of a diagonal matrix.
[0074] FIG. 21 shows a spatially distributed temporal computing paradigm.
[0075] FIG. 22 shows a temporally binned spatial computing paradigm.
[0076] FIG. 23 shows the simplest diagonal computing paradigm.
[0077] FIG. 24 shows a distributed diagonal computing paradigm.
[0078] FIG. 25 shows a temporally binned diagonal computing paradigm.
[0079] FIG. 26 shows a block-diagonal computing paradigm.
[0080] FIG. 27 shows a spatial hyper-multiplexed computing paradigm.
[0081] FIG. 28 shows an arbitrary blocked computing paradigm.
[0082] FIGs. 29A-G show experimental results. Fig. 29A - Time trace of summing computations with 500-symbols at 5 GFLOPs. Time trace of two diagonal MVMcomputations with 500-symbols at 200 MFLOPs and two channels with Fig. 29B (1-symbol delay) and Fig. 29C (5-symbol delay) between the two optical paths. Fig. 29D - Experimental apparatus. Fig. 29E - Histogram of observed compute error with 1000 uniformly distributed random symbols. Confusion matrix of the classification result for 500 MNIST-digit images implemented on a Fig. 29F - digital processor (90 %) and the 29G offset diagonal optical processor (87 %).
[0083] FIG. 30 shows batching under the computing paradigm described herein with two batches for a 8x8 matrix with 4 available compute modulators.
[0084] FIG. 31 shows embodiments of a general computing architecture comprising three fully modular subsystems: the input matrix encoder, compute cores each of which are proceeded by well-calibrated fiber length to perform coarse delays, and an optical combining module.
[0085] FIGs. 32A-D show systems with input matrix encoders. Fig. 32A -an embodiment of an integrated comb system. Fig. 32B an embodiment of an N channel laser bank plus wavelength mix plus MRRs system. Fig. 32C - an embodiment of an N channel laser bank plus electro-optic modulators system. Fig. 32D - an embodiment of an electro-optic RF comb plus MRRs system.
[0086] FIGs. 33A-D shows systems with processing cores. Fig. 33A - an embodiment of an integrated AWG output system. Fig. 33B - an embodiment of a bulk output system. Fig. 33C - an embodiment of an individual AWGs plus two-layer routing plus multi-port PDs system. Fig 33D - an embodiment of an individual AWGs plus multi-layer routing plus multiport PDs system.
[0087] FIGs. 34A-D show systems combining modules. Fig. 34A - an embodiment of an electrical combining module system. Fig. 34B - an embodiment of an optical combiningmodule system. Fig. 34C - an embodiment of a vertical stacking module system. Fig. 34D - an embodiment of a vertical plus electrical plus optical combining hierarchy module system.
[0088] FIG. 35 shows an embodiment of a system to perform computations with positive and negative weights by using dual port electro-optic weight modulators
[0089] While the present disclosure will now be described in detail, and it is done so in connection with the illustrative embodiments, it is not limited by the particular embodiments illustrated in the figures and the appended claims.DETAILED DESCRIPTION
[0090] Reference will be made in detail to certain aspects and exemplary embodiments of the application, illustrating examples in the accompanying structures and figures. The aspects of the application will be described in conjunction with the exemplary embodiments, including methods, materials and examples, such description is non-limiting and the scope of the application is intended to encompass all equivalents, alternatives, and modifications, either generally known, or incorporated here. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. One of skill in the art will recognize many techniques and materials similar or equivalent to those described here, which could be used in the practice of the aspects and embodiments of the present application. The described aspects and embodiments of the application are not limited to the methods and materials described.I. Theory of Permuted / Offset Diagonal Computing
[0091] This application describes optical artificial neural networks (ANNs) that leverage the high speed, high bandwidth, and low loss properties of light by multiplexing information in the time domain, in addition to the spatial domain with a permuted diagonal computing algorithm. In this way >billion-scale model sizes can be processed in real-time with few spatialelements, bypassing the limits of current optical computing architectures with slow reconfigurability and poor spatial scaling.
[0092] A single layer neural network maps N input neurons to N output neurons with matrix-vector multiplication (MVM) with an NxN weight-matrix through multiplication and accumulation operations ( -»= F), followed by a nonlinear activation operation (Fig. 1).
[0093] In contrast to traditional spatially distinct neurons, in this application, input and output neurons are encoded temporally in the amplitudes of a sequence of optical pulses. Multiplication and accumulation operations are implemented in time through optical modulation and addition of pulses by temporal superposition, respectively.
[0094] In the traditional spatial approach, MVMs are computed by multiplying each element of a row of the weight matrix with the corresponding element of the input vector and subsequently accumulating the products, using a distinct multiply-accumulate (MAC) unit for each row of the weight matrix (Fig. 2, Panel A). In this application’s paradigm, by contrast, MVMs are computed by decomposing the weight matrix into permuted diagonals which are multiplied by a cyclically shifted copy of the input vector with the product vectors subsequently accumulated (Fig. 2, Panel B). In particular, MVM is performed by multiplying a cyclically shifted copy of the input vector with a corresponding offset diagonal of the matrix with a distinct processing unit. Cyclic shifts can be performed with simple passive optical delay lines following optical fan-out, with each delay-line shifting a copy of the input vector by one symbol period. The partial product vectors can be accumulated spatially on photodetectors which serially read out the output vector. Matrix-matrix-multiplications (MMM) are enabled through straightforward wavelength-division multiplexing (Fig. 14, Panel C). Here for NxN MMM, N receivers are required, reducing the component count from prior approaches by a factor of N. This approach also supports maximally efficient computation of sparse matrices such as diagonalsparsity by omitting modulators corresponding to zero-values (Fig. 14D), offset diagonal sparsity by reordering the input vector (Fig. 14E), and block diagonal matrices by breaking up the input and output vectors (Fig. 14F). Zero modulation values are any values sufficiently small that the modulator is effectively idle. That is, the modulation value is so small that applying the value would make no appreciable difference to the pulse. Therefore, the modulator is unnecessary and can be physically excluded, rendered inoperable, repurposed, or otherwise omitted from the system.Permuted Diagonal Computing with All-optical In-memory accumulation
[0095] In one embodiment, a fully connected N N neural layer can be implemented with one modulating element by concatenating N cyclically shifted copies of the input vector, modulating each cyclically shifted copy with the corresponding elements of the appropriate diagonal of the weight matrix, and coherently adding the modulated vector copies within an optical cavity with a round-trip time matching the temporal length of the input vector (Fig. 3). Accumulation and cyclic shifting are performed efficiently with a cavity with a round-trip time that is one more pulse period longer than the vector length to enable the ith element to temporally overlap with (i + l)th element in a successive copy. With this form of in-memory accumulation, 107sized vectors corresponding to a 100 trillion-scale model can be readily processed with ultrashort pulses and mature low-loss fiber technology. Processing speeds beyond -100GHz bandwidth are obtained by distributing computations over several modulators.Permuted Diagonal Computing with Fan-in plus In-memory Accumulation
[0096] In another embodiment, an additional in-memory accumulator can be added to enable large vector sizes as well the highest computation speeds (Fig. 4). In this case, the fan-in performs part of the accumulations, with the remaining performed by the cavity accumulator.
[0097] In a further embodiment, the modulators can be arranged for time multiplexingwith a similar speed enhancement.
[0098] The total processing speed in optical processors (Ops) is twice the product of the number of modulators times their operation bandwidth. Notably with commercial modulators operating at 10s of GHz bandwidth with the several watts of power required to drive an efficient laser and erbium-doped fiber amplifiers, the basic compute efficiency can remarkably surpass the performance of the top supercomputers for memory -bandwidth-limited matrix-vector calculations. Moreover, the performance is significantly higher with matrix-matrix computations with large matrices which require multiple processors (GPUS for example) because of interprocessor communication latency.Permuted Diagonal Computing with Sparse Matrix Application
[0099] In an additional embodiment, as well as the described advantages regarding large computation sizes, speeds and efficiency, this platform has unique advantages for computing sparse and complex-valued problems. In-memory accumulation under permuted diagonal computing enables efficient computation of extremely sparse small-world type neural networks by simply selectively removing modulators, or by simply not using (e.g., omitting) those modulators, corresponding to specific diagonals (Fig. 5) and using reconfigurable delay lines.
[0100] In contrast, traditional architectures process the weight matrix by rows and require significant overhead in memory buffering while processing the same sparse matrices (see e.g., Fig. 2, Panel A).Permuted Diagonal Computing with Complex Value Application
[0101] A key advantage of complex value networks (CVNs) is that they allow study of leaming / training algorithms such as Spike-Timing Dependent Plasticity (STDP), which is believed to be the prominent form of synaptic change in neurons. For training more generally, while backpropagation techniques require significant computational resources and energy,because the system described herein in certain embodiments is all-optical, physical, and, in principle, reversible, this enables investigation of more advanced, energy -efficient backpropagation techniques such as physics-inspired hybrid b ackpropagation and Hamiltonian echo-backpropagation, with the potential for dramatically reduced energy costs.Offset / Permuted Diagonal Computing with Fan-in Accumulation
[0102] In an alternative embodiment, multiplications can be parallelized by fanning-out the input vector, cyclically shifting it, multiplying it with the corresponding diagonals and accumulating the products using an optical fan-in (Fig. 14a). Incredible speeds (e.g. >100TFlops) can be achieved with many integrated intensity modulators on a photonic integrated circuit (PIC) using present-day technology. This fan-in operation can be achieved with directional couplers or with electro-optical detectors.
[0103] In one embodiment, offset diagonal computing accumulation operations are achieved by routing multiple optical waveguides in close proximity, but without overlapping for simultaneous detection, on a common photodiode. In this embodiment, the detector is only measuring the total power in each waveguide.
[0104] Since coherence is not required the optical source requirements are relaxed. Coherence can also be exploited for complex valued operation by using a heterodyne detector during fan-in with a local oscillator at a distinct frequency across the spatial extent of the detector.
[0105] This architecture is represented by certain embodiments shown in Figs. 4, 7 (Panel A), and 14 (Panel A, Panel B, Panel C). The input vector stems from an optical source that is modulated with the input vector data, drawn from electrical memory (from the left, not shown in the figures). The light is fanned out with any standard coupler (no loss here) into the n waveguides, where n is the number of modulators. The speed of the processor is then given by ntimes the bandwidth of the modulators (for MVM). Each path requires a delay line of increasing delay in steps of one pulse width (determined by the modulator bandwidth).
[0106] The delay lines can be implemented on chip (silicon or silicon nitride, e.g., fiber, or even free space).
[0107] In Fig. 4, in a specific embodiment, multiplications are parallelized by fanning- out the input vector, cyclically shifting it, multiplying it with the corresponding diagonals and accumulating the products using an optical fan-in. Incredible speeds (e.g. >100TFlops) can be achieved with many integrated intensity modulators on a photonic integrated circuit (PIC) using present-day technology. This fan-in operation can be achieved with a single electro-optical detector (Fig. 14a).
[0108] In Fig. 7 (Panel A), in a particular embodiment, in the first element of a processor, periodically produced copies of the input vector pulse train are fanned out along N paths into a stack of N-l delay lines and electro-optic modulators. The combination of delay lines and electro-optic modulators implement both cyclic shifting and multiplication.
[0109] Multiplication to encode the vector elements onto an input pulse stream or with the desired matrix weights are implemented by amplitude modulation with an electro-optic modulator. A push-pull-type Mach-Zehnder intensity modulator driven with an arbitrary waveform generator encoding real-valued weights multiplies the pulse train carrying the vector. Complex weights with amplitude and phase both represented can be encoded with IQ electrooptic modulators. The modulator multiplier is validated with a normalized mode-locked laser with a calibrated fast photodiode for analysis in comparison to ideal outputs. For complex modulation, IQ modulation is validated with homodyne detection for phase and amplitude analysis.
[0110] As an alternative to delay lines with very long lengths, in certain embodiments apath can be seeded with the vector offset by the required number of delay steps at the cost of this additional memory access and modulator.
[0111] Even if in certain embodiments each path had its own modulator and memory access the power would at most be doubled from the power consumption of the entire system because the memory accessing would still be dominated by the weight matrix.
[0112] Finally, if in certain embodiments the number of modulators is fewer than the number of row diagonals in the matrix, the detector can be used as an integrator over x time steps, where x is the number of diagonals in the matrix divided by the number of modulators. In this case, the matrix to calculate is divided into banded diagonals where the thickness of the band is given by x. For example, if there were only two modulators available, the matrix is divided equally into two thick diagonals.
[0113] In one embodiment, a simple approach to the fan-in accumulation is to have all the paths come close and go through a lens which focuses the beams on one detector which adds all the intensities. In an alternative embodiment, this can also be achieved by muxing the single mode waveguides into one big multimode waveguide before hitting the detector. The detected field is translated to current with the diode which then goes to an analog to digital conversion before going to memory.
[0114] In certain embodiments, neural networks need a nonlinear activation which in this case can be achieved digitally in a standard way or, in an alternative embodiment, using the detector nonlinearity itself which is also standard. The next layer is then achieved by modulating a laser with the output vector that was stored in memory from the previous layer. These modulated lasers can also be replaced by modulated VCSELs.
[0115] In Fig. 14 (Panel C), in particular embodiments, matrix-matrix multiplication is achieved through a straightforward extension of this system. This is achieved by using a differentwavelength of light for each input vector and going through the same system. For example, in a specific embodiment, an on chip optical frequency comb can produce a broadband of 10- 100GHz separated phase-locked lasers that can each be modulated with ring or on-chip modulators with vectors corresponding to different columns of the input matrix. These can be wavelength-division multiplexed together and run through the same system described for matrixvector multiplication. Here, each wavelength is detected on a different detector which can be achieved with the wavelength spatially separated using a diffraction grating or, alternatively, wdms on chip. The output matrix is again stored in digital memory for activation before being sent to the next layer through the multi -wavelength modulators.
[0116] In light of the description herein, one of ordinary skill will understand that the spatial, temporal, and wavelength division multiplexed computing system described herein can be designed with delay -based computing in many different ways which have different practical advantages depending on what is optimized for. Given a problem size, optimization can be performed across multiple practical system metrics, such as component counts, fabrication complexity, alignment complexity, or power.
[0117] Without delays, optical computing paradigms can leverage exclusively temporal or spatial degrees of freedom (Fig. 15 and Fig. 16). In optical computing paradigms utilizing temporal integration, the weight matrix can be visualized as being broken up into rows. The input vector is optically fanned out and each optical processing element, i.e., an electro-optic modulator (EOM), multiplies a copy of the vector with each row fed into it and then subsequently summed within a detector (Fig. 15). In contrast, in the spatial paradigm (Fig. 16), the matrix is decomposed into columns, with each EOM receiving a column of the weight matrix which is multiplied by one element of the input vector. The products are subsequently summed through time-synchronized incidence on a detector.
[0118] Scaling these systems to even modest computation problems would require prohibitively large numbers of components. Additionally, neither of these frameworks are amenable to computation of sparse matrices. Herein, the study shows that these systems are in fact reductions of a much larger space of spatio-temporal -delay based optical computing systems. The addition of a third degree of freedom, namely delays, can be combined with spatial and temporal degrees of freedom to realize systems which can be tailored to nearly any dense or sparse matrix-vector multiplication (sparse-MVM) problem, as well as extended to MMM through wavelength multiplexing. Finally, the study herein develops a simple mathematical framework to describe all possible systems within this expanded design space and the resulting physical parameters.Scaling Laws
[0119] The mathematical framework underlying the systems presented herein for NxN matrices, is characterized by three independent numbers:
[0120] Horizontal subdivisions (h);
[0121] vertical subdivisions (v); and
[0122] temporal summing bin or temporal clump (c).
[0123] The horizontal and vertical subdivisions refer to the subdivisions of the matrix through horizontal and vertical cuts respectively. While these divisions are assumed to be made at equal intervals here, these can be trivially generalized for unequally spaced cuts. Horizontal and vertical cuts can be shown to be equivalent to subdividing the output or input vector.
[0124] Temporal summing bin or temporal clump refers to the number of elements, which are summed together through temporal summing, as described earlier. Below some embodiments are described for the (h,v,c) parameters for some configurations.
[0125] Scaling law equations
[0126] # of Modulators mods=Nh / c
[0127] # of detectors dets=h
[0128] # of beams / det beams=N / c
[0129] # of input mods inputs=v
[0130] Max delay required delay=N / v-c
[0131] For any given physical constraint, for e.g., max modulators or maximum allowable delay corresponding solutions (h,v,c) can be computed. This allows near arbitrary design of system-conforming physically relevant system parameters.
[0132] Constraints on Computational systems: the study also derives two constraints that any computational system within this framework must satisfy for optimal compute.
[0133] Optimal compute here is defined as the case where the total compute time, Te, is equal to the number of MACs performed divided by the combined speed of all the weight modulators, i.e. Tc=(# of MACs (multiply-accumulate operations)) / (mods * R), mods is the number of modulators in the system, R is the speed of each modulator, and # of MACs refers to the total number of MACs performed for computing the output vector. Non efficient compute then would refer to the case where the compute takes longer because weight modulators have to wait to receive appropriate vector elements.
[0134] The above requirement in addition to physical constraints, such as the physically realizable delays must be delay >0, results in a succinct equation constraining all the systems coordinates, namely, N, h, v, and c.
[0135] Constraint equation: h / v < c < N / v
[0136] To understand limitations faced by an inefficient compute system, consider the case of a system with parameters (n,h,v,c) = (8, 2, 1,1). An implementation of this system is shown in Fig. 19. Note that, this system does not satisfy the efficient compute condition, h / v < c.
[0137] Consequently, while each modulator receives all the input vector elements, it only multiplies a fraction of those elements (n / h) with the elements of the weight matrix. This results in the weight modulators waiting for excess time before performing requisite operations. The effective speed of this computing system is reduced by a factor 1 / h. Despite having sixteen modulators, this system only offers effective speed of 8R, where R is the rate of each modulator.
[0138] This system can be optimized by setting a value of v or c to satisfy the inequality condition. For example, one can change the value of v from 1 to 2, i.e. (8,2, 1,1) —> (8,2,2, 1), as shown in Fig. 17, in order to satisfy the efficient compute condition. The corresponding implementation offers an effective speed of 16R with smaller delays, at the cost of an additional input modulator.
[0139] Alternatively, one could also change c from 1 to 2, i.e. (8,2, 1,1) — (8,2, 1,2) to achieve efficient compute. In this case, the number of weight modulators is reduced to 8 from 16 while offering full effective speed of 8R. These concepts are illustrated in Fig. 19.Offset variations
[0140] The presented framework of diagonal compute is shown to span a significantly large space of matrices. This is achieved by noting that by offsetting the elements of the input vector, i .e. rearranging the order of the elements, and correspondingly rearranging the elements of the weight matrix, does not incur any additional computational penalty.
[0141] In fact, all possible offsets of the input vector given by N!, where N is the length of the input vector, are equivalent computations for this paradigm. The computation in these cases is performed by reordering the input vector elements that are input into the optical processor.
[0142] An example of a variation of a diagonal is shown in Fig. 20. While the new(permuted) matrix still has one element in each row and each column, they are arranged in adifferent order. Likewise, this input rearrangement method can be applied to any of the other variations described using the same components and at no computational cost. Only the order of the elements fed into the input modulators change, which is simple.
[0143] These offsets do not change the scaling laws. However, they enable new types of structured sparsity which may benefit and sparse computations by allowing for more efficient configurations.Designing a System for a Given Set of Coordinates
[0144] In this section, this application describes an example protocol for designing an optical computing system given the system parameters, (n,h,v,c). The only assumption the study makes while describing this recipe is that the system parameters conform to the rules for efficient compute, h / v < c < N / v. Recall the scaling laws underlying this compute paradigm:
[0145] Number of weight modulators,= Nh / c.
[0146] Number of optical detectors, Nd=h
[0147] Number of incident optical beams, Nb=N / c
[0148] Number of input modulators, Nml= vN
[0149] Maximum optical delay required, Md= - — c
[0150] Step 1 : Determine the number of input modulators,= v. Each input modulator will receive N / v elements of the input vectorNh
[0151] Step 2: Determine the number of weight modulators, N" = — . Assuming the weight matrix is divided into equal sub matrices, each sub matrix will be assigned an equal number of weight modulators.
[0152] Step 3: Determine the number of detectors, Nd=h. Next, assign equal number of weight modulators to each detector, lets call these modulators sets SlfS2> ... , SNd.
[0153] Step 4: Divide the output of each input modulator into Nd paths, with each path going to each set of weight modulators (S „ S2, ... , 5v...)- Next, determine the max delay, Ma, and further divide each path coming into each set into Md+1 additional paths followed by optical delay lines from 0 to Ma. By the end of this step, optical path exiting any one input mod would have been divided into Nd(Ma+l) paths, with each path having a unique optical delay line following it.
[0154] Step 5: Assign a weight modulator to each path coming into each modulator set
[0155] Step 6: After multiplication with weights, the optical paths exiting each modulator set should be routed onto its assigned detector. In the case of c values larger than 1, that is multiple products are needed to temporally summed the output vectors are obtained after waiting for the appropriate products to be accumulated. For example, for c=10, one must wait for 10 products to be temporally accumulated before reading out the requisite detector signal.
[0156] The system is now routed for optimal computing for the given design parameters (n,h,v,c)Possible components to realize different subsystems
[0157] Weight Multiplications: In various embodiments, weights / vectors can be encoded onto optical fields using any system which effectively changes the output optical power of the device as a function of an electrical input. In certain embodiments, this includes ultra-high-speed components such as Lithium niobate or silicon modulators, which are desirable given their fast response. However, in alternative embodiments, technologies including Spatial light modulators (SLMs), Absorptive modulators, Micro-ring resonator modulators, spectral waveshapers, optical switches (for 0 / 1) are all equally compatible with the described platform.
[0158] Optical Delay Lines: In various embodiments, implementation of time delay between optical paths can be implemented with standard optical waveguides. Lower losses andhigher-refractive index media are preferred to improve compute fidelity and reduce the required length of delay line. In exemplary embodiments, SiN-based waveguides, Si-photonic delay lines, free-space delays, and optical fiber delay lines are all compatible with the platforms discussed herein.
[0159] Wavelength multiplexer / Demultiplexer: Matrix-Matrix multiplications are performed by encoding multiple vectors onto different wavelength channels. In various embodiments, any system with the capability to combine / split optical wavelengths is compatible with the computing paradigm, including, but not limited to, free space components such as optical prisms, optical diffraction-gratings, or metamaterial surfaces or on-chip components, such as Arrayed waveguide gratings (AWG) or on-chip Bragg-gratings.
[0160] Spatial summing detectors: Spatially summing signals on detectors can be performed on any optical photodetector with the capability to accept multiple optical inputs. In various embodiments, this includes free-space detectors with optical beams incident at different angles, on-chip multi-port photodetectors, which are detectors which accept inputs from multiple optical waveguides, and multi-mode photodiodes, which accept inputs from multiple spatial modes of an optical waveguide.Batched Matrix-Vector (or -Matrix) Multiplication with limited modulators
[0161] Contemporary Machine-Learning and Al models have Billions of parameters and require trillions of multiply and accumulate operations. Typically, these are distributed over several compute cores which parallelly compute different parts of the problem. The final result is constructed by combining the outputs of all compute core. However, in certain scenarios sufficient compute cores may not be available. Here the study briefly describes computations in the scenario where the number of available modulators is less than the size of computation problem. In this situation the problem can be computed using binning as described above, or itcan be computed in batches. Each batch is computed and finally the computed batches are combined to obtain the final result.
[0162] Within this batched computing framework, calculations are batched by dividing the matrix with vertical cuts (the v-parameter described previously). Given that each available weight modulator performs computations with one offset diagonal, the number of vertical cuts required to complete the computation would depend on the number of offset diagonals that can be processed simultaneously which in turn is dictated by available modulators. Each submatrix will be computed, and the results can then be electronically summed. A pictorial illustration of this batching is shown in Fig. 30 with two batches for a 8x8 matrix with 4 available compute modulators.Modular Optical Compute Systems
[0163] A significant advantage of the architecture described herein is its inherent modularity. A general computing architecture of large-scale processor comprising of several compute cores is shown in Fig. 31. It consists of three fully modular subsystems: the input matrix encoder, which optically encodes the elements of the input matrix, compute cores each of which are proceeded by well-calibrated fiber length to perform coarse delays, and finally an optical combining module which combines the output of the compute cores to either compute the output vector or transmit the result of local computations to a subsequent combining module.II. Temporal All-optical Diagonal Computing Platform
[0164] In one embodiment, the all-optical computing platform builds off techniques developed for time-multiplexed fiber optical communication systems and leverages quantumlimited optical amplifiers in the telecommunication spectral band. Input and output neurons are encoded temporally in the amplitudes of a sequence of pulses with identical temporal envelopes. The optical processer consists of a multiplication unit, an optical in-memory accumulator, anoptical activation unit, and an optical regenerator system (Fig. 7, Panel A (overview)).Cyclic delay and multiplier
[0165] In the first element of the processor (Fig. 7, Panel B) periodically produced copies of the input vector pulse train are fanned out into a stack of M delay lines and electro-optic modulators. The combination of delay lines and electro-optic modulators implement both cyclic shifting and multiplication.
[0166] Multiplication to encode the vector elements onto an input pulse stream or with the desired matrix weights are implemented by amplitude modulation with an electro-optic modulator. A push-pull-type Mach-Zehnder intensity modulator driven with an arbitrary waveform generator encoding real-valued weights multiplies the pulse train carrying the vector. Complex weights with amplitude and phase both represented can be encoded with IQ electrooptic modulators. The modulator multiplier is validated with a normalized mode-locked laser with a calibrated fast photodiode for analysis in comparison to ideal outputs. For complex modulation, IQ modulation is validated with homodyne detection for phase and amplitude analysis.In-memory accumulator
[0167] The subsequent fan-in operation of the modulator outputs calculates partial accumulations (Fig. 7, Panel C), which are then incident into an optional optical accumulator, which is a recirculating fiber loop with an erbium gain module to compensate for cavity losses. The loop accumulates partial sums until summation is complete, and the accumulated pulses are retrieved from the loop with an optical switch.
[0168] Accumulation operations in the platform are achieved within a recirculating length-controlled fiber-loop with an input coupler and an erbium gain segment. By adjusting the total round-trip time of the cavity to match the vector length, pulses in subsequent vector copiesenter the cavity temporally aligned, resulting in coherent addition (Fig. 8, Panel A). Coherent addition through cavities has been applied previously for pulse shaping and amplification, but without arbitrary data patterns. The erbium gain in the unsaturated regime provides linear gain (^=1 / ^2) to compensate for the loss from the 3-dB input coupler. To evaluate and manage the dispersive and nonlinear effects of the fiber, the accumulator is simulated by solving nonlinear pulse propagation equations (Agrawal, G.P. Nonlinear fiber optics. (Academic press. 2007)) for 10-ps Gaussian pulses with a repetition period of T= 10 ns. Accumulation simulations with real amplitudes — 1 < Aj < 1 (Fig. 8, Panel B), and complex amplitudes | 4j | < 1 (Fig. 8, Panel C), show excellent agreement with the ideal summation operation, with relative errors ~1% (inset Fig. 8, Panel B). Nonlinear effects are negligible for nonlinear phase accumulation less than ~1%, and dispersive effects for pulses shorter than 500 fs can be suppressed through standard dispersion management techniques.
[0169] The in-memory accumulator is characterized with the GHz-electro-optic pulse source generated through strong cascaded amplitude and phase modulation of a CW -carrier followed by a segment of dispersion-compensating fiber (Fig. 9, Panel A). With this source high repetition rates and small pulse widths can be obtained with off-the shelf components as demonstrated by a 12-GHz, 25-ps source (Fig. 9, Panel B). As implemented in stabilized fiber Kerr resonator systems (Dong et al., Opt. Lett. 47, 4443 (2022); Dong et al., Phys. Rev. Lett. 125, 33902 (2020); Dong et al., Phys. Rev. Res. 3, (2021); Spiess, C. et al., Optica 8, 861 (2021); Dong et al., J. Opt. Soc. Am. B. 40, 3255-3261 (2023)), a proportional-integral-derivative (PID) system is implemented to control the round-trip phase of the loop relative to the laser, and another stabilizes the delay of the input pulse train relative to the round-trip time of the loop.
[0170] The in-line 976-nm pumped erbium amplifier operates with 1-5 dB small-signal gain, which also minimizes noise. A fast (< 100 ns) off-the-shelf optical switch ejects theaccumulated vector from the cavity to be measured through fast detection and evaluated against the ideal case and full numerical models as a function of pulse energy, duration, and cavity length, phase and group delay mismatch. This analysis is then repeated following real and complex-valued multiplication from the previous stage.All-optical activation
[0171] Finally, the pulses are incident into the all-optical nonlinear activation unit. In certain embodiments, effective all-optical activation units may include a / (3)-based fiber optical loop mirror with cubic activation, / (2)-based on-chip ReLu-like activation, and microring resonator-based sigmoid activation (Fig. 7, Panel D). One of ordinary skill in the art will understand that the type of all-optical activation unit is not limiting on the system or method described herein.
[0172] In a particular embodiment, all-optical activation can be achieved with a nonlinear optical fiber-loop mirror (NOLM) (Fig. 10), through which two counter-propagating pulses of unequal amplitude accumulate disparate nonlinear phase in a fiber loop and interfere at the output couple to yield a power dependent output response. For real -valued computations, a phase conjugation stage is applied to ensure the phase independence of the output. Full-numerical simulations of phase-preserving nonlinear activation consisting of two identical NOLMs with an intermediate phase conjugation stage (Fig. 10, Panel A) reveals a nonlinear amplitude response (Fig. 10, Panel B) with weak output phase over the range of operation. Functionality is demonstrated by measuring the nonlinear transfer function of a single NOLM and the phasepreserving system with commercial highly nonlinear fiber lengths in the range LL = 10 - 50 mm, operational peak powers of 10-100 W and pulse widths of 10-100 ps, using standard amplifiers.The phase transfer function of the devices is measured with homodyne detection. The output of this effort is a single and phase preserving NOLM based activation system optimized for fiberlengths, couplers, threshold and transfer function.Ultrashort-pulse fiber Kerr resonators for all-optical buffering and pulse regeneration
[0173] In certain embodiments, the resulting pulses representing the output of the layer can then be cleaned and regenerated in the Kerr-resonator memory for use in subsequent layers (Fig. 7, Panel E).
[0174] To mitigate noise introduced through operation from components including loss and amplification, all-optical memory and signal cleaning techniques are developed for realizing deep and recurrent neural networks. Kerr cavities support stable solitons that can be used to store information in pulses arranged around the cavity in time. The first Kerr solitons enabled storage of 15 optical bits accessed at a bit rate of lOMb / s, which with improvement to the locking circuit increased to 4536 bits of data operating at a rate of 10 Gb / s. The temporal slots are created by a simple sinusoidal phase modulation of the driving field which traps the pulses through temporal tweezing. Data can be written using a cross-phase modulation addressing beam, as well either by intensity or phase modulating the drive pulse. The accessing speed of prior results was fundamentally limited by the duration of the soliton, which was 2.6 ps. Much shorter pulses can be generated which would enable >lTb / s operation rates. The shortest pulses to date are provided by stretched-pulse Kerr resonators, a type of soliton that stretches and compresses periodically in the fiber cavity (Fig. 11, Panels A-B).
[0175] By carefully pushing the design and scaling laws 120-fs pulses are stable with excellent agreement with numerical predictions (Fig. 11, Panels A-B). Also, ‘chirped pulse’ solitons in normal dispersion cavities, which broaden available platforms, increase pulse energies and conversion efficiencies, and enables pulse generation in lossy cavities (Fig. 11, Panels C-D). These approaches provide an optical memory surpassing the performance of even the fastestmemories (optical or not) available today.
[0176] Buffer performance is characterized by speed, size and power. The speed is determined by the temporal duration of the pulse and the memory size is determined by the duration and the cavity length. To maximize memory size per unit drive power the soliton scaling laws are analyzed.
[0177] In certain embodiments, for standard single-mode fiber and intrinsic loss (-10%) coefficients, the optimum cavity length is approximately 10km. Full numerical simulations are used to determine the performance of a memory using smf-28 optical fiber with 18 ps / (nm*km) dispersion, 1.3 (W*km)_1nonlinearity and 0.18 dB / km loss at an operation wavelength of 1550.1 nm. Best performance is found with a 10km cavity with critical coupling (41%) with minimum external drive power of 26 mW and a soliton with 17.6 ps duration. The pulse is relatively long because of the large dispersion associated with the long cavity, which would limit the operation rate (the inverse of -3 times the soliton duration) to 19Gbits / s.
[0178] Higher operation rates require shorter solitons which are enabled by cavities with smaller dispersion. In a particular embodiment, a highly non-linear optical fiber (HNLF)-Zero slope fiber (OFSOptics) with small group-velocity dispersion (GVD) and a small, 0.006 ps / (nm2*km), dispersion slope near 1550nm that is well suited for this goal. To determine best performance, the wavelength (and therefore dispersion) is swept through the zero dispersion point, and optimum performance is found with 0.05 ps / (nm*km) dispersion with 1.8-km fiber and a 5% coupler (critical coupling). Stable solutions are found for a wide range of drive and detuning values (Fig. 12, Panel A), with the pulse duration varying smoothly with drive (Fig. 12, Panel B).
[0179] In one embodiment, with a 30-mW drive, the pulse envelope in time (Fig. 12, Panel C) and spectrum (Fig. 12, Panel D) are clean with very low background on the spectrum.Performance can be calculated using the cavity round trip time over 3 times the pulse duration for storage sizes. For the shortest, 320-fs pulses, a 100-mW memory of 10 Mbit is accessible at 1 Tbits / s, for an efficiency of 10 Tbit / s / W. Using 526-fs solitons with 30 mW enables a 30-mW memory of 6 Mbit accessible at 0.63 Tbits / s, for an incredible accessing efficiency of 21 Tbit / s / W.
[0180] Currently, high bandwidth memory (HBMs) dominate GPUs used for training ever increasing large models for Al applications and contribute to today’s incredible power cost problem (e.g. 1064MWh of energy to train on V100 GPUs). The newest HBMs are estimated to operate somewhere from 560 Gbit / s / W (assuming 5W and 2.8Tb / s from the partial specifications) to 206 Gbit / s / W (assuming 15.9W and 3.3Tb / s from full simulations of HBM2e using AMD Power Design Manager software).
[0181] The Fig. 12 results represent a >100x improvement compared to HBM in accessing efficiency for Mbit-scale optical memories. Larger fiber memories could be developed with the same record efficiency but at the cost of a comparable increasing in fiber length, which is impractical after the Gbit scale without future innovations in exploiting spatial degrees of freedom in fiber such as multiple cores or spatial modes. Soliton optical memories are directly applicable for constrained power budget applications such as embedded Al at the edge for training neural networks with kb-scale memories.
[0182] Finally, a memory that stores data in the optical domain reduces power consumption and latency at interconnection points while free of signal loss and interference. Most other optical memories have been demonstrated with few or single bit performance and scaling to Mbit sizes represents a daunting technical challenge, for example, with the requirement of a million lasers. An all-optical memory with a combination of speed and power surpassing currently available technology as described herein, which in addition has a sizecapacity much larger than recently demonstrated optical memories has a major significance.
[0183] For this application’s platform, memory has to accommodate the length of the vector, meaning that a 10Mbit memory is used to support operation of a brain-scale 100 trillion parameter model, the size of a fully connected model with a 10 million length vector.
[0184] Moreover, in certain embodiments, periodic cleaning of accumulated noise is all that is needed for stable long-term operation. The time required to clean pulses using Kerr solitons is determined through the simulations of the fiber cavities described above. Initial pulses with varying FWHM duration and 1% Gaussian distributed noise are input into the cavity to determine how many round trips it would take until the pulses are within 1% of the clean soliton solution. For the 1.8 km cavity 200 fs to 500 fs noisy pulses are cleaned to within 1% of the 324 fs soliton in less than 20 round trips, which is only 0.18 ms in time (Fig. 12, Panels E-F). The all- optical computer cycle time is proportional to the square of the number of elements in the vector, or for comparison, the number of round trips required for copying the vector is given by the memory size, so the cleaning time is on the order of a millionth of the computing cycle time and therefore negligible.
[0185] The present application is further illustrated by the following examples that should not be construed as limiting. The contents of all references, patents, and published patent applications cited throughout this application, as well as the Figures and Tables, are incorporated herein by reference.EXAMPLESEXAMPLE 1: A Single-layer Optical Classifier
[0186] A single-layer optical classifier is constructed to validate operation of the combined sub-systems (Fig. 13, Panel A) using MNIST digit, MNIST fashion, and CIFAR-100 datasets. Ps pulses generated with the EO pulse source are encoded first with the vector valuesand then with the matrix values through high speed (40 GHz) EO modulation. An available 25 GHz binary PRBS source is used to demonstrate high-speed multiplication and accumulation and a -100 MHz arbitrary waveform generator are used for inference with arbitrary-level coding. The weighted vector elements is optically summed within a phase-stabilized recirculating fiber loop described above optimized for the size of the network output vector (20m for MNIST and 200m for CIFAR-100) using the 100MHz AWG. The measured response from nonlinear phase, dispersion, gain saturation, and timing jitter is fed back into the numerical model to optimize the coupler, gain, and feedback mechanisms.
[0187] Phase stabilization of -1 mrad is achieved with feedback from the cavity output to drive the laser as detailed in Pi’s previous works (Fig. 13, Panel B). The delay of the input pulses are synchronized to the cavity by actively locking the frequency of the RF drive generating the input pulses to the cavity round-trip time, as demonstrated (Fig. 13, Panel C). Nonlinear activation is first implemented digitally before applying the nonlinear loop mirror described above. The weight matrix (784x10 for MNIST and 3072x100 for CIFAR-100) for the single layer classifier task is determined through standard in-silico b ackpropagation techniques, accounting for the optical nonlinear activation function in the platform. This demonstrates accurate classification as well as experimental confirmation of the accessible computation speeds, model sizes and efficiencies. Scaling computing speed is investigated by adding a second modulator as in Fig. 5, using two fiber couplers for muxing and a piezo stretcher to match the path lengths with a single pulse delay. The three models are run on the faster system for validation studies. In addition, by using the 25GHz modulator and a 1km fiber, MVM with > 105vector elements and 10 billion parameters demonstrate the scale advantage of this system.Finally, to study all-optical multi-layered networks the Kerr resonator memory system is evaluated from the first thrust for cleaning and copying the product vector for use withsubsequent layers.
[0188] The platform is demonstrated to have 100% fidelity transfer after 1 minute. To evaluate soliton cleaning, time dependent fast oscilloscope and dispersive Fourier transform evolution measurements are used.EXAMPLE 2: ALL-OPTICAL PLATFORM
[0189] A modulated laser or VCSEL produces optical pulses forming an input vector. Diagonal matrix components are multiplied by permuted copies of the cyclically shifted vector. The vector is fanned out to multiple parallel modulators on a chip. The vector is then fanned into a single mode before entering into a fiber loop memory for accumulating multiple diagonals. The signal undergoes nonlinear optical activation and is then routed to all-optical memory and / or a regenerator and / or a copy loop. The process is repeated for n layers with new weight matrices.EXAMPLE 3: FAN-IN ACCUMULATION PLATFORM
[0190] A modulated laser or VCSEL produces optical pulses forming an input vector. Diagonal matrix components are multiplied by permuted copies of the cyclically shifted vector. The vector is fanned out to multiple parallel modulators on a chip. The vector is fanned into a multi-mode waveguide or in free-space onto a common detector. The fanned-in vector is digitally activated or activated based on a detector. The signal is routed to analog-to-digital convertor (ADC) / Memory / digital-to-analog converter (DAC). The process is repeated for n layers with new weight matrices.Modular Optical Compute Systems
[0191] A significant advantage of the architecture described herein is its inherent modularity. A general computing architecture of large-scale processor comprising of several compute cores is shown in Fig. 31. It consists of three fully modular subsystems: the input matrix encoder, which optically encodes the elements of the input matrix, compute cores each of whichare proceeded by well-calibrated fiber length to perform coarse delays, and finally an optical combining module which combines the output of the compute cores to either compute the output vector or transmit the result of local computations to a subsequent combining module.
[0192] In several of the following examples, this application describes some typical examples of combining one or more degrees of freedom for an 8x8 weight matrix. Each figure shows the design layout and the corresponding matrix division, as well as the number of associated parts needed, as derived from the scaling laws discussed herein (where also h, v, c, and N are defined).
[0193] The present application is further illustrated by the following examples that should not be construed as limiting. The contents of all references, patents, and published patent applications cited throughout this application, as well as the Figures and Tables, are incorporated herein by reference.EXAMPLESEXAMPLE 1: Temporal Paradigm
[0194] In Fig. 15, a temporal computing paradigm is shown that corresponds to a case of exclusively leveraging temporal degrees of freedom. In this specific example of N=8 system the matrix is purely subdivided into 8 rows, therefore h=8. It has no additional vertical subdivisions implying v=l. Since N elements are summed at the detector temporally, it has c=8. This system can therefore be assigned the values (h,v,c)=(8,l,8).EXAMPLE 2: Spatial Paradigm
[0195] In Fig. 16, a spatial computing paradigm is shown that corresponds to the case of leveraging spatial degrees of freedom. In this specific example of N=8 the matrix is divided into 8 columns, therefore v=8. It has no additional horizontal cuts and does not temporally clump elements, therefore h=l, and c=l. This system therefore is assigned (h,v,c)=(l,8,l).EXAMPLE 3: An optical computing system design (n,h,v,c)=(10,2,5,2)
[0196] Consider an example where the problem is to design an optical computing system with the system parameters (n,h,v,c)=(10,2,5,2). This system satisfies the efficient compute condition. The study designs this system as outlined by the protocol discussed herein. These steps are visually illustrated on Fig. 17.
[0197] Step 1 : The number of requisite input modulators for this system is= 5. Each10 input modulator will receive — = 2 elements of the input vector.
[0198] Step 2: Next, the study determines the number of weight modulators, N" = 10, as shown on Fig. 16.
[0199] Step 3: Next, the study determines the number of detectors, Nd= 2. Next, assign 5 weight modulators to each detector, and we call these modulator sets SltS2.
[0200] Step 4: Next, the study splits the output of each input modulator into 2 paths, each going to each of the modulator sets, (S1,S'2)- The study then determines the max delay, Md= 0. meaning that no additional optical path delays are required for this system. Given that the max delay is 0, we do not need to further divide optical paths entering the weight modulator sets.
[0201] Step 5: The study assigns a weight modulator to each path coming into each modulator set.
[0202] Step 6: After multiplication with weights, the optical paths exiting each modulator are routed onto its assigned detector.Example 4: An optical computing system design (n,h,v,c)=(8,2,2,l)
[0203] Consider an example where the problem is to design an optical computing system with the system parameters (n,h,v,c)=(8,2,2,l). This system satisfies the efficient compute condition. The study designs this system as outlined by the protocol discussed herein. These steps are visually illustrated in Fig. 18.
[0204] Step 1 : The number of requisite input modulators for this system is= 2. Each g input modulator will now receive - = 4 elements of the input vector.
[0205] Step 2: Next, the study determines the number of weight modulators, N" = 16.
[0206] Step 3: Next, the study determines the number of detectors, Nd= 2. Therefore, assign 8 weight modulators to each detector, and call these modulator sets SltS2.
[0207] Step 4: Next, the study splits the output of each input modulator into 2 paths, each going to each of the modulator sets, (S1, S2). The study then determines the max delay, Md= 3. This implies that each path, originating from an input modulator and subsequently entering each set has to be additionally divided into 4 paths. Each of these paths is followed by optical delay lines ranging from 0 to 3. At this stage, the output of each input modulator has been fanned into 2 x 4 = 8 total optical lines.
[0208] Step 5: The study assigns a weight modulator to each path coming into each modulator set
[0209] Step 6: After multiplication with weights, the optical paths exiting each modulator are routed onto its assigned detector.EXAMPLE 5: Spatially distributed temporal compute
[0210] In Fig. 21, a spatially distributed temporal computing paradigm is shown where the input vector is divided into two halves. Each half of the vector is fanned out to a set of 8 EOMs which multiply one half of the corresponding matrix row. The two multiplied halves are spatially summed on the final detector. This configuration has the advantage of increasing the compute speed by parallelizing multiplications at the expense of additional EOMs. This system can be assigned the values (h,v,c)=(8,2,4)EXAMPLE 6: Temporally binned spatial compute
[0211] In Fig. 22, a temporally-binned spatial computing paradigm is shown where twoadjacent columns are time-interleaved and fed electrically into one modulator. The clumped columns are spatio-temporally summed at the detector. This configuration reduces the component count at the cost of lower compute speed. In this case because two columns are clumped and fed into the same detector, this can be assigned the values (h,v,c)=(l,4,2).EXAMPLE 7: Diagonal compute
[0212] In Fig. 23, a diagonal computing paradigm is shown where the weight matrix is decomposed into diagonals, and each EOM receives a cyclically shifted version of the input vector which is multiplied to the corresponding diagonal. The individual diagonals are spatially summed on a detector. This can be assigned the values (h,v,c)=( 1,1,1).EXAMPLE 8: Spatially Distributed Diagonal Compute
[0213] In Fig. 24, a spatially distributed diagonal computer paradigm, where spatial and delay DOFs can be combined by dividing the input vector and the weight matrix into two halves. The two halves of the input vector are cyclically shifted and multiplied to the corresponding diagonals of the corresponding half of the weight matrix. This combination has the advantage of reducing the maximum optical delay required while retaining compute speed. This can be assigned the values (h,v,c)=(l,2,l).EXAMPLE 9: Temporally clumped Diagonal Compute
[0214] In Fig. 25, a temporally clumped diagonal computer paradigm is shown, where temporal delay DOFs can be combined by interleaving adjacent diagonals and spatio-temporally summing the resulting products. This reduces the number of EOMs required. This can be assigned the values (h,v,c)=( 1,1,2).EXAMPLE 10: Block diagonal compute
[0215] In Fig. 26, a block diagonal computer paradigm is shown where an extension of diagonal compute is presented here whereby the weight matrix is decomposed into 4 blocks.Each block is decomposed into its corresponding diagonals. This corresponds to both subdividing the input and the output vectors. The two halves of the input vectors are fanned out to a set of 8 modulators, which multiply the input vector elements with the corresponding block diagonals. Each detector computes one half of the output vector. This can be assigned the values (h,v,c)=(2,2,l).EXAMPLE 11: Spatially hyper-multiplexed compute
[0216] In Fig. 27, a spatially hyper-multiplexed computer paradigm is shown whereas in the case of spatial compute, each vector element is encoded in a separate modulator. The vector is, however, fanned out and multiplied through sets of weight modulators, thereby further spatially multiplexing the compute. Each detector computes one half of the output vector. This paradigm is suited for computing with excess modulators with room to spatially multiplex. This can be assigned the values (h,v,c)=(2,8,l).EXAMPLE 12: Unconventional Blocked Compute
[0217] In Fig. 28, an unconventional blocked computer paradigm is shown for the sake of an example, the weight matrix is blocked into multiple rectangular blocks. While this configuration is an extension of the blocked diagonal compute (Fig. 16) it demonstrates the capability of the framework to adequately represent arbitrarily blocked matrices. This can be assigned the values (h,v,c)=(2,4,l).EXAMPLE 13: A prototypical dual-wavelength and dual-core offset-diagonal optical processor
[0218] In Figs. 29A-G, this study demonstrates a prototypical dual -wavelength and dualcore offset-diagonal optical processor with off-the-shelf optical components. The experimental apparatus (Fig. 29D) consists of two wavelength channels multiplexed and electro-optically modulated each with an input vector. Next, the vectors are fanned out and delayed by an integernumber of symbol periods (10 ns) and modulated with electro-optic modulators driven by diagonals of the weight matrix. The two wavelength channels are demultiplexed with an optical grating and spatially summed with an optical lens and high-speed free-space optical detectors. 5 GFLOPs computations are demonstrated through summing two random input vectors (Fig. 29A), and to characterize accuracy the study encodes 1000 symbols with uniformly random floatingpoint numbers between [-0.5,0.5] onto each of the modulators at 200 MFLOPS (one vector and two weight modulators) and compute the sum of the two product vectors. This operation corresponds to a MVM operation between a 1000x 1000 weight matrix with two diagonals (offset by the corresponding delay) and a 1000x 1 input vector (Fig. 29B and 29C). Typical errors of 1.7% (Fig. 29E) are observed regardless of the relative delay between the two paths (1-symbol and 5-symbol delays) shown here and the wavelength channel (where symbol means the data pulse (5 symbol spacing means a 5 pulse delay). Here, the study encodes both positive and negative floating-point numbers by upmodulating the data-stream with a local oscillator and demodulating the signal at the output (separate + and - processing and dual -output modulators are also compatible). The study also implements a two-layer MNIST-digit classifier and benchmarks its performance with 500 test images. The optical processor displays a classification accuracy of 87 % (Fig. 29G) compared to 90 % (Fig. 29F) for the same dataset with an ideal electronic classifier. While the scale and speed here is limited by available equipment, the processing speed for large MMM operations scales roughly as the number of modulators squared times the speed of modulators, corresponding to -lOOTFLOPs with 100 10GHz modulators, and 100 frequency channels, which is a realistic target for a fully integrated design.EXAMPLE 14: Input Matrix Encoders
[0219] The input matrix encoder can be physically implemented with any optical source capable of producing optical power at multiple optical wavelengths combined with a modulationtechnique to separately modulate each optical wavelength. The output of this module is fanned out to available compute cores.
[0220] The study presents several examples to illustrate this. The first example is based on optical frequency comb source and a bank of micro-ring-resonator modulators shown in Fig. 32A. An optical frequency comb is an integrated optical device which using a single-frequency pump laser to generate ultra short optical pulses through optical nonlinear interactions. The resulting output contains several wavelengths. A bank of micro-ring-resonator modulators is used to modulate the output to encode the input vector elements. Each MRR selectively addresses one wavelength channel and modulates the optical power contained in the corresponding channel. While this technique has the advantage of being compact the optical power produced by optical frequency combs are small and may require further optical amplification.
[0221] In another embodiment, a bank of single frequency lasers as an optical source. The outputs of these lasers can be combined into a single fiber with the use of standard arrayed waveguide gratings routinely used in optical communication systems. The combined output can then be modulated with an MRR modulator bank as detailed herein (see Fig. 32B).
[0222] In further embodiments, each single frequency laser can also be individually modulated with an electro-optic modulator and then subsequently multiplexed (see Fig. 32C). This technique has the capability to be versatile in the selection of laser frequencies, at the cost of being bulky requiring several lasers to be integrated simultaneously.
[0223] Alternatively, a multi -wavelength source can also be constructed by electro- optically modulating a single frequency source with a strong single-frequency RF tone. This is known as an electro-optic comb source. The wavelength channels can then be encoded with vector elements by using an MRR modulator array (see Fig. 32D).EXAMPLE 15; Processing Cores
[0224] Processing cores are subsystems which receive input vector elements, perform multiplications with weight elements and accumulate local products. These calculated results may require combining with the results of other cores to complete the required computations. The basic elements that can be within a compute cores are delay-lines to cyclically shift the received input vectors, a broad-band electro-optic modulators to perform multiplications simultaneously with all wavelength channels, a wavelength demultiplexing element to separate various wavelength channels, and spatio-temporal accumulation technique on a photodetector. The study presents four examples of possible physical implementations of a compute core, which are illustrated in Figs. 33A-D.
[0225] In the first example (see Fig. 33A), the input matrix elements from the encoder are first fanned out into several optical paths each with on-chip delay lines which perform delays to cyclically shift the input vectors. Next, a stack of broadband electro-optic modulator simultaneously performs multiplication operations across all wavelength channels. The products are then incident into a specially designed arrayed waveguide grating-component which both spatially separates the wavelength channels and focusses them on a linear detector array which performs spatial summations of the products. This component functions as follows- The input field from an electro-optic modulator containing several wavelength channels is incident into a star coupler. The star coupler divides an input signal irrespective of the wavelength into several output channels. Each output channel of the star coupler has a different optical path length, but the neighboring channels have a fixed path length difference. The structure described so far is identical to a standard optical arrayed waveguide grating. Each wavelength channel accumulates a different phase in each spatial channel of the star coupler. These channels are then incident onto a curved structure which functions as an optical lens. Each wavelength channel by the virtueof the phase difference accumulated and through optical interference exits at a different angle and as a consequence are focused onto a different detector. This implementation has the advantage of being integrable into a standard optical system at the cost of additional design requirements to ensure minimal cross-talk between various wavelength channels. This scheme (Fig. 33 A) can also be used for the arrangement from Fig. 15 by shaping the curvature of the final lens such that each beam of the same color is spatially separated for coupling into individual detectors, capacitors or Nxl mux devices.
[0226] In the examples described above, delay lines were implemented separately on each optical path preceding a modulator. However, this approach may limit the total usable device area, since longer delay lines occupy a larger footprint. To address this limitation, a cascaded delay configuration can be adopted instead (Fig. 36 demonstrates this for the case of MVM problems). As shown, each delay line adds with respect to a previous delay line. In this approach, a series of optical couplers with differing ratios are used to successively extract small portions of the input optical signal, which are fed into the weight modulators. The remaining signal is then delayed and fed into the next coupler, and so on, until the required total delay is achieved. This configuration is particularly useful when optimizing footprint is critical. However, it requires tuning the coupling ratios of the optical couplers.
[0227] While previous descriptions generally assume weight matrices are square (NXN), this is not a requirement of the system described. The techniques and embodiments here can easily be extended to use rectangular matrices. For example, as shown in Fig. 30 (batched computations), square matrices can be grouped into rectangular matrices for batched processing.However, it may be necessary to utilize all degrees of freedom, such as vertical, horizontal cuts, and temporal binning, to achieve optimal computation since a general rectangular matrix by default may not be amenable to optimal compute.
[0228] In a second example (see Fig. 33B), as before, the input matrix elements from encoder are first fanned out into several optical paths and subsequently modulated with the corresponding weight elements with the use of electro-optic modulators. The outputs are then coupled into free-space with the use of on-chip grating couplers or fiber-optic collimators. A diffraction grating is employed as a wavelength selective element which splits each wavelength channel into a corresponding angled beam. A focusing lens then focusses the beams belonging to a single wavelength channel onto an appropriate detector. This system in principle is identical to the previous example except that it used widely available inexpensive bulk components such as a gratings and lenses. However, these will require additional alignment and packaging procedures. This scheme (Fig. 33B) can also be used for the arrangement from Fig. 15 by shaping the curvature of the final lens such that each beam of the same color is spatially separated for coupling into individual detectors, capacitors orNxl mux devices.
[0229] In a third example (see Fig. 33C-D), as before, the input matrix elements from encoder are first fanned out into several optical paths and subsequently modulated with the corresponding weight elements with the use of electro-optic modulators. The outputs of the modulators are followed by a standard arrayed waveguide grating (AWG) which splits the wavelength channels into separate optical waveguide. Waveguides belonging to the same wavelength from each AWG can be routed to a multi-port photodetectors, which are on-chip photodetectors which allow for multiple input waveguides. Such a system while comprising of only standard on-chip optical components offered by several commercial optical fabrication foundries will require efficiently routing of the outputs of each AWG.
[0230] Two example techniques enable this. One, this could be performed with the use of low loss waveguide crossings or utilizing two optical waveguide layers where in each wavelength channel is successively routed to the second layer with a separate detector, whileremaining wavelength channels are allowed to continue on layer 1 till all wavelength channels are routed to layer 2 (see Fig. 33C). This technique, however, requires sufficient optical chip footprint to allow for successive wavelength drops into the other layer.
[0231] If the fabrication allows for additional optical layers the chip footprint can be significantly reduced since multiple channels can be simultaneously routed to a separate optical layer. The optimal design would likely consist of minimizing on-chip area footprint and the number of optical layers simultaneously. Alternatively, the number required routings can also be drastically reduced by utilizing cascaded AWGs, where in each electro-optic modulator is followed by several AWGs each with a finer frequency selectivity (see Fig. 33D). These schemes (Fig. 33C-D) can also be used for the arrangement from Fig. 15 by coupling the individual outputs on the final layer into individual detectors, capacitors or Nxl mux devices.
[0232] In the examples described above, delay lines were implemented separately on each optical path preceding a modulator. However, this approach may limit the total usable device area, since longer delay lines occupy a larger footprint. To address this limitation, a cascaded delay configuration can be adopted instead (Fig. 36 demonstrates this for the case of MVM problems). In this approach, a series of optical couplers with differing ratios are used to successively extract small portions of the input optical signal, which are fed into the weight modulators. The remaining signal is then delayed and fed into the next coupler, and so on, until the required total delay is achieved. This configuration is particularly useful when optimizing footprint is critical. However, it requires tuning the coupling ratios of the optical couplers.
[0233] While previous descriptions generally assume weight matrices are square (N><N), this is not a requirement of the system described. The techniques and embodiments here can easily be extended to use rectangular matrices. For example, as shown in Fig. 30 (batched computations), square matrices can be grouped into rectangular matrices for batched processing.However, it may be necessary to utilize all degrees of freedom, such as vertical, horizontal cuts, and temporal binning, to achieve optimal computation since a general rectangular matrix by default may not be amenable to optimal compute.EXAMPLE 16; Combining Modules
[0234] A combining module receives the local compute results from several compute cores and computes their sum. The output of a combining module can either be broadcast to another combining module or written into memory. A combining module within this architecture could be all-electrical, all-optical or hybrid opto-electrical architectures. This application presents three examples of such combining modules (FIGs. 34A-D).
[0235] First example (see Fig. 34A) is of an all-electrical combining technique wherein the photocurrents produced by each processing core is electrically combined by summing the currents corresponding wavelength channel photodetectors. High speed (~10 GHz) CMOS compatible current summing circuits, operational-amplifiers, and transimpedance have been demonstrated and are commercially offered by several IC manufacturers which can be adapted for this purpose. This technique has the advantage of being all electrical and requiring well established electrical circuits. However, these come at the cost of limited connectivity since electrical connections experience significant losses over long (>1 m) distances.
[0236] An all-optical combining module (see FIG. 34B) can be constructed by following every compute core with an input-encoder-like system which optically broadcast the results of the compute core. The results of a given compute core can be reencoded into optical fields with the use of a comb+MRR combination as detailed in the input matrix encoder section. Other variations such as laser banks, electro-optic comb systems can also be employed here. The optically broadcasted signals from different cores are then recombined on a combining chip with another arrayed-waveguide-grating accumulator. This is identical to the one used previouslywithin each compute core to accumulate products. In this case, this accumulation device is used to combine the results from various compute cores. As before this accumulator splits the wavelengths and accumulates them on a linear detector array. This variation while being compatible with 2D integration techniques will incur additional energy costs associated with detecting, amplifying, and broadcasting the outputs of each compute cores. However, energy calculations suggest that associated energy costs are a small fraction of total energy costs and can be minimized by increasing the number of modulators on each compute core. A variation of this scheme can also be used for the temporal summation arrangement from Fig. 5 using capacitive detectors for each output, an Nxl switching mux for each wavelength from each module, the output of which can be used to drive the MRR modulated combs. In this case, the combining receiver will require M wavelength demuxs, where M is the number of cores, followed by M*X (where X is the number of wavelengths) number of detectors and ADCs before memory.
[0237] In a further embodiment, compute cores can also be optically combined by vertically stacking compute cores. In this instance, the products generated on each core are not summed locally, but instead the outputs of several cores are summed simultaneously using a free-space grating, an optical lens and a linear detector array, similar to bulk compute core (see Fig. 34C). This allows for compact form factor while allowing several compute cores to be simultaneously summed. Alternatively, as opposed to using a free space grating, one could use an arrayed waveguide gratings to separate optical wavelengths and employ a free-space lens to accumulate the outputs on a linear array of optical detector.
[0238] Finally, a combining hierarchy may be advantageous because of the relative costs for each technique. The vertical stacking technique incurs no additional energy costs but may face alignment challenges associated with bulk optics. Electrical combining is second best in terms of energy performance and is compatible for several neighboring cores given the electricallosses that may be incurred for propagation over long lengths. Initial energy calculations suggest that electrical combining techniques retain their advantage for up to 10-30 compute cores (assuming reasonable length scales). Finally, optically combining while incurring additional cost of remodulation provides the widest reach given ultra-low loss optical propagation. About 1000 electrically combined cores can be further stacked at the cost of a new comb and rf amplification per section. Properly designed, this stack can maintain load-limited energy costs. A full combining hierarchy is illustrated in Fig. 34D.EXAMPLE 17: Negative Weights
[0239] The computing cores can also be adapted to perform computations with positive and negative weights by using dual port electro-optic weight modulators. The previously proposed arrayed-waveguide-grating accumulator can be employed here. The two output ports can be spatially separated by ensuring they occupy two separate polarization modes of the optical waveguides. In contrast to all-positive compute, each wavelength channel is now assigned two detectors which separately accumulate the computed products from the upper and lower ports of the weight modulator, respectively. The difference of the two photocurrents produces the requisite summed signal. The primary objective here is to distinguish the optical fields within the upper and lower arms of the weight modulator which the two different polarizations enable. Alternatively, two different star couplers can be designed with slightly different optical path delays one of each output of the electrooptic modulator. Changing the delay allows one to tune the angle at which different wavelengths exit the arrayed grating. By changing the angle of exit for the upper and lower arms of each weight modulator, the optical fields can separately be accumulated for each wavelength and each arm of the modulator. An example implementation is illustrated in Fig. 35. This scheme can also be used directly the temporal summation arrangement from Fig. 15.
[0240] Alternatively, positive and negative weights can also be computed by separately computing the products of positive and negative weights with input vector and subsequently accumulating them with appropriate signs digitally. In this case, the coherence requirement of the source is substantially reduced.
[0241] Negative weights can be encoded by using a frequency-detuned local oscillator. In this method, an additional optical beam called the local oscillator having a frequency offset from the main optical signal is directed onto the final photodetector. By analyzing the component of the output at the frequency difference between the local oscillator and the incident optical signals, it is possible to recover the phase information of the input optical signals. This configuration can also be extended to implement optical operations with complex numbers.
[0242] While various embodiments have been described above, it should be understood that such disclosures have been presented by way of example only and are not limiting. Thus, the breadth and scope of the subject compositions and methods should not be limited by any of the above-described exemplary embodiments but should be defined only in accordance with the following claims and their equivalents.
[0243] The above description is for the purpose of teaching the person of ordinary skill in the art how to practice the present invention, and it is not intended to detail all those obvious modifications and variations of it which will become apparent to the skilled worker upon reading the description. It is intended, however, that all such obvious modifications and variations be included within the scope of the present invention, which is defined by the following claims.
[0244] The claims are intended to cover the components and steps in any sequence which is effective to meet the objectives there intended, unless the context specifically indicates the contrary.
Claims
CLAIMSWhat is claimed is:
1. An optical matrix multiplier comprising: an optical source configured to produce a sequence of optical pulses based on an input vector; a fanout module configured to receive a sequence of optical pulses and produce a plurality of optical pulse sequences; a plurality of optical delay lines, each delay line configured to apply a delay to an associated one of the optical pulse sequences: a plurality of modulators, each modulator configured to modulate one of the delayed optical pulse sequences; and at least one accumulator configured to sum at least one modulated optical pulse sequence.
2. The optical matrix multiplier of claim 1, wherein the at least one accumulator comprises a photodetector configured to receive the modulated optical pulse sequence and produce an output.
3. The optical matrix multiplier of claim 1, wherein said at least one accumulator comprises a fan-in accumulator configured to: receive each optical pulse sequence of the plurality of modulated optical pulse sequences; and produce an output vector comprising sums of each individual optical pulse sequence.
4. The optical matrix multiplier of claim 1, wherein the accumulator comprises an optical cavity accumulator.
5. The optical matrix multiplier of claim 1, wherein each of the plurality of modulators are configured to modulate pulses of one of the delayed optical pulse sequences based on values of a weight matrix.
6. Tire optical matrix multiplier of claim 1, wherein: the optical source is configured to simultaneously produce a plurality of sequences of optical pulses, each of the sequences of optical pulses at a different wavelength, each of tire sequences of optical pulses based on a different input vector; and each of the plurality of modulators are configured to modulate each of the different wavelengths of the delayed optical pulse sequences.
7. Tire optical matrix multiplier of claim 6, wherein the at least one accumulator comprises:a wavelength-dependent fan-out configured to receive each of tire different wavelengths of the modulated optical pulse sequence and direct each different wavelength to a different photodetector.
8. The optical matrix multiplier of claim 1, further comprising an optical activation unit configured to receive an output of said at least one accumulator and produce an activation output that is a nonlinear function of the received output.
9. The optical matrix multiplier of claim 1, further comprising an optical memory coupled to the at least one accumulator and configured to store optical pulses.
10. The optical matrix multiplier of claim 9, wherein the optical memory comprises an optical resonator.
11. A modular computing system comprising: a plurality of the optical matrix multipliers of claim 1. each of the optical matrix multipliers configured to produce an output: and a combining module configured to: receive each output of tire optical matrix multipliers; and produce a combined output.
12. The modular computing system of claim 11, further comprising a plurality of coarse delay lines configured to delay to the output of each optical matrix multiplier by a different coarse delay.
13. Tire modular computing system of claim 11, wherein the combining module comprises an optical accumulator configured to sum the outputs of the optical matrix multipliers.
14. An optical matrix multiplier comprising: an optical source configured to produce a sequence of optical pulses based on an input vector; a plurality of modulators, each modulator configured to modulate one of the optical pulse sequences; and at least one accumulator configured to sum at least one modulated optical pulse sequence.
15. The optical matrix multiplier of claim 14, wherein: the optical source is configured to simultaneously produce a plurality of sequences of optical pulses, each of the sequences of optical pulses at a different wavelength, each of the sequences of opticalpulses based on a different input vector; and each of the plurality of modulators are configured to modulate each of the different wavelengths of the optical pulse sequences.
16. A method of designing and operating a modular computing system comprising: determining a number of input modulators for use in the system, wherein each input modulator includes an input and an output; determining a number of weight modulators for use in the system, wherein the weight modulators are organized in one or more sets; determining a number of detectors (Nd) for use in the system; configuring the system to divide tire output of each input modulator in Nd paths exiting the input modulators, wherein each of the Nd paths is received by a distinct and separate set of weight modulators; determining a maximum delay (Md) for the system, and further configure the system to divide each path into each set of weight modulators into Md+1 additional paths followed by optical delay lines from 0 to Md. wherein any path exiting any given input modulator becomes divided into Nd * (Md+1) paths, and wherein each path comprises a unique delay line; assigning a weight modulator to each path; routing optical paths exiting the weight modulators onto an assigned detector; and operating the system to produce an output.
17. A method of designing and operating a modular computing system comprising: receiving a weight matrix comprising modulation values for the optical matrix multiplier of claim 1; identifying a pemiuted the weight matrix and corresponding permuted input vector associated with zero modulation values for at least one modulator; omitting the at least one modulator associated with the zero modulation values; and operating the system to produce an output.
18. A method of optical matrix multiplication, the method comprising: producing a sequence of optical pulses based on an input vector; receiving, by a fanout module, a sequence of optical pulses and producing, by the fanout module, a plurality of optical pulse sequences; modulating each sequence of optical pulses based on a weight matrix; and summing, by at least one accumulator, at least one modulated optical pulse sequence.
19. Tire method of claim 18, further comprising delaying each sequence of optical pulses by a different delay prior to modulating each sequence of optical pulses.
20. The method of claim 18, wherein: producing the sequence of optical pulses comprises simultaneously producing a plurality of sequences of optical pulses, each of the sequences of optical pulses at a different wavelength, each of the sequences of optical pulses based on a different input vector; and modulating each sequence of optical pulses comprises modulating each of the different wavelengths of the optical pulse sequences.
Citation Information
Patent Citations
Dying oscillation absorption spectrum detecting and sensing device for all optical fiber cavity
CN1664517A
Optical Ising machines and optical convolutional neural networks
US11017309B2
Large-Scale Artificial Neural-Network Accelerators Based on Coherent Detection and Optical Data Fan-Out
US20210357737A1