Photon tensor accelerator for artificial neural networks
The photonic tensor accelerator, which performs multi-dimensional encoding in the optical domain using photonic units, overcomes the limitations of existing electronic accelerators in terms of scalability and computational power, achieving efficient matrix and tensor multiplication operations and possessing higher computational power and scalability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UNIVERSITY OF CENTRAL FLORIDA RESEARCH FOUNDATION INC
- Filing Date
- 2019-10-15
- Publication Date
- 2026-04-28
AI Technical Summary
Existing electronic hardware accelerators have reached their scalability limits and cannot meet the computational needs of artificial neural networks, especially in terms of computational capabilities for matrix-vector, matrix-matrix, batch matrix-matrix, and tensor-tensor multiplication.
Photonic units are used to perform vector-vector, matrix-vector, matrix-matrix, batch matrix-matrix, and tensor-tensor multiplications. Optical multiplexers, beam combiners, and optical replicators are used to perform multiplication operations in the optical domain. The results are converted into electrical signals by nonlinear optical elements, and multidimensional encoding (such as wavelength, spatial mode, polarization, etc.) is combined to improve computational efficiency.
It achieves computing power several orders of magnitude higher than traditional electronic accelerators, with greater scalability, rapid programmability and low power density, making it suitable for specific computing tasks such as MAC operations, and capable of performing complex matrix operations within one clock cycle.
Smart Images

Figure CN113518986B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application relates to and claims priority to U.S. Provisional Patent Application No. 62 / 842,771 (U.S. Attorney’s File No.: UCF34129PROV), filed May 3, 2019, entitled “Photonic Tensor Accelerator for Artificial Neural Networks,” which has been assigned to the applicant and whose entire contents are incorporated herein by reference. Background Technology
[0003] This application relates generally to optical computing, and more specifically to photon acceleration for vector-vector, matrix-vector, matrix-matrix, batch matrix-matrix, and tensor-tensor multiplication.
[0004] This application includes references enclosed in parentheses and indicated by numbers, such as [x], where x is a number. A list of numerical references can be found at the end of this application. Furthermore, these references are listed in the Information Disclosure Statement (IDS) filed together with this application. The entire teachings of each of these listed references are incorporated herein by reference.
[0005] To date, electronics and photonics have largely played their respective technological roles in the information society. Due to the fermionic nature of electrons, electronics has dominated information generation and processing technologies. Similarly, due to the bosonic nature of photons, electronics dominated communication technologies even before the invention of lasers and optical fibers, but in recent decades, photonics has dominated information transmission technologies. As long anticipated, according to Moore's Law, the processing power of electronic integrated circuits (ICs) will eventually stop growing. This expectation has spurred the optics and photonics community to explore problems in optical and photonic information processing over the past half-century. These efforts include optical transistors[1],[2] and logic gates[3],[4] for general optical computing, and Fourier optics for specialized information processing[5]. However, by the late 1980s, the misjudgment of the role of optics in computing had set the field back several times and left it almost dormant for the next two decades[6].
[0006] In recent years, ICs have indeed been unable to maintain geometric growth according to Moore's Law, mainly because the heat generated by the high-density power consumption associated with the small devices required to increase clock speeds is difficult to dissipate. Therefore, the limiting factor for the scalability of ICs in terms of computing power is not total power but power density. In the post-Moore's Law era, the industry's solution to the scalability problem is to build hardware accelerators with parallel computing architectures and optimized local memory for specific computing purposes, such as graphics processing units (GPUs) and tensor processing units (TPUs), compared to the von Neumann architecture of a single CPU [7], [8]. Thanks to these hardware accelerators, new applications based on artificial intelligence (AI) / machine learning (ML) implemented using artificial neural networks (ANNs) have proliferated in almost every corner of academia, industry and society in general, although the processing power of original ICs has remained stagnant. Summary of the Invention
[0007] This invention relates to a photonic unit for vector-vector multiplication, matrix-vector multiplication, matrix-matrix multiplication, batch matrix-matrix multiplication, and tensor-tensor multiplication. In the case of vector-vector multiplication, the photonic device includes a first optical multiplexer that receives a first optical signal representing a first vector, wherein each element of the first vector is encoded in a first degree of freedom (DOF) / dimension of the light and is non-temporally independent within one multiplication cycle to generate a first multiplexed optical signal. The photonic unit includes a second optical multiplexer that receives a second optical signal coherent with the first optical signal, the second optical signal representing a second vector, wherein each element of the second vector is encoded in an element-orthogonal degree of freedom (DOF) / dimension of the same optical mapping as the first vector to generate a second multiplexed optical signal. The photonic unit also includes a beam combiner that receives a first multiplexed optical signal from a first optical multiplexer and a second multiplexed optical signal from a second optical multiplexer, in order to combine them to produce interference between the first and second optical signals. This interference comprises the product of the first and second vectors in the total interference intensity. This accumulation does not require the entire DOF, but rather specific points or parameters in the DOF that have not yet been used for encoding.
[0008] In the case of an N×M matrix with M×1 vector multiplication, the photonic unit includes a first optical multiplexer that receives at least a first optical signal representing at least one M×1 vector having M elements, wherein each element of the M×1 vector is encoded in a first orthogonal degree of freedom (DOF) / dimension of light and is non-temporal in one multiplication cycle to generate a first multiplexed optical signal, and wherein M is a positive integer greater than or equal to 1. The photonic unit includes an optical replicator for replicating the at least first optical signal representing the M×1 vector into N copies of N additional optical signals in a second orthogonal degree of freedom (DOF) / dimension of light, wherein N is a positive integer greater than or equal to 1. The photonic unit also includes N optical multiplexers, identical to the optical replicators in the second orthogonal degree of freedom (DOF) / dimension of light. Each optical multiplexer receives M additional optical signals, each of which is coherent with the first optical signal. Each of the N additional optical signals represents an independent row of an M×N matrix, where each element in the row of the M×N matrix is encoded with the same element-orthogonal degree of freedom (DOF) / dimension of the optical mapping as the first optical signal representing an M×1 vector, to generate N additional multiplexed optical signals. The photonic unit also includes at least one beam combiner that receives the first multiplexed optical signal from the first optical multiplexer and N copies of the N additional multiplexed optical signals from the N optical multiplexers, to combine them to generate N interferences between the first optical signal and each of the N additional optical signals. The N interferences consist of the product of the M×N matrix and the M×1 vector in N total interference intensities.
[0009] In the case of an N×M matrix with M×W matrix multiplication, the photonic unit includes a first set of N optical multiplexers in the first orthogonal degree of freedom (DOF) / dimension of light, which receive N optical signals, each of which represents an independent row of the N×M matrix with M elements, wherein each element in each independent row of the N×M matrix is encoded in the second orthogonal degree of freedom (DOF) / dimension of light and is non-temporal in one multiplication cycle to produce N multiplexed optical signals, and wherein M and N are each positive integers greater than or equal to 1. The photonic unit includes a first optical replicator for replicating each of the N multiplexed optical signals representing multiple independent rows of the N×M matrix as W copies in the third orthogonal degree of freedom (DOF) / dimension of light, where W is a positive integer greater than or equal to 1. The photonic unit also includes a second set of W optical multiplexers in the third orthogonal degree of freedom (DOF) / dimension of light, identical to the first optical replicator that receives W additional optical signals, wherein each of the W additional optical signals is coherent with N optical signals, each of the W additional optical signals representing an independent column of an M×W matrix having M elements, wherein each element in each independent column of the M×W matrix is encoded with the same element-DOF / dimension of the optical mapping as each element in each independent row of an N×M matrix to produce W additional multiplexed optical signals. The photonic unit includes a second optical replicator for replicating each of the W multiplexed signals representing the independent column of the M×W matrix as N copies in the first orthogonal degree of freedom (DOF) / dimension of light, identical to the first set of N optical multiplexers. The photonic unit also includes at least one beam combiner that receives two sets of N×W multiplexed optical signals, which represent appropriate copies of rows or columns of each of the N×M and M×W matrices, in order to combine them to produce N×W interferences between each row of the N×M matrix and a column of the M×W matrix, including the result of multiplication of the total N×W interference intensity.
[0010] In the case of B-batch matrices, i.e., multiple N×M matrices × (M×W matrices), the photonic unit includes a first set of N optical multiplexers in the first orthogonal degree of freedom (DOF) / dimension of light, which are used to receive N optical signals, where each of the N optical signals represents an independent row with M elements in the first M×N matrix, where each element in each independent row of the first M×N matrix is encoded in the second orthogonal degree of freedom (DOF) / dimension of light and is non-temporal in one multiplication cycle to produce N multiplexed optical signals, where M and N are each positive integers greater than or equal to 1. The photonic unit includes a first optical replicator for replicating each of the N multiplexed optical signals representing multiple independent rows of the first N×M matrix as W copies in the third orthogonal degree of freedom (DOF) / dimension of light of multiple N×M matrices, where W is a positive integer greater than or equal to 1. The photonic unit comprises a second multiplexer in the fourth orthogonal degree of freedom (DOF) / dimension of light, which receives B optical signals, each of which contains W copies of N multiplexed optical signals in the third orthogonal degree of freedom (DOF) / dimension. The N multiplexed optical signals represent multiple independent rows of one of the B×N×M matrices, where B is a positive integer greater than or equal to 1. The photonic unit includes a third set of W optical multiplexers in the third orthogonal degree of freedom (DOF) / dimension of light, identical to those used by the first optical replicator that receives W additional optical signals. Each of the W additional optical signals is coherent with N optical signals, and each of the W additional optical signals represents an independent column with M elements in each of a plurality of M×W matrices. Each element in each independent column of the M×W matrix is encoded with the same element-DOF / dimension of the optical mapping as each element in each independent row of each of a plurality of N×M matrices to produce W additional multiplexed optical signals. The photonic unit includes a second optical replicator for replicating each of the multiplexed signals representing an independent column of the M×W matrix into N copies in the first orthogonal degree of freedom (DOF) / dimension of light, identical to the first set of N optical multiplexers. The photonic unit includes a third optical replicator in the fourth orthogonal degree of freedom (DOF) / dimension of light, which generates B identical optical signals, each containing N copies of W multiplexed signals representing independent columns of each matrix in the first orthogonal degree of freedom (DOF) / dimension. The photonic unit includes at least one beam combiner that receives two sets of B×N×W multiplexed optical signals from a second multiplexer and the third optical replicator to combine them to produce N×W interferences, which consist of the sum of the products of B distinct N×M matrices and the same M×W matrix in the total N×W interference intensity.
[0011] In the case of multiplying two tensors, where the first tensor has rank p and the second tensor has rank q, the first tensor is of the form [N1, N2, ..., N]. P-1 The second tensor is in the form [M, W1, ..., W]. q-1 N1,N2,...,N p-1 ,M,W1,...,W q-1 Each element is a positive integer greater than or equal to 1. The photonic unit encodes the elements of the first tensor from the first rank to the p-th rank onto the first to the p-th orthogonal degrees of freedom (DOF) / dimensions, respectively. The photonic unit includes a first set of optical replicators for replicating the multiplexed optical signal representing the tensor onto the (p+1)th to (p+q)th orthogonal degrees of freedom / dimensions of light as W1×W2×...W q-1 The photonic unit also encodes the elements of the second tensor from the first rank to the qth rank onto the p-th to (p+q-1)-th orthogonal degrees of freedom (DOF) / dimensions, and has the same element-orthogonal degrees of freedom (DOF) / dimensions of the light mapping as the copy mapped to the (p+1)-th (p+q)-th orthogonal degrees of freedom (DOF) / dimensions of the first tensor. The photonic unit includes a second set of optical replicators for replicating the multiplexed signal representing the second tensor into N1×N2×...×N in the first to (p-1)-th orthogonal degrees of freedom (DOF) / dimensions of light. p-1 Each photonic unit comprises a copy of the first to (p-1)th orthogonal degrees of freedom (DOF) / dimension of the light mapping, which is element-wise identical to the first to (p-1)th orthogonal degrees of freedom (DOF) / dimension of the light mapping as the first tensor. The photonic unit includes at least one beam combiner that receives two sets of [N1, N2, ..., N] optical replicas from the first and second sets of optical replicas. p-1 M, W1, ..., W q-1 Multiplex optical signals to combine them to generate [N1, N2, ..., N] p-1 W1, ..., W q-1 [Number] interferences, the interference intensity of which is contained in N1×N2×...×N p-1 ×W1×...×W q-1 The sum of the products of vector-vector multiplication of different M-element elements.
[0012] Any of the above-mentioned vector-vector, matrix-vector, matrix-matrix, batch matrix-matrix, and tensor-tensor multiplication interference signals typically enter nonlinear optical elements, or the total interference intensity is converted into an electrical signal that enters a nonlinear electronic element.
[0013] In one example, encoding or copying uses at least one of wavelength, spatial mode, polarization, orthogonality, and wave vector components. The spatial mode can be one of Hermetic-Gaussian (HG) mode, Laguerre-Gaussian mode, or discrete spatial samples forming a spatial orthogonal basis.
[0014] In another example, encoding or copying uses a hyperdimensionality consisting of a combination of two or more degrees of freedom (DOF) / dimensions of light.
[0015] In another example, for encoding or copying, at least two orthogonal degrees of freedom (DOF) / dimensions of light are non-overlapping subsets of the light's dimension or superdimensionality. Attached Figure Description
[0016] The invention, its preferred mode of use, and further objects and advantages will be best understood by referring to the following detailed description of illustrative embodiments when read in conjunction with the accompanying drawings, wherein:
[0017] Figure 1 This is a schematic diagram of an artificial neural network with convolutional hidden layers and fully connected output layers, and a schematic diagram illustrating the function of the photonic tensor accelerator of the present invention in an implementation of an artificial neural network.
[0018] Figure 2A Matrix multiplication is reduced to multiplication-accumulation operations and uses Figure 2B The CPU shown Figure 2C The GPU and shown Figure 2D The schematic diagram shown illustrates a TPU implementation with higher parallelism and improved energy efficiency for memory access.
[0019] Figure 3 This is a performance comparison table of electron accelerators and photon accelerators;
[0020] Figure 4A This is a schematic diagram of the node layer of a wavelength-coded photonic matrix accelerator.
[0021] Figure 4B This is a schematic diagram of the node layer of a pattern-coded photon matrix accelerator; and
[0022] Figure 5 This is a schematic diagram of a matrix-matrix multiplication mapping scheme. Detailed Implementation
[0023] Non-restrictive definition
[0024] The term "beam combiner" is a device that allows coherent beams of light to interfere with each other. It can be implemented by, but is not limited to, reflective optics, refractive optics, diffractive optics, fiber optic devices, or combinations of these components.
[0025] The term "beam splitter" is a device that can split a propagating beam of light into two or more paths. It can be implemented by, but is not limited to, reflective optics, refractive optics, diffractive optics, fiber optic devices, or combinations of these components.
[0026] The term "element-to-orthogonal degrees of freedom (DOF) / dimension" in a light map refers to the correspondence between matrix or vector elements and independent parameters of light, including wavelength, spatial mode, polarization, quadrature, and wave vector components.
[0027] The term "hyperdimensionality" refers to a combination of two or more degrees of freedom (DOFs) / dimensions of light.
[0028] The term "subset of dimension or hyperdimensionality" refers to a subset of independent parameters in a degree of freedom (DOF) / dimension of light or a hyperdimensionality of light.
[0029] The term "light" refers to electromagnetic radiation, which includes both the visible and invisible portions of the spectrum.
[0030] The term "multiplication cycle" refers to the number of mathematical multiplication operations included within a single calculation cycle.
[0031] The term "beam replicator" is a device that produces two or more copies of incident light, each copy having one or more specified parameters identical to the incident light, including wavelength, spatial mode, polarization, orthogonality, and wave vector. Beam replicators can be implemented using, but are not limited to, reflective optics, refractive optics, diffractive optics, fiber optic devices, or combinations thereof.
[0032] background
[0033] The significant role played by electronic hardware accelerators (TPUs and GPUs) clearly demonstrates that the future development of artificial neural networks depends on advancements in both software and hardware. However, electronic hardware accelerators have reached their limits in terms of scalability. Against this backdrop, new efforts have been made to explore the role of optics in computing
[17] , including developing optical interconnects to increasingly shorter length scales and demonstrating new computing paradigms such as optical neuromorphic computing
[18] -
[23] and optical reservoir computing
[24] -
[30] . The three main building blocks of artificial neural networks and deep neural networks (DNNs) are:
[0034] 1. Interconnection,
[0035] 2. Matrix-vector and matrix-matrix multiplication, and
[0036] 3. Nonlinear.
[0037] Because optics and photonics perform the first two functions as well as (if not better than) electronics; and because optical nonlinearities at the neuron level rather than the logic level are actually quite practical, now is the opportune time to explore the role of optics and photonics in ANNs and DNNs. We disclose a photonic tensor accelerator (PTA) with computational power several orders of magnitude higher than GPUs and TPUs, capable of performing matrix-vector multiplication, matrix-matrix multiplication, batch matrix multiplication (i.e., multiplying a 3D data cube (e.g., a batch of images) by a weight matrix), and tensor-tensor multiplication within a single clock cycle. In the remainder of this section, we first review the fundamentals of electronic ANNs and then briefly introduce relevant research on optical ANNs.
[0038] Artificial Neural Networks
[0039] There are three popular ANN models
[31] , namely, (a) multilayer perceptron, also known as a fully connected (FC) network, in which the output of each neuron is a nonlinear response of a linear combination of all neurons from the previous layer; (b) convolutional neural network (CNN), in which the output of each neuron is a nonlinear response of a specified linear combination of a subset of neurons from the previous layer (i.e., the kernel convolved with a subset of neurons from the previous layer); and (c) recurrent neural network (RNN), in which the output of each neuron is a nonlinear response of a linear combination of both neurons from the previous layer and neurons from the same layer but at a previous time. Figure 1 An ANN with convolutional hidden layers and fully connected output layers is shown.
[0040] Both convolution and linear combination can be mathematically reduced to a weight matrix W∈R M×M With a batch of input vectors {x1,…x N}∈R M×N Matrix multiplication between them. In order to improve the speed of matrix computation in neural networks, data transformation and thread parallelization schemes have been implemented in CPUs and GPUs
[32] ,
[33] . Although process parallelization and clock speed can be promoted, the performance of microprocessors is ultimately limited by on-chip power consumption. The key metric for evaluating power efficiency is the energy consumption of each multiply-accumulate operation (MAC) (a necessary operation for matrix multiplication). Figure 2A Is using Figure 2B The CPU shown Figure 2C The GPU and shown Figure 2DThe diagram illustrates how the TPU reduces matrix multiplication to a multiply-accumulate operation, which offers higher parallelism and improved energy efficiency for memory access. Figure 2B As shown, each MAC requires three memory reads (for filter weights, neuron inputs, and partial sums) and one memory write (for updating the partial sums).
[0041] In modern microprocessors, memory access consumes a significant portion of processing power. Dynamic Random Access Memory (DRAM) consumes two orders of magnitude more power per data access than small on-chip memory. Therefore, optimizing the reusability of data stored in local memory can greatly reduce overall power consumption. However, the challenge lies in the limited capacity of local memory (kilobytes (KB)) compared to DRAM (tens of gigabytes (GB)). To address this challenge in memory access, Figure 2C and 2D Application-specific integrated circuits (ASICs) are exploring new spatial architectures for computing speed. For example, Google’s TPU places its data storage space in registers close to the logic units, achieving an energy efficiency of about 1 pJ / MAC, which is 20 times lower than that of commercially available GPUs
[31] . However, the power consumption for memory access is still 3 times higher than that for logic operations.
[0042] Optical artificial neural networks
[0043] Optical artificial neural networks have been a subject of research since the 1980s. We review representative works in this field. It is evident that after a relatively long period of dormancy, the ANN field is experiencing a resurgence.
[0044] Holographic All-Optical ANN
[0045] In the 1980s, a great deal of work was undertaken to realize all-optical artificial neural networks for pattern recognition. Representatively, work on face recognition using light refraction (PR) volume holograms and nonlinear Fabry-Perot (FP) resonators
[34] ,
[35] remains the only fully-fledged all-optical ANN to date, as all functions (neural network training and pattern recognition) and all building blocks are implemented using optical devices. While this prior art demonstrates that optical devices can perform pattern recognition based on ANNs, it still suffers from the following drawbacks:
[0046] • The high power consumption required to activate the nonlinear FP resonator
[0047] Due to the limited dynamic range and scaling of light refraction holograms, and
[0048] • Slow training speed, limited by the PR carrier transmission lifetime, approximately milliseconds;
[0049] This makes it impossible for it to have meaningful practical applications.
[0050] Machine learning based on diffraction optics
[0051] In this prior art
[36] , a multiplane optical diffraction network is used, specifically as a digit classifier, to perform pattern recognition. A phase screen in the network is designed using machine learning techniques. Numerical tests were performed on 10,000 images from the MNIST (modified National Institute of Standards and Technology) handwritten digit dataset
[11] . Experimental results using 3D-printed phase screens showed a match rate of 88% between simulation and experiment. In this classifier, free-space optical diffraction is used to construct interconnects, while the phase mask is used to diversify the interconnects and establish the weight matrix for each diffraction layer.
[0052] All-optical digit classification is accomplished using a structure similar to an artificial neural network. However, the classification system is completely linear. As a result, only orthogonal inputs can be classified. Introducing nonlinearity would enable it to function like a true neural network.
[0053] Deep learning all-optical neural networks based on coherent nanophotonics
[0054] In this study, matrix-vector multiplication is performed on a coherent input optical signal array via a reconfigurable silicon photonic integrated circuit (PIC). As a result, the output optical signal becomes the product of the PIC transmission matrix and the input signal.
[0055] The results show that the transfer matrix of the silicon PIC can be set to any specified weight matrix. This is because any real-valued m×n matrix T can be decomposed using singular value decomposition (SVD). Where U and V are m×n and n×m single matrices, and ∑ is an m×n real-valued rectangular diagonal matrix. Furthermore, any unit transformation can be achieved using optical beam splitters and phase shifters
[37] .
[0056] In
[38] , the beam splitter and phase shifter were implemented in a silicon waveguide Mach-Zehnder (MZ) interferometer. To achieve an all-optical ANN, a nonlinear activation function must also be implemented in the optical domain. In
[38] , a saturable absorber was proposed to provide the optical nonlinear activation function. In actual experiments, nonlinear activation is still performed in the electrical domain.
[0057] A PIC is a coherent multiple-input multiple-output (MIMO) system in which the transmission matrix is set as the weight matrix. The advantage of this approach is that the integrated PIC can perform both matrix-vector multiplication and accumulation operations without actively consuming any power. In
[38] , the PIC has 54 MHz and occupies approximately 1.2 × 0.5 cm². 2 The area. Therefore, a 12-inch wafer can support approximately 60,500 MHz, or approximately 250 × 250 MHz. Therefore, “the footprint of the directional coupler and phase modulator makes scaling up to a large number (N > 1000) of neurons very challenging
[39] ”, while typical applications require 100,000 neurons
[38] .
[0058] based on Time Division Multiplexing (TDM) Deep learning optoelectronic artificial neural networks with coherent mixing
[0059] This is a hybrid approach where the MAC operation is performed in the optical domain and the nonlinear activation is performed in the electrical domain
[39] . Digital simulations are used to demonstrate digital classification. The power consumption per MAC is expected to be lower than that of existing electronics.
[0060] Vector-vector multiplication is performed through element-wise coherent optical mixing between the time-division multiplexed (TDM) signal and the TDM local oscillator (LO), and accumulation of the optical detection signal is performed via low-pass filtering, which is equivalent to integration. Matrix-vector multiplication is then performed by leveraging the parallelism of free space. Since the weight matrix is generated through time modulation, this configuration enables ultra-fast ANN training. It is also claimed that the power consumption per MAC can be much lower than that of electronic devices. However, in general, ANN weight matrix updates can be very slow and eventually remain in a static state. But in this configuration, even for static weights, the weight matrix always requires power-intensive high-speed modulation because accumulation is achieved through time integration. Furthermore, this construction has a direct impact on its scalability. Let's assume the integration time of the MAC is 1 ns, corresponding to a 1 GHz clock rate for electronic nonlinear activation. Assuming a maximum modulation speed of 500 GHz, the number of weights per column is limited to 500, not much more than a TPU. Secondly, this structure is incompatible with all-optical ANNs. Invention Overview
[0062] exist Figure 3The table summarizes the performance comparison between photon accelerators and existing electronic devices. As mentioned above, scalability is the most important metric. We have not listed energy efficiency here because it requires rigorous, systematic calculations, even though all optical techniques have the potential to achieve energy efficiency, since multiplication is passive. Each optical technique offers important innovations from which we can learn. As will be shown below, combining these innovations with the multidimensional approach disclosed in this paper should enable photon tensor accelerators (PTA) to ultimately surpass electronic devices in terms of scalability.
[0063] The disclosed invention utilizes optical and photonic methods that: 1) offer orders-of-magnitude scalability speeds compared to electronic devices; 2) are fast, programmable, and ideally compatible with both training and inference; and 3) reduce power density, thereby enabling ANNs to outperform purely electronic devices, at least in certain categories of AI functionalities. It is worth noting that:
[0064] • ANNs only require special operations / computations that are particularly well-suited to photon accelerators (e.g., MAC), rather than general-purpose computations.
[0065] Because ANNs are robust to changes in data dynamic range and nonlinear activation
[40] , analog photonic accelerators can achieve performance comparable to digital logic accelerators.
[0066] Embodiments of the present invention
[0067] Wavelength coding and pattern coding matrix - vector multiplication accelerator
[0068] Figure 4A and Figure 4B The matrix-vector multiplication in
[39] shares a single similarity with that in
[39] , where multiplication is performed via coherent mixing and square-law detection. The main difference between our method and
[39] is that the accumulation in
[39] is performed in the time domain, while the accumulation in our method is performed in the wavelength of light, space, and all other non-temporal / degrees of freedom / dimensions. Figure 4A and Figure 4B In this process, the input vector and weight vector are element-dependently projected onto different wavelength or spatial patterns [e.g., Hermitian-Gaussian (HG) patterns]. The wavelength-coded or pattern-coded input vector is fan-outed to the desired number of copies and mixed with the corresponding wavelength-coded or pattern-coded LOs, where the LOs contain weight vectors, which include weight matrices. Due to the orthogonality between wavelengths or spatial patterns
[41] , coherent mixing between a pair of signal and LO streams [ Figure 4A ]or[ Figure 4B It generates vector-vector multiplication, and parallelization in 2D space generates matrix-vector multiplication as a whole.
[0069] Figure 4A and Figure 4B The node layers of a photonic matrix accelerator based on (a) wavelength coding and (b) mode coding are shown. The photonic accelerator is in 2D... (x,z) In-plane parallelization is used to perform matrix-vector multiplication, forming a node layer together with (a) electronic nonlinear activation after photodetection and (b) optical nonlinear activation using a saturable absorber. Both electronic and optical nonlinear activations are compatible with wavelength-coded, mode-coded, or a combination of wavelength-coded and mode-coded photon accelerators. Here, the input data is represented in wavelength or mode dimension, the weight matrix in 2D (wavelength or mode, z) dimension, and the output port in x dimension. Accumulation is performed in the wavelength (a) or mode (b) dimension, respectively.
[0070] The output of a wavelength-coded and mode-coded photonic matrix accelerator can be converted into an electrical signal through balanced detection, which can then be used as input for electronic nonlinear activation, such as... Figure 4A As shown. Alternatively, the output can be directly input into the optical nonlinear activation unit and used as a pump for the optical nonlinear activation unit, for example... Figure 4B The saturable absorber (SA) shown [for longer wavelengths or orthogonally polarized probe waves]. Therefore, wavelength-coded and / or mode-coded matrix accelerators are compatible with all-optical ANNs or hybrid optoelectronic ANNs.
[0071] There are various methods for implementing multiplexers (and signal demultiplexers), including cascaded directional couplers
[42] ,
[43] , photonic lamps
[44] -
[46] and multiplane optical converters (MPLCs)
[47] -
[49] .
[0072] A key advantage of wavelength-coded and / or mode-coded matrix accelerators and photon tensor accelerators (PTAs) is their scalability. Using mode coding alone, our matrix accelerator can be scaled to at least 300×300. MPLC mode multiplexers have a wide operating wavelength range and are therefore wavelength-coded. Combining wavelength and mode coding, we can potentially scale matrix-vector multiplication to unprecedented sizes by combining wavelengths and modes into a superdimensional dimension, making the vector length the product of the number of wavelengths and the number of modes. With current technology, we can already achieve a vector length of 90,000 using 300 wavelengths and 300 modes with a 10 GHz channel spacing in the C-band. This is because the interference (multiplication) of the wavelength-coded and mode-coded signals and LO streams can be accumulated on a single detector. The matrix size is 90,000×[2D]. (x,z) The degree of spatial parallelization, which can easily exceed 100, makes the total MAC of the wavelength-coded and pattern-coded accelerator at least 9,000,000 and has 2D spatial parallelization.
[0073] Polarization and orthogonal coding can each double the size of matrix-vector multiplication. From here, we combine the polarization and mode dimensions into a single dimension, called the vector mode.
[0074] Matrix-matrix multiplication accelerator
[0075] This invention achieves unprecedentedly large matrix-matrix multiplications by further parallelizing the 3D space in the y-direction. In this case, as... Figure 5 As shown, the input matrix is represented in the (wavelength and / or mode, y) dimension, the weight matrix is represented in the (wavelength and / or mode, z) dimension and is copied in the y direction, and the output matrix is contained in (x, y).
[0076] Generalized tensor multiplication accelerator
[0077] In summary, our invention allows for the configuration of light in multiple dimensions (wavelength, vector mode, orthogonality, and three (3) spatial dimensions) to construct photonic accelerators. In free-space implementations, using the three spatial dimensions is natural, and for ICs or PICs, using the two spatial dimensions is natural. Any two (three) dimensions can be used to construct photonic accelerators for matrix-vector (matrix-matrix) multiplication. Multiple dimensions (e.g., the aforementioned wavelength-mode) can be combined into hyperdimensions to increase scalability. A vector mode is a combination of a spatial mode and a polarization mode. Similarly, each dimension can be used independently to construct a photonic tensor accelerator (PTA) for batch matrix multiplication operations. For example, the wavelength mode dimension can be used to represent a batch of images (3D data cubes), which can then be multiplied with a weight matrix, i.e., accelerated together once in one clock cycle. Alternatively, each dimension with a large number of parameters can be partitioned into mutually orthogonal subsets to effectively increase the number of independent / orthogonal degrees of freedom. For example, wavelength has more parameters than polarization, which has only two parameters. Space is another degree of freedom, which has a large number of parameters that can be partitioned into mutually orthogonal subsets. This is particularly useful in performing tensor-tensor multiplication.
[0078] Non-limiting embodiments
[0079] Although specific embodiments of the invention have been discussed, those skilled in the art will understand that changes can be made to the specific embodiments without departing from the scope of the invention. Therefore, the scope of the invention is not limited to the specific embodiments, and the appended claims are intended to cover any and all such applications, modifications, and embodiments within the scope of the invention.
[0080] It should be noted that some features of the invention may be used in one embodiment without using other features of the invention. Therefore, the foregoing description should be considered merely as an illustration of the principles, teachings, examples, and exemplary embodiments of the invention, and not as a limitation thereof.
[0081] Furthermore, these embodiments are merely examples of the many advantageous uses of the innovative teachings herein. Generally, the statements made in the specification of this application do not necessarily limit any of the various claimed inventions. Moreover, some statements may apply to some inventive features but not to others.
[0082] The invention has been presented for purposes of illustration and description, and is not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The embodiments were chosen and described to best explain the principles of the invention, its practical application, and to enable those skilled in the art to understand the various embodiments of the invention with various modifications suitable for the intended particular use. The terminology used herein is chosen to best explain the principles of the embodiments, their practical application, or technical improvements to technology found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.
[0083] Incorporated References
[0084] The following publications are included in whole or in part by reference:
[0085] [1]RFRutz, "Transistor-Like Device Using Optical Coupling BetweenDiffused PN Junctions in GaAs," Proc. IEEE, 1963.
[0086] [2] K.Jain and GWPratt, "Optical transistor," Appl.Phys.Lett., vol.28, no.12, pp.719-721, Jun.1976.
[0087] [3]HFTaylor, "Guided wave electrooptic devices for logic and computation," Appl.Opt., vol.17, no.10, p.1493, May 1978.
[0088] [4]M.N.Islam,“Ultrafast all-optical logic gates based on solitontrapping in fibers,”Opt.Lett.,vol.14,no.22,p.1257,Nov.1989.
[0089] [5]J.W.Goodman,“Operations Achievable with Coherent OpticalInformation Processing Systems,”Proc.IEEE,1977.
[0090] [6]D.A.B.Miller,“Are optical transistors the logical next step?,”Nat.Photonics,vol.4,no.1,pp.3-5,Jan.2010.
[0091] [7]S.A.Manavski,“CUDA compatible GPU as an efficient hardwareaccelerator for AES cryptography,”in ICSPC 2007 Proceedings-2007 IEEEInternational Conference on Signal Processing and Communications,2007.
[0092] [8]G.Quintana-Orti,F.D.Igual,E.S.Quintana-Orti,and R.A.van de Geijn,“Solving dense linear systems on platforms with multiple hardwareaccelerators,”ACM SIGPLAN Not.,2009.
[0093] [9]M.Abadi et ah,“TensorFlow:A System for Large-Scale MachineLearning This paper is included in the Proceedings of the TensorFlow:A systemfor large-scale machine learning,”Proc 12th USENIXConf.Oper.Syst.Des.Implement.,2016.
[0094]
[10] Y.Jia et al.,“Caffe:Convolutional Architecture for Fast FeatureEmbedding,”2014.
[0095]
[11] “The MNIST Database.”[Online].Available:http: / / yann.lecun.com / exdb / mnist / .
[0096]
[12] “The UCF 101 Dataset.”[Online]Available:https: / / w w w.crcv.ucf.edu / data / U CF 101.php.
[0097]
[13] “The ImageNet Dataset.”[Online].Available:http: / / www.image-net.org / .
[0098]
[14] B.Hinton,G.,Deng,L.,Yu,D.,Dahl,G.,Mohamed,A.,Jaitly,N.,...Kingsbury,“Deep Neural Networks for Acoustic Modeling in SpeechRecognition.,”IEEE Signal Processing Magazine,2012.
[0099]
[15] “IMAGENET Large Scale Visual Recognition Challenge.”[Online].Available:http: / / image-net.org / challenges / LSVRC / 2017 / results.
[0100]
[16] D.Silver et ah,“Mastering the game of Go with deep neuralnetworks and tree search,”Nature,2016.
[0101]
[17] H.J.Caulfield and S.Dolev,“Why future supercomputing requiresoptics,”Nat.Photonics,vol.4,no.5,pp.261-263,May 2010.
[0102]
[18] D.Brunner,S.Reitzenstein,and I.Fischer,“All-optical neuromorphiccomputing in optical networks of semiconductor lasers,”in 2016 IEEEInternational Conference on Rebooting Computing,ICRC 2016-ConferenceProceedings,2016.
[0103]
[19] T.Deng,J.Robertson,and A.Hurtado,“Controlled Propagation ofSpiking Dynamics in Vertical-Cavity Surface-Emitting Lasers:TowardsNeuromorphic Photonic Networks,”IEEE J.Sel.Top.Quantum Electron.,2017.
[0104]
[20] T.Ferreira de Lima,B.J.Shastri,A.N.Tait,M.A.Nahmias,andP.R.Prucnal,“Progress in neuromorphic photonics,”Nanophotonics,2017.
[0105]
[21] A.N.Tait et ah,“Neuromorphic photonic networks using siliconphotonic weight banks,”Sci.Rep.,2017.
[0106]
[22] H.T.Peng,M.A.Nahmias,T.F.De Lima,A.N.Tait,B.J.Shastri,andP.R.Prucnal,“Neuromorphic Photonic Integrated Circuits,”IEEEJ.Sel.Top.Quantum Electron.,2018.
[0107]
[23] J.K.George et ah,“Neuromorphic photonics with electro-absorptionmodulators,”Opt.Express,vol.27,no.4,p.5181,Feb.2019.
[0108]
[24] K.Vandoome et ah,“Toward optical signal processing using PhotonicReservoir Computing,”Opt.Express,vol.16,no.15,p.11182,Jul.2008.
[0109]
[25] F.Duport,B.Schneider,A.Smerieri,M.Haelterman,and S.Massar,“All-optical reservoir computing,”Opt.Express,2012.
[0110]
[26] L.Pesquera et ah,“Photonic information processing beyond Turing:an optoelectronic implementation of reservoir computing,”Opt.Express,vol.20,no.3,p.3241,2012.
[0111]
[27] Y.Paquot et ah,“Optoelectronic reservoir computing,”Sci.Rep.,2012.
[0112]
[28] A.Dejonckheere et ah,“All-optical reservoir computer based onsaturation of absorption,”Opt.Express,2014.
[0113]
[29] L.Larger,A.Baylon-Fuentes,R.Martinenghi,V.S.Udaltsov,Y.K.Chembo,and M.Jacquot,“High-speed photonic reservoir computing using a time-delay-based architecture:Million words per second classification,”Phys.Rev.X,2017.
[0114]
[30] A.Katumba,J.Heyvaert,B.Schneider,S.Uvin,J.Dambre,and P.Bienstman,“Low-Loss Photonic Reservoir Computing with Multimode Photonic IntegratedCircuits,”Sci.Rep.,2018.
[0115]
[31] N.P.Jouppi et al.,“In-Datacenter Performance Analysis of a TensorProcessing Unit,”in Proceedings of the 44th Annual International Symposium onComputer Architecture-ISCA17,2017,pp.1-12.
[0116]
[32] M.Mathieu,M.Henaff,and Y.LeCun,“Fast Training of ConvolutionalNetworks through FFTs,”pp.1-9,2013.
[0117]
[33] A.Krizhevsky,I.Sutskever,and G.E.Hinton,“ImageNet Classificationwith Deep Convolutional Neural Networks,”Adv.Neural Inf.Process.Syst.,pp.1-9,2012.
[0118]
[34] K.Wagner and D.Psaltis,“Multilayer optical learning networks,”Appl.Opt.,vol.26,no.23,p.5061,Dec.1987.
[0119]
[35] D.Psaltis,D.Brady,X.G.Gu,and S.Lin,“Holography in artificialneural networks.,Nature,vol.343,no.6256,pp.325-30,1990.
[0120]
[36] X.Lin et ah,“All-optical machine learning using diffractive deepneural networks,Science(80-.).,2018.
[0121]
[37] M.Reck,A.Zeilinger,H.J.Bernstein,and P.Bertani,“Experimentalrealization of any discrete unitary operator,”Phys.Rev.Lett.,1994.
[0122]
[38] Y.Shen et al.,“Deep learning with coherent nanophotoniccircuits,”Nat.Photonics,vol.11,no.7,pp.441-446,Jul.2017.
[0123]
[39] R.Hamerly,A.Sludds,L.Bernstein,M.Soljacic,and D.Englund,“Large-Scale Optical Neural Networks based on Photoelectric Multiplication,”arXivPrepr.arXiv 1812.07614,pp.1-18,Nov.2018.
[0124]
[40] P.Merolla,R.Appuswamy,J.Arthur,S.K.Esser,and D.Modha,“Deep neuralnetworks are robust to weight binarization and other non-linear distortions,”arXiv Prepr.arXivl606.01981,Jun.2016.
[0125]
[41] Y.Wang,N.Zhao,Z.Yang,Z.Zhang,B.Huang,and G.Li,“Few-mode SDMreceivers exploiting parallelism of free space,”IEEE Photonics Journal,2019.
[0126]
[42] B.Huang,C.Xia,G.Matz,N.Bai,and G.Li,“Structured DirectionalCoupler Pair for Multiplexing of Degenerate Modes,”2013.
[0127]
[43] Y.Gao et al.,“Non-circularly-symmetric Mode-group DemultiplexerBased on Fused-type FMF Coupler for MGM Transmission,”2018.
[0128]
[44] T.A.Birks,I.Gris-Sanchez,S.Yerolatsitis,S.G.Leon-Saval,andR.R.Thomson,“The photonic lantern,”Adv.Opt.Photonics,2015.
[0129]
[45] B.Huang et al.,“All-fiber mode-group-selective photonic lanternusing graded-index multimode fibers,”Opt.Express,2015.
[0130]
[46] B.Huang et al.,“Triple-clad photonic lanterns for mode scaling,”Opt.Express,2018.
[0131]
[47] G.Labroille,B.Denolle,P.Jian,J.F.Morizur,P.Genevaux,and N.Treps,“Efficient and mode selective spatial mode multiplexer based on multi-planelight conversion,”in 2014 IEEE Photonics Conference,IPC 2014,2014.
[0132]
[48] N.K.Fontaine,R.Ryf,H.Chen,D.T.Neilson,K.Kim,and J.Carpenter,“Scalable mode sorter supporting 210 Hermite-Gaussian modes,”2018.
[0133]
[49] S.Bade et al.,“Fabrication and Characterization of a Mode-selective 45-Mode Spatial Multiplexer based on Multi-Plane Light Conversion,”in Optical Fiber Communication Conference Postdeadline Papers,2018,p.Th4B.3.
[0134]
[50] Z.I.Borevich and S.L.Krupetskii,“Subgroups of the unitary groupthat contain the group of diagonal matrices,”J.Sov.Math.,1981.
[0135]
[51] J.-F.Morizur et ah,“Programmable unitary spatial modemanipulation,”J.Opt.Soc.Am.A,vol.27,no.11,p.2524,Nov.2010.
[0136]
[52] H.Takahashi,T.Saida,Y.Sakamaki,and T.Hashimoto,“Wavefrontmatching method:A new approach for need-oriented waveguide design,”inConference Proceedings-Lasers and Electro-Optics Society Annual Meeting-LEOS,2005.
[0137]
[53] T.Umezawa et ah,“10-GHz 32-pixel 2-D photodetector array foradvanced optical fiber communications,”2017.
[0138]
[54] D.J.Brady et ah,“Multiscale gigapixel photography,”Nature,vol.486,no.7403,pp.386-389,2012.
[0139]
[55] L.Gao,J.Liang,C.Li,and L.V Wang,“Single-shot compressed ultrafastphotography at one hundred billion frames per second,”Nature,vol.516,no.7529,pp.74-77,2014.
[0140]
[56] K.Simonyan and A.Zisserman,“Very Deep Convolutional Networks forLarge-Scale Image Recognition,”pp.1-14,2014.
Claims
1. A photonic unit for implementing multiplication of an N×M matrix with an M×1 vector, comprising: A first optical multiplexer receives and multiplexes at least one first optical signal, which represents at least one M×1 vector having M elements, wherein each element in the M×1 vector is encoded on a first orthogonal degree of freedom of light and is non-temporal in one multiplication cycle, to generate a first multiplexed optical signal, wherein M is a positive integer greater than or equal to 1. An optical replicator for replicating the first optical signal representing an M×1 vector into N copies of N additional optical signals in the second orthogonal degree of freedom of light, where N is a positive integer greater than or equal to 1; N optical multiplexers, identical to optical replicators in the second orthogonal degree of freedom of light, each receive and multiplex M additional optical signals, each of which is coherent with the first optical signal. Each of the N additional multiplexed optical signals represents an independent row of an N×M matrix, where each element in a row of the N×M matrix is encoded with an element-orthogonal degree of freedom corresponding to the first optical signal mapped to an M×1 vector, to generate the N additional multiplexed optical signals; and At least one beam combiner receives N copies of a first multiplexed optical signal from a first optical multiplexer and N additional multiplexed optical signals from N optical multiplexers, in order to combine the two to generate N interferences between the first optical signal and each of the N additional optical signals, wherein the N interferences include the result of multiplication of an M×N matrix and an M×1 vector in the total interference intensity. At least one of the first and second orthogonal degrees of freedom of the light used for encoding or copying is one of the components of wavelength, spatial mode, polarization, orthogonality, and wave vector; or, at least one of the first and second orthogonal degrees of freedom of the light used for encoding or copying is a superdimensionality consisting of a combination of two or more degrees of freedom of light. At least one of the interfering signals enters a nonlinear optical element; at least one total interference intensity is converted into an electrical signal; the electrical signal enters a nonlinear electronic element.
2. The photonic unit according to claim 1, wherein the spatial mode is at least one of the following: Hermi-Gaussian mode, Laguerre-Gaussian model, or Discrete space samples that form a spatial orthogonal basis.
3. A photonic unit for implementing multiplication of an N×M matrix and an M×W matrix, comprising: The first group of N optical multiplexers in the first orthogonal degree of freedom of light receives and multiplexes N optical signals, each of which represents an independent row of an N×M matrix with M elements, wherein each element in each independent row of the N×M matrix is encoded in the second orthogonal degree of freedom of light and is non-temporal in one multiplication period to produce N multiplexed optical signals, wherein M and N are each positive integers greater than or equal to 1; The first optical replicator is used to replicate each of the N multiplexed optical signals representing multiple independent rows of an N×M matrix as W copies of the third orthogonal degree of freedom of the light, where W is a positive integer greater than or equal to 1. The second group of W optical multiplexers in the third orthogonal degree of freedom of light, which is the same as the first optical replicator, receives and multiplexes W additional optical signals, each of which is coherent with N optical signals, and each of which represents an independent column with M elements in an M×W matrix, wherein each element in each independent column of the M×W matrix is encoded with the same element-orthogonal degree of freedom of the optical mapping as each element in each independent row of the N×M matrix to produce W additional multiplexed optical signals; A second optical replicator, used to replicate each of the W additional multiplexed optical signals representing the independent columns of an M×W matrix as N copies in the first orthogonal degree of freedom of the light, is identical to the first set of N optical multiplexers; and At least one beam combiner receives two sets of N×W multiplexed optical signals, the two sets of N×W multiplexed optical signals representing appropriately copied rows or columns of each of an N×M matrix and an M×W matrix, in order to combine them to produce N×W interferences between each row of the N×M matrix and a column of the M×W matrix, including the result of multiplication of the total N×W interference intensity. At least one of the first, second, and third orthogonal degrees of freedom of the light used for encoding or copying is one of the components of wavelength, spatial mode, polarization, orthogonality, and wave vector; or, At least one of the first, second, and third orthogonal degrees of freedom of light used for encoding or copying is a superdimensionality consisting of a combination of two or more degrees of freedom of light; At least one of the interfering signals enters a nonlinear optical element; at least one total interference intensity is converted into an electrical signal; the electrical signal enters a nonlinear electronic element.
4. The photonic unit according to claim 3, wherein the spatial mode is at least one of the following: Hermi-Gaussian mode, Laguerre-Gaussian model, or Discrete space samples that form a spatial orthogonal basis.