Photonic computing system and method of optically performing matrix vector multiplication

By using a photonic computing system and a photonic processor to perform matrix-vector multiplication, the problem of low efficiency of conventional computing processors in computationally intensive algorithms is solved, achieving efficient matrix multiplication acceleration and low-latency computing.

CN114912604BActive Publication Date: 2026-05-08LIGHT MATERIALS CO
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LIGHT MATERIALS CO
Filing Date
2019-05-14
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Conventional computing processors are inefficient at executing computationally intensive algorithms such as graphics processing, artificial intelligence, and deep learning, and electrical signal processing suffers from propagation delays and power dissipation.

Method used

A photonic computing system is employed, which utilizes a photonic processor to perform matrix-vector multiplication through optical signals. This system includes a light source, an optical encoder, a wavelength division multiplexer, and a photonic multiplier to achieve highly parallel linear transformations.

Benefits of technology

It improves the speed of matrix multiplication by two orders of magnitude compared to conventional techniques, reduces propagation delay and power dissipation, and is suitable for specific types of algorithms such as machine learning and neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114912604B_ABST
    Figure CN114912604B_ABST
Patent Text Reader

Abstract

Aspects relate to photonic computing systems and methods of optically performing matrix vector multiplication. A photonic computing system includes a light source configured to generate a plurality of input optical signals having mutually different wavelengths, a plurality of optical encoders configured to encode the plurality of input optical signals with respective input values, a first wavelength division multiplexer (WDM) configured to spatially combine the encoded plurality of input optical signals into an input waveguide, a photonic multiplier coupled to the input waveguide and configured to generate a plurality of output optical signals each encoded with a respective output value representing a product of the respective input value and a common weight value, and a second WDM coupled to the photonic multiplier and configured to spatially separate the plurality of output optical signals.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese national phase patent application with application number 201980046744.8, which was filed on January 12, 2021, after the international application PCT application with application number PCT / US2019 / 032181, international application date May 14, 2019, entitled "Photonics Processing System and Method".

[0002] Cross-references to related applications

[0003] This application claims priority to U.S. Provisional Patent Application Serial No. 62 / 671,793, filed May 15, 2018, entitled “ALGORITHMS FOR TRAINING NEURAL NETWORKS WITH PHOTONICHARDWARE ACCELERATORS”, filed under 35 U.S. SC §119(e), Agent’s Case No. L0858.70001US00, which is incorporated herein by reference in its entirety.

[0004] This application also claims priority to U.S. Provisional Patent Application Serial No. 62 / 680,557, filed June 4, 2018, entitled “PHOTONICS PROCESSING SYSTEMS AND METHODS”, Agent’s Case No. L0858.70000US00, filed under 35 U.S. SC §119(e), which is incorporated herein by reference in its entirety.

[0005] This application also claims priority to U.S. Provisional Patent Application Serial No. 62 / 689,022, filed June 22, 2018, entitled “CONVOLUTIONAL LAYERS FOR NEURAL NETWORKS USING PROGRAMMABLE NANOPHOTONICS”, filed under 35 U.S. SC §119(e), Agent’s Case No. L0858.70003US00, which is incorporated herein by reference in its entirety.

[0006] This application also claims priority to U.S. Provisional Patent Application Serial No. 62 / 755,402, filed November 11, 2018, entitled “REAL-NUMBER PHOTONIC ENCODING,” filed under 35 U.S.SC §119(e), Agent’s Case No. L0858.70008US00, which is incorporated herein by reference in its entirety.

[0007] This application also claims priority to U.S. Provisional Patent Application Serial No. 62 / 792,720, filed January 15, 2019, entitled “HIGH-EFFICIENCY DOUBLE-SLOT WAVEGUIDE NANO-OPTOELECTROMECHANICAL PHASE MODULATOR”, filed under 35 U.S.SC §119(e), Agent’s Case No. L0858.70006US00, which is incorporated herein by reference in its entirety.

[0008] This application also claims priority to U.S. Provisional Patent Application Serial No. 62 / 793,327, filed January 16, 2019, entitled “DIFFERENTIAL, LOW-NOISE HOMODYNE RECEIVER,” Agent’s Case No. L0858.70004US00, filed under 35 U.S. SC §119(e), which is incorporated herein by reference in its entirety.

[0009] This application also claims priority to U.S. Provisional Patent Application Serial No. 62 / 834,743, filed April 16, 2019, entitled “STABILIZING LOCAL OSCILLATOR PHASES IN A PHOTOCORE,” Agent’s Case No. L0858.70014US00, filed under 35 U.S. SC §119(e), which is incorporated herein by reference in its entirety. Background Technology

[0010] The processors used in conventional computing consist of circuits composed of millions of transistors to implement logic gates on bits of information represented by electrical signals. The architecture of a conventional central processing unit (CPU) is designed for general-purpose computing and is not optimized for specific types of algorithms. Some examples of computationally intensive algorithms that cannot be efficiently executed by a CPU include graphics processing, artificial intelligence, neural networks, and deep learning.

[0011] Therefore, dedicated processors with architectures better suited to specific algorithms have been developed. For example, graphics processing units (GPUs) have highly parallel architectures, making them more efficient than CPUs when performing image processing and graphics operations. After GPUs were developed for graphics processing, it was also found that GPUs are more efficient than CPUs for other memory-intensive algorithms such as neural networks and deep learning. This understanding, along with the increasing popularity of artificial intelligence and deep learning, has spurred further research into novel circuit architectures that can further improve the speed of these algorithms. Summary of the Invention

[0012] In some embodiments, a photonic computing system is provided. The photonic computing system may include: a light source configured to generate multiple input optical signals having mutually different wavelengths; multiple optical encoders configured to encode the multiple input optical signals using their respective input values; a first wavelength division multiplexing (WDM) configured to spatially combine the encoded multiple input optical signals into an input waveguide; a photonic multiplier coupled to the input waveguide and configured to generate multiple output optical signals, each output optical signal being encoded with a corresponding output value representing the product of the corresponding input value and a common weight value; and a second WDM coupled to the photonic multiplier and configured to spatially separate the multiple output optical signals.

[0013] In some embodiments, a method for optically performing matrix-vector multiplication is provided. The method may include: generating a first plurality of optical signals having different wavelengths; receiving digital representations of a plurality of input vectors; encoding the plurality of input vectors into the first plurality of optical signals using a plurality of optical encoders; combining the encoded first plurality of optical signals into an input waveguide using an input wavelength division multiplexer; performing a plurality of operations on the encoded first plurality of optical signals using a photonic processor to generate a second plurality of optical signals, the plurality of operations performing matrix multiplication of the plurality of input vectors with a matrix; dividing the second plurality of optical signals into a plurality of output waveguides using an output wavelength division multiplexer coupled to the photonic processor; and detecting the second plurality of optical signals received from the plurality of output waveguides using a plurality of photodetectors.

[0014] In some embodiments, a photonic computing system is provided. The photonic computing system may include: a light source configured to generate a plurality of input optical signals having mutually different wavelengths; a first wavelength division multiplexer (WDM) configured to spatially combine the plurality of input optical signals into an input waveguide; a photonic processor coupled to the input waveguide and configured to generate a plurality of output optical signals by performing matrix-matrix multiplication or matrix-vector multiplication using the plurality of input optical signals; a second WDM coupled to the photonic processor and configured to spatially separate the plurality of output optical signals; a plurality of optical receivers coupled to the second WDM and configured to generate a plurality of output analog signals using the plurality of output optical signals, wherein each of the plurality of optical receivers includes at least two photodetectors; and a plurality of signal combiners, each signal combiner coupled to at least two photodetectors of a corresponding optical receiver.

[0015] The foregoing apparatus and method embodiments can be implemented using any suitable combination of the aspects, features, and actions described above or in more detail below. These and other aspects, embodiments, and features of this teaching can be more fully understood from the following description taken in conjunction with the accompanying drawings. Attached Figure Description

[0016] Various aspects and embodiments of this application will be described with reference to the following accompanying drawings. It should be understood that the drawings are not necessarily drawn to scale. Items appearing in multiple drawings are indicated by the same reference numerals in all the drawings in which they appear.

[0017] Figure 1-1 This is a schematic diagram of a photonics processing system according to some non-limiting embodiments.

[0018] Figure 1-2 This is a schematic diagram of an optical encoder according to some non-limiting embodiments.

[0019] Figure 1-3 This is a schematic diagram of a photonic processor according to some non-limiting embodiments.

[0020] Figure 1-4 This is a schematic diagram of an interconnected variable beam splitter array according to some non-limiting embodiments.

[0021] Figure 1-5 This is a schematic diagram of a variable beam splitter according to some non-limiting embodiments.

[0022] Figure 1-6 This is a schematic diagram illustrating diagonal attenuation and phase shift implementation according to some non-limiting embodiments.

[0023] Figure 1-7 This is a schematic diagram of an attenuator according to some non-limiting embodiments.

[0024] Figure 1-8 This is a schematic diagram of a power tree according to some non-limiting embodiments.

[0025] Figure 1-9 This is a schematic diagram of an optical receiver according to some non-limiting embodiments.

[0026] Figure 1-10 This is a schematic diagram of a zero-difference detector according to some non-limiting embodiments.

[0027] Figure 1-11 This is a schematic diagram of a folded photon processing system according to some non-limiting embodiments.

[0028] Figure 1-12A This is a schematic diagram of a wavelength division multiplexing (WDM) photonic processing system according to some non-limiting embodiments.

[0029] Figure 1-12B According to some non-limiting embodiments Figure 1-12A A schematic diagram of the front end of a wavelength division multiplexing (WDM) photonics processing system.

[0030] Figure 1-12C According to some non-limiting embodiments Figure 1-12AA schematic diagram of the back end of a wavelength division multiplexing (WDM) photonics processing system.

[0031] Figure 1-13 This is a schematic diagram of a circuit for performing analog summation of optical signals according to some non-limiting embodiments.

[0032] Figure 1-14 This is a schematic diagram of a photonics processing system with the illustrated column-global phase according to some non-limiting embodiments.

[0033] Figure 1-15 This is a graph showing the effect of uncorrected global phase shift on zero-difference detection according to some non-limiting embodiments.

[0034] Figure 1-16 It is a graph showing the orthogonal uncertainty of the coherent state of light according to some non-limiting embodiments.

[0035] Figure 1-17 This is a diagram illustrating matrix multiplication according to some non-limiting embodiments.

[0036] Figure 1-18 This is an illustration of performing matrix multiplication by subdividing a matrix into submatrices according to some non-limiting embodiments.

[0037] Figure 1-19 This is a flowchart of a method for manufacturing a photonics processing system according to some non-limiting embodiments.

[0038] Figure 1-20 This is a flowchart of a method for manufacturing a photonic processor according to some non-limiting embodiments.

[0039] Figure 1-21 This is a flowchart of a method for performing optical calculations according to some non-limiting embodiments.

[0040] Figure 2-1 This is a flowchart of a process for training a latent variable model according to some non-limiting embodiments.

[0041] Figure 2-2 This is a flowchart of a process for configuring a photon processing system to implement a unitary transformation matrix, according to some non-limiting embodiments.

[0042] Figure 2-3 This is a flowchart illustrating the process of calculating an error vector using a photonics processing system according to some non-limiting embodiments.

[0043] Figure 2-4 This is a flowchart of a process for determining update parameters of a unitary transformation matrix according to some non-limiting embodiments.

[0044] Figure 2-5This is a flowchart of a process for updating parameters of a unitary transformation matrix according to some non-limiting embodiments.

[0045] Figure 3-1 This is a flowchart of a method for computing forward passes through a convolutional layer, according to some non-limiting embodiments.

[0046] Figure 3-2 This is a flowchart of a method for computing forward passes through a convolutional layer, according to some non-limiting embodiments.

[0047] Figure 3-3A This is a flowchart of a method suitable for computing two-dimensional convolution according to some non-limiting embodiments.

[0048] Figure 3-3B This is a flowchart of a method suitable for constructing a cyclic matrix according to some non-limiting embodiments.

[0049] Figure 3-4A The diagram illustrates a preprocessing step for constructing a filter matrix from an input filter matrix comprising multiple output channels, according to some non-limiting embodiments.

[0050] Figure 3-4B A circular matrix is ​​constructed from an input matrix comprising multiple input channels according to some non-limiting embodiments.

[0051] Figure 3-4C Two-dimensional matrix multiplication operations are illustrated according to some non-limiting embodiments.

[0052] Figure 3-4D Post-processing steps for a rotated vector row according to some non-limiting embodiments are shown.

[0053] Figure 3-4E Post-processing steps for vector row addition according to some non-limiting embodiments are shown.

[0054] Figure 3-4F The diagram illustrates how an output matrix can be reshaped into multiple output channels according to some non-limiting embodiments.

[0055] Figure 3-5 This is a flowchart of a method for performing a one-dimensional Fourier transform according to some non-limiting embodiments.

[0056] Figure 3-6 This is a flowchart of a method for performing a two-dimensional Fourier transform according to some non-limiting embodiments.

[0057] Figure 3-7 This is a flowchart of a method for performing a two-dimensional Fourier transform according to some non-limiting embodiments.

[0058] Figure 3-8This is a flowchart of a method for performing convolution using Fourier transform, according to some non-limiting embodiments.

[0059] Figure 4-1 This is a block diagram illustrating an optical modulator according to some non-limiting embodiments.

[0060] Figure 4-2A This is a block diagram illustrating an example of a photonic system according to some non-limiting embodiments.

[0061] Figure 4-2B This is a flowchart illustrating an example of a method for processing signed real numbers in the optical domain, according to some non-limiting embodiments.

[0062] Figure 4-3A This illustrates how, according to some non-limiting embodiments, it can be used with... Figure 4-2A A schematic diagram illustrating an example of a modulator used in conjunction with a photonic system.

[0063] Figure 4-3B This illustrates, according to some non-limiting embodiments, when no voltage is applied. Figure 4-3A The graph shows the intensity and phase spectrum response of the modulator.

[0064] Figure 4-3C This illustrates, according to some non-limiting embodiments, when driven at a specific voltage Figure 4-3A The graph shows the intensity and phase spectrum response of the modulator.

[0065] Figure 4-3D Based on some non-limiting embodiments and Figure 4-3A An example of an encoding table associated with a modulator.

[0066] Figure 4-3E According to some non-limiting embodiments Figure 4-3A The visual representation of the modulator's spectral response in the complex plane.

[0067] Figures 4-4A to 4-4C It is a visual representation in the complex plane of different spectral responses at the output of the light conversion unit according to some non-limiting embodiments.

[0068] Figure 4-5 This describes how optical signal detection is performed in a visual representation in a complex plane, according to some non-limiting embodiments.

[0069] Figure 4-6 This is an example of a decoding table based on some non-limiting embodiments.

[0070] Figure 4-7 This is a block diagram of an optical communication system according to some non-limiting embodiments.

[0071] Figure 4-8This illustrates a method based on some non-limiting embodiments. Figure 4-7 An example graph of the power spectral density output of the modulator.

[0072] Figure 4-9 This is a flowchart illustrating an example of a method for manufacturing a photonic system according to some non-limiting embodiments.

[0073] Figure 5-1 This is a circuit diagram illustrating an example of a differential optical receiver according to some non-limiting embodiments.

[0074] Figure 5-2 This illustrates how, according to some non-limiting embodiments, it can be used with... Figure 5-1 A schematic diagram of a photonic circuit coupled to a differential optical receiver.

[0075] Figure 5-3A This is a schematic diagram showing a substrate including a photonic circuit, a photodetector, and a differential operational amplifier according to some non-limiting embodiments.

[0076] Figure 5-3B This is a schematic diagram showing a first substrate including photonic circuitry and a photodetector and a second substrate including a differential operational amplifier, according to some non-limiting embodiments, wherein the first substrate and the second substrate are flip-chip bonded to each other.

[0077] Figure 5-3C This is a schematic diagram showing a first substrate including photonic circuitry and a photodetector and a second substrate including a differential operational amplifier, according to some non-limiting embodiments, wherein the first substrate and the second substrate are wire-bonded to each other.

[0078] Figure 5-4 This is a flowchart illustrating an example of a method for manufacturing an optical receiver according to some non-limiting embodiments.

[0079] Figures 5-4A to 5-4F An example of the manufacturing sequence of an optical receiver according to some non-limiting embodiments is shown.

[0080] Figure 5-5 This is a flowchart illustrating an example of a method for receiving optical signals according to some non-limiting embodiments.

[0081] Figure 6-1A This is a schematic top view of a nano-opto-electro-mechanical system (NOEMS) phase modulator according to some non-limiting embodiments.

[0082] Figure 6-1B This is an illustrative representation of some non-limiting embodiments. Figure 6-1A A top view of the suspended multi-slit optical structure of the NOEMS phase modulator.

[0083] Figure 6-1C This illustrates some non-limiting embodiments in Figure 6-1B A graph illustrating an example of the optical modes generated in a suspended multi-slit optical structure.

[0084] Figure 6-1D This is an illustrative representation of some non-limiting embodiments. Figure 6-1A A top view of the mechanical structure of the NOEMS phase modulator.

[0085] Figure 6-1E This is an illustrative representation of some non-limiting embodiments. Figure 6-1A A top view of the transition region of the NOEMS phase modulator.

[0086] Figure 6-2 According to some non-limiting embodiments Figure 6-1A A cross-sectional view of the NOEMS phase modulator taken along the yz plane is shown, along with the suspended waveguide.

[0087] Figure 6-3 According to some non-limiting embodiments Figure 6-1A A cross-sectional view of the NOEMS phase modulator taken along the xy plane, showing a portion of the suspended multi-slit optical structure.

[0088] Figures 6-4A to 6-4C This is a cross-sectional view illustrating, according to some non-limiting embodiments, how a suspended multi-slit optical structure can be mechanically driven to change the width of the slits between waveguides.

[0089] Figure 6-5 This is a graph showing how the effective refractive index of a suspended multi-slit optical structure, according to some non-limiting embodiments, can vary as a function of the slit width.

[0090] Figure 6-6 This is a flowchart illustrating an example of a method for manufacturing a NOEMS phase modulator according to some non-limiting embodiments. Detailed Implementation

[0091] I. Light Core

[0092] A. Overview of photonics-based processing

[0093] The inventors have recognized and understood that conventional circuit-based processors have limitations in speed and efficiency. Each wire and transistor in the circuitry of an electrical processor has resistance, inductance, and capacitance, which cause propagation delays and power dissipation in electrical signals. For example, using conductive traces with non-zero impedance to connect multiple processor cores and / or connect processor cores to memory. High impedance values ​​limit the maximum rate at which data can be transmitted through the traces at a negligible bit error rate. In applications where time delay is critical, such as high-frequency stock trading, even delays of a few hundredths of a second can render the algorithm infeasible. For processing requiring billions of transistors to perform billions of operations, these delays add up to a significant waste of time. Besides the slow speed of circuits, the heat generated by energy dissipation due to circuit impedance is also an obstacle to the development of electrical processors. The inventors have also recognized and understood that using optical signals instead of electrical signals overcomes many of the aforementioned problems with electrical computing. Optical signals travel at the speed of light in the medium in which light travels; therefore, the time delay of photonic signals is much smaller than the delay of electrical propagation. Furthermore, increasing the propagation distance of optical signals does not result in power dissipation, providing opportunities to use new topologies and processor layouts that are not feasible for electrical signals. Therefore, light-based processors (such as photonics-based processors) can achieve better speed and efficiency performance than conventional electronic processors.

[0094] Furthermore, the inventors have recognized and appreciated that light-based processors (e.g., photon-based processors) are well-suited for certain types of algorithms. For example, many machine learning algorithms, such as support vector machines, artificial neural networks, and probabilistic graphical model learning, heavily rely on linear transformations over multidimensional arrays / tensors. The simplest example is multiplying a vector by a matrix, which using conventional algorithms has an O(n) time complexity. 2 The complexity is 1 / n, where n is the dimension of the square matrix being multiplied. The inventors have recognized and appreciated that photonic processors, in some embodiments, can be highly parallel linear processors capable of performing linear transformations, such as matrix multiplication, in a highly parallel manner by propagating a specific set of input optical signals through a configurable beamsplitter array. Using this implementation, matrix multiplication of a matrix of dimension n = 512 can be completed in hundreds of picoseconds, unlike the tens to hundreds of nanoseconds using conventional processing. Using some embodiments, the speed of matrix multiplication is estimated to be two orders of magnitude faster than conventional techniques. For example, multiplications that prior art graphics processing units (GPUs) can perform in about 10 ns can be performed in about 200 ps by a photonic processing system according to some embodiments.

[0095] To realize a photonics-based processor, the inventors have recognized and understood that the multiplication of an input vector and a matrix can be achieved by propagating a coherent optical signal (e.g., a laser pulse) through a first interconnected variable beam splitter (VBS) array, a second interconnected variable beam splitter array, and a plurality of controllable optical elements (e.g., electro-optic or opto-mechanical elements) located between the two arrays, connecting a single output of the first array to a single input of the second array. Details of a photonics processing system including a photonics processor will be described below.

[0096] B. Overview of Photon Processing Systems

[0097] Reference Figure 1-1 According to some embodiments, the photonics processing system 1-100 includes an optical encoder 1-101, a photonics processor 1-103, an optical receiver 1-105, and a controller 1-107. The photonics processing system 1-100 receives an input vector represented by a set of input bit strings as input from an external processor (e.g., a CPU) and produces an output vector represented by a set of output bit strings. For example, if the input vector is an n-dimensional vector, the input vector can be represented by n separate bit strings, each bit string representing a corresponding component of the vector. The input bit strings can be received from the external processor as electrical or optical signals, and the output bit strings can be transmitted to the external processor as electrical or optical signals. In some embodiments, the controller 1-107 does not necessarily output an output bit string after each process iteration. Instead, the controller 1-107 can use one or more output bit strings to determine a new input bit stream to be fed through the components of the photonics processing system 1-100. In some embodiments, the output bit strings themselves can be used as input bit strings for subsequent iterations of the process implemented by the photonics processing system 1-100. In other embodiments, multiple output bitstreams are combined in various ways to determine a subsequent input bitstream. For example, one or more output bitstreams can be added together as part of determining the subsequent input bitstream.

[0098] Optical encoder 1-101 is configured to convert input bit strings into optically encoded information for processing by photonic processor 1-103. In some embodiments, controller 1-107 transmits each input bit string to optical encoder 1-101 in the form of an electrical signal. Optical encoder 1-101 converts each component of the input vector from its digital bit string into an optical signal. In some embodiments, the optical signal represents the value and sign of the associated bit string as the amplitude and phase of an optical pulse. In some embodiments, the phase may be limited to a binary choice of zero phase shift or π phase shift representing positive and negative values, respectively. Embodiments are not limited to real-valued input vector values. When encoding optical signals, complex vector components can be represented, for example, by using more than two phase values. In some embodiments, the bit string is received by optical encoder 1-101 as an optical signal (e.g., a digital optical signal) from controller 1-107. In these embodiments, optical encoder 1-101 converts the digital optical signal into an analog optical signal of the type described above.

[0099] Optical encoder 1-101 outputs n individual optical pulses, which are transmitted to photonic processor 1-103. Each output of optical encoder 1-101 is coupled one-to-one to a single input of photonic processor 1-103. In some embodiments, optical encoder 1-101 and photonic processor 1-103 may be disposed on the same substrate (e.g., optical encoder 1-101 and photonic processor 1-103 are located on the same chip). In such embodiments, the optical signal can be transmitted from optical encoder 1-101 to photonic processor 1-103 in a waveguide (e.g., a silicon photonic waveguide). In other embodiments, optical encoder 1-101 and photonic processor 1-103 may be disposed on separate substrates. In such embodiments, the optical signal can be transmitted from optical encoder 1-101 to photonic processor 103 in an optical fiber.

[0100] Photonic processor 1-103 performs a multiplication of the input vector with matrix M. As described in detail below, matrix M is decomposed into three matrices using a combination of singular value decomposition (SVD) and unitary matrix decomposition. In some embodiments, unitary matrix decomposition is performed with an operation similar to Givens rotation in QR decomposition. For example, SVD combined with Householder decomposition can be used. The decomposition of matrix M into three components can be performed by controller 1-107, and each component can be implemented by a portion of photonic processor 1-103. In some embodiments, photonic processor 1-103 includes three parts: a first variable beam splitter (VBS) array configured to perform a transformation on the input optical pulse array, which is equivalent to multiplying with the first matrix (see, for example...). Figure 1-3The first matrix in the array is implemented as 1-301; a set of controllable optical elements are configured to adjust the intensity and / or phase of each optical pulse received from the first array, the adjustment being equivalent to multiplying the second matrix by a diagonal matrix (see example...). Figure 1-3 The second matrix implementation (1-303); and the second variable beam splitter array, which is configured to transform the optical pulses received from the set of controllable electro-optic elements, the transformation being equivalent to multiplying with the third matrix (see, for example, the third matrix implementation 1-305 in Figure 3).

[0101] Photonic processor 1-103 outputs n individual optical pulses, which are transmitted to optical receiver 1-105. Each output of photonic processor 1-103 is coupled one-to-one to a single input of optical receiver 1-105. In some embodiments, photonic processor 1-103 and optical receiver 1-105 may be disposed on the same substrate (e.g., photonic processor 1-103 and optical receiver 1-105 are located on the same chip). In such an embodiment, the optical signal can be transmitted from photonic processor 1-103 to optical receiver 1-105 in a silicon photonic waveguide. In other embodiments, photonic processor 1-103 and optical receiver 1-105 may be disposed on separate substrates. In such an embodiment, the optical signal can be transmitted from photonic processor 103 to optical receiver 1-105 in an optical fiber.

[0102] Optical receiver 1-105 receives n optical pulses from photonic processor 1-103. Each optical pulse is then converted into an electrical signal. In some embodiments, the intensity and phase of each optical pulse are measured by a photodetector within the optical receiver. The electrical signals representing these measurements are then output to controller 1-107.

[0103] Controller 1-107 includes memory 1-109 and processor 1-111 for controlling optical encoder 1-101, photonic processor 1-103, and optical receiver 1-105. Memory 1-109 can be used to store input and output bit strings from optical receiver 1-105, as well as measurement results. Memory 1-109 also stores executable instructions that, when executed by processor 1-111, control optical encoder 1-101, execute matrix factorization algorithms, control VBS of photonic processor 103, and control optical receiver 1-105. Memory 1-109 may also include executable instructions that cause processor 1-111 to determine a new input vector to be sent to the optical encoder based on a set of one or more output vectors determined by measurements performed by optical receiver 1-105. Thus, controller 1-107 can control the iterative process of multiplying the input vector with multiple matrices by adjusting the settings of photonic processor 1-103 and feeding detection information from optical receiver 1-105 back to optical encoder 1-101. Therefore, the output vector transmitted from the photonic processing system 1-100 to the external processor can be the result of multiple matrix multiplications, rather than just a single matrix multiplication.

[0104] In some embodiments, the matrix may be too large to be encoded in a single pass within the photonic processor. In this case, a portion of the large matrix can be encoded in the photonic processor, and a multiplication process can be performed on that single portion of the large matrix. The result of this first operation can be stored in memories 1-109. Subsequently, a second portion of the large matrix can be encoded in the photonic processor, and a second multiplication process can be performed. This “blocking” of the large matrix can continue until multiplication processes have been performed on all portions of the large matrix. The results of multiple multiplication processes can be stored in memories 1-109 and then combined to form the final result of multiplying the input vector by the large matrix.

[0105] In other embodiments, the external processor uses only the collective behavior of the output vectors. In such embodiments, only collective results, such as the average or maximum / minimum values ​​of the multiple output vectors, are transmitted to the external processor.

[0106] C. optical encoder

[0107] Reference Figure 1-2 According to some embodiments, the optical encoder includes at least one light source 1-201, a power tree 1-203, an amplitude modulator 1-205, a phase modulator 1-207, a digital-to-analog converter (DAC) 1-209 associated with the amplitude modulator 1-205, and a 1-DAC 211 associated with the phase modulator 1-207. Although the amplitude modulator 1-205 and the phase modulator 1-207 are... Figure 1-2The image is shown as a single block with n inputs and n outputs (each input and each output is, for example, a waveguide), but in some embodiments, each waveguide may include a corresponding amplitude modulator and a corresponding phase modulator, such that the optical encoder includes n amplitude modulators and n phase modulators. Furthermore, a separate DAC may be available for each amplitude and phase modulator. In some embodiments, a single modulator can be used to encode both the amplitude and phase information, rather than having an amplitude modulator and a separate phase modulator associated with each waveguide. While using a single modulator to perform such encoding limits the ability to precisely tune both the amplitude and phase of each optical pulse, there are encoding schemes that do not require precise tuning of both the amplitude and phase of the optical pulse. Such schemes are described later herein.

[0108] Light source 1-201 can be any suitable coherent light source. In some embodiments, light source 1-201 can be a diode laser or a vertical-cavity surface-emitting laser (VCSEL). In some embodiments, light source 1-201 is configured to have an output power greater than 10mW, greater than 25mW, greater than 50mW, or greater than 75mW. In some embodiments, light source 1-201 is configured to have an output power less than 100mW. Light source 1-201 can be configured to emit continuous light waves or pulsed light (“light pulses”) of one or more wavelengths (e.g., C-band or O-band). The duration of a light pulse can be, for example, approximately 100 ps.

[0109] Although light source 1-201 is in Figure 1-2 The light source 1-201 is shown on the same semiconductor substrate as other components of the optical encoder, but the embodiment is not limited thereto. For example, the light source 1-201 may be a separate laser package edge-bonded or surface-bonded to the optical encoder chip. Alternatively, the light source 1-201 may be completely off-chip, and the optical pulses may be coupled to the waveguide 1-202 of the optical encoder 1-101 via optical fiber and / or grating coupler.

[0110] Light source 1-201 is shown as two light sources 1-201a and 1-201b, but embodiments are not limited thereto. Some embodiments may include a single light source. Including multiple light sources 201a-b (which may include more than two light sources) can provide redundancy in the event of a failure of one light source. Including multiple light sources can extend the lifespan of the photonics processing system 1-100. The multiple light sources 1-201a-b may each be coupled to a waveguide of the optical encoder 1-101 and then combined at a waveguide combiner configured to direct light pulses from each light source to the power tree 1-203. In such embodiments, only one light source is used at any given time.

[0111] Some embodiments may use two or more phase-locked light sources of the same wavelength simultaneously to increase the optical power entering the optical encoder system. A small portion of the light from each of the two or more light sources (e.g., obtained via a waveguide tap) can be directed to a homodyne detector, where a beat error signal can be measured. The beat error signal can be used to determine the possible phase drift between the two light sources. The beat error signal can be fed, for example, into a feedback circuit that controls a phase modulator, which phase-locks the output of one light source to the phase of another. Phase-locking can be generalized as a master-slave scheme, where N ≥ 1 slave light sources are phase-locked to a single master light source. Therefore, a total of N+1 phase-locked light sources are available for the optical encoder system.

[0112] In other embodiments, each individual light source may be associated with a different wavelength of light. Using multiple light wavelengths allows some embodiments to be reused, enabling the simultaneous execution of multiple computations using the same optical hardware.

[0113] Power trees 1-203 are configured to split a single optical pulse from light source 1-201 into an array of spatially separated optical pulses. Therefore, power trees 1-203 have one optical input section and n optical output sections. In some embodiments, the optical power from light source 1-201 is uniformly divided across n optical modes associated with n waveguides. In some embodiments, power trees 1-203 are an array of 50:50 beamsplitters 1-801, such as... Figure 1-8 As shown. The number of "depths" of power tree 1-203 depends on the number of waveguides at the output. For a power tree with n output modes, the depth of power tree 1-203 is ceil(log2(n)). Figure 1-8 Power tree 1-203 only shows trees with a depth of 3 (each level of the tree is marked at the bottom of power tree 1-203). Each level includes 2 m-1 There are 801 beam splitters, where m is the number of layers. Therefore, the first layer has a single beam splitter 1-801a, the second layer has two beam splitters 1-801b and 1-801c, and the third layer has four beam splitters 1-801d to 1-801g.

[0114] Although power tree 1-203 is shown as an array of cascaded beam splitters, which can be implemented as an evanescent waveguide coupler, the embodiments are not limited to this, as any optical device that converts a single optical pulse into multiple spatially separated optical pulses can be used. For example, power tree 1-203 can be implemented using one or more multimode interferometers (MMIs), in which case the equations for the width and depth of the control layer will be appropriately modified.

[0115] Regardless of the type of power tree 1-203 used, it may be difficult, if not impossible, to manufacture it so that the splitting ratios among the n output modes are precisely equal. Therefore, the settings of the amplitude modulator can be adjusted to correct for unequal intensities in the n optical pulses output by the power tree. For example, the waveguide with the lowest optical power can be set to the maximum power for any given pulse transmitted to the photonic processor 1-103. Thus, any optical pulse with a power higher than the maximum power can be modulated to have a lower power by the amplitude modulator 1-205, except for amplitude modulation to encode information into the optical pulse. A phase modulator can also be placed at each of the n output modes, which can be used to adjust the phase of each output mode of the power tree 1-203 so that all output signals have the same phase.

[0116] Alternatively or additionally, the power tree 1-203 may be implemented using one or more Mach-Zehnder interferometers (MZIs) that can be tuned such that the splitting ratio of each beam splitter in the power tree produces substantially equal intensity pulses at the output of the power tree 1-203.

[0117] Amplitude modulator 1-205 is configured to modify the amplitude of each optical pulse received from power tree 1-203 based on a corresponding input bit string. Amplitude modulator 1-205 may be a variable attenuator controlled by DAC 1-209 or any other suitable amplitude modulator, which may be further controlled by controller 1-107. Some amplitude modulators are known for telecommunications applications and may be used in some embodiments. In some embodiments, a variable beam splitter may be used as amplitude modulator 1-205, wherein only one output of the variable beam splitter is retained, while the other output is discarded or ignored. Other examples of amplitude modulators that may be used in some embodiments include traveling wave modulators, cavity-based modulators, Franz-Keldysh modulators, plasmon-based modulators, 2-D material-based modulators, and nano-optomechanical switches (NOEMS).

[0118] Phase modulator 1-207 is configured to modify the phase of each optical pulse received from power tree 1-203 based on a corresponding input bit string. The phase modulator may be a thermo-optical phase shifter or any other suitable phase shifter that can be electrically controlled by 1-211, which may be further controlled by controller 1-107.

[0119] Although Figure 1-2Amplitude modulator 1-205 and phase modulator 1-207 are shown as two separate components, but they can be combined into a single element controlling both the amplitude and phase of the optical pulse. However, it is advantageous to control the amplitude and phase of the optical pulse separately. That is, since amplitude shift and phase shift are connected by the Kramers-Kronenig relationship, any amplitude shift has an associated phase shift. To accurately control the phase of the optical pulse, phase modulator 1-207 should be used to compensate for the phase shift generated by amplitude modulator 1-205. By way of example, the total amplitude of the optical pulse leaving optical encoder 1-101 is A = a0a1a2, and the total phase of the optical pulse leaving optical encoder is... Where a0 is the input intensity of the input optical pulse (assuming the input of the modulator is at zero phase), a1 is the amplitude attenuation of the amplitude modulator 1-205, and Δθ is the phase shift introduced by the amplitude modulator 1-205 when modulating the amplitude. The phase shift is introduced by phase modulator 1-207, and a2 is the attenuation associated with the optical pulse passing through phase modulator 1-209. The phase is introduced into the optical signal due to the propagation of the optical signal. Therefore, setting the amplitude and phase of the optical pulse are not two independent determinations. More precisely, in order to accurately encode a specific amplitude and phase into the optical pulse output from the optical encoder 1-101, the settings of both the amplitude modulator 1-205 and the phase modulator 1-207 should be taken into account for both settings.

[0120] In some embodiments, the amplitude of the optical pulse is directly related to the bit string value. For example, a high-amplitude pulse corresponds to a high bit string value, and a low-amplitude pulse corresponds to a low bit string value. The phase of the optical pulse encodes whether the bit string value is positive or negative. In some embodiments, the phase of the optical pulse output by the optical encoder 1-101 can be selected from two phases separated by 180 degrees (π radians). For example, a positive bit string value can be encoded with a zero-degree phase shift, while a negative bit string value can be encoded with a 180-degree (π radian) phase shift. In some embodiments, the vector is intended to be complex, so the phase of the optical pulse is selected from more than two values ​​between 0 and 2π.

[0121] In some embodiments, controller 1-107 determines the amplitude and phase to be applied by both amplitude modulator 1-205 and phase modulator 1-207 based on the input bit string and the aforementioned equation relating the output amplitude and output phase to the amplitude and phase imparted by amplitude modulator 1-204 and phase modulator 1-207. In some embodiments, controller 1-107 may store a table of digital values ​​for driving amplitude modulator 1-205 and phase modulator 1-207 in memory 1-109. In some embodiments, the memory may be placed near the modulators to reduce communication latency and power consumption.

[0122] A digital-to-analog converter (DAC) 1-209, associated with and communicatively connected to amplitude modulator 1-205, receives digital drive values ​​from controller 1-107 and converts them into analog voltages to drive amplitude modulator 1-205. Similarly, a DAC 1-211, associated with and communicatively connected to phase modulator 1-207, receives digital drive values ​​from controller 1-107 and converts them into analog voltages to drive phase modulator 1-207. In some embodiments, the DAC may include an amplifier that amplifies the analog voltage to a sufficiently high level to achieve a desired extinction ratio within the amplitude modulator (e.g., the highest extinction ratio physically achievable using a particular phase modulator) and a desired phase shift range within the phase modulator (e.g., a phase shift range covering the full range from 0 to 2π). While DAC 1-209 and DAC 1-211 are in Figure 1-2 The DACs are shown as being located in the chip of the optical encoder 1-101 and / or on the chip of the optical encoder 1-101, but in some embodiments, the DACs 1-209 and DAC 1-211 may be located outside the chip while still being communicatively connected to the amplitude modulator 1-205 and the phase modulator 1-207, respectively, using conductive traces and / or wires.

[0123] After being modulated by amplitude modulator 1-205 and phase modulator 1-207, n optical pulses are transmitted from optical encoder 1-101 to photonic processor 1-103.

[0124] D. Photonic processor

[0125] Reference Figure 1-3 The photonic processor 1-103 performs matrix multiplication on an input vector represented by n input optical pulses and includes three main components: a first matrix implementation 1-301, a second matrix implementation 1-303, and a third matrix implementation 1-305. In some embodiments, as discussed in more detail below, the first matrix implementation 1-301 and the third matrix implementation 1-305 include an interconnected array of programmable, reconfigurable variable beam splitters (VBSs) configured to transform the n input optical pulses from an input vector to an output vector, the components of which are represented by the amplitude and phase of each optical pulse. In some embodiments, the second matrix implementation 1-303 includes a set of electro-optic elements.

[0126] The matrix multiplied by the input vector through the photonic processor 1-103 is called M. Matrix M is a general m×n matrix known to the controller 1-107, which should be implemented by the photonic processor 1-103. Thus, the controller 1-107 uses singular value decomposition (SVD) to decompose matrix M, such that matrix M is represented as three component matrices: M = V T ∑U, where U and V are real orthogonal n×n and m×m matrices respectively (U T U=UU T =I, and V T V = VV T =I), and ∑ is an m×n diagonal matrix with real terms. The superscript "T" in all equations denotes the transpose of the relevant matrix. The SVD of the matrix is ​​known, and the controller 1-107 can use any suitable technique to determine the SVD of matrix M. In some embodiments, matrix M is a complex matrix, in which case matrix M can be decomposed into Where V and U are complex unitary n×n and m×m matrices, respectively. and ), and ∑ is an m×n diagonal matrix with real or complex terms. The values ​​of the diagonal singular values ​​can also be further normalized so that the maximum absolute value of the singular values ​​is 1.

[0127] Once controller 1-107 has determined the matrices U, ∑, and V of matrix M, and assuming matrices U and V are orthogonal real matrices, the controller can further decompose these two orthogonal matrices U and V into a series of real-valued Givens rotation matrices. The Givens rotation matrix G(i, j, θ) is defined in component form by the following equation:

[0128] For k≠i,j,g kk =1

[0129] For k = i, j, g kk =cos(θ)

[0130] g ij =-g ji = -sin(θ),

[0131] In other cases, g kl =0,

[0132] Where g ijLet represent the elements in the i-th row and j-th column of matrix G, and θ be the rotation angle associated with the matrix. Typically, matrix G is an arbitrary 2×2 unitary matrix (SU(2) group) with a determinant of 1, and it is parameterized by two parameters. In some embodiments, these two parameters are the rotation angle θ and another phase value φ. However, matrix G can be parameterized by values ​​other than angle or phase, such as by reflectivity / transmittance or by separation distance (in the case of NOEMS).

[0133] Algorithms for representing arbitrary real orthogonal matrices as a product of sets rotated according to Givens in complex space are provided in M. Reck et al., “Experimental realization of any discrete unitary operator,” Physical Review Letters, 73, 58 (1994) (“Reck”) and WRClements et al., “Optimal design for universal multiport interferometers,” Optica, 3, 12 (2016) (“Clements”). The entire contents of these two papers are incorporated herein by reference, at least for their discussion of decomposing real orthogonal matrices according to Givens rotations. (In the event of any conflict between the terminology used herein and its use in Reck and / or Clements, the term shall be deemed to have the meaning most consistent with how it is understood by one of ordinary skill in the art for its use herein.) The resulting decomposition is given by:

[0134]

[0135] Where U is an n×n orthogonal matrix, Sk is the index set associated with the k-th Givens rotation set applied (as defined by the decomposition algorithm), and θ ij (k) Let represent the angle of the Givens rotation applied between components i and j in the k-th Givens rotation set, and D be a diagonal matrix representing the +1 or -1 entries of the global symbol on each component. Index set S k It depends on whether n is even or odd. For example, when n is even:

[0136] For odd numbers k, S k ={(1,2),(3,4),...,(n-1,n)}

[0137] For even numbers k, S k ={(2,3),(4,5),...,(n-2,n-1)}

[0138] When n is odd:

[0139] For odd numbers k, S k ={(1,2),(3,4),...,(n-2,n-1)}

[0140] For even numbers k, S k ={(2,3),(4,5),...,(n-1,n)}

[0141] By way of example rather than restriction, the decomposition of a 4×4 orthogonal matrix can be expressed as:

[0142]

[0143] A brief overview of an embodiment of an algorithm for decomposing an n×n matrix U according to n real-valued Givens rotation sets is as follows, which can be implemented using controller 1-107:

[0144]

[0145] The resulting matrix U' from the above algorithm is a lower triangular matrix and is related to the original matrix U by the following formula:

[0146]

[0147] Where S is marked L The set of two modules connected by VBS is labeled to the left of U', and labeled S. R Label the set of the two modules joined by VBS to the right of U'. Since U is an orthogonal matrix, U' is a diagonal matrix with {-1, 1} terms along its diagonal. This matrix U′ = D U It is called a "phase screen".

[0148] The next step of the algorithm is to repeatedly find which G... T jk (θ1)D U =D U G jk (θ2) is accomplished using the following algorithm, which can be implemented using controller 1-107:

[0149]

[0150] The above algorithm can also be used to decompose V and / or V T To determine the m-layer VBS value and the associated phase screen.

[0151] The above concept of decomposing an orthogonal matrix into real-valued Givens rotation matrices can be extended to complex matrices, such as unitary matrices instead of orthogonal matrices. In some embodiments, this can be accomplished by including an additional phase in the parameterization of the Givens rotation matrix. Thus, the general form of the Givens matrix with the added phase term is T(i, j, θ, φ), where

[0152] For k≠i,j,t kk =1

[0153] t ii =e iφ cos(θ),

[0154] t jj =cos(θ),

[0155] t ij = -sin(θ),

[0156] t ji =e iφ sin(θ),

[0157] In other cases, t kl =0,

[0158] Where t ij Let represent the i-th row and j-th column of matrix T, θ be the rotation angle associated with the matrix, and φ be the additional phase. Any unitary matrix can be decomposed into a matrix of type T(i, j, θ, φ). By choosing to set the phase φ = 0, we obtain the above-described conventional real-valued Givens rotation matrix. Conversely, if the phase φ = π, we obtain a set of matrices called the Householder matrix. The Householder matrix H has the following form: Where I is an n×n identity matrix and v is a unit vector. It is an outer product. The Householder matrix represents the reflection about a hyperplane orthogonal to the unit vector v. In this parameterization, the hyperplane is a two-dimensional subspace, rather than the n-1-dimensional subspace commonly seen in the Householder matrix defined by the QR decomposition. Therefore, decomposing the matrix into Givens rotations is equivalent to decomposing the matrix into Householder matrices.

[0159] Based on the above decomposition of an arbitrary unitary matrix into a finite set of Givens rotations, an arbitrary unitary matrix can be realized through a specific sequence of rotations and phase shifts. Furthermore, in photonics, rotations can be represented by variable beam splitters (VBSs), and phase shifts can be easily implemented using phase modulators. Therefore, for the n optical inputs of the photonic processor 1-103, the first matrix implementation 1-301 and the third matrix implementation 1-305 representing the unitary matrix of the SVD of matrix M can be realized through an interconnected array of VBSs and phase shifters. Due to the parallel nature of passing optical pulses through the VBS array, matrix multiplication can be performed in O(1) time. The second matrix implementation 1-303 is a diagonal matrix combining the SVD of matrix M with the diagonal matrix D associated with each orthogonal matrix of the SVD. As mentioned above, each matrix D is called a "phase screen" and can be labeled with a subscript to indicate whether it is a phase screen associated with matrix U or matrix V. Therefore, the second matrix implementation 303 is matrix ∑′=D. V ∑D U .

[0160] In some embodiments, the VBS unit of the photonic processor 1-103 associated with the first matrix implementation 1-301 and the third matrix implementation 1-305 can be a Mach-Zehnder interferometer (MZI) with an internal phase shifter. In other embodiments, the VBS unit can be a microelectromechanical system (MEMS) actuator. In some embodiments, an external phase shifter can be used to implement the additional phase required for Givens rotation.

[0161] Represents a diagonal matrix D V ∑D U The second matrix implementation 1-303 can be implemented using amplitude modulators and phase shifters. In some embodiments, VBS can be used to segment out a portion of the light that can be discarded to variably attenuate the optical pulse. Additionally or alternatively, a controllable gain medium can be used to amplify the optical signal. For example, GaAs, InGaAs, GaN, or InP can be used as an active gain medium for amplifying the optical signal. Other active gain processes can also be used, such as second harmonic generation in crystal inversion-symmetric materials (e.g., KTP and lithium niobate) and four-wave mixing processes in materials lacking inversion symmetry (e.g., silicon). Depending on the phase screen implemented, a zero or π phase shift can be applied using a phase shifter in each optical mode. In some embodiments, only a single phase shifter is used for each optical mode, rather than one phase shifter for each phase screen. This is possible because matrix D V ,∑, and D U Each of these is a diagonal matrix and therefore a commutation matrix. Thus, the value of each phase shifter in the second matrix implementation of photonic processors 1-103 (1-303) is the result of the product of the two phase screens: D V DU .

[0162] Reference Figure 1-4 According to some embodiments, the first matrix implementation 1-301 and the third matrix implementation 1-305 are implemented as arrays of VBS 1-401. For simplicity, only n = 6 input optical pulses (rows) are shown, such that the "circuit depth" (e.g., column number) is equal to the number of input optical pulses (e.g., six). For clarity, only one VBS 1-401 is labeled with the figure reference numerals. However, the VBSs are labeled with subscripts identifying which optical modes are being mixed by a particular VBS and superscripts indicating the associated columns. As described above, each VBS 1-401 implements a complex Givens rotation T(i, j, θ, φ), where i and j are equivalent to the subscripts of the VBS 1-401, θ is the rotation angle of the Givens rotation, and φ is the additional phase associated with the generalized rotation.

[0163] Reference Figure 1-5 Each VBS 1-401 can be implemented using an MZI 1-510 and at least one external phase shifter 1-507. In some embodiments, a second external phase shifter 1-509 may also be included. The MZI 1-510 includes a first evanescent coupler 1-501 and a second evanescent coupler 1-503 for mixing the two input modes of the MZI 1-510. An internal phase shifter 1-505 modulates the phase θ in one arm of the MZI 1-510 to create a phase difference between the two arms. Adjusting the phase θ causes the intensity of the light output from the VBS 1-401 to vary from one output mode of the MZI 1-510 to the other, thereby forming a controllable and variable beam splitter. In some embodiments, a second internal phase shifter may be applied to the second arm. In this case, it is the difference between the two internal phase shifters that causes the variation in output light intensity. The average value between these two internal phases introduces a global phase for the light entering modes i and j. Therefore, these two parameters θ and φ can be controlled by the phase shifters individually. In some embodiments, the second external phase shifter 1-509 can be used to correct unwanted differential phase on the output mode of the VBS due to static phase disorder.

[0164] In some embodiments, phase shifters 1-505, 1-507, and 1-509 may include thermo-optical, electro-optical, or optomechanical phase modulators. In other embodiments, an NOEMS modulator may be used instead of including an internal phase modulator 505 within the MZI 510.

[0165] In some embodiments, the number of VBSs increases with the size of the matrix. The inventors have recognized and appreciated that controlling a large number of VBSs can be challenging, and that sharing a single control circuit across multiple VBSs is advantageous. An example of a parallel control circuit that can be used to control multiple VBSs is a digital-to-analog converter (DAC) that receives as input a digital string encoding an analog signal to be introduced onto a particular VBS. In some embodiments, the circuit also receives a second input: the address of the VBS to be controlled. The circuit can then introduce the analog signal onto the addressed VBS. In other embodiments, the control circuit can automatically scan through multiple VBSs and introduce analog signals onto multiple VBSs without being actively assigned addresses. In this case, an addressing sequence is predefined such that the addressing sequence traverses the VBS array in a known order.

[0166] Reference Figure 1-6 The second matrix implements 1-303 and the diagonal matrix ∑′=D V ∑D U The multiplication of . can be accomplished by using two phase shifters 1-601 and 1-605 to implement two phase screens, and using an amplitude modulator 1-603 to adjust the intensity of the relevant optical pulses to a certain amount η. As mentioned above, in some embodiments, since the two phase screens can be combined, only a single phase modulator 1-601 can be used, because the three component matrices forming ∑′ are diagonal matrices and therefore commutative matrices.

[0167] In some embodiments, the amplitude modulator 1-603 can be implemented using attenuators and / or amplifiers. If the amplitude modulation value η is greater than 1, the optical pulse is amplified. If the amplitude modulation value η is less than 1, the optical pulse is attenuated. In some embodiments, only attenuation is used. In some embodiments, attenuation can be implemented using a series of integrated attenuators. In other embodiments, such as Figure 1-7 As shown, attenuation 1-603 can be achieved using an MZI. The MZI includes two evanescent couplers 1-701 and 1-703 and a controllable internal phase shifter 1-705 to adjust how much input light is transmitted from the input section of the MZI to the first output port 1-709 of the MZI. The second output port 1-707 of the MZI can be ignored, blocked, or discarded.

[0168] In some embodiments, controller 1-107 controls the value of each phase shifter in photonic processor 1-103. Each phase shifter discussed above may include a DAC similar to the DAC discussed in conjunction with phase modulator 1-207 of optical encoder 1-101.

[0169] Photonic processor 1-103 can include any number of input nodes, but the size and complexity of the interconnect VBS arrays 1-301 and 1-305 will increase with the number of input modes. For example, if there are n input optical modes, photonic processor 1-103 will have a circuit depth of 2n+1, where the first matrix implementation 1-301 and the second matrix implementation 1-305 each have a circuit depth of n, and the second matrix implementation 1-303 has a circuit depth of 1. Importantly, the time complexity of performing a single matrix multiplication is not linear with the number of input optical pulses; it is always O(1). In some embodiments, this low-order complexity obtained by parallelization can bring energy and time efficiencies that are unattainable using conventional electrical processors.

[0170] It should be noted that although the embodiments described herein show photonic processors 1-103 as having n inputs and n outputs, in some embodiments, the matrix M implemented by photonic processors 1-103 may not be a square matrix. In such embodiments, photonic processors 1-103 may have different numbers of outputs and inputs.

[0171] It should also be noted that, due to the interconnected topology of the VBS within 1-301 implemented by the first matrix and 1-305 implemented by the second matrix, photonic processors 1-103 can be subdivided into non-interactive row subsets, enabling the simultaneous execution of more than one matrix multiplication. For example, in Figure 1-4 In the VBS array shown, if each VBS 1-401 of the coupled optical mode 3 and optical mode 4 is configured such that optical mode 3 and optical mode 4 are not coupled at all (e.g., as if VBS 1-401 with subscript "34" is in...), Figure 1-4 If the input modes do not exist (i.e., the top three modes operate completely independently of the bottom three modes), then such subdivision can be accomplished on a larger scale using photonic processors with a large number of input modes. For example, a photonic processor with n=64 can simultaneously multiply eight octet input vectors by their corresponding 8×8 matrices (each 8×8 matrix can be programmed and controlled individually). Furthermore, photonic processors 1-103 do not need to be uniformly subdivided. For example, a photonic processor with n=64 can be subdivided into seven different input vectors with 20, 13, 11, 8, 6, 4, and 2 components respectively, each multiplied by its corresponding matrix simultaneously. It should be understood that the numerical examples above are for illustrative purposes only, and any number of subdivisions is possible.

[0172] Furthermore, although the photonic processor 1-103 performs vector-matrix multiplication, where vectors are multiplied by matrices by passing optical signals through the VBS array, it can also be used to perform matrix-matrix multiplication. For example, multiple input vectors can be passed through the photonic processor 1-103 one after another, one input vector at a time, where each input vector represents a column of the input matrix. After each of the individual vector-matrix multiplications is optically computed (each multiplication produces an output vector corresponding to a column of the output matrix), the results can be digitally combined to form the output matrix produced by the matrix-matrix multiplication.

[0173] E. Optical receiver

[0174] Photonic processor 1-103 outputs n light pulses, which are transmitted to optical receiver 1-105. Optical receiver 1-105 receives the light pulses and generates an electrical signal based on the received light signal. In some embodiments, the amplitude and phase of each light pulse are determined. In some embodiments, this is achieved using a homodyne or heterodyne detection scheme. In other embodiments, a conventional photodiode can be used to perform simple phase-insensitive light detection.

[0175] Reference Figure 1-9 According to some embodiments, the optical receiver 1-105 includes a zero-difference detector 1-901, a transimpedance amplifier 1-903, and an analog-to-digital converter (ADC) 1-905. Although in Figure 1-9 The component is shown as a single element for all optical modes, but this is for simplicity. Each optical mode may have a dedicated zero-difference detector 1-901, a dedicated transimpedance amplifier 1-903, and a dedicated ADC 1-905. In some embodiments, the transimpedance amplifier 1-903 may be omitted. Instead, any other suitable electronic circuitry that converts current to voltage may be used.

[0176] Reference Figure 1-10 According to some embodiments, the zero-difference detector 1-903 includes a local oscillator (LO) 1-1001, an orthogonal controller 1-1003, a beam splitter 1-1005, and two detectors 1-1007 and 1-1009. The zero-difference detector 1-903 outputs a current based on the difference between the currents output by the first detector 1-1007 and the second detector 1-1009.

[0177] Local oscillator 1-1001 is combined with the input optical pulse at beamsplitter 1-1005. In some embodiments, a portion of light source 1-201 is transmitted to homodyne detector 1-901 via optical waveguide and / or optical fiber. Light from light source 1-201 can itself serve as local oscillator 1-1001, or in other embodiments, local oscillator 1-1001 can be a separate light source that generates phase-matched optical pulses using light from light source 1-201. In some embodiments, MZI can replace beamsplitter 1-1005, allowing adjustment between the signal and the local oscillator.

[0178] The quadrature controller 1-1003 controls the cross-sectional angle in the phase space in which the measurement is performed. In some embodiments, the quadrature controller 1-1003 may be a phase shifter that controls the relative phase between the input optical pulse and the local oscillator. The quadrature controller 1-1003 is shown as a phase shifter in the input optical mode. However, in some embodiments, the quadrature controller 1-1003 may be in local oscillator mode.

[0179] The first detector 1-1007 detects the light output from the first output section of the beam splitter 1-1005, and the second detector 1-1009 detects the light output from the second output section of the beam splitter 1-1005. Detectors 1-1007 and 1-1009 can be photodiodes operating with zero bias.

[0180] Subtraction circuit 1-1011 subtracts the current from first detector 1-1007 from the current from second detector 1-1009. The resulting current has both amplitude and sign (positive or negative). Transimpedance amplifier 1-903 converts this current difference into a voltage, which can be positive or negative. Finally, ADC 1-905 converts the analog signal into a digital bit string. This output bit string represents the output vector result of the matrix multiplication and is an electrical-digital version of the optical output representation of the output vector output by photonic processor 1-103. In some embodiments, as described above, the output bit string can be sent to controller 1-107 for additional processing, which may include determining the next input bit string based on one or more output bit strings and / or transmitting the output bit string to an external processor.

[0181] The inventors further recognize and understand that the components of the aforementioned photonic processing system 1-100 do not necessarily need to be linked back-to-back such that the first matrix implementation 1-301 is connected to the second matrix implementation 1-303 connected to the third matrix implementation 1-305. In some embodiments, the photonic processing system 1-103 may include only a single unitary circuit for performing one or more multiplications. The output of the single unitary circuit may be directly connected to the optical receiver 1-105, wherein the result of the multiplication is determined by detecting the output optical signal. In such an embodiment, the single unitary circuit may, for example, implement the first matrix implementation 1-301. The result detected by the optical receiver 1-105 can then be digitally transmitted to a conventional processor (e.g., processor 1-111), wherein the diagonally opposite second matrix implementation 1-303 is executed in the digital domain using the conventional processor (e.g., processor 1-111). Then, controller 1-107 can reprogram individual unit circuits to execute third matrix implementation 1-305, determine the input bit string based on the result of the digital implementation of the second matrix implementation, and control the optical encoder to transmit the optical signal encoded based on the new input bit string through the individual unit circuits with the reprogrammed settings. The result of the matrix multiplication is then determined using the output optical signal detected by optical receiver 105.

[0182] The inventors also recognize and appreciate that it may be advantageous to cascade multiple photonic processors 1-103 back-to-back. For example, to implement matrix multiplication M = M1M2, where M1 and M2 are arbitrary matrices, but M2 changes more frequently based on varying input loads than M1, the first photonic processor can be controlled to implement M2, and a second photonic processor optically coupled to the first photonic processor can implement the statically held M1. In this way, only the first photonic processing system needs to be frequently updated based on varying input loads. This arrangement not only accelerates computation but also reduces the number of data bits transmitted between the controller 1-107 and the photonic processors.

[0183] F. Folded photon processing system

[0184] exist Figure 1-1In this arrangement, the optical encoder 1-101 and the optical receiver 1-105 are located on opposite sides of the photonics processing system 1-100. In applications where feedback from the optical receiver 1-105 is used to determine the input of the optical encoder 1-101 for future iterations of the process, data is electronically transmitted from the optical receiver 1-105 to the controller 1-107, and then to the optical encoder 1-101. The inventors have recognized and appreciated that reducing the distance these electrical signals need to travel (e.g., by reducing the length of electrical traces and / or wires) can save energy and reduce latency. Furthermore, it is not necessary to place the optical encoder 1-101 and the optical receiver 1-105 at opposite ends of the photonics processing system.

[0185] Therefore, in some embodiments, the optical encoder 1-101 and the optical receiver 1-105 are positioned close to each other (e.g., on the same side of the photonic processor 1-103), such that the distance the electrical signal must travel between the optical encoder 1-101 and the optical receiver 1-105 is less than the width of the photonic processor 1-103. This can be achieved by physically interleaving the components of the first matrix implementation 1-301 and the third matrix implementation 1-305 so that they are physically located in the same part of the chip. This arrangement is called a “folded” photonic processing system because light first propagates through the first matrix implementation 1-301 in a first direction until it reaches a physical portion of the chip away from the optical encoder 1-101 and the optical receiver 1-105, and then folds back, such that when the third matrix implementation 1-305 is implemented, the waveguide redirects the light to propagate in the opposite direction to the first direction. In some embodiments, the second matrix implementation 1-303 is physically located near the fold in the waveguide. This arrangement reduces the complexity of the electrical traces connecting the optical encoder 1-101, optical receiver 1-105, and controller 1-107, and also reduces the total chip area used to implement the photonics processing system 1-100. For example, the total chip area of ​​some embodiments using the folded arrangement is only a fraction of that using... Figure 1-1 The back-to-back photonic arrangement reduces the total chip area required by 65%. This can reduce the cost and complexity of photonic processing systems.

[0186] The inventors have recognized and understood that folded arrangements offer not only electrical advantages but also optical advantages. For example, by reducing the distance the optical signal must travel from the light source used as a local oscillator for zero-difference detection, time-dependent phase fluctuations in the optical signal can be reduced, resulting in higher-quality detection results. In particular, by positioning the light source and zero-difference on the same side of the photonic processor, the distance the optical signal travels for the local oscillator no longer depends on the size of the matrix. For example, in Figure 1-1In a back-to-back arrangement, the distance for optical signal transmission to the local oscillator is linearly proportional to the size of the matrix, while in a folded arrangement, the transmission distance is constant and independent of the size of the matrix.

[0187] Figure 1-11 This is a schematic diagram of a folded photonics processing system 1-1100 according to some embodiments. The folded photonics processing system 1-1100 includes a power tree 1-1101, multiple optical encoders 1-1103a to 1-1103d, multiple homodyne detectors 1-1105a to 1-1105d, multiple selector switches 1-1107a to 1-1107d, multiple U-matrix components 1-1109a to 1-1109j, multiple diagonal matrix components 1-1111a to 1-1111d, and multiple V-matrix components 1-1113a to 1-1113j. For clarity, not all components of the folded photonics processing system are shown in the figure. It should be understood that the folded photonics processing system 1-1100 may include components similar to those in the back-to-back photonics processing system 1-100.

[0188] Power tree 1-1101 is similar to power tree 1-203 in Figure 2 and is configured to transmit light from a light source (not shown) to optical encoder 1-1103. However, the difference between power tree 1-1101 and power tree 1-203 is that power tree 1-1101 directly transmits the optical signal to homodyne detector 1-1105a. In Figure 2, light source 201 transmits the local oscillator signal to the homodyne detector on the other side of the photonic processor by branching off a portion of the optical signal from the light source and guiding the optical signal using a waveguide. Figure 1-11 In the power tree 1-1101, the number of output sections is equal to twice the number of spatial modules. For example, Figure 1-11 Only four spatial modes of the photonic processor are shown, so the power tree 1-1101 has eight output modes—one output directs light to each optical encoder 1-1103, and one output directs light to each homodyne detector 1-1105. The power tree can be implemented, for example, using a cascaded beam splitter or a multimode interferometer (MMI).

[0189] The optical encoder 1-1103 is similar to the power tree optical encoder 1-101 of Figure 1 and is configured to encode information as the amplitude and / or phase of the optical signal received from the power tree 1-1101. This can be implemented, for example, as described in conjunction with the optical encoder 1-101 of Figure 2.

[0190] The zero-difference detector 1-1105 is located between the power tree 1-1101 and the U-matrix component 1-1109. In some embodiments, the zero-difference detector 1-1105 and the optical encoder 1-1103 are physically positioned in a column. In some embodiments, the optical encoder 1-1103 and the zero-difference detector 1-1105 may be interleaved in a single column. In this way, the optical encoder 1-1103 and the zero-difference detector 1-1105 are close to each other, reducing the distance of the electrical traces (not shown) used to connect the optical encoder 1-1103 and the zero-difference detector 1-1105 to the controller (not shown), which may be physically positioned adjacent to the column of the optical encoder 1-1103 and the zero-difference detector 1-1105.

[0191] Each optical encoder 1-1103 is associated with a corresponding homodyne detector 1-1105. Both the optical encoder 1-1103 and the homodyne detector 1-1105 receive optical signals from the power tree 1-1101. As described above, the optical encoder 1-1103 uses the optical signals to encode the input vector. The homodyne detector 1-1105 uses the received optical signals from the power tree as a local oscillator, as described above.

[0192] Each pair of optical encoders 1-1103 and zero-difference detectors 1-1105 is associated with and connected to selector switches 1-1107 via waveguides. Selector switches 1-1107a to 1-1107d can be implemented using, for example, conventional 2×2 optical switches. In some embodiments, the 2×2 optical switch is an MZI with an internal phase shifter to control the behavior of the MZI from crossing to bar. Switches 1-1107 are connected to a controller (not shown) to control whether the optical signal received from the optical encoders 1-1103 is directed to either the U matrix component 1-1109 or the V matrix component 1-1113. The optical switches are also controlled to direct the light received from the U matrix component 1-1109 and / or the V matrix component 1-1113 to the zero-difference detector 1-1105 for detection.

[0193] In the photonic folding photonic processing system 1-1100, the technique used to perform matrix multiplication is similar to the combination of the above. Figure 1-3 The back-to-back system described herein is a technique. The difference between the two systems lies in the physical arrangement of the matrix components and the implementation of the folds 1-1120, in which the optical signal originates from... Figure 1-11 The approximate propagation from left to right changes to an approximate propagation from right to left. Figure 1-11 In this context, the connections between components can represent waveguides. In some embodiments, solid lines represent waveguide sections where the optical signal propagates from left to right, while dashed lines represent waveguide sections where the optical signal propagates from right to left. Specifically, given this naming convention, Figure 1-11The illustrated embodiment is one in which selector switch 1-1107 first directs the optical signal to U matrix component 1-1109. In other embodiments, selector switch 1-1107 may first direct the optical signal to V matrix component 1-1113, in which case the dashed lines would represent the waveguide portion of the optical signal propagating from left to right, while the solid lines would represent the waveguide portion of the optical signal propagating from right to left.

[0194] The U matrix of the SVD of matrix M is implemented in the photonic processing system 1-1100 using U matrix component 1-1109, which is interleaved with V matrix component 1-1113. Therefore, with Figure 1-3 Unlike the back-to-back arrangement shown in the embodiment, all U matrix components 1-1109 and V matrix components 1-1113 are not physically located in corresponding self-contained arrays within a single physical region. Therefore, in some embodiments, the photonic processing system 1-1100 includes multiple columns of matrix components, and at least one column contains both U matrix components 1-1109 and V matrix components 1-1113. In some embodiments, the first column may contain only U matrix components 1-1109, such as... Figure 1-11 As shown, the U matrix component 1-1109 is implemented similarly to the first matrix implementation 1-301 in Figure 3.

[0195] Due to the interwoven structure of the U-matrix components 1-1109 and the V-matrix components 1-1113, the folded photonic processing system 1-1100 includes waveguide crosses 1-1110 at different locations between the columns of the matrix elements. In some embodiments, waveguide crosses can be constructed using adiabatic evanescent boosters between two or more layers in an integrated photonic chip. In other embodiments, the U-matrix and V-matrix may reside on different layers of the same chip, and waveguide crosses are not used.

[0196] After the optical signal propagates through all U-matrix components 1-1109, it propagates to the diagonal matrix component 1-1111, and the diagonal matrix component 1-1111... Figure 1-3 The second matrix implementation 1-303 is similarly implemented.

[0197] After the optical signal propagates through all the diagonal matrix components 1-1111, it propagates to the V matrix component 1-1113. The V matrix component 1-1113 and... Figure 1-3 The third matrix implementation 1-305 is similarly implemented. The V matrix of the SVD of matrix M is implemented in the photonic processing system 1-1100 using V matrix components 1-1113 interleaved with U matrix components 1-1109. Therefore, all V matrix components 1-1113 are not physically located in a single self-contained array.

[0198] After the optical signal propagates through all V matrix components 1-1113, the optical signal returns to the selector switch 1-1107, which guides the optical signal to the zero-difference detector 1-1105 for detection.

[0199] The inventors also recognize and appreciate that by including a selector switch after the optical encoder and before the matrix components, the folded photonics processing system 1-1100 achieves efficient bidirectionality in the circuitry. Therefore, in some embodiments, the controller, for example, combined with... Figure 1-1 The controller 1-107 described can control whether the optical signal is first multiplied by a U matrix or a V matrix. T Matrix. For a unitary matrix V configured to propagate an optical signal from left to right. T The VBS array propagates optical signals from right to left to realize a unitary matrix U. T Multiplication. Therefore, the same setup for a VBS array can achieve U... and U T The direction of propagation depends on how the optical signal travels through the array, which can be controlled using the selector switch in 1-1107. In some applications, such as backpropagation for training machine learning algorithms, it may be necessary to reverse the optical signal through one or more matrices. In other applications, bidirectionality can be used to compute the inverse matrix operation on the input vector. For example, for an invertible n×n matrix M, SVD produces M = Vinverter. T ∑U. The inverse of this matrix is ​​M. -1 =U T ∑ -1 V,,∑ -1 It is the inverse of a diagonal matrix, which can be efficiently computed by inverting each diagonal element. To multiply a vector by matrix M, the switch is configured to guide the optical signal through matrix U in the first direction, then through ∑, and then through V. T To multiply a vector by its inverse M -1 First, set singular values ​​to apply to ∑ -1 The matrix implementation is programmed. This constitutes changing only one column of the VBS instead of all 2n+1 columns of the photon processor, which is similar to... Figure 1-3 This illustrates a unidirectional photon processing system. Then, the optical signal representing the input vector propagates through matrix V in a second direction opposite to the first direction. T Then through ∑ -1 Then, via U. Using selector switch 1-1107, the folded photon processing system 1-1100 can easily start from first realizing the U matrix (or its transpose) and first realizing V. T The matrix (or its transpose) is changed.

[0200] G. Wavelength division multiplexing

[0201] The inventors further recognize and understand that there are applications where different vectors can be multiplied by the same matrix. For example, when training or using machine learning algorithms, multiple datasets can be processed using the same matrix multiplication. The inventors have recognized and understand that if the components before and after the photonic processor are wavelength division multiplexing (WDM), this can be achieved with a single photonic processor. Therefore, some embodiments include multiple front-ends and multiple back-ends, each associated with a different wavelength, while using only a single photonic processor to perform matrix multiplication.

[0202] Figure 1-12A A WDM photonics processing system 1-1200 according to some embodiments is shown. The WDM photonics processing system 1-1200 includes N front-ends 1-1203, a single photonics processor 1-1201 having N spatial modes, and N back-ends 1-1205.

[0203] Photonic processor 1-1201 can be similar to photonic processor 1-103, having N input modules and N output modules. Each of the N front-ends 1203 is connected to the corresponding input module of photonic processor 1201. Similarly, each of the N back-ends 1-1205 is connected to the corresponding output module of photonic processor 1-1201.

[0204] Figure 1-12B Details of at least one of the front-ends 1-1203 are shown. Like the photonics processing system of other embodiments, the photonics processing system 1-1200 includes optical encoders 1-1211. However, in this embodiment, there are M different optical encoders, where M is the number of wavelengths multiplexed by the WDM photonics processing system 1-1200. Each of the M optical encoders 1-1211 receives light from a light source (not shown) that generates M optical signals, each with a different wavelength. This light source can be, for example, a laser array, a frequency comb generator, or any other light source that generates coherent light of different wavelengths. Each of the M optical encoders 1-1211 is controlled by a controller (not shown) to achieve appropriate amplitude and phase modulation, thereby encoding data into optical signals. The M encoded optical signals are then combined into a single waveguide using M:1WDM 1-1213. The single waveguide is then connected to one of the N waveguides of the photonics processor 1-1201.

[0205] Figure 1-12CDetails of at least one of the back-ends 1-1205 are shown. Like the photonics processing system of other embodiments, the photonics processing system 1-1200 includes detectors 1-1223, which may be phase-sensitive or phase-insensitive detectors. However, in this embodiment, there are M different detectors 1-1223, where M is the number of wavelengths multiplexed by the WDM photonics processing system 1-1200. Each of the M detectors 1-1223 receives light from a 1:M WDM 1-1221, which divides a single output waveguide from the photonics processor 1201 into M different waveguides, each carrying an optical signal of a corresponding wavelength. Each of the M detectors 1-1223 can be controlled by a controller (not shown) to record measurement results. For example, each of the M detectors 1223 may be a homodyne detector or a phase-insensitive photodetector.

[0206] In some embodiments, the VBS in the photonic processor 1-1201 can be selected to be non-dispersive over the M wavelengths of interest. This allows all input vectors to be multiplied by the same matrix. For example, an MMI can be used instead of a directional coupler. In other embodiments, the VBS can be selected to be dispersive over the M wavelengths of interest. In some applications related to the stochastic optimization of the parameters of a neural network model, this is equivalent to adding noise when calculating the gradient of the parameters; the added gradient noise can be beneficial in accelerating optimization convergence and can improve the robustness of the neural network.

[0207] Although Figure 1-12A A back-to-back photonics processing system is shown, but a similar WDM technique can be used to form a WDM folded photonics processor, which uses the technique described in relation to folded photonics processor 1-1100.

[0208] H. Simulated summation of output

[0209] The inventors have recognized and appreciated that there are applications in which it is useful to calculate the sum or average over time of the outputs of photonic processors 1-103. For example, when photonic processing system 1-100 is used to compute a more accurate matrix-vector multiplication of a single data point, it may be desirable to run the single data point multiple times through the photonic processor to improve the statistical results of the computation. Alternatively, when computing gradients in a backpropagation machine learning algorithm, it may not be desirable for a single data point to determine the gradient; therefore, multiple training data points can be run through photonic processing system 1-100, and the average result can be used to compute the gradient. When using the photonic processing system to perform batch gradient-based optimization algorithms, this averaging can improve the quality of gradient estimation, thereby reducing the number of optimization steps required to achieve a high-quality solution.

[0210] The inventors have further recognized and understood that the output signal can be summed in the analog domain before being converted into a digital electrical signal. Therefore, in some embodiments, a low-pass filter is used to sum the output from the zero-difference detector. By performing the summation in the analog domain, the zero-difference electronics can use a low-speed ADC instead of the more expensive high-speed ADC (e.g., an ADC with high power consumption requirements) required to perform the summation in the digital domain.

[0211] Figure 1-13 A portion of an optical receiver 1-1300 according to some embodiments is shown, and how a low-pass filter 1-1305 can be used with a homodyne detector 1-1301. The homodyne detector 1-1301 performs field and phase measurements of the input optical pulses. If k is a marker of different input pulses varying over time and there are a total of K inputs, the low-pass filter 1-1305 can be used to automatically perform a summation on k in the analog domain. The optical receiver 1-1300 and... Figure 1-9 The main difference between the optical receivers 1-105 shown is that the low-pass filter is located after the transimpedance amplifier 1-1303, following the output of the zero-difference detector. If in a single slow sampling period T... s (slow) There are a total of K signals (with components y) i (k) When the zero-difference detector is reached, the low-pass filter will adjust according to y. i (k) The sign and value of the charge accumulated / removed from capacitor C. The final output of the low-pass filter is related to... Proportional, it can be represented by a sampling frequency of f s (slow) =1 / T s (slow) =f s / K's low-speed ADC (not shown) reads once, where f s This is the initial required sampling frequency. For an ideal system, the low-pass filter should have a 3dB bandwidth: f 3dB =f s (slow) / 2. For use such as Figure 1-13 The low-pass filter of the RC circuit shown in the embodiment, f 3dB = 1 / (2πRC), and the values ​​of R and C can be chosen to obtain the desired sampling frequency: f s (slow) .

[0212] In some embodiments, both high-speed and low-speed ADCs may be present. In this context, a high-speed ADC is an ADC configured to receive each individual analog signal and convert it into a digital signal (e.g., an ADC with a sampling frequency equal to or greater than the frequency at which the analog signal arrives at the ADC), while a low-speed ADC is an ADC configured to receive multiple analog signals and convert the sum or average of the multiple received analog signals into a single digital signal (e.g., an ADC with a sampling frequency less than the frequency at which the analog signal arrives at the ADC). An electrical switch can be used to switch the electrical signal from a zero-difference detector and possibly a transimpedance amplifier to a low-pass filter with a low-speed ADC or to a high-speed ADC. In this way, some embodiments of the photonics processing system can switch between performing analog summation using a low-speed ADC and measuring each optical signal using a high-speed ADC.

[0213] I. Stable phase

[0214] The inventors have recognized and understood that phase stability of the local oscillator used to perform phase-sensitive measurements (e.g., homodyne detection) is desirable to ensure accurate results. The photonic processor of the embodiments described herein performs matrix operations via interference light between N different spatial modes. In some embodiments, a phase-sensitive detector (e.g., a homodyne detector or a heterodyne detector) is used to measure the results. Therefore, to ensure accurate execution of matrix operations, the phase introduced at various parts of the photonic processor should be as accurate as possible, and the phase of the local oscillator used to perform phase-sensitive detection should be accurately known.

[0215] The inventors have recognized and understood that parallel interferometry operations (e.g., parallel interferometry operations performed within a single column of VBSs in a photonic processor) must not only employ a phase modulator that controls the relative phase within the MZI of the VBS and the phase and relative phase of the output of the MZI to introduce the correct phase, but also that each VBS in the column should introduce the same global phase shift across all spatial modes of the photonic processor. In this application, the global phase shift of a column of VBSs in a photonic processor is referred to as the “column-global phase.” The column-global phase is the phase introduced due to effects unrelated to the programming phase associated with the VBS, such as phase introduced due to waveguide propagation or phase caused by temperature variations. It is not necessary to simultaneously and precisely introduce these phases onto the VBSs in a column, but only as a result of traversing the column in question. Ensuring column-global phase consistency across different spatial modes of the column is important because the output optical signal from one column may be interfered with at one or more VBSs in subsequent columns. If the column-global phase is inconsistent at previous columns, subsequent interference will be incorrect, and therefore the accuracy of its own calculations will also be incorrect.

[0216] Figure 1-14The column-global phase and total global phase of the photonics processing system 1-1400 are shown. Similar to the above-described embodiment of the photonics processing system, the photonics processing system 1-1400 includes a U-matrix implementation 1-1401, a diagonal matrix implementation 1-1403, a V-matrix implementation 1-1405, and multiple detectors 1-1407a to 1-1407d. These implementations are similar to the first, second, and third matrix implementations described above. For simplicity, only four modes of the photonics processing system 1-1400 are shown, but it should be understood that any larger number of modes can be used. Again, for simplicity, only the VBS associated with the U-matrix implementation 1-1401 is shown. The arrangement of components in the diagonal matrix implementations 1-1403 and V-matrix implementations 1-1405 is similar to the third and fourth matrix implementations described above.

[0217] The U-matrix implementation 1-1401 comprises multiple VBSs 1-1402, although only a single VBS 1-1402 is labeled for clarity. However, the VBS labels have subscripts identifying which optical modes are being mixed by a particular VBS and superscripts indicating the associated column.

[0218] like Figure 1-14 As shown, each column is associated with a column-global phase, which is ideally consistent for each element of that column. For example, the U matrix implements column 1 of 1-1401 with a column-global phase. Relatedly, the U matrix realizes column 2 and column-global phase of 1-1401. Relatedly, the U matrix implements column 3 of 1-1401 and column-global phase. The U matrix is ​​associated with column 4 of 1-1401 and with column-global phase φ4.

[0219] In some embodiments, by implementing each VBS1-1402 as a push-pull MZI, column global phase coherence can be at least partially achieved. Alternatively or additionally, an external phase shifter can be added to the output of each MZI to correct any phase errors introduced from the internal phase elements of the MZI (e.g., phase shifters).

[0220] The inventors have further recognized and understood that even under such conditions, when each column of the photonic processing system 1-1400 provides a uniform column-global phase, phase accumulation is still possible as the signal propagates from the first column to the last. A global U-matrix phase Φ exists. U This is associated with the entire U matrix realization 1-1401 and equal to the sum of the column-global phases. Similarly, the diagonal matrix realization 1-1403 is phase Φ with the global diagonal matrix. ∑ Correlated, the V matrix realizes phase between 1-1405 and the global diagonal matrix. Correlated. Then, the total global phase Φ of the entire photonic processing system 1-1400 is given by the sum of the phases of three separate global matrices. G This total global phase can be set to be consistent across all output modes, but the local oscillator used for phase-sensitive detection does not propagate through the photon processor and is unaffected by this total global phase. If the total global phase Φ is not considered... G This could lead to errors in the values ​​read by the zero-difference detectors 1-1407a to 1-1407d.

[0221] The inventors have further realized that errors in multiplication operations may be caused by temperature variations, which alter the effective refractive index n of the waveguide. eff Therefore, in some embodiments, the temperature of each column is set to be uniform, or stabilization circuitry can be placed at each column to actively synchronize the phase of all modes introduced into a single column. Additionally, temperature differences between different parts of the system can cause errors in phase-sensitive measurements as the optical signal used for the local oscillator propagates through them. The phase difference between the signal and the local oscillator is... Where T s and T LO These are the temperatures of the signal waveguides in the photonic processor and the local oscillator waveguides, respectively, n. eff (T) is the effective refractive index as a function of temperature, λ is the average wavelength of light, and L s and L LO These are the propagation lengths of the signal waveguides in the photonic processor and local oscillator waveguides, respectively. Assume a temperature difference ΔT = T. LO -T S If the effective refractive index is small, then the effective refractive index can be rewritten as: Therefore, the phase difference between the signal and the LO can be well approximated as It increases linearly with the propagation length L. Therefore, for a sufficiently long propagation distance, a small temperature change can lead to a large phase shift (on the order of 1 radian). Importantly, L... S The value does not need to be related to L LO The values ​​are the same, and the maximum difference between the two is determined by the coherence length L of the light source. coh The decision is made. For a light source with bandwidth Δv, the coherence length can be well approximated as L. coh ≈c eff Δv, where c eff It is the speed of light in the transmission medium. As long as L... S With L LO The length difference between them is L coh Much shorter, the interference between the signal and the local oscillator could potentially be used for correction operations in photonic processing systems.

[0222] Based on the above, the inventors have discovered at least two sources of possible phase error between the photonic processor and the output signal of the local oscillator used for homodyne detection in some embodiments. Therefore, an ideal homodyne detector measures the amplitude and phase of the signal output by subtracting the outputs of both photodetectors, thereby generating a phase-sensitive intensity output measurement value I. out ∝|E s ||E LO |cos(θ s -θ LO +Φ G +Φ T ), where E s E is the electric field amplitude of the optical signal from the output of the photonic processor. LO It is the electric field amplitude of the local oscillator, θ s It is the phase shift, Φ, that is expected to be measured and introduced by the photonic processor. G It is the global phase, Φ T The phase shift is caused by the temperature difference between the local oscillator and the optical signal. Therefore, if the total global phase and the phase shift caused by the temperature difference are not taken into account, the result of zero-difference detection may be erroneous. Therefore, in some embodiments, the total system phase error ΔΦ = Φ is measured. G +Φ T The system is then calibrated based on this measurement. In some embodiments, the total system phase error includes contributions from other error sources that are not necessarily known or identified.

[0223] According to some embodiments, a zero-difference detector can be calibrated by sending a pre-calculated test signal to the detector and using the difference between the pre-calculated test signal and the measured test signal to correct the total system phase error in the system.

[0224] In some embodiments, the total global phase Φ is not used. G and the phase shift Φ caused by temperature difference T Considering the optical signals propagating through the photonic processor, they can be described as signals that do not accumulate any phase shift at all, but rather as LOs with a total system phase error of -ΔΦ. Figure 1-15 The effect of zero-difference measurement results in this case is illustrated. The signal [x, p] is rotated using a rotation matrix with parameter ΔΦ. T The original (corrected) vector of orthogonal values ​​is used to generate [x′, p′]. T The uncorrected orthogonal values.

[0225] Based on the orthogonal rotation caused by the total systematic error, in some embodiments, the value of ΔΦ is obtained as follows. First, a vector is selected using, for example, a controller 1-107. (For example, a random vector). This vector is of a type that can be generated by the optical encoder of a photonic processing system. Second, it is calculated using, for example, a controller 1-107 or some other computing device. The output value, where M is a matrix implemented by the photonic processor under the ideal condition that no unaccounted phase accumulation ΔΦ is included. As a result, Each element corresponds to x k +ip k , where k marks each output mode of the photon processor.

[0226] In some embodiments, when calculating theoretical predictions for x k +ip k In this case, the loss of propagating random vectors through a photonic processor can be considered. For example, for a photonic processor with a transmission efficiency η, x k +iP k The field signal will become

[0227] Next, random vectors are generated from the optical encoder of the actual system. The random vector is propagated through a photonic processor. And perform a biorthogonal measurement on each element of the output vector to obtain x. k ′+ip k The phase difference ΔΦ between the local oscillator and the output modulo k signal. k From the formula Given. (Generally, since the path length from LO to detector for modulus k can differ from the path length from LO to detector for modulus l, the phase difference ΔΦ for k ≠ l is...) k ≠ΔΦ l 。 )

[0228] Finally, the local oscillator phase shifter, which controls the measurement orthogonality of the zero-difference detector, introduces θ. LO,k =ΔΦ k As a result, the axis (x, p) will be aligned with the axis (x′, p′), as follows: Figure 1-15 As shown. At this stage, the vector can be propagated again. This allows us to see that the measurement results obtained when measuring two orthogonal objects are equal to the predictions. Check the calibration to ensure it is accurate.

[0229] Typically, if the field amplitude The larger the value, the more accurately ΔΦ can be determined. k The value of . For example, if field E S,k If the optical signal is considered to be a coherent signal, for example, from a laser source, then the optical signal can theoretically be modeled as a coherent state. Figure 1-16The image provided is intuitive, where the signal is the amplitude |E S,k The noise is given by the standard deviation of the Gaussian coherent states. The coherent states of modulus k |α k > is the annihilation operator a k The eigenstate, i.e., a k |α k >=α k |α k >. By The electric field of mode k with a single frequency ω is described, which is also the eigenstate of the coherent state: A zero-difference detector with a local oscillator of the same frequency ω at θ LO Orthogonal measurement is performed when = 0. And in θ LO Orthogonal measurement is performed when =π / 2 An ideal zero-difference detector will detect these measurements having intrinsic quantum noise. and This noise is related to quantum uncertainty, and it can be reduced by squeezing it on orthogonal surfaces. The angle ΔΦ can be determined. k The accuracy is directly related to the signal-to-noise ratio (SNR) of these measurements. For a total of N... ph The coherent state signal E of each photon S,k (Right now, x k and p k The upper bound of the SNR for both is: and (when θ S When = 0 or π, SNR x The upper bound is saturated, while when θ S When = π / 2 or 3π / 2, SNR p (The upper bound is saturated.) Therefore, in order to increase SNR and more accurately determine ΔΦ... k The value can be chosen from different vectors in some implementations. Propagation is performed (e.g., multiple different random vectors). In some embodiments, each time a value pair of k is propagated. Each option is selected to make the amplitude |E S,k |=N ph maximize.

[0230] Phase shifts may occur during the operation of a photonic processing system, for example, due to temperature fluctuations over time. Therefore, in some embodiments, the calibration process described above can be repeatedly performed during system operation. For example, in some embodiments, the calibration process is performed regularly on a time scale shorter than the natural time scale of the phase shift.

[0231] The inventors have further recognized and understood that signed matrix operations can be performed without requiring phase-sensitive measurements at all. Therefore, in applications, each zero-difference detector at each output mode can be replaced by a direct photodetector used to measure the intensity of light at that output mode. Since there is no local oscillator in such a system, the system's phase error ΔΦ is nonexistent and meaningless. Therefore, according to some embodiments, phase-sensitive measurements such as zero-difference detection can be avoided, making the system's phase error negligible. For example, when using unsigned matrices to compute matrix operations on signed matrices and vectors, complex matrices and vectors, and hypercomplex (quaternions, octonions, and other isomorphic (e.g., identity algebraic elements) matrices and vectors), phase-sensitive measurements are not required.

[0232] To illustrate how phase-sensitive measurements are not necessary, consider the signed matrix M and the signed vector... The case of performing matrix multiplication between them. To calculate the signed output. The value can be determined by, for example, controller 1-107 performing the following process. First, matrix M is divided into M... + and M - M + (M - ) is a matrix containing all positive (negative) terms of M. In this case, M = M + -M - Secondly, partition the vector in a similar way, such that the vector... in It includes A vector containing all positive (negative) terms. As a result of the partitioning, Each term in this final equation corresponds to a separate operation that can be performed independently by the photonics processing system. and The output of each operation is a single (positive) sign vector, and therefore can be measured using a direct detection scheme without the need for zero-difference detection. A photodetector scheme would measure the intensity, but the square root of the intensity can be determined to obtain the electric field amplitude. In some embodiments, each operation is performed individually, and the results are stored in memory (e.g., memory 1-109 of controller 1-107) until all individual operations have been performed and the results can be digitally combined to obtain the final result of the multiplication.

[0233] The above solution is effective because M + and M - Both are matrices of all positive terms. Similarly, and Both are vectors of all positive terms. Therefore, the result of multiplying them will be a vector of all positive terms, regardless of how they are combined.

[0234] The inventors have further recognized and understood that the above-mentioned segmentation technique can be extended to complex-valued vectors / matrices, quaternions, octonions, and other hypercomplex number representations. Complex numbers use two different basic units {1, i}, quaternions use four different basic units {1, i, j, k}, and octonions use eight basic units {e0≡1, e1, e2, ..., e7}.

[0235] In some embodiments, complex vectors can be multiplied by complex matrices without phase-sensitive detection by breaking down the multiplication into separate operations similar to those described above for signed matrices and vectors. In the case of complex numbers, the multiplication is broken down into 16 separate multiplications of all positive matrices and all positive vectors. The results of these 16 separate multiplications can then be numerically combined to determine the output vector result.

[0236] In some embodiments, the multiplication of a quaternion numerical vector with a quaternion numerical matrix can be performed by dividing the multiplication into separate operations similar to those described above for signed matrices and vectors, without requiring phase-sensitive detection. In the case of quaternions, the multiplication is divided into 64 separate multiplications of all positive matrices and all positive vectors. The results of these 64 separate multiplications can then be numerically combined to determine the output vector result.

[0237] In some embodiments, octagonal numerical vectors can be multiplied by octagonal numerical matrices by dividing the multiplication into separate operations similar to those described above for signed matrices and vectors, without requiring phase-sensitive detection. In the case of octagons, the multiplication is divided into 256 separate multiplications of all positive matrices and all positive vectors. The results of these 256 separate multiplications can then be numerically combined to determine the output vector result.

[0238] The inventors have further recognized and understood that the temperature-dependent phase φ can be corrected by placing a temperature sensor next to each MZI of the photonic processor. T The temperature measurement results can then be used as input to a feedback circuit that controls the external phase of each MZI. The external phase of the MZI is set to cancel out the temperature-dependent phase accumulated at each MZI. A similar temperature feedback loop can be used on the local oscillator propagation path. In this case, the temperature measurement results are used to inform the quadrature selection phase shifter settings of the zero-difference detector to cancel out the phase accumulated in the local oscillator due to the detected temperature effect.

[0239] In some embodiments, the temperature sensor may be a temperature sensor typically used in semiconductor devices (e.g., pn junction or bipolar junction transistors), or it may be a photonic temperature sensor, such as a resonator whose resonance varies with temperature. In some embodiments, an external temperature sensor, such as a thermocouple or thermistor, may also be used.

[0240] In some embodiments, the accumulated phase can be measured directly, for example, by tapping some light at each column and performing zero-difference detection using the same global local oscillator. This phase measurement can directly inform the value of the external phase used at each MZI to correct for any phase errors. In the case of directly measuring phase errors, there is no need for global error correction across the column.

[0241] J. Intermediate computing for big data

[0242] The inventors have recognized and understood that matrix-vector multiplication performed by photonic processors 1-103 and / or any other photonic processor according to other embodiments described in this disclosure can be generalized as tensor (multidimensional array) operations. For example, The core operation (where M is a matrix, An n×m matrix (where M and X are vectors) can be generalized as a matrix-matrix product: MX, where both M and X are matrices. In this particular example, the n×m matrix X is considered as a set of m column vectors, each consisting of n elements, i.e. A photonic processor can perform a matrix-matrix product MX in a manner that involves m matrix-vector products, one column at a time. This computation can be distributed across multiple photonic processors because it is a linear operation that can be performed in parallel; for example, the output of any one matrix-vector product does not depend on the results of the others. Alternatively, the computation can be performed serially over time by a single photonic processor, for example, by performing each matrix-vector product one at a time and digitally combining the results after performing all the individual matrix-vector multiplications (e.g., by storing the results in an appropriate memory configuration).

[0243] The above concept can be summarized as calculating the product (e.g., dot product) between two multidimensional tensors. A general algorithm is as follows, and can be executed at least in part by a processor such as processor 1-111: (1) take a matrix slice of the first tensor; (2) take a vector slice of the second tensor; (3) perform a matrix-vector product between the matrix slice from step 1 and the vector slice from step 2 using a photonic processor to obtain an output vector; (4) iterate over the tensor indices of the resulting matrix slice (from step 1) and the resulting vector slice (from step 2). It should be noted that multiple indices can be combined into one when taking matrix and vector slices (steps 1 and 2). For example, a matrix can be vectorized by stacking all columns into a single column vector, and a tensor can generally be matrixed by stacking all matrices into a single matrix. Since all these operations are completely linear, they can also be performed in high parallelism, with each of the multiple photonic processors not needing to know whether the other photonic processors have completed their work.

[0244] As a non-restrictive example, consider two three-dimensional tensors C. ijlm =∑ k A ijk B klm Multiplication between them. The pseudocode based on the above scheme is:

[0245] (1) Taking a slice of the matrix: A i ←A[i,:,:];

[0246] (2) Take vector slices:

[0247] (3) Calculation in as well as

[0248] (4) Iterate over indices i, l, and m to reconstruct the four-dimensional tensor C. ijlm The values ​​of all elements indexed by j are determined entirely by a single matrix-vector multiplication.

[0249] The inventors have further recognized and understood that the size of the matrix / vector to be multiplied can be greater than the number of modules supported by the photonic processor. For example, the convolution operation in a convolutional neural network structure can define a filter using only a small number of parameters, but can consist of multiple matrix-matrix multiplications between the filter and different data slices. Combining different matrix-matrix multiplications forms two input matrices larger than the original filter matrix or data matrix.

[0250] The inventors have designed a method for performing matrix operations using a photonic processor when the matrix to be multiplied is larger than the size / number of moduli of the photonic processor used to perform the computation. In some embodiments, the method involves using memory to store intermediate information during computation. The final computation result is calculated by processing the intermediate information. For example, as... Figure 1-17 As shown, considering the multiplication between an I×J matrix A and a J×K matrix B (1-1700), a new matrix C = AB with I×K elements is given using an n×n photon processing system (where n ≤ I, J, K). Figure 1-17 In the diagram, the shaded elements simply illustrate how elements 1-1701 of matrix C are calculated using elements from rows 1-1703 of matrix A and columns 1-1705 of matrix B. Figure 1-17 and Figure 1-18 The method shown is as follows:

[0251] Construct n×n submatrix blocks within matrices A and B. Use parentheses to superscript A. (ij) and B (jk) The matrix is ​​labeled with blocks where i ∈ {1, ..., ceil(I / n)}, j ∈ {1, ..., ceil(J / n)}, and k ∈ {1, ..., ceil(K / n)}. When the values ​​of I, J, or K are not divisible by n, the matrix can be zero-padded so that the new matrix has dimensions divisible by n—hence the ceil function for indices i, j, and k. Figure 1-18 In the example multiplication 1-1800 shown, matrix A is divided into six n×n submatrix blocks 1-1803, and matrix B is divided into three n×n submatrix blocks 1-1805, resulting in a final matrix C composed of two n×n submatrix blocks 1-1801.

[0252] To compute the n×n submatrix block C within matrix C (ik) Multiplication is performed in a photonic processor through, for example, the following steps:

[0253] (1) Control the photonic processor to implement submatrix A (ij) (For example, one of submatrices 1-1803);

[0254] (2) Using submatrix B (jk) One of the column vectors (e.g., one of the submatrices 1-1805) encodes the optical signal and propagates the signal through the photonic processor;

[0255] (3) Store the intermediate results of each matrix-vector multiplication in memory;

[0256] (4) Iterate over the value of j, repeating steps (a)-(c); and

[0257] (5) The final submatrix C is computed by combining intermediate results with digital electronics (e.g., a processor). (ik) (For example, one of the submatrices 1-1801).

[0258] As described above and as Figure 1-17 and Figure 1-18 As shown, the method may include using parentheses indexing notation to represent matrix multiplication, and using parentheses superscript indices instead of subscript indices (which are used to describe matrix elements in this disclosure) to perform matrix-matrix multiplication operations. These parentheses superscript indices correspond to n×n blocks of submatrices. In some embodiments, the method can be extended to tensor-tensor multiplication by decomposing the multidimensional array into slices of n×n submatrix blocks, for example, by combining the method with the tensor-tensor multiplication described above.

[0259] In some embodiments, the advantage of using a photonic processor with a smaller number of modes to process blocks of submatrices is that it provides generality regarding the shape of the multiplied matrices. For example, in the case of I >> J, performing singular value decomposition will produce a matrix of size I. 2 The first unitary matrix of size J 2 The second unitary matrix and a diagonal matrix with J parameters. Storing or processing a matrix I with a much larger number of elements than the original matrix. 2 The hardware requirements for matrix elements may be too large for the number of optical modes included in some embodiments of photonic processors. By processing submatrices one at a time instead of the entire matrix, matrices of any size can be multiplied without limiting the number of modes based on the photonic processor.

[0260] In some embodiments, submatrices of B are further vectorized. For example, matrix A may first be filled into... The matrix is ​​then divided into submatrices (each of size [n×n]). Grid, A (ij) It is the [n×n] submatrix in the i-th row and j-th column of the grid, and B has already been filled into it first. The matrix is ​​then divided into submatrices (each of size [n×K]). Grid, B (j) It is the [n×K] submatrix in the j-th row of the grid, and C has already been filled into it. The matrix is ​​then divided into submatrices (each of size [n×K]). Grid, C (i) It is the [n×K] submatrix in the i-th row of this grid. In this vectorized form, the computation is represented as:

[0261] Using the vectorization process described above, the photonic processor can... K distinct matrices are loaded into a photon array, and for each loaded matrix, K distinct vectors are propagated through the photon array to compute any GEMM. This produces There are n output vectors (each consisting of n elements), and subsets of these vectors can be summed together to produce the desired [I×K] output matrix, as defined in the equation above.

[0262] K. Precision of calculation

[0263] The inventors have recognized and understood that photonic processors 1-103 and / or any other photonic processors according to other embodiments described in this disclosure are examples of analog computers, and because in this information age, most data is stored in digital representation, the digital precision of the computations performed by the photonic processor is important for quantization. In some embodiments, the photonic processor according to some embodiments performs matrix-vector multiplication: in M is the input vector, and M is an n×n matrix. This is the output vector. In indexed annotation, this multiplication is written as... It is M ij The n elements (iterated over j) and x j The multiplication of n elements (iterwise over j) is performed, and the results are then summed. Because the photonic processor is a physical simulation system, in some embodiments, elements M are represented using a fixed-point numerical representation. ij and x j In this representation, if It is an m1-bit number and If the number of bits is m2, then a total of m1 + m2 + log2(n) bits are used to fully represent the element y of the result vector. i Typically, the number of bits used to represent the result of a matrix-vector product is greater than the number of bits required to represent the input of the operation. If the analog-to-digital converter (ADC) used cannot read the output vector with full precision, the elements of the output vector can be rounded to the precision of the ADC.

[0264] The inventors have recognized and understood that it may be difficult to construct an ADC with high bit precision at a bandwidth corresponding to the rate at which an input vector in the form of an optical signal is transmitted through a photonic processing system. Therefore, in some embodiments, the bit precision of the ADC may limit the representation of matrix elements M. ij and vector element x j The bit precision (if fully accurate calculation is required). Therefore, the inventors have devised a method to obtain the output vector with arbitrarily high full precision by calculating the sum of partial products. For clarity, it will be assumed that M needs to be represented ijor x j The number of bits is the same, i.e., m1 = m2 = m. However, this assumption can generally be excluded and does not limit the scope of embodiments of this disclosure.

[0265] According to some embodiments, as a first action, the method includes setting matrix elements M ij and vector element x j The bit string represents a partition into d partitions, each containing k = m / d bits. (If k is not an integer, zeros can be appended until m is divisible by d). Therefore, matrix element M... ij =M ij [0] 2 k(d-1) +M ij [1] 2 k(d-2) +…+M ij [d-1] 2 0 M ij [a] It is M ij The value of the k-th bit of the most significant k-bit string of the 'a'th generation. In terms of the bit string, it is written as M. ij =M ij [0] M ij [1] …M ij [d-1] Similarly, x can also be obtained. j =x j [0] 2 k(d-1) +x j [1] 2 k(d-2) +…+x j [d-1 ]2 0 In terms of its bit string, the vector element x j =x j [0] x j [1] …x j [d-1] Multiplication y i =∑ j M ij x j It can be decomposed as follows: Where set S p Let p be the set of all integer values ​​of a and b, where a + b = p.

[0266] As a second action, the method includes controlling the photonic processor to implement matrix M. ij [a]And propagate the input vector x in the form of an encoded optical signal. j [b] With a photonic processor, each input vector is only k-bit precise. The matrix-vector multiplication operation performs y... i [a,b] =∑ j M ij [a] x j [b] This method involves storing an output vector y with an accuracy of up to 2k+log2(n) bits. i [a,b] .

[0267] The method also includes in set S p Iterate over the different values ​​of a and b, repeating the second action for each different value of a and b, and storing the intermediate result y. i [a,b] .

[0268] As a third step, the method includes calculating the final result by using digital electronics (e.g., a processor) to iteratively sum a and b.

[0269] According to some embodiments of the method, the precision of the ADC used to capture fully accurate calculations is only 2k+log2(n) bits, which is less than the precision of 2m+log2(n) bits required to complete the calculation using only a single pass.

[0270] The inventors have further recognized and understood that embodiments of the aforementioned method can be extended to operations on tensors. As previously stated, photonic processing systems can perform tensor-tensor multiplication using matrix slices and vector slices of two tensors. The above method can be applied to matrix slices and vector slices to obtain output vector slices of the output tensor with full precision.

[0271] Some embodiments of the above methods use linearity in the element-wise representation of a matrix. In the above description, the matrix is ​​represented in its Euclidean matrix space form, and matrix-vector multiplication is linear in this Euclidean space. In some embodiments, the matrix is ​​represented in the phase form of the VBS, so partitioning can be performed on the bit string representing the phase, rather than directly on the matrix elements. In some embodiments, when the mapping from phase to matrix elements is linear, the relationship between the input parameters (in this case, the phase of the VBS and the input vector elements) and the output vector is linear. When this relationship is linear, the above methods still apply. However, generally, according to some embodiments, a nonlinear mapping from the element-wise representation of the matrix to the photonic representation can be considered. For example, the bit string partitioning of the Euclidean space matrix elements from its most significant k bit string to its least significant k bit string can be used to generate a series of different matrices that are decomposed into phase representations and implemented using a photonic processor.

[0272] It is not necessary to perform partitioning on both matrix elements and input vector elements simultaneously. In some embodiments, the photonic processor can propagate many input vectors for the same matrix. Because the digital-to-analog converter (DAC) used for vector preparation can operate at high bandwidth, and the DAC used for VBS can be quasi-static for multiple vectors, it may be efficient to perform partitioning only on the input vectors and keep VBS control at a set precision (e.g., full precision). Generally, including a DAC with high bit precision at higher bandwidth is more difficult than designing a DAC with high bit precision at lower bandwidth. Therefore, in some embodiments, the output vector elements can be more precise than allowed by the ADC, but the ADC will automatically perform some rounding of the output vector values ​​up to the bit precision allowed by the ADC.

[0273] L. Manufacturing method

[0274] Examples of photonic processing systems can be fabricated using conventional semiconductor manufacturing techniques. For instance, waveguides and phase shifters can be formed in a substrate using conventional deposition, masking, etching, and doping techniques.

[0275] Figure 1-19 Exemplary methods 1-1900 for manufacturing a photonics processing system according to some embodiments are illustrated. In action 1-1901, method 1-1900 includes forming an optical encoder using, for example, conventional techniques. For example, multiple waveguides and modulators may be formed in a semiconductor substrate. The optical encoder may include one or more phase and / or amplitude modulators as described elsewhere in this application.

[0276] In action 1-1903, method 1-1900 includes forming a photonic processor and optically connecting the photonic processor to an optical encoder. In some embodiments, the photonic processor is formed in the same substrate as the optical encoder, and the optical connection is accomplished using a waveguide formed in the substrate. In other embodiments, the photonic processor is formed in a substrate separate from the substrate of the optical encoder, and the optical connection is accomplished using an optical fiber.

[0277] In action 1-1905, method 1-1900 includes forming an optical receiver and optically connecting the optical receiver to a photonic processor. In some embodiments, the optical receiver is formed in the same substrate as the photonic processor, and the optical connection is accomplished using a waveguide formed in the substrate. In other embodiments, the optical receiver is formed in a substrate separate from the substrate of the photonic processor, and the optical connection is accomplished using an optical fiber.

[0278] Figure 1-20 Example method 1-2000 for forming a photonic processor is shown, such as Figure 1-19 Action 1-1903 is shown. In action 1-2001, method 1-2000 includes, for example, forming a first optical matrix implementation in a semiconductor substrate. The first optical matrix implementation may include an array of interconnected VBSs, as described in the various embodiments above.

[0279] In action 1-2003, method 1-2000 includes forming a second optical matrix implementation and connecting the second optical matrix implementation to a first optical matrix implementation. The second optical matrix implementation may include one or more optical components capable of controlling the intensity and phase of each optical signal received from the first optical matrix implementation, as described in the above embodiments. The connection between the first and second optical matrix implementations may include a waveguide formed in a substrate.

[0280] In action 1-2005, method 1-2000 includes forming a third optical matrix implementation and connecting the third optical matrix implementation to a second optical matrix implementation. The third optical matrix implementation may include an array of interconnected VBSs, as described in the embodiments above. The connection between the second and third optical matrix implementations may include a waveguide formed in a substrate.

[0281] In any of the above actions, the components of the photonic processor can be formed in the same layer of the semiconductor substrate or in different layers of the semiconductor substrate.

[0282] M. How to use

[0283] Figure 1-21Methods 1-2100 for performing optical processing according to some embodiments are illustrated. In action 1-2101, method 1-2100 includes encoding a bit string into an optical signal. In some embodiments, this can be performed using a controller and an optical encoder, as described in conjunction with various embodiments of this application. For example, complex numbers can be encoded into the intensity and phase of the optical signal.

[0284] In action 1-2103, method 1-2100 includes controlling a photonic processor to implement a first matrix. As described above, this can be achieved by having the controller perform SVD on the matrix and decompose the matrix into three separate matrix components implemented using individual parts of the photonic processor. The photonic processor may include a plurality of interconnected VBSs that control how the various modes of the photonic processor are mixed together to coherently interfere with the optical signal as it propagates through the photonic processor.

[0285] In action 1-2105, method 1-2100 includes propagating an optical signal through an optical processor such that the optical signals coherently interfere with each other in a manner that achieves a desired matrix, as described above.

[0286] In action 1-2107, method 1-2100 includes detecting the output optical signal from the photonic processor using an optical receiver. As described above, the detection can use a phase-sensitive detector or a non-phase-sensitive detector. In some embodiments, the detection result is used to determine a new input bit string to be encoded and propagated through the system. In this way, multiple computations can be performed serially, where at least one computation is based on the result of a previous computation.

[0287] II. Training Algorithm

[0288] The inventors have recognized and understood that for many matrix-based differentiable programs (e.g., neural networks or latent variable graphical models), much of the computational complexity lies in the matrix-matrix products calculated as the model layers are traversed. The complexity of a matrix-matrix product is O(IJK), where the dimensions of the two matrices are I×J and J×K, respectively. Furthermore, these matrix-matrix products are performed during both the training and evaluation phases of the model.

[0289] Deep neural networks (i.e., neural networks with more than one hidden layer) are examples of a class of matrix-based differentiable programs that can employ some of the techniques described herein. However, it should be understood that the techniques described herein for performing parallel processing can be used with other types of matrix-based differentiable programs, including but not limited to Bayesian networks, grid decoders, topic models, and hidden Markov models (HMMs).

[0290] The success of deep learning is largely attributed to the development of backpropagation techniques, which allow training on the weight matrices of neural networks. In conventional backpropagation, the error from the loss function is propagated backward through the individual weight matrix components using the chain rule of calculus. Backpropagation computes the gradients of the elements in the weight matrix and then uses them to determine updates to the weight matrix using optimization algorithms such as stochastic gradient descent (SGD), AdaGrad, RMSProp, Adam, or any other gradient-based optimization algorithm. This process is continuously applied to determine the final weight matrix that minimizes the loss function.

[0291] The inventors have recognized and appreciated that optical processors of the type described herein achieve gradient computation performance by recombining weight matrices into an alternative parameter space (referred to herein as the “phase space” or “angle representation”). Specifically, in some embodiments, the weight matrices are reparameterized as a combination of unitary transformation matrices (such as Givens rotation matrices). During this reparameterization process, training the neural network involves adjusting the angle parameters of the unitary transformation matrices. In this reparameterization process, the gradient of a single rotation angle is decoupled from other rotations, thereby allowing for parallel computation of gradients. This parallelization improves computational speed in terms of the number of computational steps required compared to conventional serial gradient determination techniques.

[0292] The above provides an example photonics processing system that can be used to implement the backpropagation technique described herein. The phase space parameters of the reparameterized weight matrix can be encoded into a phase shifter or variable beam splitter of the photonics processing system to implement the weight matrix. Encoding the weight matrix into a phase shifter or variable beam splitter can be used for both the training and evaluation phases of the neural network. While the backpropagation process is described in conjunction with the specific system described below, it should be understood that the embodiments are not limited to the specific details of the photonics processing system described in this disclosure.

[0293] As described above, in some embodiments, the photonic processing system 100 can be used to implement aspects of neural networks or other matrix-based differentiable programs that can be trained using backpropagation techniques.

[0294] Figure 2-1 An example backpropagation technique 2-100 is shown for updating the value matrix (e.g., the weight matrix for a layer in a neural network) in the Euclidean vector space of a differentiable program (e.g., a neural network or a latent variable graphical model).

[0295] In action 2-101, the value matrix in the Euclidean vector space (e.g., the weight matrix for a layer in a neural network) can be represented as an angle representation by, for example, configuring components of the photon processing system 100 to represent the value matrix. After the matrix is ​​represented as an angle representation, process 2-100 proceeds to action 2-102, where training data (e.g., a set of input training vectors and associated labeled outputs) is processed to compute an error vector by evaluating a performance metric of the model. Then, process 2-100 proceeds to action 2-103, where at least some gradients of the parameters of the angle representation required for backpropagation are determined in parallel. For example, as discussed in more detail below, the technique described herein enables the simultaneous determination of gradients of an entire column of parameters, significantly accelerating the amount of time required to perform backpropagation compared to estimating the gradients individually for each angle rotation. Then, process 2-100 proceeds to action 2-104, where the value matrix in the Euclidean vector space (e.g., the weight matrix values ​​of a layer in a neural network) is updated by updating the angle representation using the determined gradients. Further description follows. Figure 2-1 Each action shown in process 2-100.

[0296] Figure 2-2 This illustrates how it can be performed according to some embodiments. Figure 2-1 The flowchart for action 2-101 is shown. In action 2-201, the controller (e.g., controller 107) can receive the weight matrix for the layers of the neural network. In action 2-202, the weight matrix can be decomposed into a first unitary matrix V, a second unitary matrix U, and a diagonal matrix ∑ with signed singular values, such that the weight matrix W is defined as:

[0297] W = V T ∑U,

[0298] Where U is an m×m unitary matrix, V is an n×n unitary matrix, ∑ is an n×m diagonal matrix with signed singular values, and the superscript "T" indicates the transpose of the matrix. In some embodiments, the weight matrix W is first partitioned into slices, each slice being decomposed into a triple product of such matrices. The weight matrix W can be a conventional weight matrix known in the field of neural networks.

[0299] In some embodiments, the weight matrix for decomposing into the phase space is a pre-specified weight matrix, such as those provided by a random initialization process or by employing a partially trained weight matrix. If no partially specified weight matrix is ​​available for initializing the backpropagation routine, the decomposition in action 2-202 can be skipped, and the parameters of the angle representation (e.g., singular values ​​and parameters of a unitary or orthogonal decomposition) can be initialized, for example, by randomly sampling phase space parameters from a particular distribution. In other embodiments, a predetermined set of initial singular values ​​and angle parameters can be used.

[0300] At action 2-203, the two unitary matrices U and V can be represented as combinations of the first set of unitary transformation matrices and the second set of unitary transformation matrices, respectively. For example, when matrices U and V are orthogonal matrices, they can be transformed in action 2-203 into a series of real-valued Givens rotation matrices or Householder reflectors, examples of which are described in Section V above.

[0301] In action 2-204, the photonics-based processor can be configured to implement unitary transformation matrices based on the decomposed weight matrices. For example, as described above, a first set of components of the photonics-based processor can be configured based on the first set of unitary transformation matrices, a second set of components can be configured based on a diagonal matrix with signed singular values, and a third set of components can be configured based on the second set of unitary transformation matrices. Although the process described herein relates to implementing backpropagation techniques using photonics-based processors, it should be understood that backpropagation techniques can be implemented using other computing architectures that provide parallel processing capabilities, and embodiments in this regard are not limited.

[0302] Return to Figure 2-1 Process 2-100 and action 2-102 involve, for example, using a photon-based processor as described above to process the training data to compute the error vector after the weight matrix has been represented as an angle representation. Figure 2-3 A flowchart illustrating implementation details for performing action 2-102 according to some embodiments is shown. Training data can be divided into batches before processing it using the techniques described herein. The training data can take any form. In some embodiments, the data can be divided into batches in the same manner as for some conventional mini-batch stochastic gradient descent (SGD) techniques. The gradients computed in this process can be used for any optimization algorithm, including but not limited to SGD, AdaGrad, Adam, RMSProp, or any other gradient-based optimization algorithm.

[0303] exist Figure 2-3In the process shown, each vector in a specific batch of training data can be passed through a photonic processor, and the value of a loss function can be calculated for that vector. In action 2-301, an input training vector from the specific batch of training data is received. In action 2-302, the input training vector is converted into a photonic signal, for example, by encoding the vector with light pulses having amplitude and phase corresponding to the input training vector values, as described above. In action 2-303, the photonic signal corresponding to the input training vector is provided as input to a photonic processor (e.g., photonic processor 103), which has been configured as described above to implement (e.g., using an array of configurable phase shifters and beam splitters) a weight matrix to produce an output pulse vector. The light intensity of the pulses output from the photonic processor can be detected using, for example, zero-difference detection, as described above in conjunction with Figures 9 and 10, to produce a decoded output vector. In action 2-304, the value of a loss function (also called a cost function or error metric) is calculated for the input training vector. Then, repeat steps 2-301 to 2-304 until all input training vectors in a given batch have been processed and the corresponding values ​​of the loss function have been determined. In step 2-305, the total loss is calculated by aggregating these losses from each input training vector; for example, this aggregation can be in the form of an average.

[0304] Return to Figure 2-1 Process 2-100 and action 2-103 involve parallel computation of the gradients of the parameters representing the angles (e.g., the values ​​of the weights already implemented as Givens rotations using components of a photonic processor). The gradients can be computed based on the computed error vector, the input data vector (e.g., from a batch of training data), and the weight matrix implemented using a photonic processor. Figure 2-4 A flowchart is shown for performing actions 2-103 according to some embodiments. In action 2-401, for the k-th Givens rotation set G in the decomposition... (k) (For example, the columns of the decomposed matrix, referred to herein as the "derivative sequence k"), compute a block diagonal derivative matrix containing the derivative with respect to each angle in the k-th set. In action 2-402, compute the product of the error vector determined in action 2-102 with all unitary transformation matrices between the derivative sequence k and the output (referred to herein as "partial backward pass"). In action 2-403, compute the product of the input data vector with all unitary transformation matrices from the input up to and including the derivative sequence k (referred to herein as "partial forward pass"). In action 2-404, compute the inner product between consecutive element pairs of the outputs from actions 2-402 and 2-403 to determine the gradient of the derivative sequence k. The inner product between consecutive element pairs can be computed as... The superscript (k) represents the k-th column of the photon element, and i and j represent the sequence of photons with parameter θ.ij (k) The unitary transformation matrix couples the i-th and j-th optical modes, where x is the partially forward-passing output and δ is the partially backward-passing output. In some embodiments, an offset is applied before the consecutive pairings of outputs (e.g., output pairs could be (1, 2), (3, 4), etc., instead of (0, 1), (2, 3)). The determined gradient can then be appropriately used for a particular selection of optimization algorithm (e.g., SGD) that is being used for training.

[0305] The following provides a method for using, according to some embodiments, in having Figure 1-4 Example pseudocode for implementing backpropagation on a photonic processor with the topology shown from left to right:

[0306] Initialize two lists x′ and δ′ of the intermediate propagation results.

[0307] For each column of MZI, start from the last column and proceed to the first column:

[0308] Rotate the angles in this column to correspond to the derivative matrix.

[0309] ο Propagate input data vector through photonic processor

[0310] Store the result in x′.

[0311] Make the current column transparent

[0312] • Construct the transpose matrix column by column. For each new column:

[0313] The propagation error vector is transmitted through the photonic processor.

[0314] Store the results in δ′.

[0315] For each x′[i], δ′[i]

[0316] ο Calculate the inner product between consecutive pairs, and the result is the gradient of the angle in the i-th column of MZI.

[0317] According to some embodiments, instead of adjusting the weights of the weight matrix via gradient descent as in some conventional backpropagation techniques, the parameters representing the angles are adjusted (e.g., the singular values ​​of matrix ∑ and the Givens rotation angles of the orthogonal matrices U and V). To further illustrate how backpropagation works in a reparameterized space according to some embodiments, a comparison is made below of backpropagation within a single layer of a neural network using conventional techniques with methods according to some embodiments of this disclosure.

[0318] The loss function E measures the model's performance for a specific task. In some conventional stochastic gradient descent algorithms, the weight matrix is ​​iteratively adjusted such that the weight matrix at time t+1 is defined as a function of the weight matrix at time t, and the derivative of the loss function with respect to the weights of the weight matrix is ​​as follows:

[0319]

[0320] Where η is the learning rate, and (a, b) represent the items in the a-th row and b-th column of the weight matrix W, respectively. When the decomposed weight matrix is ​​used to reconstruct the iterative process, the weights w ab σ is the singular value of matrix ∑. i And the rotation angle θ of orthogonal matrices U and V ij The function. Therefore, the iterative adjustment of the backpropagation algorithm becomes:

[0321]

[0322] as well as

[0323]

[0324] To iteratively adjust the singular values ​​and rotation angles, the derivative of the loss function must be obtained. Before describing how this is implemented in a system such as the photon processing system 100, a description of backpropagation is first provided based on iteratively adjusting the weights of the weight matrix. In this case, the output measured by the system for a single layer of the neural network is represented as the output vector y. i =f((Wx) i +b i ), where W is the weight matrix, x is the data vector input to this layer, b is the bias vector, and f is a nonlinear function. The chain rule of calculus is applied to calculate the gradient of the loss function with respect to any parameter within the weight matrix (where z is defined as... for ease of representation). i =(Wx) i +b i ):

[0325]

[0326] Calculate z relative to w ab The derivative of the result yields:

[0327]

[0328]

[0329] Using this fact, the sum of the gradients of the loss function can then be written as:

[0330]

[0331] Defining the first sum as the error vector e, and x as the input vector, we obtain the following expression:

[0332]

[0333] Using the equation from conventional backpropagation, the equation can be extended to the case where the weight matrix is ​​decomposed into singular value matrices and unitary transformation matrices. Taking advantage of the fact that the weight matrix is ​​a function of the rotation angle, it can be written using the chain rule:

[0334]

[0335] Therefore, backpropagation in phase space involves the same parts as in regular backpropagation (error vector and input data), plus a term for the derivative of the weight matrix with respect to the rotation angle of the unitary transformation matrix.

[0336] To determine the derivative of the rotation angle of the weight matrix relative to the unitary transformation matrix, it should be noted that the derivative of a single Givens rotation matrix has the following form:

[0337]

[0338] It can be seen that the derivative of any term in the Givens rotation matrix that is not in the i-th row and j-th column is zero. Therefore, G (k) All derivatives of the internal angles can be grouped into a single matrix. In some embodiments, all derivatives within the columns of the unitary transformation matrix are computed. The derivative of each angle can be obtained using the two-step process described above. First, the error vector is propagated from the right (output) through the decomposed matrix until the current set of rotations is differentiated (partially in the reverse direction). Second, the input vector is propagated from the left (input) until the current set of rotations is reached (partially in the forward direction), and then the derivative matrix is ​​applied.

[0339] In some embodiments, the differentiation of singular values ​​is performed using a similar process. Regarding the singular value σ... i The derivative of the element ∑′ ii =1 and all other elements ∑′ jj The value is 0. Therefore, all derivatives of the singular values ​​can be computed together. In some embodiments, this can be done by propagating the error vector from the left (partially forward) and the input vector from the right (partially backward), and then computed the Hadamard product based on the outputs of the forward and backward passes.

[0340] exist Figure 2-4In the implementation of action 2-103 described above, all angles in column k are rotated by π / 2 in order to compute the gradient term of that column. In some embodiments, this rotation is not performed. Consider the matrix of a single MZI:

[0341]

[0342] Taking the derivative with respect to θ, we get

[0343]

[0344] Although this matrix corresponds to adding π / 2 to θ, it also corresponds to swapping the columns of the original matrix and inverting one of them. In mathematical notation, this means...

[0345]

[0346] In some embodiments, instead of rotating the angle of each MZI in a column by π / 2 and then calculating the inner product (e.g., x1δ1 + x2δ2) between consecutive element pairs output from actions 2-402 and 2-403 as described above to determine the gradient of a column of the decomposed unitary matrix, instead of rotating the angle by π / 2, the relation x1δ2 - x2δ1 is calculated to obtain the same gradient. In some embodiments, actions 2-401 to 2-404 allow the controller to obtain a unitary matrix / orthogonal matrix of size n×n when the size of matrix W (n×m) matches the size of a photon processor having a matrix U of size n×n and a matrix V of size m×m. There are several gradients. Therefore, on hardware such as the photonic processing system 100 described above, where each matrix multiplication can be computed in O(1) operations, the entire backpropagation process can be completed in O(n+m) operations when the size of the photonic processor is large enough to represent the full matrix. When the size of the photonic processor is insufficient to represent the full matrix, as mentioned above, the matrix can be segmented into pieces. Consider a photonic processor of size N. If the task is to multiply a matrix of size I×J with a vector of size J, a single matrix-vector product will have a complexity of O(IJ / N). 2 (Assuming both I and J are divisible by N), because each dimension of the matrix must be partitioned into N matrices, loaded into the processor, and used to compute a portion of the result. For a batch of K vectors (e.g., a second matrix of size J×K), the complexity of the matrix-vector product is O(IJK / N). 2 ).

[0347] As described above, embodiments of a photonic processor with n optical modes are naturally capable of computing matrix-vector products between matrices of size [n×n] and n-element vectors. This is equivalent to matrix-matrix products between matrices of size [n×n] and [n×1]. Furthermore, a sequence of K matrix-vector product operations with K distinct input vectors and a single repeating input matrix can be represented as the computation of matrix-matrix products between matrices of size [n×n] and [n×K]. However, the applications and algorithms described herein generally involve the computation of general matrix-matrix multiplication (GEMM) between matrices of arbitrary size; that is, the computation of...

[0348]

[0349] Where a ij b is the element in the i-th row and j-th column of the [I×J] matrix A. jk It is the element in the j-th row and k-th column of the [J×K] matrix B, c ik It is the element in the i-th row and k-th column of the [I×K] matrix C=AB. Due to the recursive nature of this calculation, this can be equivalently expressed as:

[0350]

[0351] A has already been filled in first. The matrix is ​​then divided into submatrices (each of size [n×n]). Grid, A ij It is the [n×n] submatrix of the i-th row and j-th column of the grid, and B has already been filled into it. The matrix is ​​then divided into submatrices (each of size [n×K]). Grid, B j It is the [n×K] submatrix of the j-th row of the grid, and C has already been filled into it. The matrix is ​​then divided into submatrices (each of size [n×K]). Grid, C i It is the [n×K] submatrix of the i-th row of the grid.

[0352] Using this process, the photonic processor can... A number of different matrices are loaded into a photon array, and for each loaded matrix, k different vectors are propagated through the photon array to compute any GEMM. This produces There are n output vectors (each consisting of n elements), and subsets of these vectors can be summed together to produce the desired [I×K] output matrix, as defined in the equation above.

[0353] exist Figure 2-4In the implementation of action 2-103 described above, it is assumed that the photonic processor used to implement the matrix has a left-to-right topology, where vectors are input to the left side of the array of optical components and output vectors are provided to the right side of the array of optical components. This topology requires the calculation of the transpose of the angle representation matrix when propagating the error vector through the photonic processor. In some embodiments, the photonic processor is implemented using a folded topology that arranges both the input and output on one side (e.g., the left side) of the array of optical components. This structure allows the use of a switch to determine the direction in which light should propagate—from the input to the output or from the output to the input. By selecting a configurable direction during operation, the propagation of the error vector can be achieved by first switching the direction to use the output as the input and then propagating the error vector through the array. This eliminates the need for inverting the phase (e.g., by rotating the angle of each photonic element in column k) and transposing the column when determining the gradient of column k, as described above.

[0354] Return to Figure 2-1 Process 2-100 and action 2-104 aim to update the weight matrix by updating the parameters represented by the angle based on the determined gradient. The previously described... Figure 2-4 The process shown is used to compute the gradient of the angular parameters of a single set (e.g., column k) of gradients in the decomposed unitary matrix based on a single input data sample. To update the weight matrix, the gradient needs to be computed for each set (e.g., column) for each input data vector in the batch. Figure 2-5 A flowchart illustrating the process for computing all the gradients required to update the weight matrix, according to some embodiments, is shown. In action 2-501, the gradient of one set of unitary transformation matrices from a plurality of sets of unitary transformation matrices (e.g., Givens rotations) is determined (e.g., using...). Figure 2-4The process is illustrated in step 2-502. Then, it is determined whether there exists an additional set of unitary transformation matrices (e.g., columns) for which gradients need to be computed for the current input vector. If it is determined that additional gradients exist, the process returns to step 2-501, where a new set (e.g., columns) is selected and gradients are computed for the newly selected set. As described below, in some computational structures, all columns of the array can be read simultaneously, making the determination in step 2-502 unnecessary. The process continues until it is determined in step 2-502 that gradients have been determined for all sets of unitary transformation matrices (e.g., columns) in the decomposed unitary matrix. The process then proceeds to step 2-503, where it is determined whether there are more input data vectors to process in the batch of training data being processed. If it is determined that additional input data vectors exist, the process returns to step 2-501, where new input data vectors from the batch are selected and gradients are computed based on the newly selected input data vectors. This process is repeated until it is determined in step 2-503 that all input data vectors in the batch of training data have been processed. In action 2-504, the gradients determined in actions 2-501 through 2-503 are averaged, and the parameters representing the angles (e.g., the angles of the Givens rotation matrix of the weight matrix decomposition) are updated based on an update rule utilizing the averaged gradients. As a non-limiting example, the update rule may include varying the magnitude of the average gradient through the learning rate, or may include "momentum" or other corrections for the gradient history of the parameters.

[0355] As briefly discussed above, while the examples are applied to real weight matrices in a single-layer neural network, the results can be generalized to networks with multiple layers and complex weight matrices. In some embodiments, the neural network consists of multiple layers (e.g., ≥50 layers in a deep neural network). To compute the gradient of the matrix of layer L, the input vector to that layer will be the output of the previous layer L-1, and the error vector to that layer will be the error backpropagated from the next layer L+1. The value of the backpropagated error vector can be computed using the chain rule of multivariable calculus as described above. Furthermore, in some embodiments, complex U and V matrices (e.g., unitary matrices) can be used by adding additional complex phase terms to the Givens rotation matrix.

[0356] While the above description is generally applied independently of hardware architecture, certain hardware architectures offer more significant computational acceleration than others. In particular, to obtain the maximum gain compared to conventional methods, it is preferable to implement the backpropagation technique described herein on a graphics processing unit, a systolic matrix multiplier, a photonic processor (e.g., photonic processing system 100), or other hardware architectures capable of parallelizing gradient computation.

[0357] As described above, the photonic processing system 100 is configured to implement any unitary transformation. Givens rotation sequences are an example of such unitary transformations, therefore the photonic processing system 100 can be programmed to compute the transformations in the above decomposition in O(1) time. As described above, the matrices can be implemented by controlling a regular array of variable beam splitters (VBS). The unitary matrices U and V... T The array can be decomposed into a VBS (Virtual Baseboard), with each VBS performing a 2×2 orthogonal matrix operation (e.g., Givens rotation). As described above, by controlling the intensity and phase of the optical pulses, the diagonal matrix ∑ and the diagonal phase screen D can be implemented in the photonics processing system 100. U and D V (using the diagonal matrix D) U ∑D V (in the form of).

[0358] Each term in the diagonal matrix ∑ corresponds to either amplification or attenuation of each photon mode. Terms with an amplitude ≥ 1 correspond to amplification, and terms with an amplitude ≤ 1 correspond to attenuation. The combination of VBS and the gain medium will allow for either attenuation or amplification. For an n×n square matrix M, the number of optical modes required to apply the diagonal matrix ∑ is n. However, if the matrix M is not square, the number of optical modes required is equal to the smaller dimension.

[0359] As described above, in some embodiments, the size of the photonic processor is the same as the size of the matrix M being multiplied and the size of the input vector. However, in practice, the size of the matrix M and the size of the photonic processor are often different. Consider a photonic processor of size N. If the task is to multiply a matrix of size I×J with a vector of size J, then a single matrix-vector multiplication will have a complexity of O(IJ / N). 2 (Assuming both I and J are divisible by N), because each dimension of the matrix must be partitioned into N matrices, loaded into the processor, and used to compute a portion of the result. For a batch of K vectors (e.g., a second matrix of size J×K), the complexity of the matrix-vector product is O(IJK / N). 2 ).

[0360] If the matrix is ​​non-square, especially if I >> J or J >> I, then the ability to work with small N×N matrix partitions may be advantageous. Assuming a non-square matrix A, a direct SVD of the matrix produces an I×I unitary matrix, a J×J unitary matrix, and an I×J diagonal matrix. If I >> J or J >> I, it means that the number of parameters required for this decomposition is much larger than that of the original matrix A.

[0361] However, if matrix A is partitioned into multiple N×N square matrices with smaller dimensions, then SVD over these N×N matrices produces two N×N unitary matrices and one N×N diagonal matrix. In this case, the number of parameters required to represent the decomposed matrices is still N. 2 —Equal to the size of the original matrix A, and can be decomposed into the total non-square matrix using approximately IJ total parameters. When IJ can be decomposed by N... 2 When divisible, they become approximately equal.

[0362] For a photonic processor with 2N+1 columns, a partial result of the backpropagation error vector for each column can be computed. Therefore, for a batch of K vectors, the complexity of backpropagation using a photonic processor of size N is O(IJK / N). By comparison, the computation of the backpropagation error using a matrix multiplication algorithm on a non-parallel processor (e.g., a CPU) would be O(IJK).

[0363] The descriptions to date have focused on the use of matrices within neural network layers with input vector data and backpropagation error vectors. The inventors have recognized and understood that the data in deep neural network computations is not necessarily vectors, but rather multidimensional tensors. Similarly, the weights describing the connections between neurons are also typically multidimensional tensors. In some embodiments, if the weight tensors are segmented into matrix slices, where each slice is independent of the others, the methods described above can be applied directly. Therefore, singular value decomposition and Givens-like rotation decompositions can be performed to obtain an efficient representation of the phase form of a particular matrix slice. The same method for calculating the gradient of the phase can then be applied with appropriate arrangement of the input tensor data and the backpropagation error data. The gradient of a particular matrix slice should be calculated using the portions of the input and error data that contribute to that particular matrix slice.

[0364] For a specific problem, consider a general n-dimensional weight tensor. From {a} which constitutes a matrix slice i} i=1,...,n Choose two indexes (assuming the choice is made by index a). b and a c (mark), then through calculation To perform decomposition to obtain phase θ ij (k) Importantly, b and c can be any values ​​between 1 and n. Now consider a general k-dimensional input tensor. For effective tensor operations to occur, the operation on the weight tensor of the input tensor must produce an (nk)-dimensional output tensor. Therefore, it can be concluded that the backpropagation error tensor, and the weight tensor of this layer, are (nk)-dimensional tensors. Therefore, the gradient to be calculated is

[0365]

[0366] For simplicity (but not a necessary condition), the indices of the weight tensor have been sorted such that the first k indices operate on x and the last (nk) indices operate on the error e.

[0367] In other embodiments, higher-order generalizations of singular value decomposition can be performed more conveniently, such as Tucker decomposition, where any n-dimensional tensor can be decomposed in this way:

[0368] Each of them They are all decomposable into orthogonal matrices of their Givens rotation phases, and It is an n-dimensional core tensor. In some cases, using a special case of Tucker decomposition known as CANDECOMP / PARAFAC (CP) decomposition, the core tensor can be chosen to be superdiagonal. The Tucker decomposition can be made analogous to a two-dimensional SVD form by multiplying some inverses (transposes or conjugate transposes) of the unitary matrices. For example, the decomposition can be rewritten as... The first m unitary matrices are pushed to the left of the core tensor. The set of unitary matrices on either side of the core tensor can be decomposed into their rotation angles, and the gradient of each rotation angle is obtained by the chain rule of calculus and the abbreviation of the gradient with respect to the input and error tensors.

[0369] The inventors have recognized and understood that the gradient of the phase (e.g., for decomposed matrices U and V) and the gradient of the signed singular values ​​(e.g., for matrix ∑) can have different upper bounds. Consider the task of computing the gradient of the scalar loss function L with respect to the parameters of a neural network. In Euclidean space, the value of the gradient is given by... Given, where W is a matrix. In the phase space, for a specific scalar phase θ k Chain rules provide:

[0370]

[0371] According to the definition of a trace, this is equivalent to:

[0372]

[0373] in and Both are matrices. It is known that the trace is a product of Frobenius norms, Tr(AB) ≤ ||A|| F ||AB|| F Bounded by, and ||A|| F =||A T|| F .therefore,

[0374]

[0375] Since the derivative with respect to θ does not change the singular values ​​of W, and therefore does not change the Frobenius norm, the following holds:

[0376]

[0377] Relative to a specific singular value σ k When we perform differentiation, all singular values ​​become zero except for the singular value that is differentiated (which becomes 1). This means that...

[0378]

[0379] therefore,

[0380]

[0381] In some embodiments, the gradient magnitudes of the phase and singular values ​​are varied, respectively, during the updating of the parameters representing the angle, for example, taking into account differences in upper bounds. By varying the gradient magnitudes, either the gradient of the phase or the gradient of the singular values ​​can be readjusted to have the same upper bound. In some embodiments, the gradient magnitude of the phase is varied by the Frobenius norm of the matrix. Depending on some update rules, varying the gradient magnitudes of the phase and singular values ​​is independently equivalent to having different learning rates for the gradients of the phase and singular values. Therefore, in some embodiments, the first learning rate used to update the component sets of the U and V matrices differs from the second learning rate used to update the component sets of the ∑ matrix.

[0382] The inventors have recognized and understood that once the weight matrix is ​​decomposed into a phase space, neither the phase nor the singular values ​​necessarily need to be updated in each iteration to obtain a good solution. Therefore, if only the singular values ​​(not the phase) are updated for a portion of the overall training time, only O(n) parameters need to be updated during these iterations instead of O(n^2). 2 This improves overall runtime by reducing the number of parameters. Updating only singular values ​​or phase during some iterations can be termed "parameter claiming." In some embodiments, parameter claiming can be performed using one or more of the following claiming techniques:

[0383] • Fixed clamping: Train all parameters for a certain number of iterations, and then only update the singular values ​​thereafter.

[0384] • Cyclic clamping: Train all parameters for M iterations, then freeze the phase for N iterations (i.e., update only the singular values). Continue training all parameters for another M iterations, then freeze the phase again for N iterations. Repeat until the desired total number of iterations has been reached.

[0385] • Warm-up clamping: Train all parameters to a certain number of iterations K, and then start cyclic clamping for the remaining iterations.

[0386] Threshold clamping: Continue updating the phase or singular values ​​until their updates are less than the threshold ε.

[0387] The inventors have recognized and understood that the architecture of a photonic processor can affect computational complexity. For example, in the architecture shown in Figure 1, the detectors are located only at one end of the photonic processor array, thus forming a column-by-column approach for gradient computation (e.g., using partial forward and backward passes), as described above. In an alternative architecture, detectors can be placed at each photonic element in the photonic array (e.g., at each MZI). For such an architecture, the column-by-column approach can be replaced with a single forward pass and a single backward pass, where the output of each column is read out simultaneously, thus providing additional computational speedup. An intermediate solution of intermittently placing detector columns throughout the array has also been considered. Arbitrarily increasing the number of detector columns will correspondingly reduce the number of partial forward and backward passes required for gradient computation.

[0388] The above technique illustrates a method for performing updates to the weight matrix parameters while preserving all computations in the phase space (e.g., using an angular representation of the matrix). In some embodiments, at least some computations can be performed in the Euclidean vector space, while others are performed in the phase space. For example, as described above, the quantities required to perform the update can be computed in the phase space, while the actual update of the parameters can occur in the Euclidean vector space. The update matrix computed in the Euclidean vector space can then be decomposed back into the weight space for the next iteration. In the Euclidean vector space, for a given layer, the update rule can be:

[0389]

[0390] The δ in this calculation can be computed in reverse through the entire photonic processor in phase space. Then, the aforementioned outer product between x and δ can be computed separately (e.g., off-chip). Once the update is applied, the updated matrix can be re-decomposed, and the decomposed values ​​can be used to set the phase of the photonic processor, as described above.

[0391] III. Convolutional Layers

[0392] Convolution and cross-correlation are common signal processing operations in many applications such as audio / video coding, probability theory, image processing, and machine learning. The terms convolution and cross-correlation generally refer to mathematical operations that take two signals as input and produce a third signal as output representing the similarity between the inputs. The inventors have recognized and understood that computing convolution and cross-correlation can be computationally resource-intensive. In particular, the inventors have developed techniques for improving the computational speed and efficiency of convolution and cross-correlation. Embodiments of these techniques include computing convolution and cross-correlation by transforming the convolution operation into a matrix-vector product and / or a product of multidimensional arrays. Embodiments of these techniques also include computing convolution based on discrete transformations.

[0393] The inventors have further recognized and understood that computation of convolution and cross-correlation can be performed in various ways depending on the intended application. Input and output signals can be discrete or continuous. The data values ​​composing the signals can be defined over various digital domains, such as real numbers, the complex plane, or finite integer rings. Signals can have any number of dimensions. Signals can also have multiple channels, a technique commonly used in convolutional neural networks (CNNs). The embodiments described herein can be implemented to adapt to these variations in any combination.

[0394] Furthermore, embodiments of these technologies can be implemented in any suitable computing system configured to perform matrix operations. Examples of such computing systems that can benefit from the technologies described herein include central processing units (CPUs), graphics processing units (GPUs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and photonic processors. While the embodiments described herein may be described in conjunction with photonic processors, it should be understood that these technologies can be applied to other computing systems, such as, but not limited to, those described above.

[0395] The various concepts and embodiments of these techniques related to the computation of convolution and cross-correlation will be described in more detail below. It should be understood that the aspects described herein can be implemented in any of a number of ways. Examples of specific implementations are provided herein for illustrative purposes only. Furthermore, the aspects described in the following embodiments can be used individually or in any combination, and are not limited to the combinations explicitly described herein.

[0396] The inventors have also recognized and understood that using optical signals instead of electrical signals overcomes many of the aforementioned problems with electrical computing. Optical signals travel at the speed of light in the medium through which light travels; therefore, the time delay of photonic signals is much smaller than the electrical propagation delay. Furthermore, increasing the propagation distance of optical signals does not cause power dissipation, revealing new topologies and processor layouts that are not feasible using electrical signals. Therefore, light-based processors (e.g., photonic processors) can have better speed and efficiency performance than conventional electronically based processors.

[0397] In order to realize a photonics-based processor, the inventors have recognized and understood that the multiplication of an input vector and a matrix can be achieved by propagating a coherent optical signal (e.g., a laser pulse) through a first array of interconnected variable beam splitters (VBS), a second array of interconnected variable beam splitters, and a plurality of controllable electro-optical elements located between the two arrays, connecting a single output of the first array to a single input of the second array.

[0398] As described herein, some embodiments of photonic processors with n optical modes naturally compute matrix-vector products between matrices of size [n×n] and vectors with n elements. This can be equivalently represented as matrix-matrix products between matrices of sizes [n×n] and [n×1]. A sequence of K matrix-vector product operations with K distinct input vectors and a single repeating input matrix can be represented as the computation of matrix-matrix products between matrices of sizes [n×n] and [n×K]. However, the applications and algorithms described herein typically involve the computation of general matrix-matrix multiplication (GEMM) between matrices of arbitrary size; that is, the computation of:

[0399]

[0400] Where a ij b is the element in the i-th row and j-th column of the [I×J] matrix A. jk It is the element in the j-th row and k-th column of the [J×K] matrix B, c ik It is the element in the i-th row and k-th column of the [I×K] matrix C=AB. Due to the recursive nature of this calculation, this can be equivalently expressed as:

[0401]

[0402] Matrix A has already been filled in. The matrix is ​​then divided into submatrices (each of size [n×n]). Grid, A ij It is the [n×n] submatrix in the i-th row and j-th column of the grid, and B has already been filled into it first. The matrix is ​​then divided into submatrices (each of size [n×K]). Grid, Bj It is the [n×K] submatrix in the j-th row of the grid, and C has already been filled into it. The matrix is ​​then divided into submatrices (each of size [n×K]). Grid, C i It is the [n×K] submatrix in the i-th row of the grid.

[0403] According to some embodiments, using this process, the photonic processor can... K distinct matrices are loaded into a photon array, and for each loaded matrix, K distinct vectors are propagated through the photon array to compute any GEMM. This produces There are n output vectors (each consisting of n elements), and subsets of these vectors can be summed together to produce the desired [I×K] output matrix, as defined in the equation above.

[0404] The inventors have recognized and understood that photonic processors can accelerate the computation of convolutions and cross-correlations; however, the embodiments described herein for computing convolutions and cross-correlations can be implemented on any suitable computing system. The embodiments described herein are discussed in the form of 2D convolutions, but can be generalized to any number of dimensions. For [I] h ×I w Input (referred to as "image" in this paper, although it should be understood that the input can represent any suitable data), G and [K] h ×K w The mathematical formula for the filter F used in two-dimensional convolution is:

[0405]

[0406] Two-dimensional cross-correlation is given by the following formula:

[0407]

[0408] in It is a function of G determined by the boundary conditions. ∠ denotes complex conjugation, and · denotes scalar multiplication.

[0409] In some implementations, convolution and cross-correlation operations can be interchangeable; for example, the cross-correlation of complex two-dimensional signals G and F can be converted into convolution using the following formula:

[0410]

[0411] The embodiments described herein will focus on the convolution case; however, it should be understood that the embodiments described herein can be used to compute convolutions and cross-correlation.

[0412] Between convolution and cross-correlation, different variations exist depending on how the boundary conditions are handled. Two boundary conditions described in some embodiments of this paper include cyclic types:

[0413]

[0414] And filled type:

[0415] If (0≤x≤I) h And 0≤y≤I w ),but Otherwise, it is 0, where a%n means amod n.

[0416] According to some embodiments, additional boundary condition variations may be used. These boundary conditions include symmetric (also known as mirror or reflective) boundary conditions, where the image is reflected on the boundary. In some embodiments, filled boundary conditions may be referred to differently as linear or filled. Cyclic boundary conditions are also known as wrapping boundary conditions.

[0417] Additionally, different output modes can be used to determine which elements interact with the boundary conditions. These output modes include valid output mode, identical (or half-filled) output mode, and full output mode. Valid output mode requires the output to consist only of elements independent of the boundary conditions. Identical output mode requires the output to be the same size as the input. Full output mode requires the output to consist of all elements not entirely dependent on the boundary conditions.

[0418] Different output modes control the number of points [x, y] on which the output is defined. Therefore, each embodiment described herein can be modified to operate with any given output mode. While the embodiments described herein focus on the case of convolutions with the same mode, it should be understood that these implementations can be extended to compute cross-correlation and / or substitute output modes to replace or supplement the embodiments described herein.

[0419] In some implementations, such as in CNNs, these operations can be generalized so that they can be applied to and / or produce multichannel data. For example, an RGB image has three color channels. For an input G with two spatial dimensions and C channels, the multichannel operation is defined as:

[0420]

[0421] in This represents convolution or cross-correlation, where M is the number of output channels and G is the three-dimensional [C×I] symbol. h ×I w The tensor, F, is a four-dimensional [M×C×I] tensor. h ×I w Tensor, ( ) is three-dimensional [M×Ih ×I w Tensors. For the above, slice indexing is used, where the spatial dimension is suppressed, so that F[m,c] accesses two dimensions of F[K]. h ×K w Spatial slice, G[c] accesses two-dimensional [I] of G h ×I w Spatial slice.

[0422] Typically, the techniques used to represent convolution as matrix operations can follow... Figure 3-1 The process is as follows: In action 3-102, preprocessing of the image and / or filter matrix can occur before matrix operations to ensure, for example, that the matrix has the correct dimensions, conforms to boundary conditions, and / or output patterns. In action 3-104, core matrix or matrix-vector operations can be applied to create the convolution output. In action 3-106, postprocessing of the convolution output can occur to, for example, reshape the output, as will be discussed in more detail herein.

[0423] Some embodiments can use a photonic processor to compute convolutions as matrix-vector products. The inventors have recognized and appreciated that arrays of variable beam splitters (VBSs), such as those that can be included in some embodiments of photonic processors as previously described herein, can be used to represent any unitary matrix. For example, these techniques can be used to represent an extended image G. mat A matrix can be decomposed into singular value decomposition.

[0424] G mat =V T ∑U.

[0425] In some embodiments of the photonic processor, the two unitary matrices U and V can then be decomposed using the previously described algorithm. The phase obtained from this calculation, along with the singular values, is programmed into the photonic array. In some embodiments, the processor decomposes filters instead of images, such that filters can be kept loaded for a batch of images.

[0426] According to some embodiments, Figure 3-2 An example of the process for computing convolution in a photonic processor is shown. In action 3-202, this process constructs matrix G from the input image matrix G. mat Matrix G mat It can be constructed in any suitable way according to the selected boundary conditions, including but not limited to constructing G. mat It is a biblock cyclic matrix or a biblock Toeplitz matrix.

[0427] In action 3-204, the decomposed matrix G matThis can then be loaded into a photon array. For each filter F in the input batch, the loop is repeated, where filter F is flattened into a column vector in action 3-206, passed through the photon array in action 3-208 to perform matrix multiplication, and then reshaped into an output with the appropriate dimensions in action 3-210. In action 3-212, it is determined whether there are any other filters F. If another filter F is to be passed through the convolutional layer, the process returns to action 3-206. Otherwise, the process ends. Due to the commutative nature of convolution, process 3-200 can be performed, where filter F is expanded into F... mat Furthermore, the image G is flattened into a column vector and passed through the photon array in action 3-208.

[0428] Photonic processors can be used to implement any suitable matrix multiplication-based algorithm. Matrix multiplication-based algorithms reorder and / or expand the input signal so that the computation can be represented as a generalized matrix-matrix multiplication (GEMM) with some preprocessing and / or post-processing. Some exemplary matrix multiplication-based algorithms that can be implemented on photonic processors include im2col, kn2row, and memory-efficient convolution (MEC).

[0429] According to some embodiments, the im2col algorithm can be implemented on a photonic processor. In the im2col algorithm, during preprocessing, the image G can be derived from [I... h ×I w The matrix is ​​expanded to [(K) h ·K w )×(I h ·I w The filter F can be derived from the matrix [K]. h ×K w The matrix flattens out to [1×(K)] h ·K w Row vectors. The output can then be generated through a matrix-vector product of the image and the filter, as this preprocessing step generates an expanded data matrix where each column contains all (K) rows that can be scaled and accumulated for each position in the output. h ×K w The im2col algorithm requires O(K) copies of each element. Therefore, the im2col algorithm may require O(K) copies. h ·K w ·I h ·I w ) data replicas and O(K h ·K w ·I h ·I w Temporary storage.

[0430] According to some embodiments, the kn2row algorithm can be implemented on a photonic processor. The Kn2row algorithm calculates the outer product of the unmodified image and the filtered signal, generating a product of size [(K h ·K w )×(I h ·I w A temporary matrix is ​​generated. Then, the kn2row algorithm sums the specific elements from each row of the outer product to produce [1×(I)]. h ·I w The output vector is kn2row. Therefore, the kn2row algorithm may also require O(K) time complexity. h ·K w ·I h ·I w ) data replicas and O(K h ·K w ·I h ·I w Temporary storage.

[0431] According to some embodiments, the MEC algorithm can be implemented on a photonic processor. The MEC algorithm can simply extend the input image by K. h or K w Multiples, instead of (K) as in the im2col algorithm h ·K w If a smaller filter dimension is chosen for expansion, the algorithm only requires O(min(K) times. h K w )·I h ·I w Temporary storage and data copies. Unlike im2col or kn2row, which compute a single matrix-vector product, the MEC algorithm computes a series of smaller matrix-vector products and concatenates the results.

[0432] In the above embodiments, due to the commutative nature of convolution, the filter matrix, rather than the image, can be expanded during preprocessing. The choice of whether the image or the filter should be tiled and reshaped into a matrix can be determined by which operations are faster and / or require less computational energy.

[0433] VII. Multidimensional Convolution Using Two-Dimensional Matrix-Matrix Multiplication

[0434] The inventors have recognized and understood that the matrix multiplication-based algorithms described above for calculating convolutions may not be suitable for some computational architectures or applications. The inventors have further recognized and understood that methods combining the computational efficiency of im2col or kn2row with the memory-efficient characteristics of MEC algorithms will be beneficial for the computation of convolutions and cross-correlation. In particular, the inventors have recognized that these benefits can be achieved by partitioning the reordering and reshaping of the input and output matrices between preprocessing and postprocessing steps, and that such methods can be generalized to N-dimensional convolutions, where N≥2.

[0435] According to some embodiments, multidimensional convolution using a two-dimensional matrix-matrix multiplication algorithm (referred to as the "cng2 algorithm" in this paper) comprises three steps. At a high level, for a non-limiting example of two-dimensional circular convolution, the preprocessing step involves copying and rotating [I h ×I w [K] is constructed from the rows of the input matrix. w ×(I h ·I w ] matrix, where in some implementations "rotation" refers to the cyclic arrangement of vector elements, such as rotation In the GEMM step, calculate [K] h ×K w The filter matrix and the [K] from the preprocessing step w ×(I h ·I w The product of matrices. In post-processing, the [K] created by GEMM is used. h ×(I h ·I w The output is constructed by rotating the rows of the matrix and summing them.

[0436] According to some embodiments, the cng2 algorithm can be modified to implement other boundary conditions. For example, in the case of padding convolutions during preprocessing and postprocessing, the vector rows are shifted instead of rotated. That is, during the rotation step, the elements that would otherwise wrap the row vectors are set to zero. Other boundary conditions that can be implemented in the cng2 algorithm include, but are not limited to, symmetric or mirror boundary conditions.

[0437] Additionally, it can be noted that, according to some embodiments, the preprocessing steps of the cng2 algorithm are not limited to being applied only to left-handed inputs (images in this paper), but can be applied to right-handed inputs (filters in this paper). For convolutions in full mode or effective mode, the operations are commutative, and the preprocessing stage can be applied to either input. For convolutions of the same mode, when I... h ≠K h or I w ·K wAt this time, the operation is non-commutative, but the preprocessing stage can still be applied to the right-hand side, although the filter must first be zero-padded and / or clipped in each dimension to match the output size.

[0438] In some implementations, the cng2 algorithm may include additional steps, such as Figure 3-3A As described above. Prior to the preprocessing stage as previously described, it may be necessary to reshape one or more input filter matrices into filter matrices F with appropriate dimensions, as shown in action 3-302. This reshaping can be accomplished by concatenating one or more input filter matrices in any suitable manner. In action 3-304, the preprocessing step of constructing the cyclic matrix H is performed, and reference is made to... Figure 3-3B A more detailed description follows. In action 3-306, the GEMM step is performed and an intermediate matrix X = F × H is created. Next, post-processing steps can be performed. In action 3-308, the vector rows of matrix X are rotated and / or shifted to form matrix X′. In action 3-310, vector row addition is performed on the rows of matrix X′ to form matrix Z. Depending on the memory layout of the specific processing system, matrix Z can be reshaped into at least one output matrix in action 3-312.

[0439] According to some embodiments, the method for constructing matrix H can depend on the desired boundary conditions, such as... Figure 3-3B The extension of action 3-304 is shown. In action 3-314, it can be determined whether the boundary condition is cyclic. If the boundary condition is determined to be cyclic, the processing system can proceed to action 3-316, where matrix H is created by copying and rotating at least one row of the input matrix. Conversely, if the boundary condition is determined to be filled rather than cyclic in action 3-314, the processing system can proceed to action 3-318. In action 3-318, matrix H is created by copying and shifting at least one row of the input matrix, as previously discussed. It should be understood that in some embodiments, boundary conditions other than filled boundary conditions may be used in action 3-318.

[0440] Alternatively, according to some embodiments, when calculating cross-correlation, it is not necessary to explicitly convert the problem into convolution as in procedures 3-300. Instead, the element-inverting step 3-302 can be omitted, and the preprocessing and postprocessing steps of the cng2 algorithm can be modified accordingly. That is, the element-inverting step can be combined with the preprocessing and postprocessing steps of the cng2 algorithm. How this is done depends on whether the preprocessing expansion is applied to the left-handed or right-handed input. If the left-handed input is expanded, the shifts or rotations in both the preprocessing and postprocessing steps can be performed in opposite directions. If the right-handed input is expanded, each cyclic matrix generated during the preprocessing stage can be transposed and concatenated in reverse order, and the i-th row of the GEMM output matrix can be shifted or rotated by (i-n+1)n elements instead of in elements in the postprocessing stage. For complex numerical data, cross-correlation still requires the complex conjugate of an input.

[0441] In some implementations, such as in CNNs, it may be desirable to generalize the above operations so that they can be applied to and / or generate multi-channel data. For a problem with C input channels and M output channels, the filter matrix takes the form [(M·K h )×(K w The input matrix takes the form [(K)], where C)] is the input matrix. w ·C)×(I h ·I w The output matrix takes the form [(M·K)]. h )×(I h ·I w )]form.

[0442] Reference Figures 3-4A to 3-4F According to some embodiments, an example of process 3-300 for multi-channel input and output is shown. For a [2×2] filter f including 4 output channels and a [3×3] image G including 3 input channels, action 3-402 is performed in... Figure 3-4A The diagram is shown below. However, filter matrices of arbitrary size, image matrices, and / or any number of input and / or output channels can be implemented. In the example of action 3-402, the filter f is reshaped into a [6×8] filter matrix F. In this example, the reshaping of filter f is accomplished by concatenating filter f without otherwise altering the order of matrix elements, although other methods, such as rotation, shifting, or otherwise changing the rows of filter f, can be used. The reshaping of filter f ensures that the filter matrix F has the appropriate dimensions for later GEMM operations. However, in some implementations where the memory is properly laid out, action 3-402 may not be necessary before action 3-404, as described below.

[0443] According to some embodiments, after action 3-402, preprocessing of image G can be performed in action 3-404, such as... Figure 3-4B As shown. In this example, image G is formed to perform convolution in the same pattern, although any output pattern can be used. In this example, the cyclic matrix H is formed based on cyclic boundary conditions, although any boundary conditions such as padding boundary conditions can be used. After action 3-404, the GEMM operation F×H=X of action 3-406 can be performed, as shown. Figure 3-4C As shown. In this example, the dimensions of the intermediate matrix X are [9×8].

[0444] Post-processing steps can occur after the GEMM operation, such as... Figures 3-4D to 3-4E As shown. In Figure 3-4D The example describes action 3-408, in which the rows of the intermediate matrix X are rotated to form matrix X′ according to the cyclic boundary conditions of this example. Other boundary conditions, such as the filling boundary conditions as a non-restrictive example, can be implemented in preprocessing and postprocessing, provided that the boundary conditions of the preprocessing and postprocessing steps are the same as each other.

[0445] Reference Figure 3-4E The description of actions 3-410 indicates that the next step in post-processing is to perform row operations on matrix X′ to form the output matrix Z. That is, in this example, x... 00 +x 16 =z 00 x 01 +x 17 =z 01 And so on. In some implementations, depending on the memory layout of the processing system, the reshaping of the output matrix Z may have to be performed in action 3-412, following action 3-410. Figure 3-4F In the exemplary illustration, the output matrix Z is reshaped into four output matrices A of [3×3] dimensions.

[0446] According to some embodiments, in addition to being generalizable to multiple input channels, the cng2 algorithm can also be extended to higher-dimensional signals (i.e., greater than two dimensions). For signals of size [K] n ×K n-1 The filter tensor and size of [×...×K1] are [I n ×I n-1 The n-dimensional convolution between images of [×...×I1] can be computed using two-dimensional matrix multiplication with steps similar to those used for two-dimensional signals to compute the desired output. During preprocessing, the input tensor can be expanded (K... a ·K a-1... K1) times, where a can be considered as the number of dimensions processed during the preprocessing stage, and can be any value in the range 1 ≤ a ≤ n-1. In the GEMM step, the data can be divided into [(K... n ·K n-1 ·...·K a+1 )×(K a ·K a-1 The filter tensor of the matrix […·K1)] is multiplied by the extended matrix from the preprocessing step. During the postprocessing step, subvectors of the matrix generated during GEMM can be rotated and accumulated.

[0447] The extended matrix generated by the preprocessing stage can be obtained from (I) i ·I n-1 ·...·I a+1 The matrix is ​​composed of horizontally concatenated submatrices, each of which is a nested Toeplitz matrix of degree a, and the innermost Toeplitz matrix is ​​defined as follows in the 2D cng2 implementation. The post-processing stage can perform (na) loops of rotation and addition, where the i-th loop partitions the matrix generated by the previous loop (or initially by GEMM operations) into submatrices of size [K]. a+i ×(I a+i ·I a+i-1 ... ... I1)] submatrix. Then, for each submatrix, perform the following operations. First, the j-th row can be rotated or shifted up to (j·(I a+i-1 ·I a+i-2 ...I1)) elements. Then, all rows can be added together.

[0448] While the above description treats dimensions sequentially—that is, the preprocessing stage expands the data along the first *a* dimensions and the postprocessing stage reduces the data along the last *na* dimensions—this is not necessarily the case in some embodiments. The preprocessing stage can expand the data along any *a* dimensions by reordering the input and output data in the same manner as described for the two-dimensional case.

[0449] The cng2 algorithm provides a flexible architecture for computing convolutions, and this paper describes several alternative implementations. In some implementations, the overlapping regions of the input signals at a given point in the output can be shifted with a constant offset. Such an offset can be applied regardless of the output mode, but it is most commonly paired with outputs of the same mode. For convolutions (cross-correlation) operating under the same mode and the definitions given above, boundary conditions can be applied along the top (bottom) edges of the input image G (K... h -1)·I w The elements and (K) along the left (right) edge of the input image G w -1)·Ih This behavior can be altered by redefining the operation with a constant offset between the filter and output positions. When calculating convolution (cross-correlation), this can be achieved by subtracting (adding) the shift or rotation offset in the preprocessing stage and by subtracting (adding) the shift or rotation offset in the postprocessing stage. w This will allow us to apply this modification to cng2.

[0450] Furthermore, according to some embodiments, the proposed methods for reducing both the time and storage requirements of the kn2row post-processing steps can be similarly applied to the cng2 algorithm. For the kn2row algorithm, the GEMM operation can be decomposed into a series of K... w ·K h This involves several smaller GEMM operations, the results of which are sequentially accumulated. This allows post-processing additions to be performed in an optimized manner and at appropriate locations relative to the storage of the final output. In the case of kn2row, this only works if boundary conditions can be ignored or if an additional (and often inefficient) so-called punching process is introduced. However, in the case of the cng2 algorithm, this process can be applied directly without sacrificing accuracy or additional processing, effectively eliminating the computational cost of the post-processing step and reducing the temporary storage required for the cng2 algorithm to O(K). w ·I h ·I w ).

[0451] In some embodiments, the spatial dimension can be processed in the reverse order of processes 3-300. The cng2 algorithm can be enhanced by applying a transpose operation to both input signals at the beginning of process 3-300 and to the final output. When the filter shape is strictly rectangular (i.e., K... w ≠K h This still produces the desired result, but the behavior changes. In this case, the input image is expanded by K. h times instead of K w The post-processing steps include O(K) times, and the post-processing steps include O(K) times. w ·I h ·I w ) addition instead of O(K) h ·I h ·I w The implementation combining this variant with the aforementioned low-memory integrated post-processing variant can further reduce the temporary storage required for the cng2 algorithm to O(min(K)). h K w )·I h ·I w ).

[0452] As an alternative implementation, rows and / or columns in the matrix passed to the GEMM operation can be reordered. If the GEMM operation is defined as C = AB, then rows and / or columns of the input matrix A or B can be reordered, provided that an appropriate permutation is applied to the other input matrix (in the case of reordering the columns of A or the rows of B) or to the output matrix (in the case of reordering the rows of A or the columns of B). In particular, in the case of multiple output channels, reordering the rows of A can reorganize the data-level parallelism available in the post-processing stage in a manner well-suited to vector processors or Single Instruction Multiple Data (SIMD) architectures.

[0453] According to some embodiments, convolution computation can also be performed using stride. For the stride S in the first dimension... x Step size S in the second dimension y The convolution operation is defined as follows:

[0454]

[0455] This definition reduces the magnitude of the output signal by S. x ·S y This is equivalent to increasing the stride of the filter across the image by a factor of 1, which is the same as increasing the stride size for each output point. This can be achieved by computed non-staggered convolutions and then downsampling the results in each dimension by an appropriate amount, but this requires more O(S) than needed. x ·S y ) Calculation steps. At least, it can be achieved by modifying preprocessing stage 3-304 to generate only each S-th cyclic matrix. x The column and the post-processing stage are modified to shift or rotate each row by i·(I) w / S x ) instead of i·I w This reduces the computational penalty to O(S) in cng2. y In some implementations, the computational penalty can be completely eliminated by making additional modifications to each stage. First, preprocessing steps 3-304 can be modified to produce S. y Instead of a single matrix, an extended matrix is ​​used, where if j = i mod S, the i-th cyclic matrix is ​​assigned to the j-th extended matrix. Then, the core processing phase must execute S. y Each GEMM operation (one GEMM operation for each expanded input matrix) uses only K of the filter matrix in each GEMM operation. w / S y Okay. Then, post-processing steps 3-308 and 3-310 can interleave the rows of the obtained matrix, and interleave each group of S yThe rows are added directly (i.e., without shifting or rotating the rows), and K is added. w / S y The row operation is performed through standard post-processing logic, and its shift or rotation amount is i·(I w / S x ).

[0456] Alternatively, according to some embodiments, the convolution can be dilated. For the dilation D in the first dimension... x Expansion D in the second dimension y The convolution operation is defined as follows:

[0457]

[0458] For each output point, dilation increases the filter's receptive field over a larger image patch and can be viewed as the insertion space between filter elements. The cng2 algorithm can be modified by increasing the rotation or shift amounts in the preprocessing and postprocessing stages, respectively. x and D y This allows us to implement dilated convolutions. Dilated convolutions can be further restricted to computation using causal output patterns.

[0459] The inventors have further recognized and understood that convolution and cross-correlation can be computed using transform-based algorithms. Transform-based algorithms alter the nature of the computational problem by first computing an equivalent representation of the input signal in an alternative digital domain (e.g., the frequency domain), performing alternative linear operations (e.g., element-wise multiplication), and then computing the inverse transform of the result to return to the original digital domain of the signal (e.g., the time domain). Examples of these transforms include the Discrete Fourier Transform (DFT), Discrete Sine Transform, Discrete Cosine Transform, Discrete Hartley Transform, Undecoded Discrete Wavelet Transform, Walsh-Hadamard Transform, Hankel Transform, and Finite Impulse Response (FIR) filters, such as Winograd's minimum filtering algorithm. This document will describe examples of DFT-based transform-based algorithms, but any suitable transform can be implemented on transform-based algorithms and photonic processors.

[0460] For unitary normalization, the Discrete Fourier Transform (DFT) of a one-dimensional signal is calculated as follows:

[0461]

[0462] The inverse of the transformation It can be calculated by taking the complex conjugate. Similarly, in two dimensions, the unitarily normalized DFT can be calculated as follows:

[0463]

[0464] Performing the one-dimensional DFT as defined above on a vector of size N can be done by computing the matrix-vector product. This is achieved by... The matrix W is called the transformation matrix and is given by the following equation:

[0465]

[0466] in The inverse transformation can be computed using a similar matrix-vector product, where W -1 The elements are complex conjugates. The DFT is a separable transformation, therefore it can be viewed as computing two one-dimensional transformations along orthogonal axes. Therefore, the two-dimensional DFT of a [M×N] (i.e., rectangular) input x can be computed using the following matrix triple product:

[0467]

[0468] Where W is the [M×M] transformation matrix associated with the columns, Y is the [N×N] transformation matrix associated with the rows, and the superscript T indicates the matrix transpose. For a square input x of size [N×N], the transformation matrix of column W is the same as the transformation matrix of row Y.

[0469] Equivalently, this can be achieved by first flattening x row-wise into a column vector x of size M·N. col And calculate the following matrix-vector product to compute:

[0470]

[0471] in It is the Kronecker product. According to some embodiments, the resulting vector X can then be... col Reorganized into a [M×N] two-dimensional array X:

[0472]

[0473] A similar process can be performed on other discrete transformations, wherein the forward transformation matrix W and the matrix W associated with the inverse transformation are defined in any suitable manner according to said other transformation. -1 .

[0474] In the case of a one-dimensional DFT, matrix W is a unitary matrix and can therefore be directly programmed into the photon array according to the previously described embodiments. For other discrete transformations, matrix W may not be a unitary matrix and therefore needs to be decomposed before being programmed into the photon array according to the previously described method. Figure 3-5 The procedure 3-500 for performing a one-dimensional transformation on a vector is illustrated according to some embodiments. In action 3-502, the matrix W is programmed into the photon array, and in action 3-504, the vector x is propagated through the array.

[0475] In some implementations, it is possible to... Figure 3-6 The two-dimensional transformation is calculated as described in procedure 3-600. For a two-dimensional input x of size N×N, procedure 3-500 can be modified to produce a two-dimensional transformation as defined above. In action 3-602, a transformation matrix W corresponding to a one-dimensional input of size N is created. small Next, in action 3-604, create a structure of size N by tiling W along the diagonal. 2 ×N 2 The block diagonal matrix B. Then, in action 3-606, a block diagonal matrix of size N is created by flattening the input x row by row. 2 column vector x col In action 3-608, multiplication X can be performed by propagating vector x through the photon array. partial =Bx. Then, in action 3-610, matrix X can be... partial Reorganized into an N×N matrix, which can be further flattened column-wise into a size of N in action 3-612. 2 The column vector. In action 3-614, X can be propagated... partial Multiplication X = BX is performed using a photon array. partial In action 3-616, the obtained vector X can be reshaped into an N×N matrix for output.

[0476] In some embodiments, the two-dimensional transformation of the [N×N] input x can then be computed by first programming the matrix W into the photon array. Next, X is computed by propagating the columns of x through the photon array. partial =Wx. Again, partial result X partial Transpose. Then, propagate X a second time. partial T The columns are calculated using an array to determine WX. partial T Finally, the result is transposed to produce X = WxW T .

[0477] Some systems (such as one embodiment of a photon-based system described herein) are limited to implementing real unitary matrices (i.e., orthogonal matrices). In such implementations, transformations can still be computed, but with additional steps. The system must keep track of the real and imaginary parts of the transformation matrix and the input vector or image separately. The embodiment defined above for computing the product can be applied to orthogonal matrices, except that the algorithm must be performed four times for each pass through the photon array as described above. Representing the real part of the variables as Re(x) and the imaginary part as Im(x), the real part of the product is Re(Wx) = Re(W)Re(x) - Im(W)Im(x), and similarly, the imaginary part of the product is Im(Wx) = Re(W)Im(x) + Im(W)Re(x). According to some embodiments, in the optical core of a photon processor that only represents real matrices, execution can be performed... Figure 3-7 Step 3-700. Depending on the input dimension, step 3-500 or 3-600 can be used in step 3-700. In step 3-702, Re(W) is loaded into the photon array. In step 3-704, depending on the input dimension, step 3-500 or 3-600 can be performed for Re(x) and Im(x). This produces Re(W)Re(x) and Re(W)Im(x). In step 3-706, Im(W) is loaded into the photon array. In step 3-708, again depending on the input dimension, step 3-500 or 3-600 can be performed for Re(x) and Im(x). This produces Im(W)Re(x) and Im(W)Im(x). In step 3-710, Re(W)Re(x) and Im(W)Im(x) are subtracted to produce Re(Wx). Furthermore, in action 3-712, Re(W)Im(x) and Im(W)Re(x) are added together to produce Im(Wx).

[0478] Using the processes described above (3-500, 3-600, and 3-700), the input matrix can be transformed into its transform. According to some embodiments, once the convolution filter F and the image G are transformed into their transform counterparts, the convolution theorem can be applied. The convolution theorem states that the convolution of two signals corresponds to the transform of the element-wise product of the transforms of the two signals. Mathematically, this can be expressed as:

[0479]

[0480] Or, equivalently,

[0481]

[0482] Where ⊙ represents element-wise multiplication. This represents the inverse transform. In some embodiments, the dimensions of the image and the filter may be different; in such cases, it should be understood that each of the forward and inverse transforms can be computed using an appropriate dimensional transformation matrix. Therefore, the matrix-multiplication equation for one-dimensional convolution, expressed in terms of the general transforms and general dimensions of the filter and image, is:

[0483]

[0484] Among them W B W is a matrix associated with the transform of filter F. D T W is a matrix associated with the transformation of image G. A T It is a matrix associated with the inverse transform of the combined signal.

[0485] Similarly, the matrix-multiplication equation for two-dimensional convolution expressed by general transformations on rectangular filters and images is:

[0486] G*F=W A T ((W B FW C T )⊙(W D T GW E ))W F ,

[0487] Among them W B and W C T W is a matrix associated with the transform of filter F. D T and W E W is a matrix associated with the transformation of image G. A T and W F It is a matrix associated with the inverse transform of the combined signal.

[0488] Reference Figure 3-8 Some embodiments of performing convolution in the optical core of a photonic processor using a transformation-based algorithm may include the following actions. In action 3-802, the transformation of image G may be performed using any of processes 3-500, 3-600, and / or 3-700.

[0489] In some embodiments, the filter F can then be zero-padded in action 3-804 to match the size of the image G, after which a transformation can be performed on the filter F in action 3-806 using any of processes 3-500, 3-600, and / or 3-700. In action 3-808, the transformed filter F can then be loaded into the element-wise multiplier of the photon array, and in action 3-810, the image G can be propagated through the photon array. In action 3-812, an inverse transformation can be performed on the previously calculated result using any of processes 3-500, 3-600, and / or 3-700. The result of action 3-812 can then be reshaped to size G and cropped in action 3-814 to produce the final convolutional image G*F.

[0490] In some embodiments, the convolution G*F can be computed in a divide-and-conquer manner, where an input is divided into a set of slices, and each slice is convolved with a second input. The results of each individual convolution can then be recombined into the desired output; however, the algorithms used to implement this divide-and-conquer approach (e.g., overlapping addition, overlapping preservation) are non-trivial. This method can be far more efficient than computed in a single operation as described above, where filters are padded to match the image size, when one input is significantly smaller than the other and a transform-based algorithm is used for the convolution operation. It is understood that such a divide-and-conquer algorithm for transform-based convolution can be implemented on a photonic processor by performing transformations on slices on a photonic array.

[0491] In some embodiments, the filter F and the image G can have multiple channels. As defined above, this means that each channel of the image is convolved with the corresponding channel of the filter tensor, and the results are summed element-wise. When multi-channel convolution is computed using a transform-based approach, the summation over the channels can be performed in either the transform domain or the output domain. In practice, it is usually chosen to perform the summation in the transform domain because this reduces the amount of data that must be subjected to the output transform. In this case, the element-wise multiplication followed by the channel-wise addition can be represented as a sequence of matrix-matrix multiplications (GEMM). Mathematically, this can be represented as follows:

[0492] Let G be the input signal of an N×N image with C data channels. Let F be the input signal of an N×N filter with M·C data channels. Let C and M be the number of input and output data channels, respectively. Let Q... m,c The transform data of the m-th output channel and c-th input channel of the filter tensor (i.e., Let R be the transformed three-dimensional [C×N×N] input tensor, R c For the c-th channel of the transformed input tensor (i.e., R),c =W D T G c W E Then, the convolutions of F and G that produce multiple output channels are:

[0493]

[0494] If S ij Let S represent a column vector consisting of the C elements at position (i, j) of each channel in a 3D [C×N×N] tensor S. This can be equivalently represented as:

[0495]

[0496] Each Q can be calculated on the photonic processor described above. m ij R ij Matrix-matrix multiplication. This can be further combined with the divide-and-conquer method described above.

[0497] Various aspects of this application provide methods, programs, and algorithms that can be executed on a processing device, such as a CPU, GPU, ASIC, FPGA, or any other suitable processor. For example, the processing device can perform the above-described processes to generate settings for a variable beam splitter and modulator for the optical core of the photonic processor described herein. The processing device can also perform the above-described processes to generate input data to be input into the photonic processor described herein.

[0498] An exemplary implementation of a computing device may include at least one processor and a non-transitory computer-readable storage medium. The computing device may be, for example, a desktop or laptop computer, a personal digital assistant (PDA), a smartphone, a tablet computer, a server, or any other suitable computing device. The computer-readable medium may be adapted to store data to be processed by the processor and / or instructions to be executed by the processor. The processor is capable of processing data and executing instructions. Data and instructions may be stored on the computer-readable storage medium and may, for example, enable communication between components of the computing device. The data and instructions stored on the computer-readable storage medium may include computer-executable instructions that implement techniques operating according to the principles described herein.

[0499] A computing device may additionally include one or more components and peripherals, including input and output devices. These devices are particularly useful for presenting a user interface. Examples of output devices that can be used to provide a user interface include a printer or display screen for visual presentation and a speaker or other sound-generating device for audible presentation. Examples of input devices that can be used for a user interface include keyboards and pointing devices such as mice, touchpads, and digitizers. As another example, a computing device may receive input information via voice recognition or in other audible formats. As yet another example, a computing device may receive input from a camera, lidar, or other device that generates visual data.

[0500] Embodiments of the computing device may also include a photonic processor, such as the photonic processor described herein. The processor of the computing device may send and receive information to and from the photonic processor via one or more interfaces. The information sent and received may include settings of the photonic processor's variable beam splitter and modulator and / or measurement results from the photonic processor's detectors.

[0501] IV. Photonic Encoder

[0502] The inventors have recognized and understood the phase-intensity relationship, in which the phase and intensity modulation of an optical signal are interdependent, posing a challenge to accurately encoding vectors in an optical field for optical processing.

[0503] Some optical modulators encode digital vectors in the optical domain, where numbers are encoded into the phase or intensity of an optical signal. In an ideal world, an intensity modulator can modulate the intensity of an optical signal while keeping its phase constant, and a phase modulator can modulate its phase while keeping its intensity constant. However, in the real world, the phase and intensity of an optical signal are interdependent. Thus, intensity modulation induces corresponding phase modulation, and phase modulation induces corresponding intensity modulation.

[0504] Consider, for example, an integrated photonic platform where intensity and phase are correlated via the Kramers-Kronig equation (a two-way mathematical relation connecting the real and imaginary parts of a complex analytic function). In these cases, the optical field can be represented by a complex function, where the real and imaginary parts of the function (and consequently, intensity and phase) are correlated. Furthermore, practical implementations of intensity and phase modulators typically suffer from dynamic losses, whereby the amount of modulation (phase or intensity) they impart depends on their current settings. These modulators experience a certain power loss when no phase modulation occurs, and different power losses when phase modulation occurs. For example, the power loss experienced without phase modulation might be L1, the power loss experienced with π / 2 phase modulation might be L2, and the power loss experienced with π phase modulation might be L3, where L1, L2, and L3 are different from each other. This behavior is undesirable because the signal undergoes intensity modulation in addition to phase modulation.

[0505] exist Figure 4-1 An example of an optical modulator (4-10) is depicted, in which ψ in Represents the input light field, ψ out Let ψ represent the optical field output by modulator 4-10, where α represents the amplitude modulation factor and θ represents the phase modulation factor. Ideally, modulator 4-10 will be able to incorporate any attenuation to modulate the optical field ψ. in The term is set to any phase, such as ψ out =αe iθ Where α ≤ 1 and θ ∈ [0, 2π]. Because modulators typically have associated Kramers-Kronig relationships and / or suffer from dynamic losses, simply setting the values ​​of α and θ is not readily feasible. Conventionally, two or more modulators with two independently controllable modulated signals are used to encode both the amplitude and phase of the optical field in a manner representing real-valued signed numbers. However, having multiple controllable modulated signals requires complex multivariate coding schemes, typically involving feedback loops for simultaneously controlling the phase and intensity of the optical signals.

[0506] Optical modulators exist that can modulate intensity without affecting phase, including modulators based on electro-optic materials. Unfortunately, electro-optic modulators are challenging to operate in large-scale commercial environments because they involve the use of impractical materials for fabrication.

[0507] Conventional computers do not perform complex calculations themselves. Instead, they use signed arithmetic, where numbers can be positive or negative, and complex number operations can be constructed by performing several calculations and combining the results. The inventors have recognized and understood that this fact can be used to greatly simplify the implementation of real number signed linear transformations, even in the presence of dynamic losses and / or non-ideal phase or intensity modulators.

[0508] Some embodiments relate to techniques for encoding signed real-number vectors using non-ideal optical modulators (modulators in which intensity modulation causes phase modulation and phase modulation causes intensity modulation). In contrast to other implementations, some such embodiments relate to using a single modulation signal to control both the phase and intensity of the optical signal. The inventors have recognized that when encoding signed real numbers into an optical field using a single modulation signal, the precise location of the optical signal in phase space (e.g., real and imaginary parts, or amplitude and phase) is not critical for decoding purposes. According to some embodiments, the problem for accurate decoding is the projection of the optical signal onto the real axis (or any other pre-selected axis as the measurement axis). To perform projection onto the pre-selected axis, coherent detection schemes are used in some embodiments.

[0509] The optical domain coding techniques described herein can be used in a variety of contexts, including but not limited to high-speed communication for short-range, medium-range, and long-range applications, on-chip phase-sensitive measurements for sensing, communication, and computation, and optical machine learning using photonic processors.

[0510] More generally, the encoding techniques of the type described herein can be used in any context where optical signals are processed according to real number transformations (the opposite of complex number transformations). Conventional computers do not directly perform complex number operations. More precisely, conventional computers use signed real number operations, where the numbers can be positive or negative. In some embodiments, complex number operations can be established in the optical domain by performing several calculations involving real number operations. For example, consider an optical linear system configured to transform an optical signal according to the following expression:

[0511] y=Mx=[Re(M)+iIm(M)][Re(x)+iIm(x)]=[Re(M)Re(x)-Im(M)Im(x)]+i[Re(m)Im(x)+Im(M)(Re(x)]

[0512] Where x represents the input vector, M represents the transformation matrix, y represents the output vector, and i represents a defined imaginary number such that i 2 = -1. Considering the real number transformation (making Im(M) = 0), y can be rewritten as follows:

[0513]

[0514] It should be noted that, based on this equation, the real and imaginary parts of the input field x contribute only to the linear independent components of the resulting field y. The inventors have recognized and understood that y can be decoded using a coherent receiver by projecting y onto any axis in the complex plane.

[0515] In some embodiments, the encoding techniques of the type described herein can be used as part of a photonic processing system. For example, in some embodiments, Figure 1-1 Optical encoder 1-101 and / or Figure 1-11 Optical encoder 1-1103 and / or Figure 1-12B The optical encoder 1-1211 can implement the type of encoding technology described in this article.

[0516] Figure 4-2A This is a block diagram of a photonic system implementing optical coding technology according to some embodiments. The photonic system 4-100 includes a light source 4-102, an encoder 4-104, an optical modulator 4-106, an optical conversion unit 4-108, a coherent receiver 4-110, a local oscillator 4-112, and a decoder 4-114. In some embodiments, the photonic system 4-100 may include... Figure 4-2A Other or alternative components not shown. In some embodiments, Figure 4-2A Some or all of the components can be disposed on the same semiconductor substrate (e.g., silicon substrate).

[0517] The light source 4-102 can be implemented in various ways, including, for example, using an optically coherent source. In one example, the light source 4-102 includes a laser configured to emit light with a wavelength of λ0. The emission wavelength can be in the visible infrared (including near-infrared, mid-infrared, and far-infrared) or ultraviolet portion of the electromagnetic spectrum. In some embodiments, λ0 can be in the O-band, C-band, or L-band. An optical modulator 4-106 is used to modulate the light emitted by the light source 4-102. The optical modulator 4-106 is a non-ideal modulator, such that phase modulation causes intensity modulation and intensity modulation causes phase modulation. In some embodiments, the phase modulation can be related to the intensity modulation according to the Kramers-Kronig equation. Alternatively or additionally, the modulator 4-106 may suffer from dynamic losses, where phase shifts cause attenuation.

[0518] In some embodiments, modulators 4-106 are driven by a single electrical modulation signal 4-105. Therefore, a single modulation signal modulates both the phase and amplitude of the optical field. This contrasts with modulators conventionally used in optical communication to encode symbols with the complex amplitude of the optical field, where each symbol represents more than one bit. In practice, in this type of modulator, multiple modulation signals modulate the optical field. Consider, for example, an optical modulator configured to provide a quadrature phase-shift keying (QPSK) modulation scheme. In these types of modulators, one modulation signal modulates the real part of the optical field, and a modulation field modulates the imaginary part of the optical field. Therefore, in a QPSK modulator, two modulation signals are used together to modulate both the phase and amplitude of the optical field.

[0519] Examples of modulators that can be used in modulators 4-106 include Mach Zehnder modulators, electro-optic modulators, ring or disk modulators or other types of resonant modulators, electroabsorption modulators, Frank-Keldysh modulators, acousto-optic modulators, Stark effect modulators, magneto-optic modulators, thermo-optic modulators, liquid crystal modulators, quantum-confined optical modulators and photonic crystal modulators, as well as other possible types of modulators.

[0520] Encoder 4-104 generates modulation signal 4-105 based on the real number to be encoded. For example, in some embodiments, encoder 4-104 may include a table that maps the real number to the amplitude of the modulation signal.

[0521] Optical conversion unit 4-108 can be configured to transform the intensity and / or phase of a received optical field. For example, optical conversion unit 4-108 may include an optical fiber, optical waveguide, optical attenuator, optical amplifier (e.g., an erbium-doped fiber amplifier), beam splitter, beam combiner, modulator (e.g., an electro-optic modulator, Franz-Keldysh modulator, resonant modulator, or Mach-Zehnder modulator, etc.), phase shifter (e.g., a thermal phase shifter or phase shifter based on plasmon dispersion), optical resonator, laser, or any suitable combination thereof. In the context of optical communication, optical conversion unit 4-108 may include an optical fiber communication channel. The optical communication channel may include, for example, an optical fiber, and optionally include an optical repeater (e.g., an erbium-doped fiber amplifier). In the context of optical processing, optical conversion unit 4-108 may include a photonics processing unit, examples of which are discussed in further detail below. In some embodiments, optical conversion unit 4-108 performs a real-number transformation such that the imaginary part of the transformation is substantially equal to zero.

[0522] The optical field output from the optical conversion unit 4-108 is provided directly or indirectly (e.g., after passing through one or more other photonic components) to the coherent receiver 4-110. The coherent receiver 4-110 may include a homodyne optical receiver or a heterodyne optical receiver. A reference signal for beating the received signal can be provided by a local oscillator 4-112, such as... Figure 4-2A As shown, or it can be provided together with the modulated optical field after passing through the optical conversion unit 4-108. The decoder 4-114 can be arranged to extract real numbers (or real number vectors) from the signal output by the coherent receiver 4-112.

[0523] exist Figure 1-1 In at least some embodiments of the photonic processing system 1-100 that implements the coding technology of the type described herein, the optical modulator 4-106 may be part of the optical encoder 1-101, the optical conversion unit 4-108 may be part of the photonic processor 1-103, and the coherent receiver 4-110 may be part of the optical receiver 1-105.

[0524] Figure 4-2B This is a flowchart illustrating a method for processing real numbers in the optical domain according to some embodiments. Method 4-150 can be used... Figure 4-2A The system or any other suitable system may be used to perform this. Method 4-150 begins with action 4-152, where a value representing a real number is provided. In some embodiments, the real number may be signed (i.e., it may be positive or negative). Real numbers may represent specific environmental variables or parameters, such as physical conditions (e.g., temperature, pressure, etc.), information associated with an object (e.g., position, motion, velocity, rate of rotation, acceleration, etc.), information associated with a multimedia file (e.g., sound intensity of an audio file, pixel color and / or intensity of an image or video file), information associated with a specific chemical / organic element or compound (e.g., concentration), information associated with a financial asset (e.g., the price of a security), or any other suitable type of information including information derived from the examples above. Information represented by signed real numbers may be useful for a variety of reasons, including, for example, training machine learning algorithms, performing predictions, data analysis, troubleshooting, or simply collecting data for future use.

[0525] In action 4-154, it is indicated that the value of the real number can be encoded onto the optical field. In some embodiments, encoding this value onto the optical field involves modulating the phase and intensity of the optical field based on the value. As a result, the phase and amplitude of the optical field reflect the encoded value. In some embodiments, action 4-154 can be performed using encoder 4-104 and modulator 4-106 (see [link to relevant documentation]). Figure 4-2AIn some such embodiments, modulating the phase and amplitude based on this value involves driving a single modulator with a single electrical modulation signal. Therefore, the single modulation signal modulates both the phase and amplitude of the optical field. (Return to reference) Figure 4-2A As an example and not a limitation, encoder 4-104 can use a single modulation signal 4-105 to drive optical modulator 4-106.

[0526] It should be noted that the fact that a single modulation signal is used to drive the modulator does not preclude the use of other control signals to control the modulator's operating environment. For example, one or more control signals can be used to control the temperature of the modulator or the temperature of a specific part of the modulator. One or more control signals can be used to power the operation of the modulator. One or more control signals can be used to bias the modulator in a certain operating manner, such as biasing the modulator in its linear region or setting the wavelength of the modulator to match the wavelength of the light source (e.g., Figure 4-2A (λ0).

[0527] According to some embodiments, in Figure 4-3A A specific type of modulator is shown. Modulator 4-206 is a ring resonant modulator. It should be noted that this document describes ring modulators only by way of example, as any other suitable type of modulator, including those listed above, may be used alternatively. Modulator 4-206 includes waveguide 4-208, ring 4-210, and phase shifter 4-212. Ring 4-210 exhibits a resonant wavelength, the value of which depends in particular on the length of the ring's circumference. When a light field ψ is emitted into waveguide 4-208... in At that time, according to ψ relative to the resonant wavelength of ring 4-210 in The wavelength of the light field can be coupled to ring 4-210 or not. For example, if ψ in If the wavelength matches the resonant wavelength, then ψ in At least a portion of the energy is transferred to ring 4-210 via ephemeris coupling. The electric current transferred to the ring will oscillate indefinitely within the ring until it is completely dispersed or otherwise dissipated. Conversely, if ψ in The wavelength does not match the resonant wavelength, ψ in It can travel directly through waveguide 4-208 without any significant attenuation.

[0528] Therefore, the light field can be determined according to ψ in The wavelength varies relative to the resonant wavelength. More specifically, the phase and intensity of the light field can be determined according to ψ. in The wavelength varies relative to the resonant wavelength. Figure 4-3B This shows how intensity (top image) and phase (bottom image) can be determined based on ψ. in The graph shows the change in wavelength relative to the resonant wavelength. The top graph shows the change as ψin The intensity spectral response α is a function of all possible wavelengths (λ). The bottom figure shows the phase spectral response θ as a function of the same wavelength. At the resonant frequency, the intensity response decreases, while the phase response shows an inflection point. Wavelengths significantly below the resonant frequency experience low intensity decay (α ~ 1) and low phase change (θ ~ 0). Wavelengths significantly above the resonant frequency experience low intensity decay (α ~ 1) and sign change (θ ~ -π). At the resonant wavelength, ψ in The intensity decays to a value corresponding to the decrease, and undergoes a phase shift equal to the inflection point value (i.e., -π / 2). Figure 4-3B In the example, ψ in The wavelength (λ0) is slightly offset relative to the resonant frequency. In this case, ψ in The intensity decreased by α V0 (where 0 < α) V0 <1), and the phase shifted by θ V0 (where -π < θ) V0 <-π / 2).

[0529] Return to reference Figure 4-3A ψ can be performed using voltage V. in The amplitude and phase modulation, in this case, the voltage V reflects Figure 4-2A The modulation signal is 4-105. When a voltage V is applied to the phase shifter 4-212, the refractive index of the phase shifter changes, and thus the effective length of the ring's circumference. This, in turn, causes a change in the ring's resonant frequency. The extent of the change in the effective circumference (and the corresponding resonant frequency) depends on the amplitude of V.

[0530] exist Figure 4-3B In the example, the voltage V is set to 0. Figure 4-3C In the example, the voltage is set to V1. For example... Figure 4-3C As shown, applying V1 to phase shifter 4-212 causes the modulator's intensity and phase spectral responses to shift along the wavelength axis. In this case, the response exhibits a redshift (i.e., a shift towards larger wavelengths). As a result, ψ in The intensity now decreases by α V1 (α V1 Unlike α V0 And the phase shift reaches θ v1 (θ V1 Unlike θ V0 ).

[0531] Therefore, the changing voltage V causes ψ in The voltage V changes in intensity and phase. In other words, V can be considered a modulation signal. In some embodiments, the voltage V can be the only modulation signal driving the modulator 4-210. In some embodiments, the voltage V can be assumed to have the following expression: V = VDC +V(t), where V DC It is a constant, and V(t) varies with time depending on the real number to be encoded. DC The ring can be pre-biased so that the resonant frequency is close to ψ. in The wavelength (λ0).

[0532] Figure 4-3D This is an encoder table (e.g., a lookup table) providing examples of how to encode real numbers into the phase and intensity of an optical field using modulator 4-206, according to some embodiments. In this case, by way of example, it is assumed that real numbers with increments of 1 between -10 and 10 will be encoded in the optical domain (see the column labeled "Real Numbers to be Encoded"). Of course, any suitable set of real numbers can be encoded using the techniques described herein. The column labeled "V (Applied Voltage)" indicates the voltage applied to phase shifter 4-212. The column labeled "α (Amplitude Modulation)" indicates ψ... in Intensity spectral response at wavelength (λ0) (see Figure 4-3B and Figure 4-3C (Top image). The column labeled "θ (phase modulation)" indicates ψ in Phase spectral response at wavelength (λ0) (see Figure 4-3B and Figure 4-3C (See bottom diagram). In this case, the real number "-10" is mapped to a voltage V1, which results in amplitude modulation α1 and phase modulation θ1; the real number "-9" is mapped to a voltage V2, which results in amplitude modulation α2 and phase modulation θ2; and so on. The real number "10" is mapped to a voltage V... N This leads to amplitude modulation α N and phase modulation θ N Therefore, different real numbers are encoded with different intensity / phase pairs.

[0533] Figure 4-3E Provided according to some embodiments Figure 4-3D The visual representation of the encoding table in the complex plane. As voltage changes, according to... Figure 4-3B and Figure 4-3C The spectral response, along line 4 of Figure 300, represents different intensity / phase pairs at different points. For example, the symbol S in The specific symbols representing intensity modulation α and phase modulation θ are used. in express Figure 4-3D A specific row in the table. Each row in the table is mapped to a different point on line 4-300.

[0534] Return to reference Figure 4-2BIn action 4-156, a transformation is applied to the modulated optical signal output by the modulator. Depending on how the encoded real numbers are to be processed, the transformation may involve intensity transformation and / or phase transformation. In some embodiments, optical transformation unit 4-108 can be used to transform the signal output by modulator 4-106. Optical transformation unit 4-108 can be configured to transform the symbol S characterized by intensity α and phase θ. in Transformed into a symbol S characterized by intensity αβ and phase θ+Δθ out In other words, the optical conversion unit 4-108 introduces intensity modulation β and phase shift Δθ. The values ​​of β and Δθ depend on the specific optical conversion unit used.

[0535] As an example, consider Figure 4-3E , Figures 4-4A to 4-4C Input symbol S in It provides an example of how the optical conversion unit performs the conversion. Figure 4-4A The transformation results in β < 1 and Δθ = 0. That is, the optical transformation unit introduces attenuation but does not change the phase of the input optical field. This could be, for example, the case of an optical fiber without an optical repeater, where the fiber length is chosen to maintain the input phase. The result is that line 4-300 is transformed to line 4-401, where line 4-401 is a compressed version of line 4-300. Output sample S out It has the same phase as Sin, but its intensity is equal to αβ (i.e., it is attenuated by β).

[0536] Figure 4-4B The transformation introduces attenuation and phase shift. This can be, for example, the case of an optical fiber of arbitrary length without an optical repeater. The result is that line 4-300 is transformed into line 4-402, where line 4-402 is a compressed and rotated version of line 4-300. Output sample S out It has a phase θ+Δθ and an intensity αβ.

[0537] Figure 4-4C The transformation involves multimode transformation. This can occur when the input optical field is combined with one or more other optical fields. Through multimode transformation, line 4-403 can be reshaped in any suitable manner. Examples of multimode transformation include optical processing units, some of which will be described in further detail below.

[0538] Return to reference Figure 4-2B In action 4-158, the coherent receiver can mix the transformed modulated optical field with the reference optical signal to obtain an electrical output signal. In some embodiments, it can use... Figure 4-2A The coherent receiver 4-110 is used to perform mixing. The reference optical signal can be generated by a local oscillator (e.g., Figure 4-2AThe signal generated by the local oscillator (4-112) can be either a signal transmitted through the optical conversion unit along with the modulated optical field. The mixing effect is... Figure 4-5 The example is shown in the text, not as a limitation. Line 4-403 (with) Figure 4-4C The same lines shown represent all possible intensity / phase pairs generated by a certain multimode transform. For example, Figure 4-5 It can describe the propagation of light signals through Figure 1-1 The effect of photon processor 1-103.

[0539] Assume S out It is a symbol generated by a transformation at a certain time. As mentioned above, the symbol S out Characterized by intensity αβ and phase θ+Δθ. When the transformed field is mixed with the reference signal, points along line 4-403 are projected onto the reference axis 4-500. Therefore, for example, the symbol S out The signal is projected onto point A on axis 4-500. Point A represents the electrical output signal generated by the mixing process and is characterized by an amplitude equidistant from 0 and A. It should be noted that the angle of shaft 4-500 relative to the real shaft (i.e., angle) This depends on the phase of the reference signal relative to the transformed modulated optical field. In this example,

[0540] Return to reference Figure 4-2B In action 4-160, the decoder can obtain the value representing the real number to be decoded based on the electrical output signal obtained in action 4-158. In some embodiments, the exact position of the symbol in the complex plane may be insignificant when decoding symbols obtained through optical transformation. According to some embodiments, the problem of accurate decoding is the projection of the optical signal onto a known reference axis. In other words, the symbol along line 4-403 can be mapped to a point along the reference axis, much like the symbol S. out It is mapped to point A. Therefore, some embodiments implement optical decoding schemes where points along the reference axis correspond to specific symbols in the complex plane.

[0541] A calibration process can be used to determine how points along a reference axis are mapped to symbols in the complex plane. During the calibration process, a set of input symbols (symbols representing the set of real numbers) with known intensity and phase are passed through an optical conversion unit, and the resulting symbols are coherently detected using a reference signal with known phase. The amplitude of the resulting electrical output signal (i.e., the amplitude of the projection along the reference axis) is recorded and stored in a table (e.g., a lookup table). This table can then be used during computation to decode real numbers based on the amplitude of the projection along the reference axis.

[0542] According to some embodiments, in Figure 4-6An example of such a table is shown below. This table includes a column for the phase of the reference signal, "Reference Phase". ", for columns projected along the reference axis ("projection") The table contains the amplitude of the electrical output signal and a column for the decoded real number ("decoded real number"). This table can be filled during the calibration process, including 1) passing a symbol of known intensity and phase through a known optical conversion unit, 2) setting the phase of the reference signal to a known value, 3) mixing the converted symbol with the reference signal to project the converted symbol onto a known reference axis, 4) determining the amplitude of the projection along the reference axis, and 5) recording the desired output real number. During computation, the user can use the table to infer the decoded real number based on the reference phase value and the measured projection.

[0543] exist Figure 4-6 In the examples, two reference phases, 0 and π, are considered only by way of example and not limitation. When φ is φ, the projections -1, -0.8, and 1 map to the real numbers -9.6, 0.2, and 8.7, respectively. When φ is φ6, the projections -0.4, 0.6, and 0.9 map to the real numbers 3.1, -5, and 10, respectively. Therefore, for example, if a user obtains a projection of -0.4 when using π as the reference phase, the user can infer that the decoded real number is 3.1.

[0544] Figure 4-7 A specific example of a photonic system 4-700, according to some embodiments, is shown, in which the optical conversion unit includes an optical fiber 4-708. This example illustrates how the type of encoding techniques described herein can be used in the context of optical communication. In this implementation, the optical fiber 4-708 separates the transmitter 4-701 from the receiver 4-702.

[0545] Transmitter 4-701 includes light source 4-102, encoder 4-104 and optical modulator 4-106 (combined with the above). Figure 4-2A (Description), Receiver 4-702 includes coherent receiver 4-110 and decoder 4-114 (also mentioned above). Figure 4-2A (Description). As described above, in some embodiments, a single modulation signal 4-105 can be used to drive the optical modulator 4-106. In some such embodiments, the optical modulator 4-106 can implement an on / off keying (OOK) modulation scheme.

[0546] It should be noted that in this implementation, the reference signal used for coherent detection is provided directly from the transmitter 4-701, rather than being generated locally at the receiver 4-702 (although other implementations may involve a local oscillator at the receiver 4-702). To illustrate this concept, consider, for example... Figure 4-8 The curve graph. Figure 4-8An example of the power spectral density of the optical field at the output of transmitter 4-701 is shown. As shown, the power spectral density includes signal 4-800 and carrier 4-801. Signal 4-800 represents information encoded using modulator 4-106. Carrier 4-801 represents the hue at the wavelength of light source 4-102 (λ0). When transmitting this optical field through optical fiber 4-708, coherent receiver 4-110 can use carrier 4-801 itself as a reference signal to perform mixing.

[0547] Some embodiments relate to methods for manufacturing photonic systems of the type described herein. According to some embodiments, in... Figure 4-9 One such method is described. Method 4-900 begins with action 4-902, in which a modulator is manufactured. This modulator can be manufactured to be driven by a single electrical modulated signal. An example of a modulator is modulator 4-106 (…). Figure 4-2A In action 4-904, a coherent receiver is manufactured. In some embodiments, the coherent receiver is manufactured on the same semiconductor substrate on which the modulator is manufactured. In action 4-906, an optical conversion unit is manufactured. The optical conversion unit may be manufactured on the same substrate on which the coherent receiver is manufactured and / or on the same semiconductor substrate on which the modulator is manufactured. The optical conversion unit may be manufactured to be coupled between the modulator and the coherent receiver. An example of an optical conversion unit is optical conversion unit 4-108 (…). Figure 4-2A ).

[0548] V. Differential receiver

[0549] The inventors have recognized and understood that some conventional optical receivers are particularly susceptible to noise generated from voltage sources, noise due to the unavoidable dark current produced by the photodetector, and other forms of noise. The presence of noise reduces the signal-to-noise ratio and thus reduces the ability of these photodetectors to accurately sense the input optical signal. This can adversely affect the performance of systems deploying these photodetectors. For example, it can negatively impact the system's bit error rate and power budget.

[0550] The inventors have developed an optical receiver with reduced sensitivity to noise. Some embodiments of this application relate to an optical receiver that performs photoelectric conversion and amplification in a differential manner. In the optical receiver described herein, two independent signal subtractions occur. First, the photocurrents are subtracted from each other to produce a pair of differential currents. Then, the resulting differential currents are further subtracted from each other to produce an amplified differential output. The inventors have recognized and appreciated that an optical receiver involving multi-stage signal subtraction can eliminate multi-stage noise, thus significantly reducing noise from the system. This can have several advantages compared to conventional optical receivers, including a wider dynamic range, a higher signal-to-noise ratio, greater output swing, and increased power supply noise immunity.

[0551] The type of optical receiver described herein can be used in a variety of settings, including, for example, telecommunications and data communications (including LANs, MANs, WANs, data center networks, satellite networks, etc.), analog applications (e.g., fiber optic radio), all-optical switching, lidar, phased arrays, coherent imaging, machine learning and other types of artificial intelligence applications, and other applications. In some embodiments, the type of optical receiver described herein can be used as part of a photonics processing system. For example, in some embodiments, the type of optical receiver described herein can be used to implement... Figure 1-1 Optical receiver 1-105 (e.g., Figure 1-9 One or more zero-difference receivers 1-901).

[0552] Figure 5-1 Non-limiting examples of an optical receiver 5-100 according to some non-limiting embodiments of this application are shown. As shown, the optical receiver 5-100 includes photodetectors 5-102, 5-104, 5-106, and 5-108, although other implementations may include more than four photodetectors. Photodetector 5-102 may be connected to photodetector 5-104, and photodetector 5-106 may be connected to photodetector 5-108. In some embodiments, the anode of photodetector 5-102 is connected to the cathode of photodetector 5-104 (at node 5-103), and the cathode of photodetector 5-106 is connected to the anode of photodetector 5-108 (at node 5-105). Figure 5-1 In the example, the cathodes of photodetectors 5-102 and 5-108 are connected to a voltage source V. DD The anodes of photodetectors 5-104 and 5-106 are connected to a reference potential (e.g., ground). In some embodiments, the reverse arrangement is also possible. The reference potential can be zero or any suitable value, such as -V. DD V DD It can have any suitable value.

[0553] The photodetectors 5-102 to 5-108 can be implemented in any of a variety of ways, including, for example, using pn junction photodiodes, pin junction photodiodes, avalanche photodiodes, phototransistors, photoresistors, etc. The photodetector can include materials capable of absorbing light at the wavelength of interest. For example, as a non-limiting example, at wavelengths in the O, C, or L bands, the photodetector can have an absorption region at least partially made of germanium. For visible light, as another non-limiting example, the photodetector can have an absorption region at least partially made of silicon.

[0554] Photodetectors 5-102 to 5-108 can be monolithically formed integrated components as part of the same substrate. In some embodiments, the substrate can be a silicon substrate, such as a bulk silicon substrate or silicon-on-insulator. Other types of substrates can also be used, including, for example, indium phosphide or any suitable semiconductor material. To reduce the variability of photodetector characteristics due to manufacturing tolerances, in some embodiments, the photodetectors can be positioned close to each other. For example, the photodetectors can be located at 1 mm. 2 Or even smaller, 0.11mm 2 Or smaller or 0.011mm 2 On the substrate, or in a smaller area.

[0555] like Figure 5-1 As further shown, photodetectors 5-102 through 5-108 are connected to the differential operational amplifier 5-110. For example, photodetectors 5-102 and 5-104 can be connected to the non-inverting input ("+") of the DOA 5-110, and photodetectors 5-106 and 5-108 can be connected to the inverting input ("-") of the DOA 5-110. The DOA 5-110 has a pair of outputs. One output is inverted, and the other is non-inverted.

[0556] In some embodiments, such as combining Figure 5-2 As described in detail, photodetectors 5-102 and 5-106 can be arranged to receive the same optical signal "t", and photodetectors 5-104 and 5-108 can be arranged to receive the same optical signal "b". In some embodiments, photodetectors 5-102 to 5-108 can be designed to be substantially equal to each other. For example, photodetectors 5-102 to 5-108 can be formed using the same process steps and the same photomask pattern. In these embodiments, photodetectors 5-102 to 5-108 can exhibit substantially the same characteristics, such as substantially the same responsivity (the ratio between photocurrent and received optical power) and / or substantially the same dark current (the current generated when no optical power is received). In these embodiments, the photocurrents generated by photodetectors 5-102 and 5-106 in response to receiving signal t can be substantially equal to each other. These photocurrents in Figure 5-1 The middle is identified as "i t It should be noted that, due to the orientation of photodetectors 5-102 and 5-106, the photocurrents generated by photodetectors 5-102 and 5-106 are directed in opposite directions. That is, the photocurrent of photodetector 5-102 points towards node 5-103, while the photocurrent of photodetector 5-106 is directed away from node 5-105. Furthermore, the photocurrents generated by photodetectors 5-104 and 5-108 in response to the received signal b can be substantially equal to each other. These photocurrents are denoted as "i". bDue to the relative orientation of photodetectors 5-104 and 5-108, the photocurrents generated by photodetectors 5-104 and 5-108 are directed in opposite directions. That is, the photocurrent of photodetector 5-108 points towards node 5-105, while the photocurrent of photodetector 5-104 is directed away from node 5-103.

[0557] Considering the orientation of the photodetector, the amplitude is i t -i b The current appears from node 5-103 and has an amplitude of i. b -i t The currents originate from nodes 5-105. Therefore, the currents have essentially the same magnitude but opposite signs.

[0558] Photodetectors 5-102 to 5-108 can generate dark current. Dark current typically arises from leakage within the photodetector, regardless of whether the photodetector is exposed to light. Because dark current occurs even when there is no input light signal, it effectively increases noise in the optical receiver. The inventors have recognized that the negative effects of these dark currents can be significantly mitigated due to the aforementioned current subtraction. Therefore, in Figure 5-1 In the example, the dark current of photodetector 5-102 and the dark current of photodiode 5-104 essentially cancel each other out (or at least essentially reduce each other), and the same is true for the dark currents of photodetectors 5-106 and 5-108. Therefore, the noise caused by the presence of dark current is greatly attenuated.

[0559] Figure 5-2 A photonic circuit 5-200, arranged according to some non-limiting embodiments, for providing two optical signals to photodetectors 5-102 to 5-108, is illustrated. The photonic circuit 5-200 may include an optical waveguide for routing the optical signals to the photodetectors. The optical waveguide may be made of a material that is transparent or at least partially transparent to light of the wavelength of interest. For example, the optical waveguide may be made of silicon, silicon oxide, silicon nitride, indium phosphide, gallium arsenide, or any other suitable material. Figure 5-2 In the example, the photonic circuit 5-200 includes input optical waveguides 5-202 and 204 and couplers 5-212, 5-214 and 5-216. As further shown, the output optical waveguide of the photonic circuit 5-200 is coupled to photodetectors 5-102 to 5-108.

[0560] exist Figure 5-2In the examples, couplers 5-212, 5-214, and 5-216 include directional couplers, where evanescent wave coupling enables the transmission of optical power between adjacent waveguides. However, other types of couplers can be used, such as Y-junctions, X-junctions, optical crossovers, anti-directional couplers, etc. In other embodiments, the photonic circuit 5-200 can be implemented using a multimode interferometer (MMI). In some embodiments, couplers 5-212, 5-214, and 5-216 can be 3dB couplers (with a 50%-50% coupling ratio), although other ratios are also possible, such as 51%-49%, 55%-45%, or 60%-40%. It should be understood that the actual coupling ratio may deviate slightly from the expected coupling ratio due to manufacturing tolerances.

[0561] Signal s1 can be provided at input optical waveguide 5-202, and signal s2 can be provided at input optical waveguide 204. Signals s1 and s2 can be provided to the respective input optical waveguides using, for example, optical fiber. In some embodiments, s1 represents a reference local oscillator signal, such as a signal generated by a reference laser, and s2 represents the signal to be detected. Thus, the optical receiver can be considered a homodyne optical receiver. In some such embodiments, s1 can be a continuous wave (CW) optical signal, while s2 can be modulated. In other embodiments, both signals are modulated or both signals are CW optical signals, as the application is not limited to any particular type of signal.

[0562] exist Figure 5-2 In the example, signal s1 has an amplitude A LO and phase The signal s2 has an amplitude A S and phase Coupler 5-212 combines signals s1 and s2 such that signals t and b appear at their respective outputs. In an embodiment where coupler 5-212 is a 3dB coupler, t and b can be given by the following expressions:

[0563]

[0564] Furthermore, the powers T and B (powers at t and b, respectively) can be given by the following expressions:

[0565]

[0566]

[0567] Therefore, in the embodiment where couplers 5-214 and 5-216 are 3dB couplers, photodetectors 5-102 and 5-106 can each receive the power given by T / 2, and photodetectors 5-104 and 5-108 can each receive the power given by B / 2.

[0568] Return to reference Figure 5-1 And assuming that the responsivity of photodetectors 5-102 to 5-108 is equal to that of each other (although not all embodiments are limited in this respect), the currents appearing from nodes 5-103 and 5-105 respectively can be given by the following expressions:

[0569]

[0570]

[0571] The DOA5-110 is configured to amplify the differential signals received at the "+" and "-" inputs and produce amplified differential outputs. Figure 5-1 The voltage V out,n and V out,p In some embodiments, the DOA 5-110, combined with impedance z, can be considered a differential transimpedance amplifier because it is based on a pair of differential currents (i... b -i t i t -i b ) generates a pair of differential voltages (V out,n V out,p In some embodiments, V out,n V out,p Each of them can be associated with current i t -i b With current i b -i t The difference between them is proportional, resulting in the following expression:

[0572] V out,p =2z(i t -i b )

[0573] V Out,n =2z(i b -i t )

[0574] This differential voltage pair can be supplied as input to any suitable electronic circuit, including but not limited to analog-to-digital converters (ADCs). Figure 5-1 (Not shown in the diagram). It should be noted that the optical receiver 5-100 provides two levels of noise suppression. The first level of noise suppression is achieved due to the subtraction of the photocurrent, and the second level of noise suppression is achieved due to the subtraction that occurs in the differential amplification stage. This results in a significant improvement in noise suppression.

[0575] exist Figure 5-1In the example, the impedances z are shown as equal to each other; however, different impedances may be used in other embodiments. These impedances may include passive electronics (such as resistors, capacitors, and inductors) and / or active electronics (such as diodes and transistors). The devices constituting these impedances may be selected to provide the desired gain and bandwidth, as well as other possible characteristics.

[0576] As described above, the optical receiver 5-100 can be monolithically integrated onto the substrate. According to some non-limiting embodiments, in... Figure 5-3A An example of such a substrate is shown. In this example, photodetectors 5-102 to 5-108, photonic circuitry 5-200, and DOA 5-110 are monolithically integrated as part of substrate 5-301. In other embodiments, photodetectors 5-102 to 5-108 and photonic circuitry 5-200 may be integrated on substrate 5-301, while DOA 5-110 may be integrated on a separate substrate 5-302. Figure 5-3B In the example, substrates 5-301 and 5-302 are flip-chip bonded to each other. Figure 5-3C In one example, substrates 5-301 and 5-302 are wire-bonded to each other. In yet another example (not shown), photodetectors 5-102 to 5-108 and photonic circuit 5-200 can be fabricated on separate substrates.

[0577] Some embodiments of this application relate to methods for manufacturing optical receivers. According to some non-limiting embodiments, in... Figure 5-4 Such a method is described in the paper. Method 5-400 begins with action 5-402, in which a plurality of photodetectors are fabricated on a first substrate.

[0578] Once manufactured, the light detector can, for example, in Figure 5-1 The components are connected together in the arrangement shown. In some embodiments, the photodetector may be located on the first substrate, within 1 mm. 2 Or even smaller, 0.1mm 2 Or even smaller, or 0.01mm 2 Or a smaller area. In action 5-404, a photonic circuit is fabricated on the first substrate. The photonic circuit can be arranged, for example, in a... Figure 5-2 The method shown provides an optical signal pair to the photodetector. In action 5-406, a differential operational amplifier can be fabricated on the second substrate. An example of a differential operational amplifier is... Figure 5-1 DOA5-110. In action 5-408, it can be achieved, for example, by inverted engagement (such as...). Figure 5-3A As shown), lead wire connection (such as Figure 5-3B (As shown) or using any other suitable bonding technique, the first substrate is bonded to the second substrate. Once the substrates are bonded, the photodetector of the first substrate can, for example, be... Figure 5-1The differential operational amplifier is electrically connected to the second substrate in the manner shown.

[0579] According to some embodiments, in Figures 5-4A to 5-4F An example of the manufacturing process is illustrated in the diagram. Figure 5-4A A substrate 5-301 is shown having an underlayer 5-412 (e.g., an oxide layer such as a buried oxide layer or other type of dielectric material) and a semiconductor layer 5-413 (e.g., a silicon layer or a silicon nitride layer, or other type of material). Figure 5-4B In this process, for example, photolithography is used to pattern the semiconductor layer 5-413 to form region 5-414. In some embodiments, region 5-414 can be arranged to form an optical waveguide. In some embodiments, the resulting pattern resembles a photonic circuit 5-200. Figure 5-2 Waveguides 5-202 and 204, and couplers 5-212, 5-214, and 5-216, are embedded in one or more regions 5-414. Figure 5-4C In this example, photodetectors 5-102, 5-104, 5-106, and 5-108 (and optionally, other photodetectors) are formed. In this example, light-absorbing material 5-416 is deposited adjacent to region 5-414. Light-absorbing material 5-416 can be patterned to form the photodetector. The material used for the light-absorbing material can depend on the wavelength to be detected. For example, germanium can be used for wavelengths in the L-band, C-band, or O-band. Silicon can be used for visible wavelengths. Of course, other materials are also possible. Light-absorbing material 5-416 can be positioned to optically couple to region 5-414 in any suitable manner, including but not limited to butt coupling, tapered coupling, and evanescent coupling.

[0580] exist Figure 5-4D In this process, DOA 5-110 is formed. In some embodiments, DOA 5-110 includes a plurality of transistors formed by ion implantation. Figure 5-4D The injection region 5-418 is depicted, which can form part of one or more transistors in the DOA 5-110. Although in Figure 5-4D Only one ion implantation is shown in the diagram, but in some embodiments, the formation of DOA 5-110 may include more than one ion implantation. Additionally, DOA 5-110 may be electrically connected to the photodetector, for example, via one or more conductive traces formed on substrate 5-301.

[0581] Figure 5-4D The arrangement allows the photonic circuit 5-200, photodetectors 5-102 to 5-108, and DOA 5-110 to be formed on a common substrate (e.g., Figure 5-3A As shown). DOA5-110 is formed on a separate substrate (e.g. Figure 5-3B or Figure 5-3C(As shown) is also possible. In one such example, DOA 5-110 is formed on a separate substrate 5-302, as shown. Figure 5-4E As shown, implantation regions 5-428 are formed through one or more ion implantations.

[0582] Next, substrate 5-301 is bonded to substrate 5-302, and photodetectors 5-102 to 5-108 are connected to DOA5-110. Figure 5-4F In the process, conductive pads 5-431 are formed and positioned to be electrically connected to the light-absorbing material 5-416, and conductive pads 5-432 are formed and positioned to be electrically connected to the injection region 5-428. The conductive pads are connected via wire bonding (e.g., ...). Figure 5-4F (as shown in the diagram) or electrically connected via flip-chip bonding.

[0583] Some embodiments relate to methods for receiving input optical signals. Some such embodiments may relate to zero-difference detection, although this application is not limited to this aspect. Other embodiments may include heterodyne detection. Still other embodiments may relate to direct detection. In some embodiments, the reception of the optical signal may involve optical receivers 5-100 (…). Figure 5-1 (in the middle), although other types of receivers can be used.

[0584] According to some embodiments, in Figure 5-5 An example of a method for receiving an input optical signal is described. Method 5-500 begins with action 5-502, wherein the input signal is combined with a reference signal to obtain a first optical signal and a second optical signal. The input signal may be encoded with data, for example, in the form of amplitude modulation, pulse width modulation, phase or frequency modulation, and other types of modulation. In some embodiments involving zero-difference detection, the reference signal may be a signal generated by a local oscillator (e.g., a laser). In other embodiments, the reference signal may also be encoded with data. In some embodiments, photonic circuit 5-200 is used ( Figure 5-2 The input signal and reference signal can be combined using a combination of optical circuits, but other types of optical combiners can be used, including but not limited to MMI, Y-junction, X-junction, optical crossover, and reverse coupler. In embodiments using photonic circuit 5-200, t and b can represent the signal obtained from the combination of the input signal and the reference signal.

[0585] In action 5-504, a first optical signal is detected using a first photodetector and a second photodetector, and a second optical signal is detected using a third photodetector and a fourth photodetector to generate a pair of differential currents. In some embodiments, action 5-504 can use an optical receiver 5-100 ( Figure 5-1This is performed by photodetectors 5-102 and 5-106, and photodetectors 5-104 and 5-108 detect the second optical signal. The resulting pair of differential currents is generated by current i. b -i t and i t -i b Common representation. Because it is differential, in some embodiments, the pair of currents may have substantially equal magnitudes but substantially opposite phases (e.g., with a π phase difference).

[0586] In action 5-506, the differential operational amplifier (e.g., Figure 5-1 The DOA5-110 uses a pair of differential currents generated in action 5-504 to produce a pair of amplified differential voltages. In embodiments using the DOA5-110, the resulting pair of differential voltages is generated by voltage V. out,n and V out,p This indicates that, because it is differential, in some embodiments the pair of voltages may have substantially equal magnitudes but substantially opposite phases (e.g., with a π phase difference).

[0587] Compared to conventional methods for receiving optical signals, method 5-500 may have one or more advantages, including, for example, a wider dynamic range, a higher signal-to-noise ratio, a larger output swing, and increased power supply noise immunity.

[0588] VI. Phase modulator

[0589] The inventors have recognized and understood that certain optical phase modulators suffer from high dynamic loss and low modulation speed, which significantly limits the range of applications that can utilize these phase modulators. More specifically, some phase modulators involve a significant trade-off between modulation speed and dynamic loss, such that an increase in modulation speed leads to an increase in dynamic loss. As used herein, the phrase "dynamic loss" refers to the optical power loss experienced by an optical signal, which depends on the degree to which the phase of the optical signal is modulated. An ideal phase modulator makes power loss independent of phase modulation. However, real-world phase modulators experience a certain power loss when no modulation occurs, and different power losses when modulation occurs. For example, the power loss experienced without phase modulation may be L1, the power loss experienced with π / 2 phase modulation may be L2, and the power loss experienced with π phase modulation may be L3, where L1, L2, and L3 are different from each other. This behavior is undesirable because, in addition to phase modulation, the signal also undergoes amplitude modulation.

[0590] Furthermore, some of these phase modulators require lengths of several hundred micrometers to provide a sufficiently large phase shift. Unfortunately, such long phase modulators are unsuitable for applications that require integrating multiple phase shifters on a single chip. A phase modulator alone can occupy most of the available space on a chip, thus limiting the number of devices that can be integrated together on the same chip.

[0591] Recognizing the aforementioned limitations of certain phase modulators, the inventors have developed a small-footprint optical phase modulator capable of providing high modulation speeds (e.g., exceeding 6-100 MHz or 1 GHz) while limiting dynamic losses. In some embodiments, the phase modulator can occupy as little as 300 μm. 2 The area. Therefore, as an example, with 1cm 2 The reticle area can accommodate up to 15,000 phase modulators while saving an additional 50mm for other components. 2 .

[0592] Some embodiments relate to nano-opto-electro-mechanical systems (NOEMS) phase modulators having multiple suspended optical waveguides positioned adjacent to each other and with multiple slits formed therebetween. The slits are small enough to form slit waveguides such that a majority (e.g., most of) of the mode energy is confined within the slits themselves. These modes are referred to herein as slit modes. With most of the mode energy in the slits, the effective refractive index of the mode can be modulated by causing a change in the size of the slits, and thus the phase of the optical signal embodying that mode can be modulated. In some embodiments, phase modulation can be achieved by applying a mechanical force that causes a change in the size of the slits.

[0593] The inventors have recognized and understood that by decoupling the mechanical driver from the region where optical modulation occurs, the modulation speed achievable with the NOEMS phase modulator described herein can be increased without significantly increasing dynamic losses. In a phase modulator where the mechanical driver is decoupled from the optical modulation region, the electrical drive signal is applied to the mechanical driver, rather than to the optical modulation region itself. This arrangement eliminates the need to make the optical modulation region conductive, thereby enabling a reduction in doping in that region. Low doping leads to a reduction in free carriers, which could otherwise lead to light absorption, thus reducing dynamic losses.

[0594] Furthermore, decoupling the mechanical driver from the optical modulation region enables greater modulation per unit length, thus allowing for a shorter modulation region. A shorter modulation region, in turn, enables a higher modulation speed.

[0595] The inventors have further recognized and appreciated that including multiple slits in the modulation region allows for a further reduction in the length of the phase modulator (and thus its size). In fact, having more than one slit significantly reduces the length of the transition region through which light couples to the modulation region. The result is a substantially more compact form factor. Therefore, NOEMS phase modulators of the type described herein can have shorter modulation regions and / or shorter transition regions. In some embodiments, phase modulators of the type described herein can have lengths as low as 20 μm or 30 μm.

[0596] As will be described in further detail below, some embodiments relate to phase modulators in which trenches are formed in a chip and arranged such that the modulation waveguide is suspended in the air and moves freely in space.

[0597] The inventors have recognized the potential drawbacks associated with using trenches due to the formation of a cladding / air interface. When a propagating optical signal enters (or leaves) a trench, it encounters the cladding / air interface (or air / cladding interface). Unfortunately, the presence of the interface causes light reflection, which in turn increases insertion loss. The inventors have recognized that the negative impact of this interface can be mitigated by reducing the physical extension of the optical mode in the region through which it passes the interface. This can be achieved in various ways. For example, in some embodiments, the extension of the optical mode can be reduced by tightly confining the mode within a ridge waveguide. The dimensions of the ridge waveguide can be determined such that only a small fraction of the mode energy (e.g., less than 20%, less than 10%, or less than 5%) is outside the edges of the waveguide.

[0598] The NOEMS phase modulators of this type described herein can be used in a variety of applications, including, for example, telecommunications and data communications (including LANs, MANs, WANs, data center networks, satellite networks, etc.), analog applications (e.g., fiber optic radio), all-optical switching, coherent lidar, phased arrays, coherent imaging, machine learning, and other types of artificial intelligence applications. Additionally, NOEMS modulators can be used as part of an amplitude modulator, for example, if combined with a Mach Zehnder modulator. For instance, a Mach Zehnder modulator can be provided in which the NOEMS phase modulator is located in one or more arms of the Mach Zehnder modulator. A variety of modulation schemes can be implemented using NOEMS phase modulators, including, for example, amplitude shift keying (ASK), quadrature amplitude modulation (QAM), phase shift keying (BPSK), quadrature phase shift keying (QPSK) and higher-order QPSK, offset quadrature phase shift keying (OQPSK), dual-polarized quadrature phase shift keying (DPQPSK), amplitude shift keying (APSK), etc. Additionally, NOEMS phase modulators can be used as phase correctors in applications where the phase of an optical signal tends to drift unpredictably. In some embodiments, NOEMS phase modulators of the type described herein can be used as part of a photonics processing system. For example, in some embodiments, NOEMS phase modulators of the type described herein can be used to implement... Figure 1-1 The phase modulator 1-207, and / or a portion thereof used to implement the variable beam splitter 1-401 of Figure 104, and / or used to implement Figure 1-5 Phase shifters 1-505, 1-507, and 1-509, and / or implementing Figure 1-6 The phase modulator 1-601, and / or implement Figure 1-6 A portion of the amplitude modulator 1-603, and / or implementation Figure 1-2 Part of the amplitude modulator 1-205.

[0599] Figure 6-1AThis is a schematic top view of a nano-opto-electro-mechanical system (NOEMS) phase modulator according to some non-limiting embodiments. The NOEMS phase modulator 6-100 includes an input waveguide 6-102, an output waveguide 6-104, an input transition region 6-140, an output transition region 6-150, a suspended multi-slit optical structure 6-120, mechanical structures 6-130 and 6-132, and mechanical actuators 6-160 and 6-162. The NOEMS phase modulator 6-100 can be fabricated using silicon photonics technology. For example, the NOEMS phase modulator 6-100 can be fabricated on a silicon substrate such as a bulk silicon substrate or a silicon-on-insulator (SOI) substrate. In some embodiments, the NOEMS phase modulator 6-100 may also include electronic circuitry configured to control the operation of the mechanical actuators 6-160 and 6-162. The electronic circuitry can be fabricated in a housing... Figure 6-1A The components are either on the same substrate or fabricated on separate substrates. When mounted on separate substrates, the substrates can be joined to each other in any suitable manner, including 3D bonding, flip-chip bonding, wire bonding, etc.

[0600] At least a portion of the NOEMS phase modulator 6-100 is formed in trench 6-106. As will be described in further detail below, trenches of the type described herein can be formed by etching a portion of the overlay. Figure 6-1A In the example, trench 6-106 has a rectangular shape, although any other suitable trench shape can be used. In this example, trench 6-106 has four sidewalls. Sidewalls 6-112 and 6-114 are spaced apart from each other along the z-axis (referred to here as the propagation axis), and the other two sidewalls (in...) Figure 6-1A (Not shown in the image) are spaced apart from each other along the x-axis.

[0601] In some embodiments, the spacing along the z-axis between sidewalls 6-112 and 6-114 can be less than or equal to 50 μm, less than or equal to 30 μm, or less than or equal to 20 μm. Therefore, the modulation region of this NOEMS phase modulator is significantly shorter than other types of phase modulators that require hundreds of micrometers to modulate the phase of the optical signal. This relatively short length is achieved through one or more of the following factors. First, the multiple slits improve coupling with the optical modulation region, which in turn allows for a reduction in the length of the transition region. The improved coupling may be a result of enhanced mode symmetry in the multi-slit structure. Second, decoupling the mechanical actuator from the optical modulation region enables greater modulation per unit length, thus achieving a shorter modulation region.

[0602] During operation, an optical signal can be supplied to the input waveguide 6-102. In one example, the optical signal can be a continuous wave (CW) signal. Phase modulation can occur in the suspended multi-slit optics 6-120. The phase-modulated optical signal can exit the NOEMS phase modulator 6-100 from the output waveguide 6-104. The transition region 6-140 ensures lossless or near-lossless optical coupling between the input waveguide 6-102 and the suspended multi-slit optics 6-120. Similarly, the transition region 6-150 ensures lossless or near-lossless optical coupling between the suspended multi-slit optics 6-120 and the output waveguide 6-104. In some embodiments, transition regions 6-140 and 6-150 may include tapered waveguides, as described in further detail below. As discussed above, the length of the transition regions can be shorter relative to other embodiments.

[0603] The input optical signal can have any suitable wavelength, including but not limited to wavelengths in the O-band, E-band, S-band, C-band, or L-band. Alternatively, the wavelength can be in the 850 nm band or in the visible light band. It should be understood that the NOEMS phase modulator 6-100 can be made of any suitable material, provided that the material is transparent or at least partially transparent at the wavelength of interest and that the refractive index of the core region is greater than that of the surrounding cladding. In some embodiments, the NOEMS phase modulator 6-100 can be made of silicon. For example, the input waveguide 6-102, the output waveguide 6-104, the input transition region 6-140, the output transition region 6-150, the suspended multi-slit optical structure 6-120, and the mechanical structures 6-130 and 6-132 can be made of silicon. Given the relatively low optical bandgap of silicon (approximately 1.12 eV), silicon is particularly suitable for use in conjunction with near-infrared wavelengths. In another example, the NOEMS phase modulator 6-100 can be made of silicon nitride or diamond. Given the relatively high optical band gaps of silicon nitride and diamond (approximately 5 eV and 5.47 eV, respectively), these materials are particularly suitable for use with visible light wavelengths. However, other materials are also possible, including indium phosphide, gallium arsenide, and / or any suitable III-V or II-VI alloy.

[0604] In some embodiments, the dimensions of the input waveguide 6-102 and the output waveguide 6-104 can be configured to support single-mode operation at the operating wavelength (although multimode waveguides can also be used). For example, if the NOEMS phase modulator is designed to operate at 1550 nm (of course, not all embodiments are limited to this), the input waveguide 6-102 and the output waveguide 6-104 can support single-mode operation at 1550 nm. This enhances mode confinement within the waveguides, thereby reducing optical loss due to scattering and reflection. Waveguides 6-102 and 6-104 can be ridge waveguides (e.g., having a rectangular cross-section) or can have any other suitable shape.

[0605] As described above, a portion of the NOEMS phase modulator 6-100 can be formed within the trench 6-106, such that the waveguide in the modulation region is surrounded by air and moves freely in space. A disadvantage of including the trench is the formation of a cladding / air interface and an air / cladding interface along the propagation path. Therefore, the input optical signal passes through the cladding / air interface (corresponding to sidewall 6-112) before reaching the modulation region and through the air / cladding interface (corresponding to sidewall 6-114) after the modulation region. These interfaces can introduce reflection loss. In some embodiments, reflection loss can be reduced by positioning the transition region 6-140 inside rather than outside the trench 6-106 (e.g., ...). Figure 6-1A (As shown). Thus, mode expansion associated with the transition region occurs where the optical signal has already passed through the cladding / air interface. In other words, the mode is strictly confined as it passes through the cladding / air interface, but is expanded in the trench using a transition region in order to couple to the suspended multi-slit structure 6-120. Similarly, a transition region 6-150 can be formed within the trench 6-106 to spatially re-define the mode before it reaches the sidewall 6-114.

[0606] Figure 6-1B A suspended multi-slit optical structure 6-120 according to some non-limiting embodiments is shown in more detail. Figure 6-1B In the example, the multi-slit optical structure 6-120 includes three waveguides (6-121, 6-122, and 6-123). Slit 6-124 separates waveguide 6-121 from waveguide 6-122, and slit 6-125 separates waveguide 6-122 from waveguide 6-123. The widths of the slits (d1 and d2) can be smaller than the critical widths (at the operating wavelength) used to form the slit mode, so that most of the mode energy (e.g., greater than 40%, greater than 50%, greater than 60%, or greater than 75%) is within the slits. For example, each of d1 and d2 can be equal to or less than 200 nm, equal to or less than 6-150 nm, or equal to or less than 6-100 nm. The minimum width can be set by the lithographic resolution.

[0607] Figure 6-1C This is a graph illustrating examples of optical modes supported by waveguides 6-121, 6-122, and 6-123 according to some non-limiting embodiments. More specifically, the graph illustrates the mode (e.g., electric field E) x E y or E z Or magnetic field H x H y or H z The amplitude of the mode is as follows. As shown, most of the total energy is confined within the slit, where the mode exhibits a peak amplitude. In some embodiments, the light energy in one slit is greater than the light energy in any single waveguide. In some embodiments, the light energy in one slit is greater than the light energy in all waveguides considered together. Outside the outer wall of the outer waveguide, the mode energy decays (e.g., exponentially).

[0608] Widths d1 and d2 can be equal or different. The widths of the slit and waveguide can be constant along the z-axis (e.g., ...). Figure 6-1B (As shown) or may vary. In some embodiments, the widths of waveguides 6-121, 6-122, and 6-123 may be smaller than the width of input waveguide 6-102. In some embodiments, when the operating wavelength is in the C-band, the widths of waveguides 6-121, 6-122, and 6-123 may be between 200 nm and 400 nm, between 250 nm and 350 nm, or in any other suitable range, whether within or outside such a range.

[0609] Although Figure 6-1B The example illustrates a suspended multi-slit optical structure 6-120 with three waveguides and two slits, but any other suitable number of waveguides and slits can be used. In other examples, the suspended multi-slit optical structure 6-120 may include five waveguides and four slits, seven waveguides and six slits, nine waveguides and eight slits, etc. In some embodiments, the structure includes an odd number of waveguides (and therefore an even number of slits) such that only symmetric modes are excited, while antisymmetric modes remain unexcited. The inventors have recognized that enhancing the symmetry of the modes enhances the coupling to the slit structure, thereby enabling a significant reduction in the length of the transition region. However, implementations with an even number of waveguides are also possible.

[0610] As will be described in further detail below, by making the outer waveguide ( Figure 6-1B (6-121 and 6-123 in the middle) relative to the central waveguide ( Figure 6-1BPhase modulation occurs when waveguide 6-122 moves along the x-axis. As waveguide 6-121 moves relative to waveguide 6-122 on the x-axis, the width of slit 6-124 changes, and the shape of the mode supported by the structure changes accordingly. This results in a change in the effective refractive index of the mode supported by the structure, thus causing phase modulation. Mechanical structures 6-130 and 6-132 can be used to induce movement of the outer waveguide.

[0611] According to some non-limiting embodiments, examples of mechanical structures 6-130 are shown in... Figure 6-1D As shown in Figure 6-132. Mechanical structure (see Figure 6-132) Figure 6-1A (They) can have a similar arrangement. Figure 6-1D In the example, mechanical structure 6-130 includes beams 6-133, 6-134, 6-135, and 6-136. Beam 6-133 connects mechanical driver 6-160 to beam 6-134. Beams 6-135 and 6-136 connect beam 6-134 to the outer waveguide. To limit optical loss, beams 6-135 and 6-136 may be located in the transition regions 6-140 and 6-150, respectively, instead of in the modulation region (as discussed below). Figure 6-1E (As shown) Attached to the outer waveguide. However, it is also possible to attach beams 6-135 and 6-136 to the outer waveguide for the modulation region. Beams of different shapes, sizes, and orientations can be used to replace or supplement them. Figure 6-1D The beam shown.

[0612] Mechanical structure 6-130 can transmit the mechanical force generated at mechanical actuator 6-160 to waveguide 6-121, thereby causing waveguide 6-121 to move relative to waveguide 6-122. Mechanical actuators 6-160 and 6-162 can be implemented in any suitable manner. In one example, the mechanical actuator may include a piezoelectric device. In one example, the mechanical actuator may include conductive fingers. When a voltage is applied between adjacent fingers, the fingers may experience acceleration, thus applying a mechanical force to the mech...

Claims

1. A photon processing system, comprising: An optical encoder configured to encode an input vector into a first plurality of optical signals; Photonic processor, the photonic processor being configured to: The first plurality of optical signals are received, each of which is received by a corresponding input spatial mode among the plurality of input spatial modes of the photonic processor; Multiple operations are performed on the first plurality of optical signals, wherein the plurality of operations implement matrix multiplication of the input vector and the matrix; and The output represents a second plurality of optical signals representing an output vector, each of which is transmitted by a corresponding output spatial mode among a plurality of output spatial modes of the photonic processor; An optical receiver configured to detect the second plurality of optical signals and output an electrical-digital representation of the output vector. A plurality of front ends, wherein each of the plurality of front ends is associated with one of the plurality of input space modes of the photonic processor, wherein each of the plurality of front ends includes: A plurality of optical encoders, each configured to encode a corresponding component of an input vector into an optical signal, wherein each optical encoder is configured to output an optical signal with a wavelength different from that output by the other optical encoders; and An input wavelength division multiplexer, configured to receive each optical signal in a separate spatial mode from each of the plurality of optical encoders, and to output each optical signal in a single spatial mode of a corresponding input spatial mode of the plurality of input spatial modes connected to the photonic processor; and A plurality of back-ends, wherein each of the plurality of back-ends is associated with one of the plurality of output spatial modes of the photonic processor, wherein each of the plurality of back-ends includes: An output wavelength division multiplexer, configured to receive optical signals of different wavelengths from a corresponding output spatial mode of the plurality of output spatial modes of the photonic processor, and to output each of the different wavelength optical signals in a corresponding spatial mode of the plurality of output spatial modes of the wavelength division multiplexer; and A plurality of optical receivers, each configured to determine a corresponding component of the output vector by detecting a corresponding optical signal associated with a corresponding output spatial mode of the wavelength division multiplexer.

2. The photon processing system according to claim 1, wherein, The photonic processor includes: The first interconnected variable beam splitter array includes a first plurality of optical input sections and a first plurality of optical output sections corresponding to the first plurality of input spatial modes; The second interconnected variable beam splitter array includes a second plurality of optical input sections and a second plurality of optical output sections corresponding to the plurality of output spatial modes; and A plurality of controllable optical elements, each of the plurality of controllable optical elements coupling one of the first plurality of optical output sections of the first array to a corresponding optical input section of the second plurality of optical input sections of the second array.

3. The photon processing system according to claim 2 further includes a controller, the controller being configured to: Perform singular value decomposition on the matrix to determine a first singular value decomposition matrix, a second singular value decomposition matrix, and a third singular value decomposition matrix; The first singular value decomposition matrix is ​​realized by controlling the first plurality of interconnected variable beam splitters; The second plurality of interconnected variable beam splitters are controlled to realize the second singular value decomposition matrix; as well as The third singular value decomposition matrix is ​​achieved by controlling the plurality of controllable optical elements, wherein the third singular value decomposition matrix is ​​a diagonal matrix.

4. The photon processing system according to claim 3, wherein, The controller further includes at least one digital-to-analog converter to adjust one or more parameters of the first plurality of interconnected variable beam splitters and the second plurality of interconnected variable beam splitters.

5. The photon processing system according to claim 4, wherein: Each of the first plurality of interconnected variable beam splitters and each of the second plurality of interconnected variable beam splitters is associated with a corresponding address; and The at least one digital-to-analog converter includes a single digital-to-analog converter that uses the address to control a plurality of variable beams in the first plurality of interconnected variable beams and / or the second plurality of interconnected variable beams.

6. The photon processing system according to claim 1, wherein, The optical receiver includes a low-pass filter configured to perform analog summation on a plurality of subsequent signals associated with each of the plurality of output spatial modes of the photonic processor.

Citation Information

Patent Citations

  • Apparatus and Methods for Optical Neural Network

    US20170351293A1

  • Methods, systems, and apparatus for programmable quantum photonic processing

    US9791258B2