Photonic neural network
The method and apparatus for training PNNs using unitary matrices and nonlinear optical elements in photonic integrated circuits address the challenges of optical backpropagation and loss, enabling stable and accurate optical training of deep learning models.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LUMAI LTD
- Filing Date
- 2026-01-13
- Publication Date
- 2026-07-23
Smart Images

Figure EP2026050664_23072026_PF_FP_ABST
Abstract
Description
[0001] PHOTONIC NEURAL NETWORK
[0002] The present invention relates to systems and methods for optical computation, in particular for training a neural network using optical computation.
[0003] BACKGROUND
[0004] The rapid emergence and increasing implementation of large-scale artificial intelligence (Al) models places intense demand on existing computational electronics.1In particular, machine vision and the real-time management of vast sensory data streams call for an alternative approach to both training and inference.2
[0005] Optoelectronic hardware and optical computing architectures have the potential to decrease compute latency and energy consumption,3and the development of optical processor layers of appropriate size and complexity is accelerating. The recent realisation of a full optically trained optical neural network (ONN)4validates the progress of the optical computing field and invites questions over which paths it should take to guarantee an impactful future.
[0006] Photonic neural networks (PNNs) are a class of ONN with distinct properties. Photonic neural networks marry the efficiency and latency benefits of optical computation with a compact and hardware-compatible platform that may support future visual input-based artificial intelligence.
[0007] PNNs support optical signal transmission and computation via a configuration of wafer-embedded photonic waveguides and modulators, making them intrinsically compact and readily compatible with electronics. Integrated photonic platforms can support coherent optical signals without the stability issues associated with optical fibres, and are less afflicted by problems of heat dissipation and signal attenuation than electronic circuits. Photonic processors computing linear matrix-vector multiplication (MVM) have been deployed for neural network acceleration5and PNNs have been successfully trained by digital, hybrid and in situ means.6Despite these advances, PNNs must evolve if they are to play a useful role in future Al, one step of which will be to reconcile with the requirements of deep learning models.
[0008] The present invention is concerned with providing further improvements in PNNs, as well as providing various other advantages as discussed herein.
[0009] SUMMARY
[0010] In accordance with a first aspect, there is provided a method of training a neural network using an optical neural network unit, the optical neural network unit comprising: anoptical data processing unit configured to implement a linear layer of the neural network; and a non-linear optical element configured to implement a non-linear activation function of the neural network, the method comprising: performing a forward propagation to determine an error vector by transmitting an optical input signal through the optical neural network unit to multiply with a unitary matrix an input vector encoded by the optical input signal; performing an error backpropagation by transmitting an optical error signal encoding the error vector through the optical neural network unit; updating the unitary matrix based on the error backpropagation; and repeating the steps of performing a forward propagation, performing an error backpropagation, and updating the unitary matrix until a pre-determined convergence criterion is met.
[0011] This configuration has a number of advantages. Using a unitary matrix provides improved stability in the values of the weights. Further, this system allows a nonlinear activation function and backpropagation to be provided in a fully optical manner. This closes the reality gap that exists when parts of the neural network are simulated in silico as is commonly done in existing approaches, thereby making the training more accurate.
[0012] Optionally, performing the error backpropagation comprises reconfiguring the optical data processing unit such that the error vector is multiplied by an inverse matrix of the unitary matrix used during the forward propagation. This allows light to propagate through the system in the same direction for both forward and backward propagation, which may make configuring the apparatus more straightforward. This is also enabled by the choice of a unitary matrix for the linear weight matrix.
[0013] Optionally, performing the error backpropagation further comprises: transmitting a probe signal into the optical neural network unit, the probe signal having a lower intensity than the optical input signal; and determining a derivative of the non-linear activation function based on a change in intensity of the probe signal caused by the non-linear optical element. This is particularly advantageous where a saturable gain element is used for the non-linear optical element, because the probe signal has a lower intensity that does not significantly disturb the effective gain of the non-linear element.
[0014] Optionally, the probe signal differs in one or both of wavelength and polarisation from the optical input signal. This makes the probe signal easily distinguishable from the optical input signal.
[0015] Optionally, the optical error signal is transmitted through the optical data processing unit in a direction opposite to a direction in which the optical input signal is transmittedthrough the optical data processing unit. This can make the error and input signals and their operations more easily distinguishable.
[0016] In accordance with a second aspect, there is provided an apparatus for training a neural network comprising an optical neural network unit, the optical neural network unit comprising: an optical data processing unit configured to implement a linear layer of the neural network by performing multiplication with a unitary matrix of an input vector encoded by an optical input signal; and a non-linear optical element configured to implement a nonlinear activation function of the neural network, wherein the apparatus is configured to: perform a forward propagation to determine an error vector by transmitting the optical input signal through the optical neural network unit; perform an error backpropagation by transmitting an optical error signal encoding the error vector through the optical neural network unit; update the unitary matrix based on the error backpropagation; and repeat the steps of performing a forward propagation, performing an error backpropagation, and updating the unitary matrix until a pre-determined convergence criterion is met.
[0017] The apparatus permits the implementation of the advantageous method described above, using a unitary matrix in the linear operations and a non-linear optical element to implement an optical activation layer.
[0018] Optionally, the optical data processing unit is configured to produce an output optical signal encoding the result of the multiplication with the unitary matrix of the input vector; the optical data processing unit is configured to transmit a reference signal having a constant phase relative to a phase of the optical input signal; and the apparatus is configured to determine a phase of the output optical signal by comparison of the output optical signal and the reference signal. By transmitting the reference signal through the optical data processing unit, the phase can be more accurately determined because the reference signal is subjected to the same propagation effects as the optical input signal, with the exception of any processing operations to perform the matrix-vector multiplication.
[0019] Optionally, the optical data processing unit comprises a photonic integrated circuit, for example a photonic waveguide mesh. Optionally, the optical data processing unit comprises a plurality of Mach-Zender interferometers each having one or more reconfigurable phase shifters. This provides a compact and reconfigurable implementation.
[0020] Optionally, the non-linear optical element is an optical amplifier having a non-linear response function, for example an erbium-doped fibre amplifier. Optical losses can limit the depth of an optically implemented neural network. This allows the amplifier to provide gainas well as a non-linear effect to counteract loss of amplitude as the optical signals propagate through the system.
[0021] Optionally, the apparatus further comprises an encoder unit configured to receive one or more input beams and produce one or both of a) the optical input signal during the forward propagation, and b) the optical error signal during the error backpropagation. This reduces the need to provide phase controls on separately-provided external input signals.
[0022] Optionally, the optical input signal comprises a plurality of input sub-signals; and the optical data processing unit comprises a plurality of input channels, each input channel configured to receive one input sub-signal. This can allow the system to encode different elements of the input vector as different input sub-signals, allowing an increase in the parallelisation by processing multiple vector elements simultaneously.
[0023] LIST OF FIGURES
[0024] Embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings in which corresponding reference symbols represent corresponding parts, and in which
[0025] Fig. 1 shows an apparatus fortraining a neural network comprising an optical neural network (ONN) unit;
[0026] Fig. 2 shows a deep ONN formed by a chain of ONN units;
[0027] Fig. 3 shows a recurrent ONN;
[0028] Fig. 4 shows a block diagram of an ONN unit during optical backpropagation;
[0029] Fig. 5 shows an example of the internal structure of an ONN unit comprising a photonic waveguide mesh;
[0030] Fig. 6 shows a block diagram representation of a pump-probe strategy for measuring the derivative of the optical nonlinearity;
[0031] Fig. 7 provides a summary of an example photonic neural network (PNN);
[0032] Fig. 8 shows detail of the photonic waveguide mesh of the PNN of Fig. 7;
[0033] Fig. 9 shows further detail of an example training scheme applied to the PNN of Fig. 7; Fig. 10 shows the electric field amplitude response for erbium-doped fibre amplifiers with varying pump power;
[0034] Fig. 11 shows scatter plots of test results for coherent optical matrix-vector multiplication with the PNN of Fig. 7;
[0035] Fig. 12 shows optical training results on agglomerative clustering datasets for the PNN of Fig. 7; andFig. 13 shows optical training results for radial classification datasets with varying separation parameter for the PNN of Fig. 7.
[0036] DETAILED DESCRIPTION
[0037] Deep learning constructs such as convolutional neural networks (CNNs) and transformers are demonstrating outstanding and continuously increasing capabilities on a wide range of computational tasks.7,8Optical neural networks (ONNs) have the potential to decrease compute latency and energy consumption compared to electronic models.
[0038] Physical demonstrations of ONNs contain few layers4in comparison with state-of-the-art deep learning models, though diffractive neural networks with up to 5 layers have been demonstrated albeit without an optical nonlinearity.9The assembly of a deep ONN requires a scalable design and innovative strategy to enable optical training.
[0039] ONNs suffer from appreciable optical loss which can be fatal to performance through deep or numerous layers. Moreover, with increasing ONN layer size, the threat of gradient decay or explosion increases, causing training instability. This is not limited to optical networks and relates to the lack of constraint on weight values within fully connected layers, during training. Another problem with ONNs has been the difficulty of realising an optical nonlinear activation function.
[0040] In practice, the solution to many of these problems has been to enact hybrid ONNs which might be trained in silico or contain some electronic hardware10, such as an electronic activation layer.
[0041] Any neural network must be trained before being utilised for inference. Three training methods for optical neural networks are:
[0042] a) Digital Training: The ONN structure is modelled in silico and trainable parameters (e.g., weights and biases) are determined from training simulations. Parameters are then applied to the physical ONN for optical inference. Physical setup is simple, not requiring optical back propagation. Accuracy is low since the trained digital model is not an accurate representation of the physical model used for inference.
[0043] b) Hybrid Training: Forward propagation during training is carried out on the physical ONN while backpropagation for gradient calculations is carried out in silico on the ONN model. Physical setup is simple and optical characteristics that affect inference accuracy are learned during training. Digital back propagation is reliable. These combined benefits yield high accuracy but the alternation between physical and simulated steps during training can be time and energy inefficient.c) Optical Training: Both forward and back propagation are performed optically. More physical hardware is required but optical characteristics are accounted for during training. This fully optical configuration can be used to maximise optical advantage. Thus, is it beneficial to be able to conduct optical training on an optical neural network, to maximise performance, prevent systematic errors associated with a “reality gap” between training and inference scenarios, and fully exploit energy efficiency benefits of all-optical computing.
[0044] Key challenges in implementing fully optical models include implementing optical backpropagation to optimise training and ultimately improve inference accuracy, realising deep optical neural networks for increased capability, and establishing optical linear and nonlinear layer structures that can form the building blocks of various network models.
[0045] The present disclosure provides a method for optically training an ONN, such as a photonic neural network (PNN), and a corresponding apparatus comprising an optical neural network unit whose structure offers a scalable framework for the development of deep optical learning models.
[0046] Fig. 1 shows an apparatus 1 for training a neural network comprising an optical neural network unit 2, which can be used for a corresponding method of training a neural network using the optical neural network unit 2.
[0047] Photonic Processor
[0048] The optical neural network unit 2 comprises an optical data processing unit 4 (ODPU) configured to implement a linear layer of the neural network by performing multiplication with a unitary matrix (I7n) of an input vector encoded by an optical input signal. The optical data processing unit 4 may comprise a photonic integrated circuit, preferably a reconfigurable photonic integrated circuit. For example, the optical data processing unit 4 may comprise a photonic waveguide mesh such as a plurality of Mach-Zender interferometers each having one or more reconfigurable phase shifters. The array of Mach-Zender Interferometers (MZI) may be provided in a triangular or rectangular configuration. An example of such a photonic waveguide mesh is shown in Fig. 7 and Fig. 8 and discussed in more detail below.
[0049] The ODPU 4 performs as a fully connected linear layer and can either be used repeatedly (i.e., so that a single ODPU 4 acts as plural different linear layers in the network) as in the recurrent network of Fig. 3, or in sequence with other ODPUs in the case of a deep PNN structure, as in the deep network of Fig. 2.The input vector may be a complex input vector (xn) encoded using the amplitude and phase of the optical input signal. The optical input signal may comprise a plurality of input sub-signals and the optical data processing unit 4 may comprise a plurality of input channels. Each input channel may be configured to receive one input sub-signal. The input sub-signals may, for example, correspond to different elements of the input vector or to different bits of an n-bit input vector element. Each input sub-signal might optionally have a distinct photon energy (wavelength). When the input vector is a complex input vector, the complex value of respective vector elements of the complex input vector might be stored within the amplitude and phase of respective input sub-signals.
[0050] During training and inference, the complex input vector might be reconfigured in various time steps to represent different complex input vectors in an input batch. During training, the unitary matrix might be reconfigured between training steps to reflect weight updates calculated via backpropagation.
[0051] The unitary matrix is the weight matrix of the linear layer, and has a matrix determinant which lies on the unit circle of the complex plane. The significance of this is that individual elements of the unitary weight matrix are prevented from exploding or vanishing throughout optimisation and training. Weight gradients are therefore stable during training - a crucial property for viable deep networks. In the case of an MZI mesh with an equal number of input and outputs, where every output is connected to every input, the matrix formed is necessarily unitary.
[0052] The ODPU 4 performs optical matrix-vector multiplication on the input vector. The multiplication results in an output vector, which may be a complex output vector (zn). The optical data processing unit 4 may be configured to produce an output optical signal encoding the output vector, representing the result of the multiplication with the unitary matrix of the input vector. Where the output vector is complex, the output vector may be encoded in both the amplitude and phase of the output optical signal. Similarly, as for the input optical signal, the optical data processing unit 4 may be configured to produce a plurality of output subsignals, which may, for example, correspond to different elements of the output vector or to different bits of an n-bit output vector element. The ODPU 4 may comprise a plurality of output channels, and one output vector element may be provided at each output channel.
[0053] The output vector may be detected using an optical detector. For example, the apparatus 1 may comprise one or more photodetectors 12 configured to detect an output of theoptical data processing unit 4. The optical intensity of the output signal and / or each output sub-signal might be directly detected using an optical detector such as a photodiode interfaced with an electrical amplifier circuit, and the optical amplitude inferred from this.
[0054] Where the input vector and the output vector are complex, the optical data processing unit 4 may be configured to transmit a reference signal having a known or constant phase relative to a phase of the optical input signal. The apparatus 1 may then be configured to determine a phase of the output optical signal by comparison of the output optical signal and the reference signal.
[0055] Any suitable comparison of the reference signal and the output signal may be used to determine the phase. For example, the optical phase of the output signal and / or each output sub-signal might be determined by intensity measurements made after the ODPU or after the non-linear optical element 6, the latter particularly if the non-linear optical element 6 has no effect (or a consistent, known effect) on phase. Specifically, the output signal or each output sub-signal signal and the reference signal may be mixed (for example within the photonic mesh where the ODPU 4 comprises a mesh), such that the mixed signal exhibits phasedependent optical interference. The resulting combined optical amplitude with varying reference phase yields phase information for the signal channel. The phase of the reference signal may also be varied to observe the change in intensity of the mixed signal.
[0056] Amplitude and phase measurements might be conducted during separate time steps. The measured output vector might be arbitrarily sent to another network layer via an optical or electronic interconnection to form the input vector for that layer (n+i) or might form the result of the desired calculation.
[0057] Power is conserved through the linear layer of the ODPU 4 (ignoring attenuation losses from propagation through the ODPU 4 itself). The output vector from the linear layer, which is the product of the unitary matrix and the input vector, contains the same total power as the input vector. A learning model containing these complex unitary layers is capable of handling complex values from input to output, by storing real and imaginary components within the amplitude and phase of each optical input and output signal, respectively. In some embodiments, a learning model containing these complex unitary layers can be trained using a unitary stochastic gradient descent (USGD) optimiser. This technique helps in preventing vanishing or exploding weight gradients by constraining weight parameters. This type of optimizer has been proposed as a suitable training procedure for recurrent neural networks(RNNs).
[0058] An optical source 10 sends optical signals into the apparatus 1. The source 10 may be part of the apparatus 1, or may be external to the apparatus. The apparatus 1 may comprise an encoder unit 8 configured to receive one or more input beams from the source 10 and produce the optical input signal. This is particularly relevant to the case where the optical input signal comprises a plurality of input sub-signals, and provides an alternative to each input sub-signal being generated by separate sources and / or being generated externally to the apparatus 1.
[0059] The encoder unit 8 may comprise an optical modulator array or photonic mesh.
[0060] Producing the optical input signal may comprise producing each of the plurality of input subsignals. The source 10 might provide exactly one continuous-wave optical signal to the apparatus 1, from which the input optical signal is produced by the encoder unit 8. Producing the input optical signal using the encoder unit 8 reduces the need to provide phase controls on separately-provided external input signals because, for example, the relative phase of the input sub-signals can be controlled by the encoder unit 8.
[0061] Non-linear element
[0062] The optical neural network unit 2 further comprises a non-linear optical element 6 configured to implement a non-linear activation function (or nonlinear layer) of the neural network. The non-linear optical element 6 can either be used repeatedly (i.e., so that a single non-linear optical element 6 acts as plural different nonlinear layers in the network) or in sequence with other non-linear optical elements 6, in the case of a deep PNN structure.
[0063] The output vector from the ODPU 4 may pass directly to the non-linear optical element 6. The non-linear optical element 6 nonlinearly transforms the output vector from the ODPU 4 depending on the form of a transfer function ( ) of the non-linear optical element 6 and the properties of the output vector. The output from the non-linear optical element 6 might be known as the activated output vector (an= f(zn)). The activated output vector might optionally be detected using an optical detector in a similar manner as discussed for the output vector above. For example, the apparatus 1 may comprise one or more photodetectors 12 configured to detect an output of the optical neural network unit 2, i.e. the activated output vector.
[0064] Where the ODPU 4 comprises a plurality of output channels, plural non-linear optical elements 6 may be provided, where each non-linear optical element 6 implements a non-linear activation function for one or more of the output channels. In this way, different outputchannels may have different activation functions applied depending on the requirements of the ONN.
[0065] The non-linear layer is realised optically. This negates the need for an electronic activation function, as is often used in exiting ONNs. In some embodiments, the non-linear optical element 6 has a nonlinear response function with respect to optical input intensity. Any suitable optical element with a non-linear response may be used, for example the nonlinear optical element 6 may be a saturable optical absorber.
[0066] In some embodiments, the non-linear optical element 6 may be provided as a physically separate component to the ODPU 4, for example being connected to the ODPU 4 via optical fibre. Alternatively, the non-linear optical element 6 may be integrated with the ODPU 4, for example on a single compact photonic platform.
[0067] In some embodiments, the non-linear optical element 6 is an optical amplifier having a non-linear response function. The optical amplifier may be a fibre amplifier, booster amplifier, semiconductor amplifier. The optical amplifier may exhibit gain saturation / compression such that input optical signals experience reduced gain with increasing power / intensity. In particular, the optical amplifier may be an erbium-doped fibre amplifier (EDFA), an ytterbium-doped fibre amplifier.
[0068] The choice of an optical amplifier as the non-linear optical element 6 has the advantage that the non-linear optical element 6 acts to compensate optical losses by augmenting the optical signal at each nonlinear layer between linear layers, to compensate for optical losses experienced elsewhere in the network, particularly through the ODPU 4. In some embodiments, the optical amplifier can be selected for its gain and noise properties. In some embodiments, the optical amplifier may have tuneable gain e.g., via its input current setting.
[0069] In some embodiments, the nonlinear response function can be described by the saturable gain relationship, with the gain compression effect limiting gain at high input power. Gain saturation is a useful strategy for simultaneously implementing optical nonlinear layers11while amplifying signals between linear layers using a saturable gain amplifier to offset loss. This yields a nonlinear activation function similar in nature to the tanh(x) function, which is known for use in learning models.
[0070] Deep PNN
[0071] Deep ONN structures can be provided with an arrangement of one or more optical neural network (ONN) units 2. Any arrangement of linear layers provided by an ODPU 4 and nonlinear layers provided by a non-linear optical element 6 can form a unit as part of a largernetwork or can be reutilised through different timesteps to perform as a full ONN or a optical recurrent neural network (ORNN). The overall ONN framework may be integrated on a single photonic chip to provide a photonic neural network (PNN) or photonic recurrent neural network (PRNN), with optical sources and detectors co-packaged to yield a compact system.
[0072] Optionally, the apparatus 1 may comprise a plurality of optical neural network units 2 implementing plural layers of the neural network. Fig. 2 shows a deep ONN formed by a chain of ONN units 2. For example, this deep PNN might form the whole or part of a multilayer perceptron, convolutional neural network, or transformer model.
[0073] Alternatively (or additionally in other parts of the same ONN), Fig. 3 shows a recurrent ONN containing one ONN unit 4. The output of the ONN unit 4 is detected and fed back into the ONN unit 4 as the input in a subsequent timestep. The ONN unit 4 may be reconfigured between these timesteps, for example to change the unitary matrix. The recurrent ONN could optionally contain more than one ONN unit 4.
[0074] Training
[0075] The full network structure is compatible with optical training. The method of training a neural network comprises performing a forward propagation to determine an error vector by transmitting an optical input signal through the optical neural network unit 2. This results in multiplying the input vector with the unitary matrix.
[0076] The apparatus 1 is configured to perform the forward propagation. The configuration for forward propagation is substantially as shown in Fig. 1. For example, a controller (not shown) may be used to set the unitary matrix of the ODPU 4 and / or the properties of the nonlinear optical element 6 appropriately for the current step of the training procedure. Where the apparatus 1 comprises an encoder unit 8, the controller may also control the encoder unit 8 to set the optical input signal appropriately.
[0077] The error vector may be determined by comparing the output of the ONN unit 2 during forward propagation to an expected (or “ground truth”) output. The error vector might be the derivative of an appropriate loss function with respect to the output vector from the ONN unit 2. For example, the error vector might be defined as where L is the loss function,
[0078]
[0079] znis the output vector from the ODPU 4 of the nth ONN unit in the network, and anis the activated output vector from the non-linear optical element 6 of the nth ONN unit in the network.
[0080] For ONN training, weight updates might be calculated via derivatives of the lossfunction (L) with respect to the weights in question. Weight updates for the final ONN unit in a given model might be trivial to calculate. However, weight updates for ONN units deeper within the model, or that require knowledge of the derivative of an optical nonlinearity for example, call for optical backpropagation steps to enable their determination.
[0081] The method further comprises performing an error backpropagation by transmitting an optical error signal encoding the error vector through the optical neural network unit. This allows the ONN unit to conduct unitary matrix-vector multiplication (MVM) of the error vector and a unitary matrix to calculate loss gradients that enable determination of weight updates for prior network layers
[0082]
[0083] Un-2...).
[0084] The method may further comprise updating the unitary matrix used in forward propagation based on the error backpropagation. The steps of performing a forward propagation, performing an error backpropagation, and updating the unitary matrix may be repeated until a pre-determined convergence criterion is met.
[0085] The apparatus 1 is configured to perform the error backpropagation by transmitting an optical error signal encoding the error vector through the optical neural network unit 2. The optical error signal is transmitted through at least some of the optical components of the optical neural network unit 2 that carry the optical input signal through the optical neural network unit 2, optionally all of the optical components of the optical neural network unit 2 that carry the optical input signal through the optical neural network unit 2. The optical error signal may have substantially the same wavelength as the optical input signal.
[0086] Fig. 4 shows a block diagram of an ONN unit 2 during the first stage of optical backpropagation for optical training. The optical error signal encoding the error vector is generated at the input stage of the ONN unit 2 and / or by the optical source, similarly as for forward propagation. For example, where the apparatus 1 comprises an encoder unit 8, a controller of the apparatus 1 may control the encoder unit 8 to set the optical error signal appropriately according to the error vector calculated during forward propagation.
[0087] Performing the error backpropagation may comprise reconfiguring the optical data processing unit 4 such that the error vector is multiplied by an inverse matrix of the unitary matrix used during the forward propagation. In particular, for backpropagation, the configured unitary matrix by which the ODPU 4 multiplies the error vector may be the conjugate (Hermitian) transpose of the unitary weight matrix present during forward propagation (
[0088]
[0089] . This means that evaluation of the backpropagation unitary MVM can yieldderivatives - - or - - required for subsequent determination of the loss gradient with dzn_ dan_
[0090] •
[0091] respect to weights,
[0092]
[0093] During b ackpropagation, the optical error signal may be transmitted through the optical data processing unit 4 in a direction opposite to a direction in which the optical input signal is transmitted through the optical data processing unit. The optical error signals encoding the error vectors may physically propagate through the ONN unit 2 in either the same or opposite direction as that used for the optical input signal during forward propagation. The direction is defined by the choice of source and detector.
[0094] The optical error signal may propagate in either direction if the optical components used within the ONN unit 2 are reciprocal in nature. Fig. 5 shows an example of the internal structure of the ONN unit 2 where the ODPU 4 comprises a photonic waveguide mesh with numerous Mach-Zender Interferometers (MZIs) for carrying out unitary transformations. In this example, the presence of two external phase shifters along with the internal phase shifter enable the MZI mesh to be reconfigured in an identical manner, regardless of physical propagation direction. The presence of just one such MZI unit at the input and output of the photonic mesh would have a similar effect of enabling any arbitrary unitary transform to be configured, but with different phase shifter configurations for each physical propagation direction.
[0095] Propagating the optical error signal through the ONN unit 2 in the same direction as the optical input signal has the benefit of requiring fewer optical sources and detectors in the overall system and might yield more representative weight updates given the optical signals in and out of the PNN unit will encounter common paths.
[0096] A second stage of optical backpropagation might be used to determine the derivative of the optical nonlinearity of the non-linear optical element 6 with respect to the optical input signal it receives
[0097]
[0098] = '(zn-i))- For example, the optical nonlinearity may be the gain where the non-linear optical element is an amplifier, or the absorption where the non-linear optical element 6 is a saturable absorber.
[0099] A “pump-probe” scheme can be used to physically extract the derivative, where the optical input signal represents the pump signal. In particular, performing the error backpropagation may further comprise transmitting a probe signal into the optical neural network unit 2, the probe signal having a lower intensity than the optical input signal, and determining a derivative of the non-linear activation function based on a change in intensityof the probe signal caused by the non-linear optical element 6.
[0100] The “pump” signals representing the output vector zn-xare activated by the non-linear activation function , yielding the activated output vector, an-x, as discussed during forward propagation for inference or training. Separately, “probe” signals with optical intensity significantly smaller than the pump signals, are generated and sent through the non-linear optical element 6. The low optical intensity of the probe signal with respect to the pump signal is crucial for its ability to sample the response function of the non-linear optical element 6 without influencing the response beyond a minimal degree.
[0101] The non-linear activation function experienced by each low intensity probe is dependent on the intensity of the corresponding pump signal. The derivative of the optical nonlinearity is determined by measurement of the pump output, and the probe input and output with respect to the optical nonlinearity.
[0102] For the case where the non-linear optical element 6 is a saturable gain amplifier, the response function can be written as:
[0103]
[0104] Where E1 inis the electric field of the pump signal into an amplifier with small signal gain gssand saturation power P, and E1 outis the electric field of the amplified (activated) pump signal. The derivative of this response function is given by:
[0105] 0.5
[0106] f'(Elin) = gss
[0107]
[0108] 7777
[0109]
[0110] p
[0111]
[0112] Where the first bracketed term is simply the gain term from Eq. 1. Thus Eq. 2 can be rewritten as:
[0113]
[0114] Where E2tnis the electric field of the probe signal into the nonlinearity and E2 out is the electric field of the amplified probe signal. The constant term c is equal to gssP.
[0115] Therefore, the derivative of the optical nonlinearity can be determined by measuring the relative amplification of the probe signal, and the final amplified pump signal, without knowledge of the input pump signal.
[0116] In addition to this scheme using generic, low intensity probe signals, the probe signalsthemselves might constitute an error vector originating from another ONN unit 2. This strategy might be useful for detecting the error vector and probe input vector simultaneously, enabling the calculation of subsequent weight gradients in fewer time steps. Fig. 6 shows a block diagram representation of this pump-probe strategy, using an error vector as the low intensity probe input.
[0117] The pump-probe method is valid depending on the interaction of the pump and probe signals with the non-linear optical element 6, and with each other. Specifically, the pump and probe should remain separable from one another before and after interactions with the nonlinear optical element 6. For example, the probe signal may differ in one or both of wavelength and polarisation from the optical input signal. This prevents interference and allows separation of the “probe” signal from the original “pump” signal via an appropriate optical filter (e.g. band-pass, short-pass, long-pass, notch filters etc.) or polarisation-sensitive optics (e.g. a polarising beam splitter or reflective polariser). However, this is not essential and the probe signal may be separated from the optical input signal in other ways. For example, the probe signal and the optical input signal may be propagated through the nonlinear optical element 6 in opposing directions, to simplify the separation of pump and probe.
[0118] Non-linear optical elements 6 to which this strategy is applicable include linear and non-linear optical amplifiers such as semiconductor optical amplifiers and doped fibre amplifiers, and linear and non-linear absorbers such as saturable absorbers.
[0119] RESULTS
[0120] We build and optically train a two-layer PNN to perform both fully optical and hybrid-optical training on nonlinear classification datasets and showcase strong experimental learning capabilities in comparison with network simulations. Optical backpropagation is realised by propagating known loss gradients forwards through the PNN unit to determine unknown loss gradients and subsequent weight updates.
[0121] Our model achieves >90% training and test accuracy on classification of spherical quadrants and outperforms realistic imperfect network simulations on several tasks, indicating that the simple PNN is capable of learning and compensating for its own systematic errors.
[0122] Fig. 7 provides a summary of the PNN, which has four input, hidden, and output neurons, respectively, as illustrated schematically in Fig. 7(a). The PNN consists of two complex linear layers, each supporting a 4x4 complex unitary weight matrix, either side of one activation layer.
[0123] The central component is the optical data processing unit 4 or photonic processor. Inthis example, this is provided by a 12-channel (with 12-input and output ports) Mach-Zender interferometer (MZI) mesh fabricated within a low-loss SisN4 wafer and designed for classical or quantum information processing.12The array of configurable Mach-Zender interferometer units arranged in the Clements design13are manipulated to carry out arbitrary unitary transforms on C-band input signals.
[0124] Laser input from a source 10 is coupled via polarisation maintaining fibres to layer 1 of the photonic processor and split into a four-channel input vector and reference channel. The input vector is multiplied with a unitary weight matrix and resulting four vector elements are each coupled to four erbium-doped fibre amplifiers (EDFAs) respectively, for saturable gain amplification. This hidden layer vector is detected (red arrow; PD 1) and its phase values determined by mixing with the reference channel.
[0125] The hidden layer vector is regenerated in layer 2 of the photonic processor, using a second laser source having a different wavelength (X = 1553 nm) where the second matrixvector multiplication (MVM) is carried out. The complex PNN output vector is detected (blue arrow; PD 2) and its four elements are compared with the ground truth for evaluation of the loss function.
[0126] For optical ’backpropagation’ to determine the weight update for layer 1, the layer 2 weight matrix is switched to its own conjugate transpose and the gradient vector is az2
[0127] optically generated (see Fig. 9). The output of this MVM yields the -y- gradient vector, which is colinearly propagated through the EDFA with the layer 1 signal to sample gain saturation (measured at the error detector; PD 3), informing the weight updates.
[0128] The MZI mesh is shown in more detail in Fig. 8. The system is constructed to receive arbitrary complex input vectors. All signals travel left to right on the 12-port interferometer mesh where black lines are waveguides. Each MZI waveguide unit consists of back-to-back 3 dB beam splitters with an internal phase shifter (9) between to control amplitude mixing, and an external phase shifter ((|>) to control the relative output phase.
[0129] For the example 4x4x4 network, the 12-channel photonic processor is partitioned to support both complex linear layers. The first linear layer is fed at input port 10 and the second linear layer is fed at input port 4. At each input, the beam is split equally between the ‘Input vector’ and ‘Reference’ regions.
[0130] To avoid necessary external phase controls on each of the five inputs, MZI units in columns B and C are configured to act as an encoder unit to set the input vector amplitudes.We inject the input laser light through a single channel and split this into a 1 *4 vector with normalised amplitude and arbitrary phase, and a reference channel. Column D sets the combined output phase of the vector and the input phase of the unitary weight matrix.
[0131] Columns E-H apply the unitary transform. Columns I-K are configured to diagonally transmit the column H signals to the outputs (i.e., the path of signal at H12 is II 1-J10-K9-L9).
[0132] Outputs 9-12 are coupled to the non-linear optical element 6. Outputs 3-6 are coupled to photodetector array 2. For phase measurements, Reference from rows 8 and 2 is bled into the signal channels at columns G, I and K (5:95 split at each MZI). External phase shifter E8 (E2) is ramped through 1 OJT / 6 radians in five steps. The resulting amplitude modulation of each signal channel is used to determine their phase. Note that signals at LIO and LI 1 (L4 and L5 for the second linear layer) are intermixed with a single reference beam.
[0133] Returning to Fig. 7, the PNN uses multiple detection stages during operation. There are two arrays of four photodetectors in the hidden layer (PD array 1 and PD array 3), and one array of four photodetectors at the output layer (PD array 2). Optical signals are detected by InGaAs PIN photodiodes. Photodetector circuits were custom designed and are two-stage photocurrent amplifiers with adjustable gain up to 4* 105.
[0134] For forward propagation, each layer (n) contains a block to generate vectors (x and a in layer 1 and 2, respectively) and a block to generate unitary weight matrices (I7n). The resulting vector (zn) from the MVM is measured at the output of each layer. Reference channels provide a means to calculate the phase of each output vector element, through interference with each signal channel during a separate measurement step.
[0135] Each of the layer 1 output (zx) elements are coupled to an individual erbium-doped fibre amplifier (EDFA) which supplies the nonlinear amplification function, (z- via saturable gain amplification. The resulting amplified vector (a- is detected (PD 1 in Fig. 7(b)) and its phase determined by mixing with the reference channel. The vector is then regenerated in layer 2 of the processor using a second laser source where the second matrixvector multiplication (MVM) is carried out. Finally, the output of the second linear layer (z2) is detected (PD 2 in Fig. 7(b)) as the resulting inference. The complex PNN output is detected and compared with the ground truth for evaluation of the loss function.
[0136]
[0137] For backward propagation during training, the loss gradient is generated at the input az2
[0138] block of layer 2 and is multiplied by the conjugate transpose of the weight matrix,
[0139]
[0140] to yieldthe gradient - — . Phases for - — are calculated digitally to increase the throughput of training da-^ da-^
[0141]
[0142] data. Layer 1 remains unchanged such that both the
[0143]
[0144] and signals are sent through the da^
[0145] EDFA. Crucially, laser sources with distinct wavelengths are used within layer 1 (pump; X = 1559 nm) and layer 2 (probe; X = 1550 nm) of the PNN, enabling independent sampling of the signals after the EDFA through spectral separation using a band-pass filter. The pumpdependent amplifier gain, '(zi)< is sampled by measuring the amplification of the comparatively small probe signal. From this, the final loss gradient (-^-) is calculated and dz
[0146] weight matrices are updated.
[0147] The training scheme is illustrated in more detail in Fig. 9.
[0148] As shown in Fig. 9(a), input power from the pump laser source (X = 1559nm) is stabilised using a variable optical attenuator (VOA) controlled by a responsive proportional-integral-derivative (PID) algorithm receiving input from a power meter, which receives an optical signal via a 90:10 fibre beam splitter.
[0149] The first complex linear layer, first PNN unit (Layer 1), is implemented on the photonic processor as described above in connection with Fig. 8 and each of the four, layer 1 output vector elements is coupled to an individual erbium-doped fibre amplifier.
[0150] Each non-linearly amplified channel is coupled into free space where an optical bandpass filter reflects pump signals to photodetector (PD) array 1.
[0151] As shown in Fig. 9(b), to complete forward propagation, the detected signals (amplitude and phase) are regenerated in the second layer of the photonic processor, which is fed by the probe laser source via a second VOA.
[0152] Output from the second complex linear layer, second PNN unit (Layer 2), is detected by PD array 2. This is the final PNN output, a2.
[0153] As shown in Fig. 9(c), the error signal for back propagation is generated in the second PNN unit (Layer 2) of the photonic processor, along with the conjugate transpose of the second unitary weight matrix ([L )• The gradient dL / da^ is measured by PD array 2. Both the pump and probe signals pass through the EDFA and are separated at the bandpass filter. Finally, PD array 3 measures the amplified probe signal, such that the amplifier gain may be discerned, and the gradient dL / dztcalculated.
[0154] Fig. 10 shows the electric field amplitude response for each EDFA with varying pump power. Fig. 10(a) shows the pump response curve for each of the four neural networkchannels. Solid lines show the measured power output at photodetector array 1 over the full pump input range (0 - 200 pW). Dashed lines are fitted response curves using the model of equation 1 in the main text.
[0155] Fig. 10(b) shows unwanted power leakage into the pump output, from the probe channel and from the amplified spontaneous emission (ASE) of the EDFA. Fig. 10(c) shows the probe response curve for each channel showing power output at photodetector array 3 over the full probe range (0 - 20 pW). Fig. 10(d) shows unwanted power leakage into the probe output, from the pump channel and from the ASE of the EDFA.
[0156] The response curves account for coupling loss contributions. The difference in response between EDFAs 1-4 does not affect the learning capability of the PNN given the curves remain constant in time.
[0157] System Characterisation
[0158] Having established the PNN composition, we now characterise the performance of the linear layers in carrying out coherent MVM.
[0159] Fig. 11 shows scatter plots of test results for coherent optical matrix-vector multiplication with the photonic processor against ground truth for various calculations. In particular, Fig. 11(a) to Fig. 11(c) show Resulting vs. Target output amplitudes for implementation on layer 1 of the photonic processor of (a) random normalised 1x4 vector inputs, (b) random unitary weight matrices multiplied by single channel inputs, and (c) matrix-vector multiplication of random unitary matrices with random 1x4 vector inputs. Fig.
[0160] 11(d) shows Resulting vs. Target output phase for implementation of random unitary matrices multiplied by single channel inputs.
[0161] We also determine the mean residuals and root-mean-square errors for each calculation to quantify the fidelity of various regions of the processor. Mean Residual and Root-meansquare error are shown in Fig. 11(a) to Fig. 11(c). Median Residual is shown in Fig. 11(d), to avoid mean contribution of the modulo 2K data errors at the top left and bottom right of the plot.
[0162] Classification Tasks and Network Simulations
[0163] We train the complex unitary PNN on nonlinear classification tasks that are suitable for its limited size and shape. We generate two different agglomerative clustering datasets with four equally populated classes in the positive real quadrant of the four-dimensional unit sphere. These datasets have classification boundaries which are in direct contact and are not linearly separable. We create a further task by defining four data classes separated radially intwo dimensions on the 4-D unit sphere. Separation of these classes is parameterised to vary the learning difficulty in a controlled manner.
[0164] Simulations are carried out on a digital replica of our PNN to determine maximum performance on the classification tasks. For each of the cluster datasets, mean training and validation accuracies converge rapidly at approximately 90% over the first 10 epochs (batch size = 100, batches per epoch = 6). Mean training and validation accuracies for the radial classification tasks trend from 85% to >99 % for increasing class separation. Convergence is slower for these tasks, taking up to 30 epochs.
[0165] We perform the following experiments on the physical PNN:
[0166] 1. Digital training and optical inference: weights determined from simulated training are applied to the physical model. Physical setup is simple, not requiring optical back propagation. Accuracy is low since the trained digital model is not an accurate representation of the physical model used for inference.
[0167] 2. Hybrid training and optical inference: forward propagation is performed optically (as in Fig. 9(a) & Fig. 9(b)) with digital back propagation. Physical setup is simple but optical characteristics that influence inference accuracy are learned during training. Digital back propagation is reliable. These combined benefits yield high accuracy. 3. Optical training and optical inference: both forward and back propagation are performed optically. Full physical setup is utilised (Fig. 9(a) to Fig. 9(c)) and optical characteristics are learned during training. This fully optical configuration can be used to maximise optical advantage. Optical back propagation does not significantly depreciate performance in training or inference, yielding high accuracy.
[0168] In summary, the PNN inference accuracy is low after digital training but high after hybrid or optical training. Optical inference results from digital training on the cluster (average) dataset show the resulting accuracy of 51% is drastically lower than the simulated performance of 91%.
[0169] Similar trends are observed for digital training on the radial (y = 0.3) dataset. These results confirm that the accumulation of errors through the PNN - particularly the dominant error from layer 2 of the photonic processor - considerably degrades inferencing accuracy when no error correction is handled during training.
[0170] Fig. 12 shows optical training results on each agglomerative clustering dataset.
[0171] Fig. 12(a) shows loss (blue) and training accuracy (orange) by optical training iteration on an agglomerative clustering dataset with four classes merged by average linkage. Quotedaccuracy and loss values are averaged over final five iterations and illustrated by dashed lines. Fig. 12(b) shows validation accuracy by optical training epoch on the same average linkage dataset. Quoted validation accuracy is the final epoch validation accuracy and is illustrated by the dashed line. Inset shows test data confusion plot with test accuracy value.
[0172] Fig. 12(c) shows loss (blue) and optical training accuracy (orange) for another agglomerative clustering dataset merged by Ward-type linkage. Fig. 12(d) shows validation accuracy by optical training epoch on the Ward-type linkage dataset. Inset shows test data confusion plot.
[0173] Training trajectories for each dataset are similar, with accuracies converging at > 85% within 60 iterations (10 epochs). Final test accuracies of > 84% are achieved; overall performance is comparable to simulation (training accuracy within 6 percentage points). Training trajectories are stable using the unitary optimizer (learning rate = 0.01, batch size = 100) and convergence is fast. Optical training yields far superior optical inferencing accuracy than digital training.
[0174] Hybrid training of the PNN on the cluster datasets is also successful, with test accuracies exceeding 87%, 3 percentage points higher than achieved by optical training.
[0175] Accuracy and loss trajectories are comparable between the hybrid and optical trainings.
[0176] Importantly, the optical backpropagation which is introduced for optical training but is not employed during hybrid training, does not significantly degrade the learning capability or stability of the PNN.
[0177] Optical training results for the radial classification datasets with varying separation parameter are summarised in Fig. 13. Here, the separation parameter (y) is tuned to control the complexity of the task.
[0178] Fig. 13(a) and Fig. 13(b) show two-dimensional t-distributed stochastic neighbour embedding (t-SNE) representation of four classes in the positive real quadrant of the 4-D unit sphere separated by radial distance in two dimensions, with a separation parameter of (a) y = 0.0 and (b) y = 0.3.
[0179] Fig. 13(c) to Fig. 13(f) show the evolution of optical training accuracy and loss for radially separated classification tasks. Quoted accuracy and loss values are averaged over final five iterations and illustrated by dashed lines.
[0180] Fig. 13(g) shows the evolution of the real and imaginary parts of the 4x4 dW weight gradient throughout optical training, for the y = 0.3 dataset.Fig. 13(h) shows the evolution of the complex weight elements throughout optical training, for the y = 0.3 dataset. Crosses and circles mark the initial and final weight values, respectively.
[0181] As captured by the t-distributed stochastic neighbour embedding (t-SNE) plots of Fig.
[0182] 13(a) and Fig. 13(b), with increasing y, the classes are more clearly separable, and the degree of nonlinearity required to resolve a boundary between classes reduces. The PNN shows impressive optical training results across the datasets (Fig. 13(c) to Fig. 13(f)), with training accuracy exceeding 91% within 15 epochs (90 iterations) for the dataset with largest separation parameter (y = 0.3). All training runs converge within the experimental window, with stable trajectories throughout. Fig. 13(g) and Fig. 13(h) survey the unitary weight evolution during optical training. Weight gradients decrease quite uniformly as the model converges, with individual weight elements (constrained by the unitary requirement) moving in the complex plane towards their optima.
[0183] On the four radial datasets, hybrid training produces comparable or worse training results than optical training. Crucially, both hybrid and optical training show strong learning capability and tolerance to the errors which afflict digital inference results. Finally, the importance of the saturable gain amplification nonlinearity is verified by comparing our results with simulations of the PNN model without any nonlinear layer. Average training and test accuracies on the radial datasets are low (50 - 57%) without the nonlinearity. We note that this simulation retains the \z212operation at the output layer, representing the optical intensity measurement.
[0184] Conclusion
[0185] This invention demonstrates an optical training solution for a photonic neural network structure that provides a potential pathway towards deep optical learning models. The ability of the PNN to learn its own systematic errors, which include optical loss and crosstalk, is key to surpassing the performance limitations which afflict the optical inference accuracy after digital training.
[0186] The realisation of large-scale PNNs remains partly limited by the footprint of photonic chips. While PNN layers might eventually be utilised as sections of transformer-like models, more useful applications lie in smaller machine vision models which are heavily deployed in resource-constrained edge devices, and which would benefit from the low latency and power consumption of a PNN (e.g., RepViT and MobileViT).The following numbered clauses define further optional configurations of the present disclosure. These are not the claims of the present application, which follow under the heading “CLAIMS”.
[0187] 1. A method of training a neural network using an optical neural network unit, the optical neural network unit comprising: an optical data processing unit configured to implement a linear layer of the neural network; and a non-linear optical element configured to implement a non-linear activation function of the neural network, the method comprising: performing a forward propagation to determine an error vector by transmitting an optical input signal through the optical neural network unit to multiply with a unitary matrix an input vector encoded by the optical input signal; and performing an error backpropagation by transmitting an optical error signal encoding the error vector through the optical neural network unit. 2. The method of clause 1, further comprising updating the unitary matrix based on the error backpropagation.
[0188] 3. The method of clause 2, further comprising repeating the steps of performing a forward propagation, performing an error backpropagation, and updating the unitary matrix until a pre-determined convergence criterion is met.
[0189] 4. The method of any of clauses 1 to 3, wherein performing the error backpropagation comprises reconfiguring the optical data processing unit such that the error vector is multiplied by an inverse matrix of the unitary matrix used during the forward propagation. 5. The method of any of clauses 1 to 4, wherein performing the error backpropagation further comprises: transmitting a probe signal into the optical neural network unit, the probe signal having a lower intensity than the optical input signal; and determining a derivative of the non-linear activation function based on a change in intensity of the probe signal caused by the non-linear optical element.
[0190] 6. The method of clause 5, wherein the probe signal differs in one or both of wavelength and polarisation from the optical input signal.
[0191] 7. The method of any of clauses 1 to 6, wherein the optical error signal is transmitted through the optical data processing unit in a direction opposite to a direction in which the optical input signal is transmitted through the optical data processing unit.
[0192] 8. An apparatus for training a neural network comprising an optical neural network unit, the optical neural network unit comprising: an optical data processing unit configured to implement a linear layer of the neural network by performing multiplication with a unitary matrix of an input vector encoded by an optical input signal; and a non-linear optical elementconfigured to implement a non-linear activation function of the neural network, wherein the apparatus is configured to: perform a forward propagation to determine an error vector by transmitting the optical input signal through the optical neural network unit; and perform an error backpropagation by transmitting an optical error signal encoding the error vector through the optical neural network unit.
[0193] 9. The apparatus of clause 8, wherein performing the error backpropagation comprises reconfiguring the optical data processing unit such that the error vector is multiplied by an inverse matrix of the unitary matrix used during the forward propagation.
[0194] 10. The apparatus of clause 8 or 9, wherein the input vector is a complex input vector encoded using the amplitude and phase of the optical input signal.
[0195] 11. The apparatus of any of clauses 8 to 10, wherein: the optical data processing unit is configured to produce an output optical signal encoding the result of the multiplication with the unitary matrix of the input vector; the optical data processing unit is configured to transmit a reference signal having a constant phase relative to a phase of the optical input signal; and the apparatus is configured to determine a phase of the output optical signal by comparison of the output optical signal and the reference signal.
[0196] 12. The apparatus of any of clauses 8 to 11, wherein the optical data processing unit comprises a photonic integrated circuit, for example a photonic waveguide mesh.
[0197] 13. The apparatus of clause 12, wherein the optical data processing unit comprises a plurality of Mach-Zender interferometers each having one or more reconfigurable phase shifters.
[0198] 14. The apparatus of any of clauses 8 to 13, wherein the non-linear optical element is an optical amplifier having a non-linear response function, for example an erbium-doped fibre amplifier.
[0199] 15. The apparatus of any of clauses 8 to 13, wherein the non-linear optical element is a saturable optical absorber.
[0200] 16. The apparatus of any of clauses 8 to 15, wherein the apparatus further comprises an encoder unit configured to receive one or more input beams and produce one or both of a) the optical input signal during the forward propagation, and b) the optical error signal during the error backpropagation.
[0201] 17. The apparatus of any of clauses 8 to 16, wherein: the optical input signal comprises a plurality of input sub-signals; and the optical data processing unit comprises a plurality of input channels, each input channel configured to receive one input sub -signal.18. The apparatus of any of clauses 8 to 17, wherein the apparatus comprises one or more photodetectors configured to detect an output of one or both of a) the optical neural network unit, and b) the optical data processing unit.
[0202] 19. The apparatus of any of clauses 8 to 18, wherein the apparatus comprises a plurality of optical neural network units implementing plural layers of the neural network.
[0203] REFERENCES
[0204] 1. Garcia-Martin, E., Rodrigues, C. F., Riley, G. & Grahn, H. Estimation of energy consumption in machine learning. Journal of Parallel and Distributed Computing 134, 75-88 (2019).
[0205] 2. Wetzstein, G. et al. Inference in artificial intelligence with deep optics and photonics. Nature 588, 39-47 (2020).
[0206] 3. Guo, X., Barrett, T. D., Wang, Z. M. & Lvovsky, A. I. B ackpropagation through nonlinear units for the all-optical training of neural networks. Photon. Res., PRJ 9, B71-B80 (2021).
[0207] 4. Spall, J., Guo, X. & Lvovsky, A. I. Training neural networks with end-to-end optical backpropagation. Submitted (2023).
[0208] 5. Shen, Y. et al. Deep learning with coherent nanophotonic circuits. Nature Photon 11, 441-446 (2017).
[0209] 6. Pai, S. et al. Experimentally realized in situ backpropagation for deep learning in photonic neural networks. Science 380, 398-404 (2023).
[0210] 7. Dhillon, A. & Verma, G. K. Convolutional neural network: a review of models, methodologies and applications to object detection. Prog Ar tif Intell 9, 85-112 (2020).
[0211] 8. Touvron, H. et al. LLaMA: Open and Efficient Foundation Language Models. Preprint at https: / / doi.org / 10.48550 / arXiv.2302.13971 (2023).
[0212] 9. Lin, X. et al. All-optical machine learning using diffractive deep neural networks. Science 361, 1004-1008 (2018).
[0213] 10. Mengu, D., Luo, Y., Rivenson, Y. & Ozcan, A. Analysis of Diffractive Optical Neural Networks and Their Integration With Electronic Neural Networks. IEEE Journal of Selected Topics in Quantum Electronics 26, 1-14 (2020).
[0214] 11. Duport, F., Schneider, B., Smerieri, A., Haelterman, M. & Massar, S. All-optical reservoir computing. Opt. Express, OE 20, 22783-22795 (2012).12. Taballione, C. etal. A universal fully reconfigurable 12-mode quantum photonic processor. Mater. Quantum. Technol. 1, 035002 (2021).
[0215] 13. Clements, W. R., Humphreys, P. C., Metcalf, B. J., Kolthammer, W. S. & Walmsley, I. A. Optimal design for universal multiport interferometers. Optica, OPTICA 3, 1460-1465 (2016).
Claims
CLAIMS1. A method of training a neural network using an optical neural network unit, the optical neural network unit comprising:an optical data processing unit configured to implement a linear layer of the neural network; anda non-linear optical element configured to implement a non-linear activation function of the neural network,the method comprising:performing a forward propagation to determine an error vector by transmitting an optical input signal through the optical neural network unit to multiply with a unitary matrix an input vector encoded by the optical input signal;performing an error backpropagation by transmitting an optical error signal encoding the error vector through the optical neural network unit;updating the unitary matrix based on the error backpropagation; andrepeating the steps of performing a forward propagation, performing an error backpropagation, and updating the unitary matrix until a pre-determined convergence criterion is met.
2. The method of claim 1, wherein performing the error backpropagation comprises reconfiguring the optical data processing unit such that the error vector is multiplied by an inverse matrix of the unitary matrix used during the forward propagation.
3. The method of claim 1 or 2, wherein performing the error backpropagation further comprises:transmitting a probe signal into the optical neural network unit, the probe signal having a lower intensity than the optical input signal; anddetermining a derivative of the non-linear activation function based on a change in intensity of the probe signal caused by the non-linear optical element.
4. The method of claim 3, wherein the probe signal differs in one or both of wavelength and polarisation from the optical input signal.
275. The method of any of claims 1 to 4, wherein the optical error signal is transmitted through the optical data processing unit in a direction opposite to a direction in which the optical input signal is transmitted through the optical data processing unit.
6. An apparatus for training a neural network comprising an optical neural network unit, the optical neural network unit comprising:an optical data processing unit configured to implement a linear layer of the neural network by performing multiplication with a unitary matrix of an input vector encoded by an optical input signal; anda non-linear optical element configured to implement a non-linear activation function of the neural network,wherein the apparatus is configured to:perform a forward propagation to determine an error vector by transmitting the optical input signal through the optical neural network unit;perform an error backpropagation by transmitting an optical error signal encoding the error vector through the optical neural network unit;update the unitary matrix based on the error backpropagation; andrepeat the steps of performing a forward propagation, performing an error backpropagation, and updating the unitary matrix until a pre-determined convergence criterion is met.
7. The apparatus of claim 6, wherein performing the error backpropagation comprises reconfiguring the optical data processing unit such that the error vector is multiplied by an inverse matrix of the unitary matrix used during the forward propagation.
8. The apparatus of claim 6 or 7, wherein the input vector is a complex input vector encoded using the amplitude and phase of the optical input signal.
9. The apparatus of any of claims 6 to 8, wherein:the optical data processing unit is configured to produce an output optical signal encoding the result of the multiplication with the unitary matrix of the input vector;the optical data processing unit is configured to transmit a reference signal having a constant phase relative to a phase of the optical input signal; andthe apparatus is configured to determine a phase of the output optical signal by comparison of the output optical signal and the reference signal.
10. The apparatus of any of claims 6 to 9, wherein the optical data processing unit comprises a photonic integrated circuit, for example a photonic waveguide mesh.
11. The apparatus of claim 10, wherein the optical data processing unit comprises a plurality of Mach-Zender interferometers each having one or more reconfigurable phase shifters.
12. The apparatus of any of claims 6 to 11, wherein the non-linear optical element is an optical amplifier having a non-linear response function, for example an erbium-doped fibre amplifier.
13. The apparatus of any of claims 6 to 11, wherein the non-linear optical element is a saturable optical absorber.
14. The apparatus of any of claims 6 to 13, wherein the apparatus further comprises an encoder unit configured to receive one or more input beams and produce one or both of a) the optical input signal during the forward propagation, and b) the optical error signal during the error backpropagation.
15. The apparatus of any of claims 6 to 14, wherein:the optical input signal comprises a plurality of input sub-signals; andthe optical data processing unit comprises a plurality of input channels, each input channel configured to receive one input sub-signal.
16. The apparatus of any of claims 6 to 15, wherein the apparatus comprises one or more photodetectors configured to detect an output of one or both of a) the optical neural network unit, and b) the optical data processing unit.
17. The apparatus of any of claims 6 to 16, wherein the apparatus comprises a plurality of optical neural network units implementing plural layers of the neural network.