Modular programmable n-channel interferometer
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2026-08-13
Smart Images

Figure RU2025000045_13082026_PF_FP_ABST
Abstract
Description
MODULAR PROGRAMMABLE N-CHANNEL INTERFEROMETER AREA OF TECHNOLOGY
[0001] The invention relates to methods for creating optical devices that perform linear multiplication operations of large matrices by vectors. The invention can be used in classical optical computing devices, for example, those solving problems in machine learning and artificial intelligence. The invention can be used in the implementation of individual elements of communication and computing networks serving a large number of subscribers and computing nodes; these elements and networks can be both classical and quantum. Furthermore, the invention can be used to create devices that analyze and synthesize multimode electromagnetic fields. LEVEL OF TECHNOLOGY
[0002] Transmitted or processed information can be encoded into optical signals. Optical elements are capable of performing linear operations on electromagnetic field amplitudes. Therefore, the use of optical methods for certain computational operations can improve the speed and / or energy efficiency of computations. Matrix-vector multiplication is one of the most computationally expensive and frequently used operations in modern computers, particularly in neural networks. To improve the efficiency of matrix-vector multiplication, it has been proposed to use optical devices—multichannel programmable interferometers.
[0003] The prior art discloses methods for creating universal N-channel programmable interferometers, as described in the works of C. W. Clements et al., "Optimal Design for Universal Multiport Interferometers," and M. Reck et al., "Experimental Realization of Any Discrete Unitary Operator." Interferometers created using these methods are capable of multiplying N x 1 vectors, encoded in the field amplitudes arriving at the interferometer inputs, by arbitrary N x N unitary matrices.
[0004] Fig. 1 illustrates the well-known Clements architecture. Fig. 1a and Fig. 1b show diagrams of 4- and 8-channel programmable interferometers. The interferometers consist of sequentially connected layers of Mach-Zehnder interferometers (MZIs). Fig. 1a shows one MZI layer 2 and one MZI element 3 within this layer.
[0005] In addition, the interferometers have a phase shift layer 4, which is not related to the MZI layers, and is located at the outputs of the interferometers. In Fig. 1a, one phase shift element 5 is also indicated, related to the phase shift layer 4. To perform the multiplication of vectors by the matrix set in the 4-channel programmable interferometer, the input field amplitudes 6, encoding the vector, are fed to its inputs 7. The field propagates through a 4-channel interferometer and the field amplitudes a) emerge from the output channels 8 out ^ 9, which are equal to the vector a^ п) , multiplied by the transfer matrix of the 4-channel interferometer.
[0006] The transfer matrix of an N-channel interferometer with the Clements architecture is set by programming all the components of the MZI and the phase shift layer 4. An example of MZI 3 is shown in Fig. 1c. This MZI contains two phase shifters (5a and 5b) and two identical balanced dividers 10. The MZI is programmed by setting the phase shift values 5a (the phase shift value β) and 5b (the phase shift value <ρ). In turn, by specifying the phase shift in each MZI, it is possible to program the transfer matrix of a multichannel interferometer, which consists of multiple MZIs.
[0007] In general, N-channel interferometers based on the Clements architecture are capable of multiplying N x 1 vectors by arbitrary N x N unitary matrices. This requires that the number of MZI layers be at least N, i.e., the depth of the interferometer circuit (the number of phase shift layers) be at least 2N. In addition to the MZI layers, the multichannel interferometer circuit also contains a layer of N or N – 1 phase shifts located at the inputs or outputs of the M-channel interferometer (in the examples shown in Fig. 1a and Fig. 1b, this layer is located at the output). This layer is not related to any MZI layer. In some cases, it can be removed from the circuit without losing the required functionality. A method for reducing the depth of N-channel interferometer circuits is known, proposed in the work of B. A. Bell and L. A. Walmsley “Further compactifying linear optical units” / / APL Photonics 6, 070804 (2021).This paper proposes a method for placing phase-shifting elements more densely in fewer layers while maintaining the interferometer's functionality. This method allows the depth of an N-channel interferometer with a Clements architecture to be reduced to D. CL = N + 2. The minimum number of phase shift elements in N-channel Clements interferometers that need to be programmed to set unitary matrices is P CL = N 2 , with an additional layer of phase shifts at the input or output, or P CL = N(N - 1) without an additional layer of phase shifts.
[0008] There are a number of other architectures of programmable M-channel interferometers capable of implementing a wide class of unitary matrices, for example, those proposed in the works of M. Reck et al., “Experimental realization of any discrete unitary operator” / / Phys. Rev. Lett. 73, 58 (1994) and S. A. Fldzhyan, M. Yu. Saygin and S. P. Kulik, “Optimal design of error-tolerant reprogrammable multiport interferometers” / / Opt. Lett. 45, No. 9, pp. 2632 (2020) and patent RU 2734454 "N-channel linear converter of electromagnetic signals" (Application: 2019132796, October 16, 2019). These devices exhibit scaling laws for the depth D and the number of parameters P, similar to those of the Clements architecture.
[0009] A disadvantage of these methods is the large number of programmable phase-shift elements in the interferometer circuit, which increases proportionally to the square of the number of channels N. This makes optical multipliers, even for relatively small values of N, large, complex to calibrate and program, and difficult to manufacture. Another disadvantage is the linear increase in circuit depth with increasing number of channels N. As a result, even for relatively small N, N-channel universal interferometers introduce high levels of loss into optical signals.
[0010] The prior art discloses methods for creating N-channel programmable interferometers capable of performing multiplications by non-unitary matrices, as described in the works of S. A. Fldzhyan, M. Yu. Saygin, and S. S. Straupe, "Low-depth, compact, and error-tolerant photonic matrix-vector multiplication beyond the unitary group" / / Optics Express https: / / doi.org / 10.1364 / OE.539666 and R. Tang et al., "Lower-depth programmable linear optical processors" / / Physical Review Applied v.21, 014054 (2024). All these architectures do not allow creating N-channel programmable interferometers capable of performing multiplications by arbitrary matrices of a given size, without a quadratic increase in the number of programmable phase shifts in N and without a linear increase in the depth of the circuits in N. [UN] The disadvantage of these methods is the large number of programmable phase-shift elements in the interferometer circuit, which increases proportionally to the square of the number of channels N. This makes optical multipliers, even for relatively small values of N, large, complex to calibrate and program, and difficult to manufacture. Another disadvantage is the linear increase in the circuit depth with increasing number of channels N. As a result, even for relatively small N 3N-channel universal interferometers introduce high levels of loss into optical signals.
[0012] Fig. 2 illustrates the implementation of matrix-vector multiplication using programmable multichannel converters. Multiplication by an M x M matrix W can be performed in N-channel converters for which N > M. The procedure for multiplying an M x 1 vector matrix W using an optical M-channel converter (interferometer) consists of the following sequence of steps: 1. in the N-channel interferometer 17, from the set of N input channels 12, a subset of N input channels 13 is selected into which the multiplied vector of optical amplitudes a^ will be fed in From the set N of output channels 14, a subset M of output channels I is selected, from which the resulting vector a^ will emerge. out 2. The interferometer is adjusted so that the M x M submatrix W of its transfer matrix U of size N x N (W is obtained from U by selecting columns and rows corresponding to the inputs and outputs selected in point 1) best corresponds to the given matrix IA°. For this, the values of the phase shifts described by the vector of values 0 applied to the optical signals in the interferometer are set to values that best correspond to the given matrix. The calculations of the phase shift values are performed on a traditional digital computer 15. The calculated phase shifts in the form of electrical signals are fed to the phase shift elements via electrical lines 16. 3. vector of optical amplitudes a^ 1П which encodes the input vector to be multiplied by the given matrix fed to a selected subset of the converter inputs. The result of multiplying this vector by a given matrix is obtained after the optical fields pass through the converter. The output vector of optical amplitudes a. ut = andA°)a( 1П ) is the result of multiplying the input vector by the given matrix. 4. If the next stage of the calculation, for example, the application of nonlinear activation functions in a neural network, is performed optically, then the amplitude vector a (out ^ is fed to the inputs of the next optical converter, which implements the corresponding stage of the calculation. If the next stage of the calculation is performed on an electronic computer, then the optical amplitudes a are measured. out and conversion of measurement results into digital format with sending data to a digital electronic computer.
[0013] Generalization of the procedure for multiplication by non-square matrices of size M out x M in with M in M out and max(Min , M out ) < is quite simple. For this, only part of the inputs and outputs of the N-channel interferometer can be used.
[0014] A technical challenge in creating optical multipliers using standard N-channel programmable interferometers is the large number of programmable phase-shift elements required to implement matrix multiplication even for relatively small values of N. For example, the largest programmable interferometer circuits have N = 64 inputs and outputs (https: / / hc32.hotchips.org / assets / program / conference / day2 / HotChips2020_ML_Inference_Lightmater.pdf). At the same time, modern computers solving, for example, artificial intelligence problems, including machine learning, require operations of multiplication by matrices of larger dimensions. For example, linear layers performing multiplication by matrices of size 1024 x 1024 and larger are often found in neural network algorithms (https: / / pytorch.org / tutorials / beginner / blitz / cifar 10_tutorial.html).The use of optical multiplication by such matrices using traditional N-channel interferometer architectures mentioned above leads to the need to create circuits containing P > 10. 6 programmable phase shifts. The depths of such circuits are D > 10 3 . This is beyond the current level of optical technology, namely, due to unattainable characteristics: 1) large circuit sizes, 2) high propagation losses, 3) complex programming of the interferometer, and 4) high power consumption of converters from electronic digital to optical analog formats and back (consumption of digital-to-analog and analog-to-digital converters) (see, for example, J T. Meech et al., “The data conversion bottleneck in analog computing accelerators” / / arxiv:2308.01719 (2023)). ESSENCE OF THE INVENTION
[0015] The claimed invention solves three problems of optical matrix-vector multipliers: 1) the technical complexity of creating universal programmable interferometers with a large number of channels capable of performing a wide class of linear transformations - operations of multiplying matrices by vectors, 2) the complexity of programming / tuning these interferometers due to the large number of programmable elements, and 3) high losses introduced into optical signals during propagation through these interferometers, due to the large depth of the circuits, which increases linearly with the number of channels N.
[0016] The technical result of the invention is a reduction in losses introduced into optical signals as they pass through the interferometer by reducing the depth of the optical multiplier circuits. Furthermore, the modular architecture of the interferometer contributes to the reduction in losses, as the optical signal passes through fewer layers of optical elements (e.g., phase shift layers). An additional technical result is a reduction in the number of programmable elements in optical multipliers compared to traditional universal interferometers.
[0017] The claimed technical result is achieved by implementing a modular programmable N-channel interferometer consisting of at least two layers of modules, wherein each layer of modules consists of n modules, and the number of channels of the modular programmable interferometer N = n 2 , where n > 3; each module has the same number of inputs and outputs, equal to at least and, moreover, the number of inputs and outputs in all modules is the same; each module from the layer of modules is connected to each module from the subsequent layer of modules, wherein the inputs of the modules of the first layer of modules are the inputs of the modular programmable interferometer, and the outputs of the modules of the last layer of modules are the outputs of the modular programmable interferometer; each module is a programmable interferometer that performs linear transformations on at least n field amplitudes supplied to the interferometer inputs; Each module, which is a programmable interferometer, consists of at least 2 layers of phase shifts.
[0018] In one particular implementation example, each module is programmed with a number of phase shifts equal to at least n 2 or and 2 — 1, or lying in the range from n log2n to n 2— 1, depending on the type of programmable interferometers used as modules. BRIEF DESCRIPTION OF DRAWINGS
[0019] Fig. 1 shows traditional programmable interferometer circuits used to implement matrix-vector multiplication: a) 4-channel interferometer with Clements architecture, b) 8-channel interferometer with Clements architecture, c) 2-channel Mach-Zehnder interferometer.
[0020] Fig. 2 shows a diagram explaining the implementation of matrix-vector multiplication using programmable interferometers.
[0021] Fig. 3 shows a diagram of a programmable interferometer with a modular architecture proposed in the present invention.
[0022] Fig. 4 shows examples of the implementation of modules in the form of programmable interferometers with a logarithmic dependence of the circuit depth on the number of channels.
[0023] Fig. 5 illustrates a possible method for increasing the number of programmable parameters in a programmable interferometer with the proposed modular architecture by serially connecting the modular interferometers shown in Fig. 3.
[0024] Fig. 6 illustrates a possible way to increase the number of programmable parameters in a modular interferometer by adding layers of modules.
[0025] Fig. 7 shows a diagram of an optoelectronic neural network that solves the problem of classifying handwritten digits.
[0026] Fig. 8 shows a graph of the dependence of the test accuracy of the neural network shown in Fig. 7 on the training epoch. IMPLEMENTATION OF THE INVENTION
[0027] The invention concerns converters of electromagnetic fields, which can belong to different wavelength ranges - from radio to optical.
[0028] For a more unambiguous understanding of the essence of the claimed solution, the main terms and definitions used within the framework of the present invention are presented below.
[0029] A conversion channel is any degree of freedom of an electromagnetic field that can be unambiguously assigned to an independent signal characterized by a complex amplitude. For clarity, we will separately identify the following degrees of freedom that can serve as a conversion channel: 1) A spatial degree of freedom of a field can be a mode of a single-mode waveguide through which an electromagnetic signal can propagate. In this case, several single-mode waveguides form multiple channels. A spatial channel can also be a single spatial mode of a multimode waveguide or a single mode of free space. In this case, several spatial modes of a single multimode waveguide or several modes of free space form multiple independent channels. 2) The frequency degree of freedom of the field is represented by individual lines of the frequency spectrum. A set of non-overlapping spectral lines forms a set of channels. PC17RU2025 / 000045 3) A temporal degree of freedom is a given time interval, characterized by the timing of its beginning and end. The presence of an electromagnetic signal in this time interval is interpreted as the presence of a signal in this channel. Thus, a set of time intervals forms a set of channels. Depending on which of the designated degrees of freedom of the electromagnetic field is used to encode information, we speak of spatial, frequency, and temporal channels.
[0030] A linear N-channel transform is a transform performed between N channels whose operation can be described by a linear law. Without loss of generality, the following description will focus on multiplying vectors by square matrices in which the input and output vectors have the same dimensions. In this case, a linear N-channel transform is written as: (oi) where N is the number of conversion channels, - complex amplitudes at the input of the transformation, a^ ut ^ - complex amplitudes at the output of the transformation. Here the indices j and k take values from 1 to N. In (0.1) the complex coefficients U k j form a matrix U of dimensions N x N, which determines a specific linear transformation. Expression (0.1) can be represented in matrix form: a (0Ut) = Ua (in), (0.2) where and a^ оиС>- columns composed of the amplitudes of the signals at the input and output of the conversion, respectively. The number of conversion channels N characterizes the dimension of the conversion. The invention relates to the case where the number of conversion channels N > 3. For multiplication by matrices whose size is less than N x N (specified by the number of inputs and outputs of the programmable N-channel converter) and for expanding the class of matrices available for multiplication (see, for example, S. A. Fldzhyan, M. Yu. Saygin, S. S. St. aipe, "Low-depth, compact and error-tolerant photonic matrix-vector multiplication beyond the unitary group" / / Optics Express https: / / doi.org / 10.1364 / OE.539666) only a part of the N inputs and N outputs are used. The number of used inputs M in and outputs M out is given by the size M out x M in multiplied matrix W. In what follows, without loss of generality, we assume that M in = M out = N.
[0031] The transformation transfer matrix, or simply the transformation matrix, is the matrix U that relates the column of amplitudes at the transformation output to the column at its input (see expression (0.2)).
[0032] An N-channel interferometer (N-channel linear converter or linear N-channel device or N-channel converter) is any device that performs linear N-channel conversion of electromagnetic signals.
[0033] The transfer matrix of an interferometer is the transformation transfer matrix implemented by a given interferometer.
[0034] A two-channel conversion unit (independent two-channel unit or simply conversion unit) is a two-channel element used to construct a multi-channel interferometer.
[0035] A divider is a two-channel element that performs a static transformation of two incoming field amplitudes, which is specified during the design and manufacturing process. The divider element is described by a 2 x 2 transfer matrix: / l 1t\ «»s = (h p)' < L3 > where p is the splitter's reflection coefficient, and m is the splitter's transmission coefficient. The splitter's reflection and transmission coefficients are related to each other by the ratio: p 2 + t 2 = 1, thus the transformation of one element of the divider is described by one parameter, for example, its transmittance coefficient t.
[0036] A phase shifter (phase shift element or phase shift element) is a single-channel element that adds a given phase b to the amplitude of the field a input to it. As a result, the field at the output of this element has an amplitude b = e 1вa. The present invention concerns interferometers that consist of static two-channel dividers and programmable phase shift elements.
[0037] A Mach-Zehnder interferometer (MZI) is a two-channel element consisting of two series-connected static divider elements with a balanced division ratio and one variable / programmable phase shift element located in one of the channels between the dividers. The MZI acts as a tunable divider, in which the division ratio can be changed by varying the phase shift. It serves as the core component of the conversion units in known methods for implementing universal and non-universal converters. An arbitrary transmission coefficient in the interferometer can only be achieved by precisely balancing the components of the static dividers.
[0038] The phase-shift layer of an N-channel interferometer (converter) is the portion of the conversion circuit that adds phase shifts to the signals entering the channels of this layer. The phase-shift layer allows for independent setting of phase-shift values in the N-1 channels of the layer. In N-channel interferometer implementations, the phase-shift layers are programmable, meaning that the phase-shift values of the devices can be changed after manufacture.
[0039] The static layer of an N-channel interferometer (converter) is the portion of the conversion circuit located between two successive layers of programmable elements, such as phase shift layers. In N-channel converter implementations, the static layers are not changed after fabrication.
[0040] An n-channel module, or simply an interferometer (converter) module, is a single-type programmable interferometer that forms the basis of a modular n-channel converter circuit, implemented using the proposed architecture. Each module is described by an n x n transfer matrix. The transfer matrix of each module can be set independently of the transfer matrices of other modules. The description below assumes that N = n. 2 .
[0041] A layer of n-channel modules, or a layer of modules, is a set of modules of an N-channel interferometer (converter) that act independently on a set of N input signals. The description below assumes that N = n. 2 , from which it follows that one layer of modules consists of n n-channel modules.
[0042] A cascade of interferometers (converters) is a series connection of these interferometers (converters), in which the outputs of the preceding interferometer (converter) are connected to the inputs of the following interferometer (converter). The inputs of the cascade are the inputs of the first interferometer (converter), and the outputs of the cascade are the outputs of the last interferometer (converter).
[0043] The circuit depth of an optical circuit is the minimum number of layers of elementary programmable elements, such as phase shifters, through which optical signals pass as they propagate from the inputs to the outputs of the circuit. The concept of circuit depth can be applied to any type of programmable linear optical circuit that contains programmable elements. In particular, we can talk about the circuit depth of a module, or the depth of the entire modular circuit. The depth of a modular circuit is equal to the sum of the depths of all the layers that make up the circuit.
[0044] A training epoch or learning epoch of a neural network is one complete cycle of training the neural network on the entire training data set.
[0045] Fig. 3 illustrates the proposed modular architecture of programmable interferometers implementing a linear optical layer for multiplying matrices by vectors. The circuit of the modular interferometer 17 consists of modules 18a, 18b,... 18b and 19a, 19b,... 19b, located in the first layer of modules 20a and in the second layer of modules 20b. The number of modules in each layer of modules is n.
[0046] Each module is a programmable interferometer that performs linear transformations on n field amplitudes fed to its inputs. Fig. 3 shows the inputs of the modules in the first layer of modules 20a: • inputs 21aa, 21ab,... 21av in the first module 18a of this layer, • inputs 21 ba, 2166,... 21bv in the second module 186 of this layer, • inputs 21va, 21vb,... 21vv in the last n-th module 18v of this layer.
[0047] Also in Fig. 3 the outputs of the modules in the second layer of modules 206 are marked: • outputs 22aa, 22ab,... 22av in the first module 19a of this layer, • outputs 22ba, 2266,... 22bv in the second module 196 of this layer, • outputs 22VA, 22VB,... 22VV in the last n-th module 19V of this layer.
[0048] The transformation of each module is described by transfer matrices of size n x n, where the superscript I denotes the layer to which the module belongs (I = 1 for modules in the module layer 20a and I = 2 for modules in the module layer 20b), and the subscript j denotes the module number in the layer (J = 1... n).
[0049] Total number of inputs N in and N outputs out in a modular interferometer is equal to n = ou t = N = p 2 . Thus, the transformation of a modular interferometer is described by a matrix U of size N x N = n. 2 x p 2 .
[0050] The outputs of the modules located in the first layer of modules 20a are connected to the inputs of the modules located in the second layer of modules 206 by a connecting layer 23. The connecting layer consists of optical connections that redirect the output optical signals from the outputs of the modules of the first layer of modules 20a to the inputs of the second layer of modules 206. This layer is described by a transfer matrix of permutations P of size N x N. One output of a module of the first layer of modules 20a corresponds to its connection with one input of the second layer of modules 206. In this case, the connections of the outputs of all modules of the first layer of modules 20a with the inputs of all modules of the second layer of modules 206 are organized in such a way that optical signals are received from all modules of the layer 206 to each module of the layer 206. For example, in Fig. 3 separately designated two connections lioutputs of the first module 18a of the first layer of modules 20a - 23a and 23b - connecting this module with the first module 19a and the last (n-th) module 19b of the second layer of modules.
[0051] The transfer matrix of a modular N-channel converter can be written as: U = W PW, (1.4) where = diagtW^,...,И^ 1 ' ) ) and = diag(W^ 2 \ И^ 2 *) - block-diagonal matrices describing the transformations of the first and second layers of modules.
[0052] The transfer matrix of the N-channel converter (1.4) is programmed by specifying the transfer matrices of the programmable modules.
[0053] In the modular architecture shown in Fig. 3, the number of programmable parameters and the circuit depth increase with increasing dimension N of the N x N multiplier transfer matrix more slowly than in traditional interferometers. Therefore, in modular optical circuits, the above-mentioned problems that hinder the creation of large-scale optical multipliers arise at higher matrix dimensions.
[0054] To estimate the number of parameters and the depth of the modular circuit, we will assume that each module in the modular circuit shown in Fig. 3 is implemented using the same architecture. If each n x n module contains P о programmable parameters (phase shifts), then the total number of parameters in the modular circuit is P mO duiar = 2пР0. In the case where universal unitary programmable n-channel interferometers are used as modules, for example, Clements interferometers (examples of circuits of 4- and 8-channel Clements interferometers are shown in Fig. 1a and Fig. 16), the number of parameters in the module is equal to P о = p 2 . Thus, the total number of programmable phase shifts in the circuit is: « And? = 2« 3 = (1.5)
[0055] To evaluate the advantages of a modular architecture composed of Clements interferometer modules, let's compare its characteristics with those of a single non-modular N-channel Clements interferometer. If a single N-channel Clements interferometer were used instead of a modular N-channel design, the number of programmable phase shifts in it would be equal to pClements > ^4 = at 2
[0056] The number of parameters in (1.6) turns out to be [N / 2 times greater than in a modular architecture with the same number of inputs and outputs N.
[0057] It is worth noting that the number of elements in universal Clements interferometers can be less by 1, since out of N 2 programmable shifts one is responsible for programming the overall phase of the transfer matrix, which often does not affect the observed values.
[0058] For specific values of N of interest to computing systems, particularly for matrix-vector multiplication, the reduction in the number of programmable parameters in modular designs is significant. For example, when multiplying by 1024 x 1024 matrices (N = 1024), the modular architecture contains 16 times fewer phase shifts than a non-modular universal architecture.
[0059] In addition to the number of programmable parameters, an important characteristic of multichannel interferometers is the depth of the circuit. The depth of the circuit determines the level of losses introduced during the propagation of optical signals through the circuit. The depth of the N-channel Clements interferometer (the minimum number of phase-shift layers) is (B. A. Bell and L. A. Walmsley, "Further compactifying linear optical unitaries" / / APL Photonics 6, 070804 (2021): D Clements = N + (1.7)
[0060] The depth of the modular circuit composed of n-channel (modules) Clements interferometers is equal to: = 2(n + 2) = 26 / N + 2). (1.8)
[0061] Thus, for N » 1, the depth of the modular circuit is ~ V / 2 times smaller than the depth of the N-channel Clements interferometer circuit.
[0062] Similar estimates can be made for other multichannel interferometer architectures discussed above (in the prior art). Depending on the choice of architecture for the universal modules that make up the modular N-channel converter circuit, the estimates for the number of programmable parameters (phase shifts) may differ by different coefficients, but the asymptotic behavior with increasing matrix dimension N remains the same as in the example considered.
[0063] The number of parameters and the depth of the modular circuit can be further reduced by choosing more economical architectures for the optical circuits of the modules. For this purpose, non-universal architectures of n-channel interferometers with fewer than n parameters can be used in modular circuits. 2 . As an example, in Fig. 4a and PC17RU2025 / 000045 Fig. 46 shows the optical schemes of 8- and 16-channel interferometers with the FFTUnitary architecture (K Tian et al., “Scalable and compact photonic neural chip with low learning-capability-loss” / / Nanophotonics 11(2), pp. 329 (2022)). These interferometers consist of 2 MZI layers. In an n-channel interferometer of this type, the number of programmable phase shifts is P FFTU = n log2n, and the depth of the circuit is D FFTU = ]og2п. In a modular scheme with such modules, the number of parameters is equal to Pmoduiar = 2n 2 log2л; the depth of the circuit is ^ modular =? log2П.
[0064] In n-channel interferometers that make up the modules, the number of phase layers and the number of parameters can be even lower than in the previous example of the FFTUnitary architecture, due to an even narrower class of possible transformations. For example, the interferometers could have only one MZI layer 2, rather than n layers, as in the case shown in Fig. 1.
[0065] Reducing the number of parameters in a linear layer that performs matrix-vector multiplication, such as in optical neural networks, can reduce their accuracy. This may necessitate increasing the number of parameters in the optical linear layer.
[0066] Fig. 5 shows an approach to increasing the number of parameters in an optical modular interferometer based on adding layers of modules to the interferometer. Each added layer of modules is connected to the previous one in the same way 23 as the second layer of modules is connected to the first. The outputs of the last layer of modules 20b are the outputs of the modular interferometer 14.
[0067] An alternative approach to increasing the accuracy of the linear optical layer is a cascade connection of modular interferometers, shown in Fig. 6, to form a new linear layer with a larger number of programmable parameters than in a single modular interferometer 17. In this case, the total number of programmable parameters in the cascade can still be less than in universal N-channel interferometers with N 2 parameters.
[0068] To illustrate the possibility of using the proposed modular architecture of multichannel programmable interferometers as matrix-vector multipliers, the problem of recognizing handwritten digits from 0 to 9 from the MNIST dataset (https: / / en.wikipedia.org / wiki / MNIST_database) by a neural network was considered, the block diagram of which is shown in Fig. 7. The input of the neural network is images of handwritten digits measuring 28 by 28 pixels. The images are stretched into a one-dimensional vector x = (% 1( ...,X 784) with dimensions of 784 x 1, where the value Xj encodes the gray level in the corresponding pixel of the image.PC17RU2025 / 000045
[0069] In the simulation, it was assumed that the image vectors (the index i denotes the image number in the data set) are encoded into the amplitudes of the optical signals arriving at the inputs of an optical programmable interferometer with a modular architecture 17, shown in Fig. 3, with the number of inputs and outputs N = 784 (n = 28). Each module was an n-channel universal Clements interferometer with n 2programmable phase shifts. At the outputs of the modular interferometer 17, the AbsReLU function was applied, which is the element-by-element application of the modulus to the complex field amplitudes output by the modular interferometer. After this, the vectors are fed to a digital real layer with 784 inputs and 10 outputs. This digital layer, assumed to be implemented on a traditional digital computer, performs matrix multiplication j ( dt 0 ltal ) of size 10 x 784 into its input vectors of size 784 X 1. After passing through the digital linear layer, the Softmax function (htps: / / en.wikipedia.org / wiki / Softmax_function) was applied to the vectors. The position of the maximum element in the vector (of length 10 X 1) at the output of the digital layer after applying the Softmax function gives the answer.
[0070] The neural network with a modular optical linear layer was trained using the PyTorch software package (https: / / pytorch.org / ). During training, the phase shift values of the modular interferometer were optimized, while the digital layer parameters remained constant and were randomly set before training. Figure 8 shows the dependence of the recognition accuracy of the neural network with an optical linear layer 17 on the training epoch. The graph also shows the dependence of the accuracy of the same neural network, but with a digital layer instead of an optical one. In this network, the multiplication performed in the optical linear layer is replaced by the standard multiplication of real matrices by vectors. This digital layer has 784 2= 614,656 trainable parameters, while the modular optical layer is 14 times smaller – 43,904 phase shift parameters. As can be seen from comparing the two dependences, the neural network with an optical layer achieves the same accuracy as the neural network with a digital layer.
[0071] The submitted application materials disclose preferred examples of the implementation of the technical solution and should not be interpreted as limiting other, particular examples of its implementation that do not go beyond the scope of the requested legal protection, which are obvious to specialists in the relevant field of technology.
Claims
FORMULA 1. A modular programmable N-channel interferometer consisting of at least two layers of modules, wherein each layer of modules consists of n modules, and the number of channels of the modular programmable interferometer N = n 2 , where n > 3; each module has the same number of inputs and outputs, equal to at least n, and the number of inputs and outputs in all modules is the same; each module from the layer of modules is connected to each module from the subsequent layer of modules, wherein the inputs of the modules of the first layer of modules are the inputs of the modular programmable interferometer, and the outputs of the modules of the last layer of modules are the outputs of the modular programmable interferometer; each module is a programmable interferometer that performs linear transformations on at least n field amplitudes supplied to the interferometer inputs; Each module, which is a programmable interferometer, consists of at least 2 layers of phase shifts.
2. A modular interferometer according to claim 1, characterized in that each module is programmable with a number of phase shifts equal to at least n 2 or p 2 — 1, or lying in the range from nlog n to p 2 — 1, depending on the type of programmable interferometers used as modules.