Optical circuit, forward propagation neural network, and regression neural network

WO2026203287A1PCT designated stage Publication Date: 2026-10-01NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/012770
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2026-10-01

Smart Images

  • Figure JP2025012770_01102026_PF_FP_ABST
    Figure JP2025012770_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This optical circuit has a structure in which a plurality of array layers (2-1 to 2-3), each of which includes a PSBS (1) that is a basic constituent unit, are cascade-connected, and the plurality of array layers (2-1 to 2-3) are butterfly-connected by waveguides (3-1, 3-2). Each PSBS (1) is configured from a two-input two-output beam splitter (11) and a programmable phase shifter (10) inserted into any one of four signal paths of the two inputs and the two outputs of the beam splitter (11).
Need to check novelty before this filing date? Find Prior Art

Description

Optical circuits, feedforward neural networks, and recurrent neural networks

[0001] The present invention relates to an optical circuit that uses light as an information medium, and more particularly to a circuit that performs analog multiplication.

[0002] Conventionally, in neural networks using silicon photonics, when multiplying a complex vector by a unitary matrix, the complex vector was encoded into a signal using light as the information medium and input into a unitary matrix circuit that could be set to any value to perform analog vector-matrix multiplication. This unitary matrix circuit is a mesh structure circuit with a Mach-Zehnder interferometer (MZI) as its basic building block (Non-Patent Documents 1-3). Hereafter, the Clements-type mesh structure described in Non-Patent Document 2 will be used as an example of a mesh structure. A butterfly structure with an MZI as its basic building block has also been reported (Non-Patent Document 4).

[0003] Figure 25 shows an example of the MZI configuration. The 2-input, 2-output MZI9 has two programmable phase shifters (PS) 10-1, 10-2 and two beam splitters (BS) 11-1, 11-2 to configure a variable interferometer. 1 , x 2 is the input complex vector, y 1 , y 2 is the output complex vector. The phase shift amount of PS10-1 and 10-2 is φ, and the branching ratio of BS11-1 and 11-2 is 0.5:0.5.

[0004] A mesh structure with MZI9 as the basic building block requires N layers of N-input MZI arrays to realize an arbitrary N×N unitary matrix. Figure 26 shows the configuration of a linear circuit of a mesh structure that realizes a 4×4 unitary matrix W when the number of inputs N=4. 200-1 to 200-4 are MZI array layers, and 201 is a unitary diagonal matrix layer. X is the input light, and Z is the output light.

[0005] In order to reduce the number of elements constituting a circuit and the number of parameters related to variability, it is conceivable to reduce the number of MZI array layers. However, when the number of MZI array layers is reduced, there has been a problem that the reachability that an arbitrary input signal can reach all output terminals cannot be satisfied. Conventional problems will be described with reference to FIGS. 27 and 28. FIG. 27 shows the configuration of a linear circuit having a mesh structure that implements an 8×8 unitary matrix W when the number of inputs N=8. On the other hand, FIG. 28 shows the configuration of a linear circuit having a mesh structure when the number of MZI array layers is reduced. 200-1 to 200-8 are MZI array layers, and 201 is a unitary diagonal matrix layer.

[0006] In the configuration of FIG. 27, an optical signal can propagate only from the left side to the right side, and when an optical signal can propagate to either of the two outputs in MZI 9, any input terminal (signal x 1 to x 8 ) can reach all output terminals (signal y 1 to y 8 ). Reference numeral 202 in FIG. 27 indicates a path through which signal x 1 reaches the terminal of signal y 8 . On the other hand, in the configuration of FIG. 28 in which the number of MZI array layers is reduced to 4, there are output terminals that the optical signal cannot reach. For example, signal x 1 can reach the terminals of signal y 1 to y 5 , but cannot reach the terminals of signal y 6 to y 8 .

[0007] M.Reck, A.Zeilinger, HJBernstein, and P.Bertani, “Experimental realization of any discrete unitary operator”, Physical Review Letters, vol.73, no.1, pp.58-63, 1994 W.R.Clements, PC Humphreys, BJMetcalf, WSKolthammer, and IAWalmsley, “Optimal design for universal multiport interferometers”, Optica, vol.3, no.12, pp.1460-1465, 2016 Y. Shen, NC Harris, S. Skirlo, M. Prabhu, T. Baehr-Jones, M. Hochberg, X. Sun, S. Zhao, H. Larochelle, D. Englund, and M. Soljacic, “Deep learning with coherent nanophotonic circuits”,Nature Photonics, vol.11, pp.441-447, 2017 M.Y.-S.Fang, S.Manipatruni, C.Wierzynski, A.Khosrowshahi, and MRDeWeese, “Design of optical neural networks with component imperfections”, Optics Express, vol.27, no.10, pp.14009-14029, 2019

[0008] The present invention was made to solve the above problems, and aims to provide an optical circuit, a feedforward neural network, and a recurrent neural network that can satisfy signal reachability while reducing the number of elements and parameters.

[0009] The optical circuit of the present invention has a structure in which a plurality of array layers, each containing one or more basic structural units, are connected in cascading order, and the plurality of array layers are butterfly-connected by waveguides, and the basic structural unit is characterized by comprising a beam splitter with two inputs and two outputs, and a programmable phase shifter inserted at any one of the four signal paths, namely the two inputs and two outputs of the beam splitter.

[0010] According to the present invention, multiple array layers, each containing one or more basic structural units, are butterfly-connected by waveguides. Each basic structural unit consists of a 2-input, 2-output beam splitter and a programmable phase shifter inserted at one of the four signal paths (2 inputs and 2 outputs) of the beam splitter. This reduces the number of elements and parameters while satisfying the requirement for signal reachability.

[0011] Figure 1 is a block diagram showing the configuration of the basic component unit of the present invention. Figure 2 is a block diagram showing the configuration of the optical circuit according to the present invention. Figure 3 is a block diagram showing the configuration of the optical circuit according to the present invention. Figure 4 is a block diagram showing the configuration of the optical circuit according to the present invention. Figure 5 is a block diagram showing the configuration of the optical circuit according to the present invention. Figure 6 is a block diagram showing the configuration of a feedforward neural network according to the first embodiment of the present invention. Figure 7 is a diagram showing the relationship between training accuracy and the number of training iterations in a conventional configuration and the configuration of the first embodiment of the present invention. Figure 8 is a diagram showing the relationship between the loss function value and the number of training iterations in a conventional configuration and the configuration of the first embodiment of the present invention. Figure 9 is a diagram showing the relationship between test accuracy and the total number of parameters in a conventional configuration and the configuration of the first embodiment of the present invention. Figure 10 is a diagram showing the results of training accuracy, test accuracy and loss function value in a conventional configuration. Figure 11 is a diagram showing the results of training accuracy, test accuracy and loss function value in the configuration of the first embodiment of the present invention. Figure 12 is a block diagram showing the configuration of a regressive neural network according to the second embodiment of the present invention. Figure 13 is a diagram illustrating the processing of the patch image generation unit according to the second embodiment of the present invention. Figure 14 is a diagram showing the relationship between training accuracy and the number of training iterations in a conventional configuration and the configuration of the second embodiment of the present invention. Figure 15 shows the relationship between the loss function value and the number of training iterations in a conventional configuration and the configuration of the second embodiment of the present invention. Figure 16 shows the relationship between the test accuracy and the total number of parameters in a conventional configuration and the configuration of the second embodiment of the present invention. Figure 17 shows the relationship between the test accuracy and the total number of parameters in another conventional configuration. Figure 18 is a block diagram showing the configuration of an optical circuit according to the third embodiment of the present invention. Figure 19 is a block diagram showing the configuration of an optical circuit according to the fourth embodiment of the present invention. Figure 20 is a block diagram showing the configuration of an optical circuit according to the fifth embodiment of the present invention. Figure 21 is a block diagram showing another configuration of the optical circuit according to the fifth embodiment of the present invention. Figure 22 is a block diagram showing the configuration of a feedforward neural network according to the sixth embodiment of the present invention. Figure 23 is a block diagram showing the configuration of an optical circuit according to the seventh embodiment of the present invention. Figure 24 is a block diagram showing another configuration of the optical circuit according to the seventh embodiment of the present invention.Figure 25 is a block diagram showing the configuration of a Mach-Zehnder interferometer. Figure 26 is a block diagram showing the configuration of a mesh-structured linear circuit. Figure 27 is a block diagram showing another configuration of a mesh-structured linear circuit. Figure 28 is a block diagram showing the configuration of a mesh-structured linear circuit when the number of MZI array layers is reduced.

[0012] [Principle of the Invention] In this invention, in order to reduce the number of elements, the basic structural unit (hereinafter referred to as PSBS) is a pair of PS10 and BS11 with two inputs and two outputs, as shown in Figure 1. PS10 is inserted in the middle of one of the four signal paths (arm waveguides 12-1 to 12-4) of the two inputs and two outputs of BS11. In the following description, the configuration in which PS10 is inserted in the middle of arm waveguide 12-1 of the two inputs of BS11 will be described. 1 , x 2 is the input complex vector, y 1 , y 2 is the output complex vector. The phase shift amount of PS10 is φ, and the branch ratio of BS11 is 0.5:0.5. The 2-input, 2-output PSBS1 is no longer an interferometer on its own. Thus, in this invention, circuits that are not interferometers on their own are used as the basic configuration units.

[0013] The circuit structure, which combines the basic constituent units, uses a butterfly structure rather than a mesh structure. Examples of optical circuits using the butterfly structure of the present invention are shown in Figures 2 to 5. The example in Figure 2 shows a Decimation-In-Time (DIT) connection of an 8-input, 8-output butterfly structure. The optical signals passing through PSBS1 are paired. The configuration in Figure 2 is the same as that of a radix 2 DIT fast Fourier transform (FFT). Decimation-In-Frequency (DIF) can also be used in the same way. In the case of DIF, the input and output of the DIT are inverted.

[0014] Figure 3 shows an example layout of an 8-input, 8-output butterfly structure. When PSBS array layers 2-1 to 2-3 are arranged regularly, and PSBS array layers 2-1 and 2-3 are butterfly-connected by waveguide 3-1 according to the connection description in Figure 2, and PSBS array layers 2-2 and 2-3 are butterfly-connected by waveguide 3-2, the circuit layout becomes as shown in Figure 3. Each of the cascaded PSBS array layers 2-1 to 2-3 consists of one or more PSBS1. Pairs of optical signals are processed by PSBS1. In the configuration of Figure 3, any input signal can reach all output terminals.

[0015] Multiple combinations of signal pairs and PSBS1 exist based on the connection description in Figure 2. Figure 4 shows an example in which the arrangement of PSBS1 within PSBS array layer 2-2 is swapped while maintaining the connection between PSBS array layer 2-2 and PSBS array layer 2-1 in Figure 3. Figure 5 shows an example in which the arrangement of PSBS1 within PSBS array layer 2-3 is swapped while maintaining the connection between PSBS array layer 2-2 and PSBS array layer 2-1 in Figure 3. Hereinafter, in this invention, a butterfly structure with PSBS1 as the basic constituent unit will be referred to as a PSBS-butterfly structure.

[0016] Conventional mesh structures, such as the triangular structure in Non-Patent Document 1 and the rectangular structure in Non-Patent Document 2, are polygonal meshes with MZIs (Multiple Waveguides) placed at the vertices of the polygons. Waveguides connecting the MZIs are placed at the sides of the polygons and do not intersect with other waveguides. On the other hand, in the PSBS-butterfly structure of the present invention, some of the waveguides connecting the PSBS1s intersect. The present invention uses a PSBS-butterfly structure or a hybrid structure of the PSBS-butterfly structure and an existing structure.

[0017] [First Embodiment] Figure 6 is a block diagram showing the configuration of a feedforward neural network (FNN) according to the first embodiment of the present invention. In this embodiment, an example is described in which the FNN identifies an image. The input image 103 is a grayscale image with a size of 28 pixels in the ax direction (horizontal direction) x 28 pixels in the ay direction (vertical direction), where one pixel is represented by 256 gradations. The FNN comprises a preprocessing unit 100, a complex number processing unit 101, and a real number processing unit 102.

[0018] The preprocessing unit 100 includes a vector transformation unit 1000, a normalization unit 1001, an FFT unit 1002, and a vector generation unit 1003. The complex number processing unit 101 includes a hidden layer unit 1010 that takes a complex number output from the preprocessing unit 100 as input, a nonlinear function unit 1011 that takes a complex number output from the hidden layer unit 1010 as input, a hidden layer unit 1012 that takes a complex number output from the nonlinear function unit 1011 as input, a nonlinear function unit 1013 that takes a complex number output from the hidden layer unit 1012 as input, and a hidden layer unit 1014 that takes a complex number output from the nonlinear function unit 1013 as input. The real number processing unit 102 includes a vector transformation unit 1020, a softmax function unit 1021, and an error calculation unit 1022.

[0019] The vector conversion unit 1000 of the preprocessing unit 100 converts the input image 103 into a 784-dimensional non-negative integer vector. The normalization unit 1001 normalizes the integer elements of the non-negative integer vector output from the vector conversion unit 1000, from 0 to 255, to real values ​​from 0 to 1. The FFT unit 1002 converts the real values ​​output from the normalization unit 1001 into a 393-dimensional complex vector using FFT. When converting real numbers, half of the values, excluding the DC component, are complex conjugates.

[0020] The vector generation unit 1003 acquires the power of the complex vector output from the FFT unit 1002 and reduces the dimensionality by selecting k nodes (vector dimensions) in descending order of power. The vector generation unit 1003 then uses the complex vector consisting of the outputs of the selected nodes as the feature vector for input to the complex number processing unit 101. Here, k = 256.

[0021] In this embodiment, a PSBS-butterfly structure is used as the optical circuit constituting each of the hidden layer sections 1010, 1012, and 1014 of the complex number processing unit 101. All three hidden layer sections 1010, 1012, and 1014 have the same configuration. The optical circuit of the PSBS-butterfly structure has one, two, three, or four blocks per hidden layer. Here, a block is, in the case of the number of inputs N, log 2 This treats the PSBS thin layer of (N) as a single unit, and one block can satisfy the requirement for signal reachability. The phase φ of PS10 of each basic constituent unit of the optical circuit is a parameter to be optimized. The phase φ is optimized by learning through backpropagation. In this invention, a layer with PSBS1 as the basic constituent unit (PSBS array layer) is called a thin layer to distinguish it from one with MZI as the basic constituent unit.

[0022] The nonlinear function units 1011 and 1013 are inputs of the complex number z output from the hidden layer units 1010 and 1012, respectively. The nonlinear function units 1011 and 1013 output a value modReLU(z) as shown in equation (1). Here, the bias value, a real number b, is the parameter to be optimized. The real number b is optimized through learning.

[0023]

[0024] The vector conversion unit 1020 of the real number processing unit 102 converts complex numbers into real number vectors by acquiring the power of the complex numbers output from the hidden layer unit 1014. The softmax function unit 1021 applies the softmax function to the real number vectors output from the vector conversion unit 1020. The error calculation unit 1022 calculates the cross-entropy loss between the correct image recognition result (Target label) and the image recognition result by the FNN.

[0025] The results of this embodiment are described below. First, Figure 7 shows the relationship between training accuracy and the number of training iterations (epochs) in the conventional configuration and the configuration of this embodiment. In the conventional configuration, a mesh structure with MZI as the basic structural unit (hereinafter referred to as the MZI-mesh structure) or a butterfly structure with MZI as the basic structural unit (hereinafter referred to as the MZI-butterfly structure) was used as the optical circuit constituting each of the hidden layers 1010, 1012, and 1014. All three hidden layers 1010, 1012, and 1014 had the same configuration. The number of MZI layers per hidden layer was set to one of 4, 32, 64, 96, 128, 192, or 256. However, when the number of MZI layers was 256, a unitary diagonal matrix layer consisting of a phase shifter was added to the MZI layer to realize an arbitrary unitary matrix, and the number of variable parameters per hidden layer was set to N. 2 = 65536.

[0026] Figure 7, section 700, shows the results when the number of blocks per hidden layer is 8 in this embodiment. The total number of parameters is 25088, and the training accuracy per 100 training iterations is 0.934. Although not shown in Figure 7, when the number of blocks is 7, the total number of parameters is 22016, and the training accuracy per 100 training iterations is 0.931.

[0027] Figure 7, 701 shows the results when an MZI-butterfly structure is used for the hidden layers 1010, 1012, and 1014. The number of MZI layers per hidden layer is 4, the total number of parameters is 25088, and the training accuracy per 100 training iterations is 0.913. Figures 7, 702, and 703 show the results when an MZI-mesh structure is used for the hidden layers 1010, 1012, and 1014. In the case of 702, the number of MZI layers per hidden layer is 256, the total number of parameters is 197120, and the training accuracy per 100 training iterations is 0.815. In the case of 703, the number of MZI layers per hidden layer is 32, the total number of parameters is 24993, and the training accuracy per 100 training iterations is 0.533.

[0028] According to the present example, high training accuracy can be achieved with a smaller number of parameters for a configuration using an MZI-mesh structure. Also, in the present example, high training accuracy can be achieved with the same number of parameters for a configuration using an MZI-butterfly structure. Furthermore, learning according to the present example was more stable compared to a configuration using an MZI-mesh structure.

[0029] FIG. 8 shows the relationship between the loss function value and the number of training epochs in the conventional configuration and the configuration of the present example. Reference numeral 800 in FIG. 8 indicates the result when the number of blocks per hidden layer is set to 8 in the present embodiment. The loss function value per 100 training epochs is 0.191. Although not shown in FIG. 8, when the number of blocks is 7, the loss function value per 100 training epochs is 0.196.

[0030] Reference numeral 801 in FIG. 8 indicates a result when an MZI-butterfly structure is used for the hidden layers 1010, 1012, and 1014. The number of MZI layers per hidden layer is 4, and the loss function value per 100 training epochs is 0.242. Reference numerals 802 and 803 in FIG. 8 indicate results when an MZI-mesh structure is used for the hidden layers 1010, 1012, and 1014. In the case of 802, the number of MZI layers per hidden layer is 256, and the loss function value per 100 training epochs is 0.509. In the case of 803, the number of MZI layers per hidden layer is 32, and the loss function value per 100 training epochs is 1.358.

[0031] According to the present embodiment, the value of the loss function, which is the optimization objective function, can be reduced compared to the case where an MZI-butterfly structure is used. Also, in the present embodiment, the value of the loss function can be significantly reduced compared to the case where an MZI-mesh structure is used.

[0032] Figure 9 shows the relationship between test accuracy and the total number of parameters in the conventional configuration and the configuration of this embodiment. Figure 900 shows the results for this embodiment. Figure 901 shows the results when an MZI-butterfly structure is used for the hidden layers 1010, 1012, and 1014. Figures 902 to 907 show the results when an MZI-mesh structure is used for the hidden layers 1010, 1012, and 1014.

[0033] In the case of 902, there are 256 MZI layers per hidden layer, a total of 197,120 parameters, and a test accuracy of 0.820 per 100 training iterations. In the case of 903, there are 192 MZI layers per hidden layer, a total of 147,392 parameters, and a test accuracy of 0.829 per 100 training iterations. In the case of 904, there are 128 MZI layers per hidden layer, a total of 98,432 parameters, and a test accuracy of 0.818 per 100 training iterations. In the case of 905, there are 96 MZI layers per hidden layer, a total of 73,952 parameters, and a test accuracy of 0.784 per 100 training iterations. In the case of 906, there are 64 MZI layers per hidden layer, a total of 49,472 parameters, and a test accuracy of 0.716 per 100 training iterations. In the case of 907, the number of MZI layers per hidden layer is 32, the total number of parameters is 24,993, and the test accuracy per 100 training iterations is 0.545.

[0034] According to the present embodiment, high test accuracy can also be achieved with a small number of parameters. Among the configurations using the MZI-mesh structure, a configuration with 256 MZI layers per hidden layer section has the same number of parameters as the degree of freedom of a 256×256 unitary matrix. A configuration with 192, 128, 96, 64, or 32 MZI layers per hidden layer section has a smaller number of parameters than the degree of freedom of a 256×256 unitary matrix. A configuration having one hidden layer and 128, 96, 64, or 32 MZI layers in the hidden layer cannot satisfy reachability, meaning any input signal can reach all output terminals. A configuration having 64 or 32 MZI layers in the hidden layer cannot satisfy reachability even when three hidden layers are connected. On the other hand, in the present embodiment, reachability can be satisfied even when the number of blocks per hidden layer section is one.

[0035] Figure 10 shows the results of training accuracy, test accuracy, and loss function values for a configuration using an MZI-butterfly structure as the three hidden layer sections 1010, 1012, and 1014. Figure 11 shows the results of training accuracy, test accuracy, and loss function values for the configuration of the present embodiment. In the present embodiment, when the number of parameters is large, test accuracy comparable to or higher than that of a configuration using an MZI-butterfly structure can be achieved. In addition, in the present embodiment, since the number of blocks per hidden layer section can be adjusted with a granularity that is 1 / 2 that when using the MZI-butterfly structure, the number of parameters can be finely controlled.

[0036] [Second Embodiment] FIG. 12 is a block diagram showing the configuration of a recurrent neural network (RNN) according to a second embodiment of the present invention. In the present embodiment, an example in which an RNN identifies an image will be described. The input image 103 is a grayscale image having a size of 28 pixels in the ax direction (horizontal direction) × 28 pixels in the ay direction (vertical direction), with one pixel represented by 256 gray levels. The RNN includes a preprocessing unit 100a, a complex number processing unit 101a, and a real number processing unit 102.

[0037] The preprocessing unit 100a includes a normalization unit 1004, a patch image generation unit 1005, a vector generation unit 1006, an FFT unit 1007, a vector generation unit 1008, and a time series data generation unit 1009. The complex number processing unit 101a includes an input layer unit 1015 that takes a complex number output from the preprocessing unit 100a as input, a hidden layer unit 1016 that takes a complex number output from a nonlinear function unit (described later) as input, an adder unit 1017 that adds the outputs of the input layer unit 1015 and the hidden layer unit 1016, a nonlinear function unit 1018 that takes a complex number output from the adder unit 1017 as input, and an output layer unit 1019 that takes a complex number output from the nonlinear function unit 1018 as input. The configuration of the real number processing unit 102 is the same as in the first embodiment.

[0038] The normalization unit 1004 of the preprocessing unit 100a normalizes the value of each pixel in the input image 103 to a real value between 0 and 1. As shown in Figure 13, the patch image generation unit 1005 generates six patch images in the ax direction and six in the ay direction by sliding a window WI of a predetermined size (8 pixels in the ax direction x 8 pixels in the ay direction) 4 pixels at a time on the input image 103 and extracting images of a predetermined size. As a result, 36 patch images with a size of 8 pixels in the ax direction x 8 pixels in the ay direction are generated.

[0039] The vector generation unit 1006 performs a scan of one line in the ax direction of the patch image, acquiring the value of each pixel from the left end to the right end. After completing the scan in the ax direction, it performs the same scan on a line shifted by one pixel in the ay direction. By performing this process for each line in the ax direction, the patch image is made one-dimensional and a 64-dimensional vector is generated.

[0040] Furthermore, the vector generation unit 1006 performs a scan of one line in the ay direction of the patch image, acquiring the value of each pixel from the top to the bottom. After the scan in the ay direction is completed, it performs the same process on a line shifted by one pixel in the ax direction. By performing this process for each line in the ay direction, the patch image is made one-dimensional and a 64-dimensional vector is generated. The vector generation unit 1006 performs the above 64-dimensional vector generation process for each patch image.

[0041] The FFT unit 1007 converts the 64-dimensional vector obtained by scanning in the ax direction and the 64-dimensional vector obtained by scanning in the ay direction into 33-dimensional complex vectors using FFT. The FFT unit 1007 performs the FFT processing for each patch image.

[0042] The vector generation unit 1008 acquires the power of the complex vector output from the FFT unit 1007 and reduces the dimensionality by selecting 32 nodes (vector dimensions) in descending order of power. The vector generation unit 1008 then uses the complex vector consisting of the outputs of the selected nodes as a 32-dimensional feature vector. The vector generation unit 1008 performs the 32-dimensional feature vector generation process for each 64-dimensional vector and for each patch image.

[0043] Next, the vector generation unit 1008 combines the 32-dimensional feature vector obtained by scanning in the ax direction and selecting the vector dimension with the 32-dimensional feature vector obtained by scanning in the ay direction and selecting the vector dimension for each identical patch image to generate a 64-dimensional feature vector. The vector generation unit 1008 performs the 64-dimensional feature vector generation process for each patch image. Through the above process, 36 64-dimensional feature vectors are generated from the input image 103.

[0044] The time-series data generation unit 1009 generates one of the following from the 64-dimensional feature vector output from the vector generation unit 1008: a 256-dimensional feature vector with a time-series length of 9, a 128-dimensional feature vector with a time-series length of 18, or a 64-dimensional feature vector with a time-series length of 36.

[0045] In this embodiment, a PSBS-butterfly structure is used as the optical circuit constituting the input layer 1015, the hidden layer 1016, and the output layer 1019 of the complex number processing unit 101a. In the PSBS-butterfly structure optical circuit, the number of blocks in the input layer 1015, the hidden layer 1016, and the output layer 1019 is set to 1, 2, 3, or 4. Here, a block is defined as log in when the number of inputs is N. 2 The PSBS layer of (N) is treated as one unit, and the signal reachability can be satisfied with one block. However, for the hidden layer 1016, log 2(N) A unitary diagonal matrix layer consisting of a phase shifter was added after the thin layer. The phase φ of PS10 of each basic component unit of the optical circuit is a parameter to be optimized. The phase φ is optimized by learning through backpropagation.

[0046] The addition unit 1017 adds the complex number output from the input layer unit 1015 and the complex number output from the hidden layer unit 1016. The nonlinear function unit 1018 takes the complex number z output from the addition unit 1017 as input and outputs a value modReLU(z) as shown in equation (1). The hidden layer unit 1016 takes modReLU(z) output from the nonlinear function unit 1018 as input. The operation of the real number processing unit 102 is the same as in the first embodiment.

[0047] The results of this embodiment are described below. First, Figure 14 shows the relationship between training accuracy and the number of training iterations (epochs) in the conventional configuration and the configuration of this embodiment. In the conventional configuration, an MZI-mesh structure was used as the optical circuit constituting the input layer 1015, the hidden layer 1016, and the output layer 1019, respectively. In both the conventional configuration and the configuration of this embodiment, the number of dimensions of the input vector to the input layer 1015 was 256, the time series length of the input vector was 9, and each layer of the input layer 1015, the hidden layer 1016, and the output layer 1019 was configured to realize a 256 × 256 unitary matrix. When using the MZI-mesh structure, the input layer 1015, the hidden layer 1016, and the output layer 1019 all had the same configuration. The input layer 1015, the hidden layer 1016, and the output layer 1019 each have 256 MZI layers, and a unitary diagonal matrix layer consisting of a phase shifter is added to the MZI layers to realize any unitary matrix.

[0048] Figure 14, section 1400 shows the results in this embodiment when the number of blocks in the input layer 1015, hidden layer 1016, and output layer 1019 is set to 4. The total number of parameters is 9728, and the training accuracy per 100 training iterations is 0.896. Figure 14, section 1401 shows the results in this embodiment when the number of blocks is set to 1. The total number of parameters is 3584, and the training accuracy per 100 training iterations is 0.850. Although not shown in Figure 14, when the number of blocks is set to 2, the total number of parameters is 5632, and the training accuracy per 100 training iterations is 0.887.

[0049] Figure 14, section 1402, shows the results when an MZI-mesh structure is used for the input layer 1015, the hidden layer 1016, and the output layer 1019. The number of MZI layers in the input layer 1015, the hidden layer 1016, and the output layer 1019 is 256, the total number of parameters is 196,864, and the training accuracy per 100 training iterations is 0.829. According to this embodiment, a high training accuracy could be achieved with a smaller number of parameters compared to when an MZI-mesh structure is used.

[0050] Figure 15 shows the relationship between the loss function value and the number of training iterations (epochs) in the conventional configuration and the configuration of this embodiment. Figure 1500 shows the results when the number of blocks in the input layer 1015, hidden layer 1016, and output layer 1019 is set to 4 in this embodiment. The loss function value per 100 training iterations is 0.351. Figure 1501 shows the results when the number of blocks is set to 1 in this embodiment. The loss function value per 100 training iterations is 0.401. Although not shown in Figure 15, the loss function value per 100 training iterations when the number of blocks is set to 2 is 0.377.

[0051] Figure 15, section 1502, shows the results when an MZI-mesh structure is used for the input layer 1015, the hidden layer 1016, and the output layer 1019. The number of MZI layers in the input layer 1015, the hidden layer 1016, and the output layer 1019 is 256, and the loss function value per 100 training iterations is 0.524. According to this embodiment, the value of the loss function can be reduced compared to when an MZI-mesh structure is used.

[0052] Figure 16 shows the relationship between test accuracy and the total number of parameters in the conventional configuration and the configuration of this embodiment. Figures 1600 to 1611 in Figure 16 show the results for this embodiment. Figures 1600 to 1603 show the results when the number of dimensions of the input vector to the input layer 1015 is 256 and the time series length of the input vector is 9. In the case of 1600, the number of blocks in the input layer 1015, the hidden layer 1016, and the output layer 1019 is 4, the total number of parameters is 7680, and the test accuracy per 100 training iterations is 0.861.

[0053] Results 1604-1607 show the case where the input vector to the input layer 1015 has 128 dimensions and a time series length of 18. In case 1604, the number of blocks in the input layer 1015, hidden layer 1016, and output layer 1019 is 4, the total number of parameters is 4288, and the test accuracy per 100 training iterations is 0.855. Results 1608-1611 show the case where the input vector to the input layer 1015 has 64 dimensions and a time series length of 36. In case 1608, the number of blocks in the input layer 1015, hidden layer 1016, and output layer 1019 is 4, the total number of parameters is 1472, and the test accuracy per 100 training iterations is 0.843. In cases 1601, 1605, and 1609, the number of blocks is 3. In the cases of 1602, 1606, and 1610, the number of blocks is 2. In the cases of 1603, 1607, and 1611, the number of blocks is 1.

[0054] Figures 1612 to 1614 show the results when an MZI-mesh structure is used for the input layer 1015, the hidden layer 1016, and the output layer 1019. In the case of 1612, the number of dimensions of the input vector to the input layer 1015 is 256, the time series length of the input vector is 9, the total number of parameters is 196,864, and the test accuracy per 100 training iterations is 0.822. In the case of 1613, the number of dimensions of the input vector to the input layer 1015 is 128, the time series length of the input vector is 18, the total number of parameters is 49,280, and the test accuracy per 100 training iterations is 0.837. In the case of 1614, the number of dimensions of the input vector to the input layer 1015 is 64, the time series length of the input vector is 36, the total number of parameters is 12,352, and the test accuracy per 100 training iterations is 0.843. In both the conventional configuration and the configuration of this embodiment, the matrices constituting the input layer 1015, the hidden layer 1016, and the output layer 1019 are unitary matrices corresponding to the number of input dimensions. For example, in the cases of 1608 to 1611 and 1614, these are 64 × 64 unitary matrices.

[0055] According to this embodiment, when the number of dimensions of the input vector is small, high test accuracy can be achieved with approximately one order of magnitude fewer parameters compared to a configuration using an MZI-mesh structure. Furthermore, in this embodiment, when the number of dimensions of the input vector is large, high test accuracy can be achieved with two or more orders of magnitude fewer parameters compared to a configuration using an MZI-mesh structure.

[0056] Figure 17 shows the relationship between test accuracy and the total number of parameters in a conventional configuration. Figures 1700 to 1711 show the results when an MZI-butterfly structure is used as the optical circuit constituting the input layer 1015, the hidden layer 1016, and the output layer 1019, respectively. In cases 1700 to 1703, the dimensionality of the input vector to the input layer 1015 is 256, and the time series length of the input vector is 9. In cases 1704 to 1707, the dimensionality of the input vector to the input layer 1015 is 128, and the time series length of the input vector is 18. In cases 1708 to 1711, the dimensionality of the input vector to the input layer 1015 is 64, and the time series length of the input vector is 36.

[0057] [Third Embodiment] Figure 18 is a block diagram showing the configuration of an optical circuit according to the third embodiment of the present invention. The optical circuit using the PSBS of the present invention as the basic structural unit is also useful in structures other than butterfly structures. The optical circuit of this embodiment includes a plurality of cascaded PSBS array layers 2-10 to 2-17, a unitary diagonal matrix layer 5, a waveguide 6-1 connecting PSBS array layers 2-10 and 2-11, a waveguide 6-2 connecting PSBS array layers 2-11 and 2-12, a waveguide 6-3 connecting PSBS array layers 2-12 and 2-13, a waveguide 6-4 connecting PSBS array layers 2-13 and 2-14, a waveguide 6-5 connecting PSBS array layers 2-14 and 2-15, a waveguide 6-6 connecting PSBS array layers 2-15 and 2-16, a waveguide 6-7 connecting PSBS array layers 2-16 and 2-17, and a waveguide 7 connecting PSBS array layer 2-17 and the unitary diagonal matrix layer 5. The unitary diagonal matrix layer 5 consists of PS50 inserted in the middle of each of the four waveguides 7.

[0058] In this embodiment, a mesh structure is adopted as the structure of the circuit combining PSBS1. The mesh structure is a polygonal network, with PSBS1 positioned at the vertices of the polygons. Waveguides 6-1 to 6-7 connecting the PSBS1 are positioned at the sides of the polygons. In this embodiment, any 4-input 4-output unitary matrix can be realized. Hereinafter, in this invention, a mesh structure with PSBS1 as the basic constituent unit will be referred to as a PSBS-mesh structure.

[0059] The degrees of freedom of a 4x4 unitary matrix are N, where N is the number of inputs. 2Therefore, the result is 16. The circuit shown in Figure 18 has 16 PS10, 50 for adjusting parameters, so any unitary matrix can be realized. Unlike the MZI-mesh structure, when the number of inputs N is even, the PSBS-mesh structure can satisfy reachability, meaning that any input signal can reach all output terminals, even when the number of fine layers (number of PSBS array layers) is reduced to N-1. In other words, although eight PSBS array layers are used in Figure 18, reachability can be satisfied with three PSBS array layers 2-10 to 2-12, so the number of PSBS array layers can be reduced, i.e., the number of parameters can be reduced. Such a configuration of PSBS array layers 2-10 to 2-12 can be used as the optical circuit of the complex number processing units 101, 101a.

[0060] [Fourth Embodiment] To realize an arbitrary N×N unitary matrix with an MZI-mesh structure, N layers of MZI arrays and one layer of unitary diagonal matrix are required, and the reachability that any input signal can reach all output terminals can be satisfied with (N-1) layers of MZI arrays. On the other hand, PSBS-butterfly structure, log 2 The signal reachability can be satisfied with the PSBS thin layer of (N).

[0061] While it is not possible to realize an arbitrary unitary matrix, the (N-1) layer MZI array and log 2 An intermediate configuration between the PSBS layers of (N) may also exist. As such an intermediate configuration, an example of a hybrid configuration of PSBS-mesh structure and PSBS-butterfly structure is shown in Figure 19.

[0062] The optical circuit in Figure 19 comprises a PSBS-butterfly structured block 20-1, a PSBS-mesh structured block 20-2, a PSBS-butterfly structured block 20-3, a waveguide 21 connecting blocks 20-1 and 20-2, and a waveguide 22 connecting blocks 20-2 and 20-3.

[0063] Block 20-1 consists of PSBS array layers 2-20 to 2-22, a waveguide 3-20 connecting PSBS array layers 2-20 and 2-21, and a waveguide 3-21 connecting PSBS array layers 2-21 and 2-22. Block 20-2 consists of PSBS array layers 2-23 to 2-26, a waveguide 6-20 connecting PSBS array layers 2-23 and 2-24, a waveguide 6-21 connecting PSBS array layers 2-24 and 2-25, and a waveguide 6-22 connecting PSBS array layers 2-25 and 2-26. Block 20-3 consists of PSBS array layers 2-27 to 2-29, a waveguide 3-22 connecting PSBS array layers 2-27 and 2-28, and a waveguide 3-23 connecting PSBS array layers 2-28 and 2-29. The configuration shown in Figure 19 can be used as the optical circuit for the complex number processing units 101 and 101a.

[0064] [Fifth Embodiment] Figure 20 is a block diagram showing the configuration of an optical circuit according to the fifth embodiment of the present invention. The optical circuit of this embodiment comprises PSBS-butterfly structured blocks 30-1 and 30-2, a waveguide 31 connecting blocks 30-1 and 30-2, and a nonlinear function layer 32 inserted in the middle of the waveguide 31.

[0065] Block 30-1 consists of PSBS array layers 2-30 to 2-32, a waveguide 3-30 connecting PSBS array layers 2-30 and 2-31, and a waveguide 3-31 connecting PSBS array layers 2-31 and 2-32. Block 30-2 consists of PSBS array layers 2-33 to 2-35, a waveguide 3-32 connecting PSBS array layers 2-33 and 3-34, and a waveguide 3-33 connecting PSBS array layers 3-34 and 3-35. The nonlinear function layer 32 consists of nonlinear function sections 320 inserted in the middle of each of the eight waveguides 31. The nonlinear function section 320 performs processing such as that shown in equation (1).

[0066] In this embodiment, nonlinear function layers are inserted between blocks, but as shown in Figure 21, nonlinear function layers may also be inserted between PSBS array layers within a block. The block 40 shown in Figure 21 consists of PSBS array layers 2-40 to 2-42, a waveguide 3-40 connecting PSBS array layers 2-40 and 2-41, a waveguide 3-41 connecting PSBS array layers 2-41 and 2-42, a nonlinear function layer 32-1 inserted in the middle of waveguide 3-40, and a nonlinear function layer 32-2 inserted in the middle of waveguide 3-41. The nonlinear function layer 32-1 consists of nonlinear function sections 320-1 inserted in the middle of each of the eight waveguides 3-40. The nonlinear function layer 32-2 consists of nonlinear function sections 320-2 inserted in the middle of each of the eight waveguides 3-41. The configurations shown in Figures 20 and 21 can be used as the optical circuits for the complex number processing units 101 and 101a.

[0067] [Sixth Embodiment] Figure 22 is a block diagram showing the configuration of an FNN according to the sixth embodiment of the present invention. In this embodiment, a complex number processing unit 101b is provided instead of the complex number processing unit 101 of the FNN described in the first embodiment. The complex number processing unit 101b comprises hidden layers 1010, 1012, and 1014, a loss compensation layer 1030 inserted between the output of hidden layer 1010 and the input of hidden layer 1012, and a nonlinear function layer 1031 with loss compensation function inserted between the output of hidden layer 1012 and the input of hidden layer 1014.

[0068] In this embodiment, a loss compensation layer 1030 and a nonlinear function layer 1031 with loss compensation function are provided to compensate for optical loss between devices such as PS and BS that constitute the optical circuit and the waveguide. This improves information processing performance. As the loss compensation layer 1030 and the nonlinear function layer 1031 with loss compensation function, optical amplification circuits such as semiconductor optical amplifiers (SOAs) and photoelectric conversion circuits such as electro-absorption modulators can be used.

[0069] In this embodiment, the case of FNN was described, but instead of the nonlinear function section 1018 of the RNN described in the second embodiment, a nonlinear function layer with loss compensation function may be provided.

[0070] [Seventh Embodiment] Figure 23 is a block diagram showing the configuration of an optical circuit according to the seventh embodiment of the present invention. The optical circuit of this embodiment comprises PSBS-butterfly structured blocks 50-1 to 50-3, a waveguide 51 connecting blocks 50-1 and 50-2, a waveguide 52 connecting blocks 50-2 and 50-3, a nonlinear function layer 53 with loss compensation function inserted in the middle of waveguide 51, and a loss compensation layer 54 inserted in the middle of waveguide 52.

[0071] Block 50-1 consists of PSBS array layers 2-50 to 2-52, a waveguide 3-50 connecting PSBS array layers 2-50 and 2-51, and a waveguide 3-51 connecting PSBS array layers 2-51 and 2-52. Block 50-2 consists of PSBS array layers 2-53 to 2-55, a waveguide 3-52 connecting PSBS array layers 2-53 and 2-54, and a waveguide 3-53 connecting PSBS array layers 2-54 and 2-55. Block 50-3 consists of PSBS array layers 2-56 to 2-58, a waveguide 3-54 connecting PSBS array layers 2-56 and 2-57, and a waveguide 3-55 connecting PSBS array layers 2-57 and 2-58.

[0072] The loss-compensating nonlinear function layer 53 consists of loss-compensating nonlinear function sections 530 inserted in the middle of each of the eight waveguides 51. The loss-compensating layer 54 consists of loss-compensating sections 540 inserted in the middle of each of the eight waveguides 52.

[0073] In this embodiment, a nonlinear function layer is inserted between blocks, but as shown in Figure 24, a nonlinear function layer with loss compensation function and a loss compensation layer may be provided between PSBS array layers within a block. The block 60 shown in Figure 24 comprises PSBS array layers 2-60 to 2-62, a waveguide 3-60 connecting PSBS array layers 2-60 and 2-61, a waveguide 3-61 connecting PSBS array layers 2-61 and 2-62, a nonlinear function layer 61 with loss compensation function inserted in the middle of waveguide 3-60, and a loss compensation layer 62 inserted in the middle of waveguide 3-61.

[0074] The loss-compensated nonlinear function layer 61 consists of loss-compensated nonlinear function sections 610 inserted in the middle of each of the eight waveguides 3-60. The loss-compensated layer 62 consists of loss-compensated sections 620 inserted in the middle of each of the eight waveguides 3-61. The configurations shown in Figures 23 and 24 can be used as the optical circuits for the complex number processing units 101, 101a, and 101b.

[0075] Some or all of the above examples may also be described as follows, but are not limited to the following:

[0076] (Note 1) The optical circuit of the present invention has a structure in which a plurality of array layers, each containing one or more basic structural units, are connected in cascading order, and the plurality of array layers are butterfly-connected by waveguides, and the basic structural unit consists of a beam splitter with two inputs and two outputs, and a programmable phase shifter inserted at any one of the four signal paths, the two inputs and two outputs of the beam splitter.

[0077] (Note 2) The optical circuit of the present invention has a structure in which a plurality of array layers, each containing one or more basic structural units, are connected in cascading order, and the plurality of array layers are mesh-connected by waveguides, and the basic structural unit consists of a 2-input 2-output beam splitter and a programmable phase shifter inserted at any one of the four signal paths, the 2 inputs and 2 outputs of the beam splitter.

[0078] (Note 3) The optical circuit of the present invention includes both a structure in which a plurality of array layers, each containing one or more basic structural units, are connected in cascading order, and the plurality of array layers are butterfly-connected by waveguides, and a structure in which the plurality of array layers are mesh-connected by waveguides, wherein the basic structural unit consists of a 2-input 2-output beam splitter and a programmable phase shifter inserted at any one of the four signal paths, the 2 inputs and 2 outputs of the beam splitter.

[0079] (Note 4) The optical circuit described in any one of Notes 1 to 3 further comprises a nonlinear function layer inserted in the middle of the waveguide connecting the plurality of array layers.

[0080] (Note 5) The optical circuit described in any one of Notes 1 to 3 further comprises a loss compensation layer inserted in the middle of the waveguide connecting the plurality of array layers.

[0081] (Note 6) The optical circuit described in any one of Notes 1 to 3 further comprises a nonlinear function layer with loss compensation function inserted in the middle of the waveguide connecting the plurality of array layers.

[0082] (Note 7) The feedforward neural network of the present invention uses the optical circuit described in any one of Notes 1 to 3 as the hidden layer.

[0083] (Note 8) The recurrent neural network of the present invention uses the optical circuit described in any one of Notes 1 to 3 as the input layer, the hidden layer, and the output layer.

[0084] 1...PSBS, 2-1 to 2-3, 2-10 to 2-17, 2-20 to 2-29, 2-30 to 2-35, 2-40 to 2-42, 2-50 to 2-58, 2-60 to 2-62...PSBS array layer, 3-1, 3-2, 3-20 to 3-23, 3-30 to 3-33, 3-40, 3-41, 3-50 to 3-55, 3-60, 3-61, 6-1 to 6-7, 6-20 to 6-22, 7, 2 1, 22, 51, 52... Waveguides, 5... Unitary diagonal matrix layer, 10, 50... Programmable phase shifter, 11... Beam splitter, 12-1 to 12-4... Arm waveguides, 20-1 to 20-3, 30-1, 30-2, 40, 50-1 to 50-3, 60... Blocks, 32, 32-1, 32-2... Nonlinear function layers, 53, 61, 1031... Nonlinear function layers with loss compensation, 54 ,62...Loss compensation layer, 100,100a...Preprocessing unit, 101,101a,101b...Complex number processing unit, 102...Real number processing unit, 320,320-1,320-2...Nonlinear function unit, 530,610...Nonlinear function unit with loss compensation function, 540,620...Loss compensation unit, 1000...Vector transformation unit, 1001,1004...Normalization unit, 1002,1007...FFT unit, 1003, 1006, 1008... Vector generation unit, 1009... Time series data generation unit, 1010, 1012, 1014, 1016... Hidden layer unit, 310, 1011, 1013, 1018... Nonlinear function unit, 1015... Input layer unit, 1017... Addition unit, 1019... Output layer unit, 1020... Vector transformation unit, 1021... Softmax function unit, 1022... Error calculation unit, 1030... Loss compensation layer.

Claims

1. An optical circuit characterized in that a plurality of array layers, each containing one or more basic structural units, are connected in cascading order, and the plurality of array layers are butterfly-connected by waveguides, wherein the basic structural units consist of a 2-input, 2-output beam splitter and a programmable phase shifter inserted at one of the four signal paths (2 inputs and 2 outputs) of the beam splitter.

2. An optical circuit characterized in that a plurality of array layers, each containing one or more basic structural units, are connected in series, and the plurality of array layers are mesh-connected by waveguides, wherein the basic structural units consist of a 2-input 2-output beam splitter and a programmable phase shifter inserted at one of the four signal paths (2 inputs and 2 outputs) of the beam splitter.

3. An optical circuit comprising both a structure in which multiple array layers, each containing one or more basic structural units, are connected in cascading order, and the multiple array layers are butterfly-connected by waveguides, and a structure in which the multiple array layers are mesh-connected by waveguides, wherein the basic structural units consist of a 2-input 2-output beam splitter and a programmable phase shifter inserted at any one of the four signal paths (2 inputs and 2 outputs) of the beam splitter.

4. An optical circuit according to any one of claims 1 to 3, further comprising a nonlinear function layer inserted in the middle of a waveguide connecting the plurality of array layers.

5. An optical circuit according to any one of claims 1 to 3, further comprising a loss compensation layer inserted in the middle of a waveguide connecting the plurality of array layers.

6. An optical circuit according to any one of claims 1 to 3, further comprising a nonlinear function layer with loss compensation function inserted in the middle of a waveguide connecting the plurality of array layers.

7. A feedforward neural network characterized by using the optical circuit described in any one of claims 1 to 3 as a hidden layer.

8. A recurrent neural network characterized by using the optical circuit described in any one of claims 1 to 3 as an input layer, a hidden layer, and an output layer.