Transform neural network-oriented photon calculation device
By designing a photonic computing device for Transformer neural networks, using optical encoding and linear transformation modules to perform multi-input and multi-output matrix operations, and combining adjustable optical nonlinear modules to realize nonlinear computing, the bandwidth and energy efficiency bottleneck problems of traditional electron accelerators are solved, low-energy consumption, high-parallelism photonic computing is achieved, and the all-optical reconstruction of Transformer neural networks is supported.
Patent Information
- Application Number
- CN202510855228.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional electron accelerators face bandwidth bottlenecks, energy efficiency ratio and thermal power consumption bottlenecks when processing large model inference tasks, making it difficult to meet the computing requirements of low latency and high energy efficiency. Existing optical computing devices lack the inter-layer interaction and nonlinear computing capabilities for Transformer neural networks.
A photonic computing device for Transformer neural networks is designed, including an optical encoding module, an optical linear transformation module, an adjustable optical nonlinear module, and a multi-path optical output module. The optical encoding and linear transformation modules are used to perform multi-input and multi-output matrix operations, and the adjustable optical nonlinear module is combined to realize nonlinear computing, forming an all-optical computing path to support the complex structure of Transformer neural networks and the transmission of information flow between network layers.
It realizes low-energy, high-parallel matrix and vector operations, improves the accuracy and reliability of photonic computing, supports all-optical reconstruction of Transformer neural networks, and has high scalability and system stability.
Smart Images

Figure CN120806010A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence and photonic integrated circuit, and particularly relates to a photonic computing device for a Transformer neural network. BACKGROUND
[0002] The conventional electronic accelerator in the prior art, such as a graphics processing unit, a tensor processing unit, an application-specific integrated circuit, etc., has high parallel computing capability, but has a bandwidth bottleneck in electrical signal transmission, and has a physical bottleneck in energy efficiency and thermal power consumption, and is difficult to meet the demand of low delay and high energy efficiency of data processing of large model inference tasks. SUMMARY
[0003] Therefore, the present disclosure provides a technical scheme of a photonic computing device for a Transformer neural network.
[0004] According to an aspect of the present disclosure, a photonic computing device for a Transformer neural network is provided, comprising: an optical encoding module, an optical linear transformation module, an adjustable optical nonlinear module, and a multi-channel optical output module; the optical encoding module is configured to perform word embedding encoding on a plurality of input optical signals, and determine an encoded optical signal corresponding to each input optical signal; the optical linear transformation module is configured to perform a multiple-input multiple-output matrix operation on the encoded optical signals corresponding to the plurality of input optical signals based on a multi-head attention model according to a preset parameter, and determine a plurality of modulated optical signals, wherein the preset parameter is determined by pre-training a Transformer neural network according to a preset algorithm; the adjustable optical nonlinear module is configured to perform feature dimension adjustment on the plurality of modulated optical signals according to a linear rectified activation function, and determine an optical processing result; and the multi-channel optical output module is configured to determine a target output result corresponding to the plurality of input optical signals according to the optical processing result.
[0005] In a possible implementation manner, the device further comprises a multi-channel input module configured to: perform matrix block processing on high-dimensional input data to determine a plurality of low-dimensional vectors; and perform electro-optical conversion on each low-dimensional vector to determine the plurality of input optical signals.
[0006] In a possible implementation manner, the multi-channel input module comprises a plurality of input channels, and any one input channel comprises an electrically controllable variable attenuator.
[0007] In a possible implementation manner, the optical encoding module comprises a plurality of modulators.
[0008] In a possible implementation, the preset algorithm comprises any one of a stochastic gradient descent method, a back propagation method, and a dual adaptive training method.
[0009] In a possible implementation, the preset parameter comprises a weight matrix of wavelength coding; and the optical linear transformation module comprises a plurality of matrix multiplication operator modules and a plurality of element-wise multiplication operator modules; any one matrix multiplication operator module is configured to perform matrix-vector multiplication in the optical domain on any two coded optical signals according to the weight matrix, to determine eigenvector optical signals of the two coded optical signals; and any one element-wise multiplication operator module is configured to perform element-wise vector-vector multiplication in the optical domain on any two eigenvector optical signals, to determine modulation optical signals of the two eigenvector optical signals.
[0010] In a possible implementation, any one matrix multiplication operator module comprises a field programmable gate array (FPGA) unit and a Mach-Zehnder interferometer (MZI) array unit; the FPGA unit is configured to determine, for any two input optical signals, first driving signals corresponding to the two input optical signals according to the weight matrix and real-time feedback signals corresponding to the multi-channel optical output module; and the MZI array unit is configured to perform linear transformation on coded optical signals corresponding to the two input optical signals according to the first driving signals corresponding to the two input optical signals, to determine eigenvector optical signals of the two coded optical signals.
[0011] In a possible implementation, any one element-wise multiplication operator module comprises a photodetector, a transimpedance amplifier, and a phase shifter; the photodetector is configured to perform photoelectric conversion on any one eigenvector optical signal, to determine an electrical signal corresponding to the eigenvector optical signal; the transimpedance amplifier is configured to determine, according to the electrical signal corresponding to any one eigenvector optical signal, a second driving signal corresponding to the eigenvector optical signal; and the phase shifter is configured to perform linear transformation on another eigenvector optical signal according to the second driving signal corresponding to the eigenvector optical signal, to determine modulation optical signals of the two eigenvector optical signals.
[0012] In a possible implementation, the adjustable optical nonlinear module comprises a plurality of nonlinear layers based on clipping-type nonlinear materials.
[0013] In a possible implementation, the multi-channel optical output module is configured to perform photoelectric conversion on the optical processing result, to determine an electrical signal corresponding to the optical processing result; and perform normalization processing and value weighting processing on the electrical signal corresponding to the optical processing result, to determine the target output result.
[0014] The photon computing device for the Transformer neural network provided by the present disclosure utilizes an optical coding module to perform word embedding coding on a plurality of input optical signals, determines the coded optical signals corresponding to each input optical signal, and can realize multi-head parallel processing in combination with wavelength multiplexing technology, thereby improving the efficiency of photon computing. Then, the optical linear transformation module performs multi-input multi-output matrix operation on the coded optical signals corresponding to the plurality of input optical signals based on the multi-head attention model and the preset parameters determined by pre-training the Transformer neural network according to a preset algorithm, determines a plurality of modulated optical signals, and realizes simulation of the attention mechanism network layer of the Transformer neural network. The photon computing device not only supports the traditional linear transformation function, but also can solve the inter-network layer interaction and nonlinear calculation problem of the Transformer neural network, complete the matrix and vector operation with low energy consumption and high parallelism in the optical domain, and improve the accuracy and reliability of photon computing. In combination with the adjustable optical nonlinear module, the feature dimension of the plurality of modulated optical signals can be adjusted according to the linear rectified activation function, and the optical processing result is determined, so as to simulate the complex structure network layer such as the activation function network layer and the residual connection layer of the Transformer neural network, form a full-optical computing path, and solve the nonlinear calculation and inter-network layer information flow transmission problem of the Transformer neural network. According to the optical processing result, the multi-channel optical output module can determine the target output result corresponding to the plurality of input optical signals, and realize full-optical reconstruction of the Transformer neural network. On the other hand, the photon computing device for the Transformer neural network of the present disclosure has a modular structure, all module structures can be manufactured based on a standard silicon photon platform, and a cross-module cooperative regulation function is realized through waveguide interconnection to realize a function closed loop on a chip, support a complex structure of the Transformer neural network, be easy to integrate and deploy, and have high scalability and system stability.
[0015] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0016] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and serve to explain the principles of the present disclosure.
[0017] Figure 1 A structural schematic diagram of a Transformer neural network according to the prior art is shown;
[0018] Figure 2 A block diagram of a photon computing device for a Transformer neural network according to an embodiment of the present disclosure is shown;
[0019] Figure 3 A system schematic diagram of a photonic computing device for a Transformer neural network is shown according to an embodiment of the present disclosure;
[0020] Figure 4 A structural schematic diagram of a photonic computing device for a Transformer neural network is shown according to an embodiment of the present disclosure;
[0021] Figure 5 A schematic diagram of a matrix multiplication operator module is shown according to an embodiment of the present disclosure;
[0022] Figure 6 A schematic diagram of an element-wise multiplication operator module is shown according to an embodiment of the present disclosure;
[0023] Figure 7 A schematic diagram of a limiting-amplitude type nonlinear material is shown according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0024] Various exemplary embodiments, features, and aspects of the present disclosure will be described herein below with reference to the accompanying drawings. The same reference numbers in different drawings indicate the same or similar elements / functionality. Although various aspects of embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically indicated.
[0025] As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having,” or the like are open-ended, and include one or more stated features, integers, elements, steps, components or functions but do not preclude the presence or addition of one or more other features, integers, elements, steps, components, functions or groups thereof.
[0026] When an element is referred to as being “connected,” “coupled,” “responsive,” or “correlated” to another element, it can be directly connected, coupled, or responsive to the other element, or one or more intervening elements can exist.
[0027] Although the terms first, second, third, and the like can be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another element / operation. Thus, a first element / operation in some embodiments could be termed a second element / operation in other embodiments without departing from the teachings of the present inventive concept.
[0028] The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any implementation described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other implementations.
[0029] The term "and / or", as used herein, merely describes association between associated objects, and can indicate three possible cases: A and / or B, which can mean that A exists alone, A and B exist together, and B exists alone. In addition, the term "at least one" herein means any one of a plurality or at least any combination of at least two of a plurality, for example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0030] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the specific embodiments below. Those skilled in the art should understand that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, elements and circuits that are well known to those skilled in the art are not described in detail in order to highlight the main ideas of the present disclosure.
[0031] With the development of deep learning technology, neural network models based on the Transformer structure gradually replace traditional convolutional neural networks and recurrent neural networks, and become the current mainstream general AI modeling framework. Specifically, the Transformer structure takes the self-attention mechanism (Self-Attention) and the multi-head attention mechanism (Multi-Head Attention) as the core, does not need the sequence dependence of the traditional recurrent neural network (RNN), and has higher parallelism, and is widely used in natural language processing, image recognition and speech understanding technical fields, and evolves to form deep large models such as the bidirectional encoder based on Transformer (Bidirectional Encoder Representations from Transformers, BERT), the generative pre-training Transformer model (Generative Pre-trained Transformer, GPT), the visual Transformer model (Vision Transformer, ViT), and the Pathways language model (Pathways Language Model, PaLM).
[0032] Figure 1 A structural schematic diagram of a Transformer neural network according to the prior art is shown. As shown in FIG. 1, the Transformer neural network includes an input layer, a self-attention layer, a multi-head attention layer, a fully connected layer, and an output layer. Figure 1As shown, the Transformer neural network includes an encoder 101, a decoder 102, and an output layer 103. The encoder 101 and the decoder 102 each include a multi-head attention mechanism network layer, a residual connection layer and a normalization layer, a feedforward layer, and the like. The output layer 103 includes a linear layer and a network layer based on a soft activation function. The multi-head attention mechanism includes a linear layer for implementing linear transformation, an attention calculation layer based on an activation function, a residual connection layer, and the like. By calculating linear transformation of a query vector (Q vector), a key vector (K vector), and a value vector (V vector), and calculating an attention score based on a scaling factor, weighted summation can be performed to extract key information in the input sequence.
[0033] In addition, the input data of the input multi-head attention mechanism network layer can also be projected into multiple subspaces for parallel attention calculation, thereby enhancing the expression ability of the Transformer neural network.
[0034] However, as the parameters of the neural network model based on the Transformer structure expand from hundreds of millions to tens of billions and hundreds of billions, the computing power required to implement model inference also increases exponentially, and only computing hardware with higher computing power can meet the use requirements.
[0035] Traditional electronic accelerators, such as a graphics processing unit (GPU), a tensor processing unit (TPU), and an application specific integrated circuit (ASIC), have bandwidth bottlenecks in electrical signal transmission and physical bottlenecks in energy efficiency and heat dissipation, which limit the ability of computing power expansion and are difficult to meet the low-latency and high-energy efficiency computing requirements of large model inference tasks.
[0036] While the optical signal has the advantages of high parallelism, high bandwidth and low delay, the photonic computing technology for realizing matrix operation in an optical manner can have the information processing characteristics of high bandwidth, high energy efficiency and high parallelism, realize the fusion of data processing and communication through optical signal transmission, significantly improve the computing efficiency, and can be fused with the existing electronic computing hardware platform to realize the complementary advantages in performance. On this basis, the prior art provides an optical matrix multiplier constructed based on a Mach-Zehnder interferometer (MZI) array, an optical matrix processing system based on a Micro-Ring Resonator (MRR) array, a computing platform using a Spatial Light Modulator (SLM) array, and the like.
[0037] Specifically, the optical matrix multiplier constructed based on the MZI array can realize new mapping of optical intensity output and input vectors through the MZI and behavior control array, and constitute a matrix-form neural network linear layer. However, this optical computing device only has good programmable linear mapping capability for general matrix-vector multiplication, lacks dynamic weight loading capability, does not support nonlinear operation, is difficult to realize all-optical vector-vector multiplication operation, and lacks specific design for the Transformer neural network. On the other hand, this optical computing device does not take into account the inevitable noise interference and manufacturing errors in the optical system. In the MZI array, the small error of phase shift will be accumulated with the increase of the number of neural network layers, causing the overall output of the optical computing device to deviate from the expected value, and the error can be more than 10%.
[0038] The optical matrix processing system based on the MRR array can utilize the MRR to realize wavelength division multiplexing technology (WDM), and then perform multiplication operation between vectors and matrices in an optical manner. However, the system integration of this optical computing device is limited by the temperature drift sensitivity of the MRR, and cannot stably perform multi-layer stacked neural network computing tasks. On the other hand, this optical computing device focuses on realizing the linear transformation function in a small-scale feedforward neural network, and is mainly used for simple linear transformation scenarios, and is difficult to solve the network layer interaction and nonlinear calculation problems of the Transformer neural network, and does not support the functions corresponding to the complex network layers of the Transformer neural network such as residual connection, activation function and normalization, and has great functional limitations.
[0039] The computing platform based on the SLM array can convert the dot product operation of the Transformer neural network into an optical interference problem, adopt a data-driven error modeling method, fit a hardware error model through simulation modeling and table calibration methods, combine singular value decomposition (SVD) and an electro-optical hybrid training strategy, map the linear layer of the Transformer neural network on the interference array for calculation, realize the linear transformation function, and theoretically has a high energy efficiency ratio. However, this optical computing device still needs electronic computing devices to complete part of the matrix multiplication calculation, activation function, and normalization in the attention mechanism network layer of the Transformer neural network, and does not have an end-to-end pure optical reasoning capability. On the other hand, due to the slow response speed of the SLM array, this optical computing device is only suitable for offline experimental simulation environment, and the calculation accuracy depends on the use environment and initial setting, which cannot meet the demand for calculation speed and system stability of real-time reasoning, and lacks practicality and robustness.
[0040] Therefore, the present disclosure provides a photonic computing device for a Transformer neural network, which can solve the interaction between network layers of the Transformer neural network, the non-linear calculation problem and the information flow transmission problem between network layers, complete low-energy and high-parallel matrix and vector operations in the optical domain, and realize the full-optical reconstruction of the Transformer neural network, thereby improving the efficiency and accuracy of photonic computing. The photonic computing device for the Transformer neural network of the present disclosure will be described in detail below.
[0041] Figure 2 A block diagram of a photonic computing device for a Transformer neural network according to an embodiment of the present disclosure is shown. As shown in Figure 2 The device 200 includes an optical encoding module 201, an optical linear transformation module 202, an adjustable optical non-linear module 203, and a multi-channel optical output module 204.
[0042] The optical encoding module 201 is configured to perform word embedding encoding on a plurality of input optical signals, and determine an encoded optical signal corresponding to each input optical signal.
[0043] The optical linear transformation module 202 is configured to perform a multiple-input multiple-output matrix operation on the encoded optical signals corresponding to the plurality of input optical signals according to a preset parameter based on a multi-head attention model, and determine a plurality of modulated optical signals, wherein the preset parameter is determined by pre-training the Transformer neural network according to a preset algorithm.
[0044] An adjustable optical nonlinear module 203 is configured to perform feature dimension adjustment on the plurality of modulated optical signals according to a linear rectified activation function, to determine an optical processing result.
[0045] A multi-channel optical output module 204 is configured to determine a target output result corresponding to the plurality of input optical signals according to the optical processing result.
[0046] The specific form of the device 200 can be flexibly set according to actual use requirements, for example, it can be designed as a photonic chip, and the present disclosure does not make specific limitations thereto.
[0047] The plurality of input optical signals herein can represent laser pulses obtained by performing dimension reduction processing and electro-optical conversion on any one high-dimensional data requiring a Transformer neural network to perform calculation and processing; the specific form of any one input optical signal and the specific number of input optical signals can be flexibly set according to actual use requirements, and the present disclosure does not make specific limitations thereto.
[0048] The specific method for determining the input optical signal can refer to the implementation in the related art, and the present disclosure does not make specific limitations thereto.
[0049] In one possible implementation, the device 200 further includes a multi-channel input module 205 configured to perform matrix blocking on the high-dimensional input data to determine a plurality of low-dimensional vectors, and perform electro-optical conversion on each low-dimensional vector to determine a plurality of input optical signals.
[0050] For example, as shown in Figure 2 For example, as shown in Figure 2 The device 200 further includes a multi-channel input module 205 connected with the optical encoding module 201, which can perform matrix blocking on any one high-dimensional input data to obtain a plurality of low-dimensional vectors. The specific form of the high-dimensional input data can be flexibly set according to actual use requirements, and the present disclosure does not make specific limitations thereto; the specific method for performing matrix blocking on the high-dimensional input data can refer to the implementation in the related art, and depends on the specific form of the high-dimensional input data, and the present disclosure does not make specific limitations thereto; the specific form of the low-dimensional vector depends on the specific form of the high-dimensional input data, and the present disclosure does not make specific limitations thereto.
[0051] Figure 3 A system principle schematic diagram of a photonic computing device for a Transformer neural network according to an embodiment of the present disclosure is shown. As Figure 3As shown, the high-dimensional input data is an image, and the image is directly divided into blocks to obtain a plurality of image blocks as low-dimensional vectors corresponding to the image, so as to realize matrix blocking of the high-dimensional input data. Further, the multi-channel input module 205 can perform electro-optical conversion on each low-dimensional vector respectively, and obtain an input optical signal corresponding to the low-dimensional vector based on an electrical signal corresponding to any low-dimensional vector; the input optical signal corresponding to each low-dimensional vector is independently input to the optical coding module 201, and sequentially passes through the optical linear transformation module 202, the adjustable optical nonlinear module 203, and the multi-channel optical output module 204, and finally obtains a target output result corresponding to the image.
[0052] The specific method of electro-optical conversion of the multi-channel input module 205 can refer to the implementation in the related art, for example, the electrical signal corresponding to any low-dimensional vector can be input to a laser to obtain the input optical signal corresponding to the low-dimensional vector, and the present disclosure does not make a specific limitation hereon.
[0053] The optical coding module 201 includes a plurality of input channels, which can perform word embedding coding on each input optical signal respectively to determine the encoded optical signal corresponding to each input optical signal for representing the input tensor corresponding to the Transformer neural network. The specific method of word embedding coding of the optical coding module 201 on any input optical signal and the specific form of the optical coding module 201 can be flexibly set according to actual use requirements, and the present disclosure does not make a specific limitation hereon.
[0054] In one possible implementation, the optical coding module 201 includes a plurality of modulators.
[0055] Each input channel of the optical coding module 201 is provided with a modulator for mapping a word embedding vector to light field amplitude information or light field phase information and performing light field modulation on the input optical signal of the input channel to realize word embedding coding. The specific form of the word embedding vector can refer to the implementation in the related art, and the present disclosure does not make a specific limitation hereon; the specific number of modulators can be flexibly set according to actual use requirements and depends on the number of input channels of the optical coding module 201, and the present disclosure does not make a specific limitation hereon; the specific form of any modulator can refer to the implementation in the related art, for example, a high-speed digital-to-analog converter (DAC) driven electro-optical modulator can be used, and the present disclosure does not make a specific limitation hereon.
[0056] Figure 4 Fig. 1 shows a structural schematic diagram of a photonic computing device for a Transformer neural network according to an embodiment of the present disclosure. As shown in Fig. 1, the photonic computing device for a Transformer neural network includes an optical coding module 201, an optical linear transformation module 202, an adjustable optical nonlinear module 203, a multi-channel optical output module 204, and a multi-channel input module 205. Figure 4As shown, the optical encoding module 201 includes 8 independent input channels, each of which is provided with a modulator, which are modulators 2011 to 2018 respectively. The 4 input optical signals corresponding to vector X1 are input to modulators 2011 to 2014, and the 4 input optical signals corresponding to vector X2 are input to modulators 2015 to 2018.
[0057] With the optical encoding module 201 including multiple modulators, high-speed parallel input greater than 50 Gbps can be achieved, and different input optical signals can be assigned independent wavelengths to improve the overall throughput of the device 200 in combination with wavelength division multiplexing technology.
[0058] The optical linear transformation module 202 can be used to simulate the attention mechanism network layer corresponding to the Transformer neural network, and based on the multi-head attention model, performs multiple-input multiple-output matrix operations on the encoded optical signals corresponding to the multiple input optical signals according to the preset parameters to determine multiple modulated optical signals. The number of modulated optical signals is usually less than or equal to the number of encoded optical signals.
[0059] The attention mechanism network layer here can be a self-attention mechanism network layer or a multi-head attention mechanism network layer, which can be flexibly set according to actual use requirements, and the present disclosure does not make specific limitations on this; the multiple-input multiple-output matrix operation can include the matrix-vector multiplication operation and the vector-vector element-wise multiplication operation included in the attention mechanism network layer corresponding to the Transformer neural network, and the specific form thereof can refer to the implementation in related technologies, for example, the matrix-vector multiplication operation for generating Q vectors and K vectors, the vector-vector element-wise multiplication operation between Q vectors and K vectors, etc., and the present disclosure does not make specific limitations on this.
[0060] The preset parameters corresponding to the optical linear transformation module 202 are determined according to the pre-training of the Transformer neural network based on a preset algorithm. After the network layer parameters satisfying the use requirements are determined by pre-training the Transformer neural network, the preset parameters can be determined according to the actual structure of the optical linear transformation module 202 in combination with the network layer parameters. The specific form of the preset algorithm can refer to the implementation in related technologies, and the present disclosure does not make specific limitations on this.
[0061] In one possible implementation, the preset algorithm includes any one of a stochastic gradient descent method, an error back propagation method, and a dual adaptive training method. The specific form of any one of the preset algorithms here can refer to the implementation in related technologies, and the present disclosure does not make specific limitations on this.
[0062] The optical linear transformation module 202 will be described in detail below in combination with possible implementation manners of the present disclosure, and thus no further description is provided herein.
[0063] The adjustable optical nonlinear module 203 can simulate a linear rectified activation function (ReLU) network layer corresponding to a Transformer neural network, and perform feature dimension adjustment on the plurality of modulated optical signals according to the ReLU activation function, so as to realize feature dimension increase or feature dimension decrease on the attention output features of the optical linear transformation module 202, thereby determining the optical processing result and reducing the loss of photoelectric digital-to-analog conversion. The feature dimension adjustment can include residual connection processing and multi-layer stacking characteristics processing facing the Transformer neural network, and specific forms thereof can be referred to the implementation manners in the related art, which are not limited in the present disclosure.
[0064] The optical processing result can include at least one optical signal, and the corresponding optical intensity information and phase information thereof can reflect the operation of the Transformer neural network on the input optical signal, and specific forms thereof can be flexibly set according to actual use requirements, which are not limited in the present disclosure.
[0065] The adjustable optical nonlinear module 203 will be described in detail below in combination with possible implementation manners of the present disclosure, and thus no further description is provided herein.
[0066] The multi-channel optical output module 204 can simulate an output layer in the Transformer neural network for realizing normalization processing and the like, extract the corresponding optical intensity information and phase information from the optical processing result, and perform post-processing according to actual use requirements to determine the target output result corresponding to the plurality of input optical signals. Specific forms of the target output result can be flexibly set according to actual use requirements, which are not limited in the present disclosure.
[0067] In an example, in a case where the optical processing result includes at least one optical signal, and the target output result needs to be further input to other optical computing devices, for example, another photonic computing device facing the Transformer neural network, for processing, the target output result can be set to include a normalized optical signal obtained by performing optical power ratio operation on the optical processing result.
[0068] In an example, in a case where the optical processing result includes at least one optical signal, and the target output result needs to be further input to an electronic computing device for processing, the target output result can be set to include a voltage signal obtained by performing weighted combination and electro-optical conversion on the optical processing result.
[0069] The multi-path light output module 204 will be described in detail later in conjunction with possible implementations of the present disclosure, and will not be elaborated here.
[0070] The photonic computing device for Transformer neural networks provided by the present disclosure uses an optical coding module to perform word embedding coding on multiple input optical signals, determine the coded optical signal corresponding to each input optical signal, and can be combined with wavelength multiplexing technology to realize multi-head parallel processing, thereby improving the efficiency of photonic computing; then, through an optical linear transformation module, based on a multi-head attention model and preset parameters determined by pre-training the Transformer neural network according to a preset algorithm, multi-input and multi-output matrix operations are performed on the coded optical signals corresponding to the multiple input optical signals, multiple modulated optical signals are determined, and the simulation of the attention mechanism network layer of the Transformer neural network is realized. It not only supports traditional linear transformation functions, but also can solve the network layer interaction and nonlinear calculation problems for Transformer neural networks, complete low-energy, high-parallelism matrix and vector operations in the optical domain, and improve the accuracy and reliability of photonic computing. Combined with the adjustable optical nonlinear module, the characteristic dimensions of multiple modulated optical signals can be adjusted according to the linear rectification activation function to determine the optical processing results, thereby simulating the activation function network layer, residual connection layer and other complex structure network layers of the Transformer neural network to form an all-optical computing path, which can solve the nonlinear computing and information flow transmission problems between network layers for the Transformer neural network; according to the optical processing results, the multi-channel optical output module can determine the target output results corresponding to multiple input optical signals, and realize the all-optical reconstruction of the Transformer neural network. On the other hand, the photonic computing device for the Transformer neural network of the embodiment of the present disclosure adopts a modular structure. All module structures can be manufactured based on the standard silicon photonic platform, and the on-chip functional closed loop is realized through waveguide interconnection. It has cross-module collaborative control functions, supports complex structure Transformer neural networks, is easy to integrate and deploy, and has high scalability and system stability.
[0071] In a possible implementation, the multi-channel input module 205 includes multiple input channels, and any input channel includes an electrically controlled variable attenuator.
[0072] The number of input channels included in the multi-channel input module 205 is the same as the number of input channels included in the optical encoding module 201 .
[0073] In order to make the device 200 have an adjustable optical power function, improve its reliability and stability, an electrically controllable variable attenuator can be arranged at each input channel of the multi-channel input module 205 to independently adjust the power of each input optical signal obtained through electro-optical conversion. The specific form of any one electrically controllable variable attenuator can refer to the implementation in the related art, and the present disclosure does not make specific limitations thereto.
[0074] In a possible implementation, the preset parameters include a wavelength-encoded weight matrix; the optical linear transformation module 202 includes a plurality of matrix multiplication operator modules and a plurality of element-wise multiplication operator modules; any one matrix multiplication operator module is configured to perform a matrix-vector multiplication operation in the optical domain on any two encoded optical signals according to the weight matrix to determine a feature vector optical signal of the two encoded optical signals; and any one element-wise multiplication operator module is configured to perform a vector-vector element-wise multiplication operation in the optical domain on any two feature vector optical signals to determine a modulation optical signal of the two feature vector optical signals.
[0075] The wavelength-encoded weight matrix can include a matrix required by an attention mechanism network layer in a Transformer neural network. The number of weight matrices can be flexibly set according to actual use requirements, and the present disclosure does not make specific limitations thereto.
[0076] In an example, the weight matrix can include a weight matrix a weight matrix a weight matrix a weight matrix The specific content of each weight matrix can be flexibly set according to actual use requirements, and the present disclosure does not make specific limitations thereto.
[0077] The optical linear transformation module 202 can include a plurality of matrix multiplication operator modules and a plurality of element-wise multiplication operator modules to respectively implement the matrix-vector multiplication operation and the vector-vector element-wise multiplication operation of the attention mechanism network layer in the Transformer neural network. The specific number of matrix multiplication operator modules and the specific number of element-wise multiplication operator modules can be flexibly set according to actual use requirements, and the present disclosure does not make specific limitations thereto.
[0078] Specifically, any one matrix multiplication operator module can receive any two different encoded optical signals, and perform matrix-vector multiplication operation on the two encoded optical signals in the optical domain according to a weight matrix, to determine an eigenvector optical signal of the two encoded optical signals; wherein the eigenvector optical signal of any two encoded optical signals can be an optical signal for representing a Q vector of the two encoded optical signals, or an optical signal for representing a K vector of the two encoded optical signals, depending on the type of the weight matrix used when performing the matrix-vector multiplication operation in the optical domain.
[0079] For example, as shown in FIG. 2, the optical linear transformation module 202 includes four matrix multiplication operator modules: the matrix multiplication operator module 2021, the matrix multiplication operator module 2022, the matrix multiplication operator module 2023, and the matrix multiplication operator module 2024. Figure 4 For example, as shown in FIG. 2, the optical linear transformation module 202 includes four matrix multiplication operator modules: the matrix multiplication operator module 2021, the matrix multiplication operator module 2022, the matrix multiplication operator module 2023, and the matrix multiplication operator module 2024. Figure 4 The matrix multiplication operator module 2021 and the matrix multiplication operator module 2022 are used to perform matrix-vector multiplication operation on the encoded optical signals corresponding to the vector X1; the matrix multiplication operator module 2023 and the matrix multiplication operator module 2024 are used to perform matrix-vector multiplication operation on the encoded optical signals corresponding to the vector X2.
[0080] Any one encoded optical signal corresponding to the vector X1 is input into the matrix multiplication operator module 2021 and the matrix multiplication operator module 2022 after being split by the beam splitter. The matrix multiplication operator module 2021 can perform matrix-vector multiplication operation on any two encoded optical signals input, to obtain an eigenvector optical signal capable of representing a Q vector of the two encoded optical signals, based on a weight matrix The matrix multiplication operator module 2022 can perform matrix-vector multiplication operation on any two encoded optical signals input, to obtain an eigenvector optical signal capable of representing a K vector of the two encoded optical signals, based on a weight matrix The matrix multiplication operator module 2022 can perform matrix-vector multiplication operation on any two encoded optical signals input, to obtain an eigenvector optical signal capable of representing a K vector of the two encoded optical signals, based on a weight matrix
[0081] Similarly, any one encoded optical signal corresponding to the vector X2 is input into the matrix multiplication operator module 2023 and the matrix multiplication operator module 2024 after being split by the beam splitter. The matrix multiplication operator module 2023 can perform matrix-vector multiplication operation on any two encoded optical signals input, to obtain an eigenvector optical signal capable of representing a Q vector of the two encoded optical signals, based on a weight matrix The matrix multiplication operator module 2024 can perform matrix-vector multiplication operation on any two encoded optical signals input, to obtain an eigenvector optical signal capable of representing a K vector of the two encoded optical signals, based on a weight matrix The matrix multiplication operator module 2024 can perform matrix-vector multiplication operation on any two encoded optical signals input, to obtain an eigenvector optical signal capable of representing a K vector of the two encoded optical signals, based on a weight matrix
[0082] For example, as shown in FIG. 2, the optical linear transformation module 202 includes four matrix multiplication operator modules: the matrix multiplication operator module 2021, the matrix multiplication operator module 2022, the matrix multiplication operator module 2023, and the matrix multiplication operator module 2024.Figure 4 For example, as shown in Figure 4 The optical linear transformation module 202 includes 12 element-wise multiplication operator modules. Any one element-wise multiplication operator module can receive any one eigenvector optical signal for representing a K vector of any two encoded optical signals, and any one eigenvector optical signal for representing a Q vector of any two encoded optical signals, and perform vector-vector element-wise multiplication operation on the two eigenvector optical signals in the optical domain to determine the modulation optical signal of the two eigenvector optical signals, and output to the adjustable optical nonlinear module 203.
[0083] In a possible implementation, any one matrix multiplication operator module includes: a field programmable gate array (FPGA) unit, a Mach-Zehnder interferometer (MZI) array unit; the FPGA unit is configured to determine, for any two input optical signals, a first driving signal corresponding to the two input optical signals according to a weight matrix and a real-time feedback signal corresponding to the multi-channel optical output module 204; and the MZI array unit is configured to perform linear transformation on encoded optical signals corresponding to the two input optical signals according to the first driving signal corresponding to the two input optical signals to determine eigenvector optical signals of the two encoded optical signals.
[0084] The specific form of the field programmable gate array (FPGA) unit can refer to the implementation in the related art, and the present disclosure does not make a specific limitation hereon.
[0085] The MZI array unit is determined according to a plurality of MZIs, and a weight matrix is used as a target unitary matrix to construct an optical path. The specific form thereof can be flexibly set according to actual use requirements, and the present disclosure does not make a specific limitation hereon. The specific form of any one MZI in the MZI array unit can refer to the implementation in the related art, for example, a MZI based on the thermo-optic effect or a MZI based on the electro-optic effect, and the present disclosure does not make a specific limitation hereon.
[0086] The FPGA unit can determine a control signal according to the weight matrix and the real-time feedback signal corresponding to the multi-channel optical output module, and then generate the first driving signal corresponding to the two input optical signals, and transmit the first driving signal corresponding to the two input optical signals to the MZI array unit.
[0087] The specific form of the first driving signal corresponding to the two input optical signals can refer to the implementation in the related art, for example, the shunt switching driving signal can be set as a driving voltage in the range of 0-3V, and the specific form depends on the specific form of the FPGA unit, and the present disclosure does not make a specific limitation hereon.
[0088] The MZI array unit can perform optical path switching according to the first driving signal after receiving the encoded optical signals corresponding to any two input optical signals and the first driving signal corresponding to the two input optical signals, thereby performing linear transformation on the two encoded optical signals, realizing matrix-vector multiplication operation in the optical domain of the two encoded optical signals, and obtaining the eigenvector optical signals of the two encoded optical signals. The eigenvector optical signals of the two encoded optical signals can be output to the multi-channel optical output module 204, and corresponding real-time feedback signals are generated and fed back to the FPGA unit to perform secondary modulation on the eigenvector optical signals of the two encoded optical signals, thereby improving the accuracy and reliability thereof and realizing closed-loop regulation and control of the optical domain matrix-vector multiplication operation.
[0089] The specific method for determining the real-time feedback signal by the multi-channel optical output module 204 and the specific form of the real-time feedback signal can be flexibly set according to actual use requirements, and the present disclosure does not make specific limitations in this regard.
[0090] In an example, the multi-channel optical output module 204 can perform photoelectric conversion and analog-to-digital conversion on the eigenvector optical signals of the two encoded optical signals when receiving the eigenvector optical signals of the two encoded optical signals, to determine the real-time feedback signal in the form of an electrical digital signal.
[0091] In addition, the FPGA unit can also be used to control each modulator in the optical encoding module 201 to realize unified scheduling of the device 200 and realize comprehensive cooperation between electrical domain control and optical domain calculation.
[0092] Figure 5 A principle schematic diagram of a matrix multiplication operation sub-module according to an embodiment of the present disclosure is shown. As shown in Figure 5 The field programmable gate array unit 501 cooperates with the scheduling of each modulator of the optical encoding module 201 through high-speed digital-to-analog conversion to perform word embedding encoding on each input optical signal to obtain the encoded optical signal corresponding to each input optical signal. After any two encoded optical signals are input to the Mach-Zehnder interferometer array unit 502, linear transformation is performed on the two encoded optical signals under the control of the field programmable gate array unit 501 to determine the eigenvector optical signals of the two encoded optical signals. The eigenvector optical signals of the two encoded optical signals are output to the multi-channel optical output module 204, and corresponding real-time feedback signals are generated, which are fed back to the field programmable gate array unit 501 after analog-to-digital conversion to form a closed-loop regulation and control.
[0093] Through the matrix multiplication operation sub-module in the above form, an all-optical computing path for matrix-vector multiplication operation can be realized, and combined with the intelligent scheduling capability of FPGA photoelectric signal hybrid processing, low-power and high-parallel optical calculation can be realized, thereby improving the accuracy and reliability of the matrix-vector multiplication operation.
[0094] In a possible implementation, the element-wise multiplication operation module includes: a photodetector, a transimpedance amplifier, and a phase shifter; the photodetector is configured to perform photoelectric conversion on the arbitrary feature vector optical signal to determine an electrical signal corresponding to the feature vector optical signal; the transimpedance amplifier is configured to determine a second driving signal corresponding to the feature vector optical signal according to the electrical signal corresponding to the arbitrary feature vector optical signal; and the phase shifter is configured to perform linear transformation on another feature vector optical signal according to the second driving signal corresponding to the arbitrary feature vector optical signal to determine a modulated optical signal of the two feature vector optical signals.
[0095] The specific forms of the photodetector, the transimpedance amplifier, and the phase shifter can refer to the implementation in the related art, and the present disclosure does not make a specific limitation.
[0096] Figure 6 A principle schematic diagram of an element-wise multiplication operation module according to an embodiment of the present disclosure is shown. As shown in the figure, Figure 6 The element-wise multiplication operation module 600 includes: a photodetector 601, a transimpedance amplifier 602, and a phase shifter 603. When the feature vector optical signal A and the feature vector optical signal B are input into the element-wise multiplication operation module 600, the photodetector 601 can perform photoelectric conversion on the feature vector optical signal A to determine an electrical signal corresponding to the feature vector optical signal A and transmit the electrical signal to the transimpedance amplifier 602. The transimpedance amplifier 602 amplifies the electrical signal corresponding to the feature vector optical signal A to determine a second driving signal corresponding to the feature vector optical signal A; and controls the phase shifter 603 to perform linear transformation on the feature vector optical signal B through the second driving signal to determine a modulated optical signal of the feature vector optical signal A and the feature vector optical signal B.
[0097] The specific form of the second driving signal corresponding to the arbitrary feature vector optical signal can be flexibly set according to actual use requirements, and the present disclosure does not make a specific limitation.
[0098] Based on the combination of the matrix multiplication operation module and the element-wise multiplication operation module, the complete process calculation of the attention mechanism network layer can be completed, and all calculation processes are performed in the optical domain to form a full connection reasoning path. Moreover, the matrix multiplication operation module and the element-wise multiplication operation module can both adopt modular design, so that the structure of the device 200 can be flexibly expanded, can adapt to a complex structure of a Transformer neural network, and can reduce the overall system complexity and manufacturing cost of the device 200 in the network layer growth scenario, facilitate the integrated design and manufacturing of the photonic chip, and improve the adaptability of the photonic chip in industrial deployment.
[0099] In a possible implementation, the tunable optical nonlinear module 203 includes: a plurality of nonlinear layers based on limiting nonlinear materials.
[0100] Specifically, the optical properties (eg, refractive index or absorption coefficient) of the limiting nonlinear material may vary with the intensity of the input light, exhibiting limiting characteristics such as linear loss and saturation absorption.
[0101] In one example, when the intensity of the input light is less than the limiting threshold of the material, the limiting nonlinear material exhibits linear optical properties, so that the intensity of the output light is approximately proportional to the intensity of the input light; when the intensity of the input light is greater than the limiting threshold of the material, the limiting nonlinear material excites the limiting effect, limiting the intensity of the output light to near the limiting threshold.
[0102] Figure 7 FIG. 1 is a schematic diagram showing a limiting nonlinear material according to an embodiment of the present disclosure. Figure 7 As shown in (a) of FIG. 1 , the limiting nonlinear material 700 includes an input port 701 and an output port 702. Figure 7 As shown in (b) of FIG. 1 , the horizontal axis of the light intensity detection image represents the horizontal position distribution of the limiting nonlinear material 700, and the vertical axis represents the vertical position distribution of the limiting nonlinear material 700. The light intensity detection image shows the linear loss characteristics of the limiting nonlinear material 700, and the light intensity corresponding to the output light is less than the light intensity corresponding to the input light. Figure 7 As shown in (c) in FIG. 5 , the light intensity detection image shows the saturation absorption characteristics of the limiting nonlinear material 700, and the light intensity corresponding to the output light is close to 0; Figure 7 As shown in (d), the horizontal axis is the light intensity corresponding to the input light, and the vertical axis is the light intensity corresponding to the output light passing through the limiting nonlinear material 700. It can be seen from the curve that the light intensity between the output light and the input light presents a nonlinear relationship.
[0103] Based on the parameters corresponding to the ReLU activation function, a structure of a nonlinear layer based on a limiting nonlinear material is designed. This property of the limiting nonlinear material can be used to simulate the nonlinear input-output relationship of the ReLU activation function, thereby performing feature dimensionality upscaling and dimensionality reduction operations on any modulated optical signal input to the adjustable optical nonlinear module 203. The specific number of nonlinear layers of the adjustable optical nonlinear module 203 can be flexibly set according to actual usage requirements, and this disclosure does not specifically limit this. The specific form of the limiting nonlinear material used in any nonlinear layer can refer to the implementation methods in the relevant technology. For example, a germanium-silicon composite material can be used, and this disclosure does not specifically limit this.
[0104] With the above Figure 4 For example,Figure 4 As shown, the adjustable optical nonlinear module 203 includes 4 nonlinear layers: nonlinear layer 2031 to nonlinear layer 2034. Taking the nonlinear layer 2031 as an example, after receiving the modulated optical signals from the three element-wise multiplication sub-modules, the nonlinear layer 2031 adjusts the feature dimensions of the three modulated optical signals respectively, and then combines and outputs an optical signal, so as to obtain the subsequent optical processing result.
[0105] Further, the adjustable optical nonlinear module 203 can also include at least one photodetector (PD) to convert the optical signal output by any nonlinear layer into an electrical signal, so as to perform corresponding weight distribution, simulate the attention mechanism network layer based on the Relu function, further improve the computing performance of the photonic computing device, and improve the accuracy and reliability of the target output result.
[0106] In a possible implementation, the multi-channel optical output module 204 is configured to: perform photoelectric conversion on the optical processing result to determine an electrical signal corresponding to the optical processing result; perform normalization processing and value weighting processing on the electrical signal corresponding to the optical processing result to determine the target output result.
[0107] The multi-channel optical output module 204 can perform photoelectric conversion on each optical signal included in the optical processing result, so as to extract the optical intensity information and phase information corresponding to each optical signal, and obtain the electrical signal corresponding to the optical processing result. The specific method of performing photoelectric conversion on the optical processing result by the multi-channel optical output module 204 can be flexibly set according to actual use requirements, for example, any optical signal included in the optical processing result can be output to a PD, and the present disclosure does not make a specific limitation in this regard.
[0108] By performing normalization processing and value weighting processing on the electrical signal corresponding to the optical processing result, the multi-channel optical output module 204 can determine the target output result, and realize the context-aware output for multiple input optical signals on the basis of ensuring the stability of the device 200. The specific method of normalization processing and value weighting processing can refer to the implementation in the related art, and the present disclosure does not make a specific limitation in this regard.
[0109] The photon computing device for the Transformer neural network provided by the disclosure utilizes the optical coding module to perform word embedding coding on the plurality of input optical signals, determines the coded optical signal corresponding to each input optical signal, and can realize multi-head parallel processing in combination with the wavelength multiplexing technology, thereby improving the efficiency of photon computing. Then, the optical linear transformation module is used to perform multi-input multi-output matrix operation on the coded optical signals corresponding to the plurality of input optical signals based on the multi-head attention model and the preset parameters determined by pre-training the Transformer neural network according to a preset algorithm, to determine a plurality of modulated optical signals, thereby realizing simulation of the attention mechanism network layer of the Transformer neural network. The photon computing device not only supports the traditional linear transformation function, but also can solve the inter-network layer interaction and nonlinear calculation problem of the Transformer neural network, complete the matrix and vector operation with low energy consumption and high parallelism in the optical domain, and improve the accuracy and reliability of photon computing. In combination with the adjustable optical nonlinear module, the feature dimension of the plurality of modulated optical signals can be adjusted according to the linear rectified activation function, and the optical processing result is determined, so as to simulate the activation function network layer, the residual connection layer and other complex structure network layers of the Transformer neural network, form a full-optical computing path, and solve the nonlinear calculation and inter-network layer information flow transmission problem of the Transformer neural network. According to the optical processing result, the multi-channel optical output module can determine the target output result corresponding to the plurality of input optical signals, and realize full-optical reconstruction of the Transformer neural network. On the other hand, the photon computing device for the Transformer neural network of the embodiment of the disclosure adopts a modular structure, all module structures can be manufactured based on a standard silicon photon platform, and the on-chip function closed loop is realized through waveguide interconnection, has the cross-module cooperative regulation and control function, supports the complex structure of the Transformer neural network, is easy to integrate and deploy, and has high scalability and system stability.
[0110] It should be noted that although the photon computing device for the Transformer neural network and the structure of part of the modules are exemplarily introduced as above, Figure 4 、 Figure 5 、 Figure 6 and Figure 7 , the person skilled in the art can understand that the disclosure should not be limited thereto. In fact, the user can completely flexibly set the specific structure form of the photon computing device according to personal preferences and / or actual application scenarios, as long as the above-mentioned principles can be used to solve the inter-network layer interaction and nonlinear calculation problem of the Transformer neural network, and to construct the full-optical computing path for the Transformer neural network with low energy consumption and high parallelism.
[0111] Having described above several embodiments of the disclosure, any modifications and variations that fall within the scope of the described embodiments are also intended to be within the scope of the disclosure. As will be apparent to those skilled in the art, some modifications and variations to the embodiments described above can be practiced while staying within the scope and spirit of the described embodiments. The foregoing description of the described embodiments has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the described embodiments to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. It is intended that the disclosed embodiments be limited only by the claims.
Claims
1. A photonic computing device for Transformer neural networks, characterized in that: The device comprises: an optical encoding module, an optical linear conversion module, an adjustable optical nonlinear module, and a multi-path light output module; The optical encoding module is used to perform word embedding encoding on multiple input optical signals and determine the encoded optical signal corresponding to each input optical signal; The optical linear transformation module is configured to perform a multi-input multi-output matrix operation on the coded optical signals corresponding to the multiple input optical signals according to preset parameters based on a multi-head attention model to determine multiple modulated optical signals, wherein the preset parameters are determined by pre-training a Transformer neural network according to a preset algorithm; The adjustable optical nonlinear module is used to adjust the characteristic dimensions of the multiple modulated optical signals according to the linear rectification activation function to determine the optical processing results; The multi-path optical output module is used to determine target output results corresponding to the multiple input optical signals according to the optical processing results.
2. The device according to claim 1, characterized in that The device further comprises a multi-channel input module, configured to: Perform matrix partitioning on high-dimensional input data to determine multiple low-dimensional vectors; Perform electrical-to-optical conversion on each low-dimensional vector to determine the multiple input optical signals.
3. The device according to claim 2, characterized in that The multi-channel input module includes multiple input channels, and any input channel includes an electrically controlled variable attenuator.
4. The device according to any one of claims 1 to 3, characterized in that The optical encoding module includes a plurality of modulators.
5. The device according to any one of claims 1 to 3, characterized in that The preset algorithm includes: any one of a stochastic gradient descent method, an error back propagation method, and a dual adaptive training method.
6. The device according to any one of claims 1 to 3, characterized in that The preset parameters include: a wavelength-encoded weight matrix; The optical linear transformation module includes: a plurality of matrix multiplication operator modules, a plurality of element-by-element multiplication operator modules; Any matrix multiplication operator module, configured to perform a matrix-vector multiplication operation in the optical domain on any two coded optical signals according to the weight matrix to determine the eigenvector optical signals of the two coded optical signals; Any one of the element-by-element multiplication operation submodules is used to perform a vector-vector element-by-element multiplication operation in the optical domain on any two eigenvector optical signals to determine the modulated optical signals of the two eigenvector optical signals.
7. The device according to claim 6, characterized in that Any matrix multiplication operator module includes: a field programmable gate array FPGA unit, a Mach-Zehnder interferometer MZI array unit; The FPGA unit is configured to determine, for any two input optical signals, first drive signals corresponding to the two input optical signals based on the weight matrix and the real-time feedback signals corresponding to the multi-channel optical output modules; The MZI array unit is used to perform linear transformation on the coded optical signals corresponding to any two input optical signals according to the first driving signals corresponding to the two input optical signals, and determine the characteristic vector optical signals of the two coded optical signals.
8. The device according to claim 6, characterized in that Any element-by-element multiplication operator module includes: a photodetector, a transimpedance amplifier, and a phase shifter; The photoelectric detector is used to perform photoelectric conversion on any eigenvector optical signal to determine the electrical signal corresponding to the eigenvector optical signal; The transimpedance amplifier is configured to determine a second driving signal corresponding to any one of the eigenvector optical signals based on the electrical signal corresponding to the eigenvector optical signal; The phase shifter is used to perform linear transformation on any one eigenvector optical signal according to the second driving signal corresponding to the other eigenvector optical signal, so as to determine the modulated optical signal of the two eigenvector optical signals.
9. The device according to any one of claims 1 to 3, characterized in that The adjustable optical nonlinear module includes: a plurality of nonlinear layers based on limiting nonlinear materials.
10. The device according to any one of claims 1 to 3, characterized in that The multi-channel light output module is used for: Performing photoelectric conversion on the optical processing result to determine an electrical signal corresponding to the optical processing result; Normalization processing and value weighting processing are performed on the electrical signal corresponding to the optical processing result to determine the target output result.