Optical neural network system adaptive to multi-class neural network architecture
By designing an optical neural network system including lasers, amplitude-type and phase-type spatial light modulators, 4f systems and imaging devices, the problem that existing systems are difficult to adapt to multiple neural network architectures is solved, and direct transplantation of mainstream neural network architectures and high-speed and low-power inference tasks are realized.
Patent Information
- Application Number
- CN202411882548.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-06
AI Technical Summary
Existing optical neural network systems are difficult to adapt to multi-class neural network architectures, especially when dealing with complex artificial intelligence tasks, and face the challenges of computing requirements.
An optical neural network system including a laser, an amplitude-type spatial light modulator, a 4f system, a phase-type spatial light modulator and an imaging device is designed. The system can realize the computing functions of multiple types of neural network architectures through a variety of optical components and system components.
This system enables mainstream neural network architectures (such as convolutional neural networks, fully connected neural networks, Vision Transformer networks, etc.) to be directly transplanted into the optical system, achieving high-speed and low-power inference tasks without redesigning the optical system or performing additional on-optical training.
Smart Images

Figure CN119940439A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of optical neural network technology, and in particular to an optical neural network system adapted to multiple types of neural network architectures. Background Art
[0002] Artificial intelligence (AI) technology is rapidly changing the way people live, but at the same time, its training and reasoning also put tremendous pressure on digital processors. Optics, with its high bandwidth and low power consumption, has become one of the ideal candidate carriers for accelerating neural networks; in recent years, a variety of optical neural network (ONN) systems and solutions have emerged. However, the characteristics of optical systems determine that they can often only realize one or a class of computing functions, such as convolution, Fourier transform, etc. For example, a simple convolutional neural network (CNN) includes a multi-channel convolution layer and a fully connected layer. In order to implement it using existing optical systems, separate architectures need to be built for these two layers, and multiple optical propagations are required for convolution. This is why existing ONN systems usually adopt a hybrid optoelectronic approach, using optical systems to implement only one type of computing to avoid overly complex structures.
[0003] However, the field of artificial intelligence is developing rapidly, and new neural network architectures are constantly emerging. If existing ONN structures have difficulty matching the latest neural network designs and meeting computational requirements, they will inevitably face challenges when applied to more complex AI tasks. For example, the recent surge in large-scale models led by Transformer has pushed the application of neural networks to another level, not only in language models such as ChatGPT, but also in the field of image and video processing called VisionTransformer (ViT). The basic ViT includes complex computational processes such as multi-channel convolutional layers, multi-head self-attention layers, and multi-sample fully connected layers, which brings difficulties to porting it to optics using the current ONN architecture. Summary of the invention
[0004] The purpose of the present invention is to provide an optical neural network system that is adaptable to multiple types of neural network architectures.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] An optical neural network system adapted to multiple types of neural network architectures, the system comprising: a laser, a first amplitude-type spatial light modulator, a 4f system, a phase-type spatial light modulator, a second amplitude-type spatial light modulator and an imaging device, wherein the laser is used to emit a coherent linearly polarized light signal, the coherent linearly polarized light signal is collimated, expanded and polarized to form a parallel light beam that irradiates the surface of the first amplitude-type spatial light modulator, the first amplitude-type spatial light modulator modulates the input matrix of each layer of the neural network onto the parallel light beam in the form of amplitude, the 4f system images the light field distribution on the surface of the first amplitude-type spatial light modulator onto the surface of the phase-type spatial light modulator, the light beam after phase modulation by the phase-type spatial light modulator reaches the surface of the second amplitude-type spatial light modulator, the second amplitude-type spatial light modulator modulates the weight matrix of the neural network onto its surface light beam in the form of amplitude, the modulated light beam reaches the surface of the imaging device, and the multiplication operation of the input matrix and the weight matrix is completed.
[0007] The system also includes a first polarization wave plate and a first polarization beam splitter cube. The first polarization wave plate converts its input light into P polarized light. The optical properties of the first polarization beam splitter cube are P light transmission and S light reflection. The coherent linear polarized light signal emitted by the laser is converted into parallel light after collimation and expansion and enters the first polarization wave plate and is converted into P polarized light. The P polarized light passes through the first polarization beam splitter cube and reaches the surface of the first amplitude-type spatial light modulator.
[0008] The first amplitude-type spatial light modulator modulates the input matrix of each layer of the neural network in the form of amplitude onto the parallel light beam on its surface, and at the same time converts the P-polarized light into S-polarized light and reflects it. The reflected light is reflected by the first polarization beam splitter cube and then enters the 4f system.
[0009] The system also includes a depolarizing beam splitter cube, which has a splitting ratio of 50:50 and an optical characteristic of P light transmission and S light reflection. The 4f system includes a first spherical lens, a second polarizing wave plate and a second spherical lens. The second polarizing wave plate converts its input light into P polarized light. The light beam entering the 4f system passes through the first spherical lens, the second polarizing wave plate and the second spherical lens in sequence, and then passes through the depolarizing beam splitter cube to reach the surface of the phase-type spatial light modulator.
[0010] The system also includes a first cylindrical lens group, a second polarization beam splitter cube and a second cylindrical lens group. The optical characteristics of the second polarization beam splitter cube are P light transmission and S light reflection. The light beam modulated by the phase-type spatial light modulator passes through the first cylindrical lens group and the second polarization beam splitter cube in sequence and reaches the surface of the second amplitude-type spatial light modulator. The second amplitude-type spatial light modulator modulates the weight matrix of the neural network in the form of amplitude onto the surface light beam, and at the same time converts the P polarized light into S polarized light and reflects it. The reflected light is reflected by the second polarization beam splitter cube and then passes through the second cylindrical lens group to reach the surface of the imaging device.
[0011] The wavelength of the laser is 532 nm.
[0012] The pixel size of the amplitude-type spatial light modulator is 8 μm, the number of pixels is 1200×1920, the modulation accuracy is 8 bits, the maximum refresh rate is 180 Hz per second, and the zero-order diffraction efficiency is 95%;
[0013] The pixel size of the phase-type spatial light modulator is 8μm, the number of pixels is 1200×1920, the modulation accuracy is 10bits, the maximum refresh rate is 180Hz per second, and the zero-order diffraction efficiency is 95%.
[0014] The focal lengths of the first spherical lens and the second spherical lens are 150 mm.
[0015] The first cylindrical lens group and the second cylindrical lens group each include three cylindrical lenses arranged in parallel, wherein the focal lengths of the first and third cylindrical lenses are 100 mm, and the focal length of the second cylindrical lens is 200 mm.
[0016] The imaging device is a qCMOS camera with a quantization accuracy of 16 bits.
[0017] Compared with the prior art, the present invention has the following beneficial effects:
[0018] The present invention constructs an optical system that is adapted to multiple types of neural network architectures, so that mainstream neural network architectures (such as convolutional neural networks, fully connected neural networks, Vision Transformer networks, etc.) can be directly transplanted into the optical system described in the present invention without the need to redesign and build different optical systems. At the same time, the system does not require additional optical training or optical simulation training. The standard neural network architecture trained based on the GPU system can be directly transplanted to light, and the reasoning task can be completed with the high speed and low power consumption characteristics of light. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic diagram of the system structure of the present invention;
[0020] Figure 2 An example of a CNN based on an optical system in the MNIST dataset in one embodiment;
[0021] Figure 3 An example of a ViT network based on an optical system in the MNIST dataset in one embodiment;
[0022] The reference numerals in the figure are: 1-laser, 2-first amplitude-type spatial light modulator, 3-4f system, 4-phase-type spatial light modulator, 5-second amplitude-type spatial light modulator, 6-imaging device, 7-first polarization wave plate, 8-first polarization beam splitter cube, 9-depolarization beam splitter cube, 10-first cylindrical lens group, 11-second polarization beam splitter cube, 12-second cylindrical lens group, 301-first spherical lens, 302-second polarization wave plate, 303-second spherical lens. DETAILED DESCRIPTION
[0023] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0024] This embodiment provides an optical neural network system that is adaptable to multiple types of neural network architectures, such as Figure 1 As shown, the system includes: a laser 1 (Laser), a first amplitude spatial light modulator 2 (Amplitude spatial light modulator 1, ASLM1), a 4f system 3, a phase spatial light modulator 4 (Phase spatial light modulator, PSLM), a second amplitude spatial light modulator 5 (Amplitude spatial light modulator 2, ASLM2) and an imaging device 6. In addition, a first polarization wave plate 7 (Wave Plate 1, WP1), a first polarizing beam splitter cube 8 (Polarizing Beam Splitter 1, PBS1), a depolarizing beam splitter cube 9 (Non-Polarizing Beam Splitter, NPBS), a first cylindrical lens group 10, a second polarizing beam splitter cube 11 (Polarizing Beam Splitter 2, PBS2) and a second cylindrical lens group 12. Among them, the 4f system 3 includes a first spherical lens 301 (Lens1), a second polarization wave plate 302 (Wave Plate2, WP2) and a second spherical lens 303 (Lens2).
[0025] In this embodiment, the wavelength of the laser 1 is 532nm; the pixel size of the amplitude-type spatial light modulator is 8μm, the number of pixels is 1200×1920, the modulation accuracy is 8bits, the maximum refresh rate is 180Hz per second, and the zero-order diffraction efficiency is 95%; the focal length of the first spherical lens 301 and the second spherical lens 303 is 150mm; the pixel size of the phase-type spatial light modulator 4 is 8μm, the number of pixels is 1200×1920, the modulation accuracy is 10bits, and the maximum refresh rate is 1 per second. 80Hz, the zero-order diffraction efficiency is 95%; the first cylindrical lens group 10 is composed of three cylindrical lenses arranged in parallel, which are denoted as CL1, CL2, and CL3 respectively, and the second cylindrical lens group 12 is also composed of three cylindrical lenses arranged in parallel, which are denoted as CL4, CL5, and CL6 respectively, wherein the focal lengths of CL1, CL3, CL4, and CL6 are 100mm, and the focal lengths of CL2 and CL5 are 200mm; the imaging device 6 is a quantitative complementary metal oxide semiconductor camera (qCMOS camera) with a quantization accuracy of 16bits.
[0026] The first polarization wave plate 7 and the second polarization wave plate 302 convert their input light into P polarized light; the optical characteristics of the first polarization beam splitter cube 8 and the second polarization beam splitter cube 11 are P light transmission and S light reflection; the depolarization beam splitter cube 9 has a splitting ratio of 50:50, and its optical characteristics are P light transmission and S light reflection.
[0027] like Figure 1As shown, the laser 1 emits a coherent linear polarized light signal, which is converted into parallel light after collimation and beam expansion, and enters the first polarization wave plate 7 to be converted into P polarized light. The P polarized light passes through the first polarization beam splitter cube 8 to reach the surface of the first amplitude-type spatial light modulator 2. The first amplitude-type spatial light modulator 2 modulates the input matrix of each layer of the neural network in the form of amplitude to the parallel light beam on its surface, and at the same time converts the P polarized light into S polarized light and reflects it. The reflected light is reflected by the first polarization beam splitter cube 8 and enters the 4f system 3. The light beam entering the 4f system 3 passes through the first spherical lens 301, the second polarization wave plate 302 and the second spherical lens 303 in sequence, and then changes from S polarized light to P polarized light again, and then passes through the depolarization beam splitter cube 9 to reach the surface of the phase-type spatial light modulator 4. The light beam modulated by the phase-type spatial light modulator 4 passes through the first cylindrical lens group 10 (CL1, CL2, CL3) and the second polarization beam splitter cube 11 in sequence and reaches the surface of the second amplitude-type spatial light modulator 5. The second amplitude-type spatial light modulator 5 modulates the weight matrix of the neural network in the form of amplitude onto its surface light beam, and at the same time converts the P-polarized light into S-polarized light and reflects it. The reflected light is reflected by the second polarization beam splitter cube 11 and then passes through the second cylindrical lens group 12 (CL4, CL5, CL6) to reach the surface of the imaging device 6.
[0028] like Figure 2, this embodiment uses a most basic single-layer CNN network to describe the principle of the system. The network structure is: a 50-channel convolution layer; a standard fully connected layer, and a nonlinear activation function (rectified linear unit). For the MNIST and Fashion MNIST datasets, the original image size is [28,28]. After padding [1,2,1,2], it becomes a matrix of size [31,31]. The convolution kernel size is [6,6] and the step size is 5. Therefore, one convolution requires 6×6=36 dot products between an image slice of size [6,6] and a convolution kernel of size [6,6] to cover an input matrix (image) of size [31,31]. This layer will obtain a convolution result of size [6,6]. For 50 channels, this is a process of converting a [1,28,28] tensor to a [50,6,6] tensor. The dot product between the image slice of size [6,6] and the convolution kernel of size [6,6] can be regarded as the dot product between the image slice vector of size [1,36] and the convolution kernel vector of size [1,36], and the 36 dot products can be regarded as the vector-matrix multiplication between the vector of size [1,36] and the matrix of size [36,36]. Therefore, the 50-channel convolution can be abstracted as the matrix-matrix multiplication between the matrix of size [36,36] and the matrix of size [36,50]. In fact, this is also the way it is implemented on the GPU. This process can be implemented by the optical system of the present invention, and the nonlinear function can be implemented by setting the photosensitivity curve of qCMOS, which means that the result of qCMOS has been processed by ReLU. Then, the result is flattened to a vector of size [1,1800], and after the fully connected layer, a vector of size [1,10] is obtained, and the process can be abstracted as the vector-matrix multiplication between the vector of size [1,1800] and the matrix of size [1800,10]. The [1,10] size vector represents the ten-class result, passes through the Softmax function, uses the vector with the image label, calculates the loss through the cross entropy loss function, and then performs backpropagation to update the gradient and complete the training. Therefore, the system implements all CNN processes except flattening through two light propagations.
[0029] like Figure 3For the ViT network, the present invention uses a standard single-layer ViT structure: an embedding layer, an encoder layer, and a classifier layer. In the Embedding layer, an embedding block (patch embedding) is generated by 50-channel convolution according to the same process as the multi-channel convolution layer in the above optical CNN. The difference is that the convolution kernel size is [4,4] and the step size is 4. Therefore, a convolution kernel requires 7×7=49 dot products to cover an input image of size [28,28]. This process can also be abstracted as a matrix-matrix multiplication between a matrix of size [49,16] and a matrix of size [16,50] to obtain a matrix of size [49,50]. Then, the class token of size [1,50] should be added to the patch embedding to form a matrix of size [50,50]. In the Encoder layer, multi-head self-attention is described as mapping a query (Q) and a set of key-value (KV) pairs: attention = QK T V. Q, K, and V are the Encoder layer input matrix and weight matrix W using [50,50] size respectively. Q , W K , and W V Obtained by matrix-matrix multiplication, the size of each weight matrix is also [50,50]. The input of the Encoder layer is then added to the output of the self-attention, and then processed through the multi-sample fully connected layer to obtain the [50,50] size output matrix of the Encoder layer. The multi-sample fully connected layer can be abstracted as a matrix-matrix multiplication between a matrix of size [50,50] and a matrix of size [50,50]. The Classifier layer retrieves a [1,50] size class token from the standardized output of the Encoder layer. The class token then passes through two simple fully connected layers to obtain a [1,10] size classification result vector. The loss is also calculated through the Softmax function and the cross entropy loss function, and then backpropagation is performed to update the gradient.
[0030] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. An optical neural network system adapted to multiple types of neural network architectures, characterized in that: The system comprises: a laser (1), a first amplitude-type spatial light modulator (2), a 4f system (3), a phase-type spatial light modulator (4), a second amplitude-type spatial light modulator (5) and an imaging device (6), wherein the laser (1) is used to emit a coherent linear polarized light signal, the coherent linear polarized light signal is collimated, expanded and polarized to form a parallel light beam that irradiates the surface of the first amplitude-type spatial light modulator (2), the first amplitude-type spatial light modulator (2) modulates the input matrix of each layer of the neural network onto the parallel light beam in the form of amplitude, the 4f system (3) images the light field distribution on the surface of the first amplitude-type spatial light modulator (2) onto the surface of the phase-type spatial light modulator (4), the light beam after phase modulation by the phase-type spatial light modulator (4) reaches the surface of the second amplitude-type spatial light modulator (5), the second amplitude-type spatial light modulator (5) modulates the weight matrix of the neural network onto its surface light beam in the form of amplitude, the modulated light beam reaches the surface of the imaging device (6), and the multiplication operation of the input matrix and the weight matrix is completed.
2. The optical neural network system adapted to multiple types of neural network architectures according to claim 1, characterized in that: The system further comprises a first polarization wave plate (7) and a first polarization beam splitter cube (8); the first polarization wave plate (7) converts its input light into P polarized light; the optical characteristics of the first polarization beam splitter cube (8) are P light transmission and S light reflection; the coherent linear polarized light signal emitted by the laser (1) is converted into parallel light after collimation and beam expansion and enters the first polarization wave plate (7) and is converted into P polarized light; the P polarized light passes through the first polarization beam splitter cube (8) and reaches the surface of the first amplitude-type spatial light modulator (2).
3. The optical neural network system adapted to multiple types of neural network architectures according to claim 2, characterized in that: The first amplitude-type spatial light modulator (2) modulates the input matrix of each layer of the neural network in the form of amplitude onto the parallel light beam on its surface, and at the same time converts the P-polarized light into S-polarized light and reflects it. The reflected light is reflected by the first polarization beam splitting cube (8) and then enters the 4f system (3).
4. The optical neural network system adapted to multiple types of neural network architectures according to claim 1, characterized in that: The system further comprises a depolarizing beam splitter cube (9), the depolarizing beam splitter cube (9) having a light splitting ratio of 50:50 and an optical characteristic of P light transmission and S light reflection. The 4f system (3) comprises a first spherical lens (301), a second polarizing wave plate (302) and a second spherical lens (303), the second polarizing wave plate (302) converting its input light into P polarized light. The light beam entering the 4f system (3) passes through the first spherical lens (301), the second polarizing wave plate (302) and the second spherical lens (303) in sequence, and then passes through the depolarizing beam splitter cube (9) to reach the surface of the phase-type spatial light modulator (4).
5. The optical neural network system adapted to multiple types of neural network architectures according to claim 1, characterized in that: The system further comprises a first cylindrical lens group (10), a second polarization beam splitting cube (11) and a second cylindrical lens group (12); the optical characteristics of the second polarization beam splitting cube (11) are P light transmission and S light reflection; the light beam modulated by the phase-type spatial light modulator (4) passes through the first cylindrical lens group (10) and the second polarization beam splitting cube (11) in sequence and then reaches the surface of the second amplitude-type spatial light modulator (5); the second amplitude-type spatial light modulator (5) modulates the weight matrix of the neural network in the form of amplitude onto the light beam on its surface, and at the same time converts the P polarized light into S polarized light and reflects it; the reflected light is reflected by the second polarization beam splitting cube (11) and then passes through the second cylindrical lens group (12) to reach the surface of the imaging device (6).
6. The optical neural network system adapted to multiple types of neural network architectures according to claim 1, characterized in that: The wavelength of the laser (1) is 532 nm.
7. The optical neural network system adapted to multiple types of neural network architectures according to claim 1, characterized in that: The pixel size of the amplitude-type spatial light modulator is 8 μm, the number of pixels is 1200×1920, the modulation accuracy is 8 bits, the maximum refresh rate is 180 Hz per second, and the zero-order diffraction efficiency is 95%; The pixel size of the phase-type spatial light modulator (4) is 8 μm, the number of pixels is 1200×1920, the modulation accuracy is 10 bits, the maximum refresh rate is 180 Hz per second, and the zero-order diffraction efficiency is 95%.
8. The optical neural network system adapted to multiple types of neural network architectures according to claim 4, characterized in that: The focal lengths of the first spherical lens (301) and the second spherical lens (303) are 150 mm.
9. The optical neural network system adapted to multiple types of neural network architectures according to claim 5, characterized in that: The first cylindrical lens group (10) and the second cylindrical lens group (12) each include three cylindrical lenses arranged in parallel, wherein the focal lengths of the first and third cylindrical lenses are 100 mm, and the focal length of the second cylindrical lens is 200 mm.
10. The optical neural network system adapted to multiple types of neural network architectures according to claim 1, characterized in that: The imaging device (6) is a qCMOS camera with a quantization accuracy of 16 bits.