All-optical convolutional neural network design method for image recognition

By designing a full-optical convolutional neural network, using a multi-core metasurface convolutional device and an optical full-connection layer for full-optical processing, the existing optical neural networks have solved the problems of low accuracy and high crosstalk in image recognition, and achieved high recognition accuracy and low energy consumption image recognition effect.

CN119990218AActive Publication Date: 2025-05-13UNIV OF SCI & TECH OF CHINA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510057557.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

In the image recognition, existing optical neural networks have problems such as low recognition accuracy, poor signal-to-noise ratio of the output plane and large crosstalk in various energy areas, especially when the digital connection layer is required after the convolution operation, resulting in partial delay of electronic parts.

Method used

A fully optical convolutional neural network is designed, including a lens, a multi-core metasurface convolutional device and an optical fully connected layer. Image recognition is achieved through full optical processing, a multi-core metasurface convolutional operation is used for convolution operations, and diffraction modulation is performed through an optical fully connected layer. Only one diffraction layer is needed to achieve high recognition accuracy.

Benefits of technology

The accuracy of image recognition and output plane signal-to-noise ratio are significantly improved, the crosstalk of each energy area is reduced, and multi-task recognition is realized, which only needs to change the phase distribution of the optical fully connected layer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990218A_ABST
    Figure CN119990218A_ABST
Patent Text Reader

Abstract

The invention discloses an all-optical convolutional neural network design method for image recognition, and relates to the technical field of image recognition, and the method comprises the steps: inputting a to-be-detected image into an all-optical convolutional neural network which comprises a lens, a multi-core metasurface convolver and an optical full connection layer; the target convolution kernel is flatly laid on one face and placed in the front focal plane of the lens, a multi-kernel metasurface convolver is obtained on the rear focal plane of the lens, and the multi-kernel metasurface convolver is used for conducting convolution operation on the to-be-detected image to obtain a convolved array image; the optical full connection layer is an array image after diffraction modulation processing convolution, and outputs an identification result of the image to be detected; according to the image recognition method, all-optical processing is carried out between input and output, a multi-kernel super-surface convolver is utilized to realize multi-kernel convolution operation, and the convolved array image can realize optical full-connection operation through an optical diffraction layer; due to the fact that image features are extracted through convolution operation, differentiation of different types of images is increased, and compared with a traditional diffraction neural network, especially when full connection layers are single layers, the image recognition accuracy and the output face signal-to-noise ratio are obviously improved, and crosstalk of all energy areas is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to a method for designing an all-optical convolutional neural network for image recognition. Background Art

[0002] Convolutional Neural Networks (CNN) is a deep feedforward neural network with the characteristics of local connection and weight sharing. It is also one of the representative algorithms of deep learning. It is good at processing images, especially image recognition and other related machine learning problems. For example, it has a significant improvement effect in various visual tasks such as image classification, target detection, and image segmentation. It is one of the most widely used models.

[0003] However, as the amount of neural network tasks and data increases, the central processing unit in the computer designed by the traditional von Neumann architecture is very time-consuming and cannot continuously process large amounts of data. Therefore, some customized computing platforms such as graphics processing units, field programmable logic gate arrays, and tensor processing units and other advanced electronic devices have been proposed to accelerate computing. However, Moore's Law predicts that the number of transistors on a single chip in an electronic device will double every two years. As the density of transistors increases, their feature size is getting closer and closer to the boundary between macroscopic physics and quantum physics. In addition, the speed and energy consumption of hardware accelerators are limited by the heat generated by the charging and discharging of parasitic capacitances, electronic crosstalk, electromagnetic radiation, and the movement of electrons.

[0004] Optical computing uses the high speed, low power consumption, high bandwidth, inherent parallelism and compatibility with the semiconductor industry of photons to replace electrons for computing. It can overcome the inherent limitations of electronics and is expected to improve energy efficiency by 6 orders of magnitude and computing speed by 3 orders of magnitude.

[0005] At present, most optical neural networks are all-optical diffraction neural networks based on digital fully connected neural networks. They are based on a series of diffraction layers to realize the recognition and imaging of the input light field, while the underlying feature convolution operation of extracting the image is not reflected, so more diffraction layers are required. In addition, although most convolution accelerators realize convolution operations, they require digital connection layers behind them, so the delay of the electronic part will affect the speed of recognition or imaging. Summary of the invention

[0006] Based on the technical problems existing in the background technology, the present invention proposes an all-optical convolutional neural network design method for image recognition, which improves the accuracy of image recognition.

[0007] The present invention proposes an all-optical convolutional neural network design method for image recognition, comprising:

[0008] Inputting the image to be detected into an all-optical convolutional neural network, which includes a lens, a multi-core metasurface convolution device, and an optical fully connected layer;

[0009] The target convolution kernel is flattened on a surface and placed in the front focal plane of the lens. A multi-core hypersurface convolver is obtained at the back focal plane of the lens. The multi-core hypersurface convolver is used to perform a convolution operation on the image to be detected to obtain an array image after convolution.

[0010] The optical fully connected layer is the array image after the diffraction modulation processing convolution, and outputs the recognition result of the image to be detected.

[0011] Furthermore, when there is only one optical fully connected layer, its accuracy, output surface signal-to-noise ratio, and crosstalk in each energy region are significantly better than those of existing diffraction neural networks.

[0012] Furthermore, the target convolution kernel is selected based on the difference in the structural similarity index between images of different categories before and after convolution. The judgment is based on selecting a convolution kernel with a small structural similarity index and a large differentiation, so as to facilitate the recognition of images of different categories.

[0013] Furthermore, the multi-core metasurface convolver adopts standard micro-nano processing technology.

[0014] Furthermore, the training process of the all-optical convolutional neural network is as follows:

[0015] Construct a training set based on the MNIST handwriting dataset and the label value corresponding to each data;

[0016] Randomly select N data in the training set, illuminate the image with parallel light, and then sequentially enter the lens, multi-core metasurface convolution device, and optical fully connected layer, and then calculate the total energy of N target areas on the output surface;

[0017] The energy of each target area and the label value matrix of the corresponding data are calculated to construct a loss function, and the phase of the diffraction layer in the optical fully connected layer is updated using the gradient descent method until the set number of iterations is completed.

[0018] Furthermore, the loss function is a cross entropy loss function.

[0019] Furthermore, the light field distribution E3 of the array image after convolution on the convolution plane is as follows:

[0020]

[0021] Among them, (x3, y3) is the position coordinate of the convolution surface, is the result without modulation by the multi-core metasurface convolver. is the target convolution kernel, PSF is the electric field distribution of the image to be detected on the convolution surface, U(x0,y0) is the light field distribution of the input image, is the equivalent aperture on the lens spectrum plane, h is the amplitude and phase distribution of the multi-core metasurface convolver, Represents a convolution operation.

[0022] An image recognition system based on an all-optical convolutional neural network, which includes a lens, a multi-core metasurface convolution device and an optical fully connected layer;

[0023] The target convolution kernel is flattened on a surface and placed in the front focal plane of the lens. A multi-core hypersurface convolver is obtained at the back focal plane of the lens. The multi-core hypersurface convolver is used to perform a convolution operation on the image to be detected to obtain an array image after convolution.

[0024] The optical fully connected layer is used to perform diffraction modulation processing on the convolved array image and output the recognition result of the image to be detected.

[0025] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the all-optical convolutional neural network design method as described above when executing the computer program.

[0026] A computer-readable storage medium having stored thereon a plurality of classification programs, wherein the plurality of classification programs are used to be called by a processor and execute the all-optical convolutional neural network design method as described above.

[0027] The advantages of the all-optical convolutional neural network design method for image recognition provided by the present invention are as follows: the all-optical convolutional neural network design method for image recognition provided in the structure of the present invention is all-optical processing between input and output, and a multi-core metasurface convolver is used to implement a multi-core convolution operation, the purpose of which is to extract the underlying features of the image, so that the structural similarity differences of input images of different categories become larger, and the array image after convolution can realize an optical full connection operation through an optical full connection layer, and only one diffraction layer is required to achieve a high recognition accuracy, and multi-task recognition can be achieved by only changing the phase distribution of the optical full connection layer; because the convolution operation extracts image features, the differentiation of images of different categories becomes larger, compared with traditional diffraction neural networks, especially when the fully connected layers are all single layers, the image recognition accuracy and output surface signal-to-noise ratio are significantly improved, and the crosstalk of each energy region is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a schematic diagram of the optical path of the present invention;

[0029] Figure 2 A schematic diagram of the process of processing silicon-based supersurfaces using electron beam lithography;

[0030] Figure 3 Schematic diagram of the optimization process for the optical fully connected layer;

[0031] Figure 4 The intensity distribution of the all-optical convolutional neural network and the diffractive neural network on the output surface when the digital image '6' is incident;

[0032] Figure 5 The simulated and experimental intensity distribution of the all-optical convolutional neural network on the output surface when the digital images from '0' to '9' are incident;

[0033] Figure 6 Simulated and experimental intensity distributions of the all-optical convolutional neural network on the output surface when the gesture images ‘0’ to ‘9’ are incident. DETAILED DESCRIPTION

[0034] Below, the technical solution of the present invention is described in detail through specific embodiments. Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific implementation disclosed below.

[0035] like Figures 1 to 6 As shown, the present invention proposes an all-optical convolutional neural network design method for image recognition, comprising:

[0036] Inputting the image to be detected into an all-optical convolutional neural network, which includes a lens, a multi-core metasurface convolution device, and an optical fully connected layer;

[0037] The target convolution kernel is flattened on a surface and placed in the front focal plane of the lens. A multi-core hypersurface convolver is obtained at the back focal plane of the lens. The multi-core hypersurface convolver is used to perform a convolution operation on the image to be detected to obtain an array image after convolution.

[0038] The optical fully connected layer is used to perform diffraction modulation processing on the convolved array image and output the recognition result of the image to be detected.

[0039] In this embodiment, all-optical processing is performed between input and output, and a multi-core metasurface convolver is used to implement a multi-core convolution operation, the purpose of which is to extract the underlying features of the image and increase the structural similarity of input images of different categories. The array image after convolution can be subjected to an optical fully connected layer to implement an optical fully connected operation, and only one diffraction layer is required to achieve a high recognition accuracy. In addition, multi-task recognition can be achieved by only changing the phase distribution of the diffraction layer in the optical fully connected layer.

[0040] 1. Multi-core Hypersurface Convolver

[0041] Image features are parts of an image that are easily recognizable, unique, and different from other areas. They often contain rich information, such as the location of important objects in the image or edge changes. The basis for selecting convolution kernels is to use the difference in the structural similarity index (SSIM) of different categories of images before and after convolution to make judgments.

[0042] The convolution kernel of the all-optical convolutional neural network selects the Sobel operator for extracting features in the digital convolutional neural network. The Sobel operator is a one-dimensional first-order operator commonly used for edge detection. The difference between different categories of digital images before and after convolution becomes larger. The convolution kernel is a target convolution kernel tiled with Sobel operators in different rotation directions to ensure that edges in different directions are used. The target convolution kernel is placed on the front focal plane of the lens, and after optical simulation transmission, a multi-core metasurface convolver is obtained on the back focal plane of the lens (i.e., the spectrum plane). Figure 1 As shown, the input image passes through the lens and the multi-core metasurface convolver to obtain an array image after convolution with different kernels on the imaging surface, and then the convolved array image passes through a series of diffraction layers to reach the output surface.

[0043] In this embodiment, the multi-core metasurface convolver adopts standard micro-nano processing technology, such as Figure 2 As shown. The silicon-based super surface first uses chemical vapor deposition to deposit a 300nm thick single-crystal silicon film on a sapphire substrate, which determines the height of the designed silicon nanostructure. Then, a positron beam photoresist with a thickness of 100nm is coated on the silicon film and baked. Subsequently, an electron beam lithography machine is used to form a pattern on the photoresist with a mask of device structure information under an acceleration voltage of 100kV. To retain the desired pattern, 15nm thick chromium is deposited. Therefore, the pattern of the nanostructure is transferred to the chromium film, and the remaining photoresist is peeled off. Next, an inductively coupled plasma etcher is used to etch a 300nm thick silicon layer. Finally, a dry etching technique is used to remove the residual chromium mask to obtain the expected super surface structure.

[0044] 2. Optical Fully Connected Layer

[0045] The optical fully connected layer includes one or more diffraction layers. All diffraction layers act as optical fully connected layers to perform optical diffraction modulation on the convolved array image. After the design of the optical fully connected layer is completed, it needs to be trained before deployment. The training process of the optical fully connected layer is detailed as follows (III. Training of the all-optical convolutional neural network). The all-optical convolutional neural network composed of the optical fully connected layer and the multi-core metasurface convolution device can realize the recognition of image categories.

[0046] 3. Training of all-optical convolutional neural networks;

[0047] The training process of the all-optical convolutional neural network is as follows:

[0048] S1, build a training set based on the MNIST handwriting data set and the label value corresponding to each data;

[0049] S2, randomly select N data in the training set, illuminate the input image with parallel light, and then enter the lens, multi-core metasurface convolution device, and optical fully connected layer in sequence, and then calculate the total energy of N target areas on the output surface;

[0050] If a point source is incident, passes through the lens, and then passes through the multi-core metasurface convolver placed on the rear focal plane (i.e., the spectrum plane), an electric field distribution is formed on the convolution surface, namely:

[0051]

[0052] Among them, PSF (x3, y3) is the point spread function distribution of the convolution layer, (x0, y0) is the position coordinate of the input surface, (x1, y1) is the position coordinate of the lens surface, (x2, y2) is the position coordinate of the multi-core hypersurface convolver surface, (x3, y3) is the position coordinate of the convolution array image surface, f1 is the distance from the eyepiece to the front focal plane or to the back focal plane, d′2 is the distance from the lens to the convolution surface, d′1 is the distance between the input image and the lens, M is the ratio between d′2 and d′1, that is, M=d′2 / d′1, M1 is the ratio between f1 and d′1, that is, M1=f1 / d′1, M2 is the ratio between (d′2-f1) and d′2, that is, M2=(d′2-f1) / d′2. Among them, A is a constant, i is an imaginary unit, is the equivalent aperture on the lens spectrum plane, represents the convolution operation, h(x2,y2) is the amplitude and phase distribution of the multi-core metasurface convolver, λ is the wavelength of light, and f x and f y is the spatial frequency of the (x2, y2) plane. The equivalent aperture on the x2-y2 plane can be approximated as: is the tiled distribution of the multi-core convolution operator, M1=f1 / d′1, M2=(d′2-f1) / d′2, M=d′2 / d′1, and PSF(x3,y3) is the point spread function distribution of the convolution layer.

[0053] Therefore, any input image E0(x0,y0) is distributed on the convolution surface after passing through the convolution layer:

[0054]

[0055] is the result without modulation by the multi-core metasurface convolution device (only the result of the lens input image on the convolution surface). It is regarded as the target multi-convolution kernel, that is, the multi-convolution kernel operator.

[0056] This embodiment designs a multi-core metasurface convolver through the above (i). After the multi-core metasurface convolver is fixed, in order to achieve the category recognition effect of the all-optical convolutional neural network on the image, the phase of the diffraction layer in the optical fully connected layer is optimized, and the existing diffraction integral formula is used to calculate the light field distribution of the input image through the convolution layer and then through the diffraction layer to the target surface, and the energy of each preset target area is calculated, and then the loss function is calculated with the label value matrix, and the phase of the diffraction layer is optimized using a suitable algorithm according to the loss function.

[0057] S3, calculate the energy of each target area and the label value matrix of the corresponding data to construct a loss function, and use the gradient descent method to update the phase of the diffraction layer in the optical fully connected layer until the set total number of iterations is reached;

[0058] Specifically: Figure 3 As shown in the figure, the input recognition image passes through the convolution layer to obtain the convolution array image on the convolution surface, and the energy of ten regions is calculated on the output surface after passing through the diffraction layer, and the cross entropy loss function is calculated with the label value matrix. Then, the diffraction layer phase is updated using the gradient descent method according to the loss function. If the required maximum number of iterations is not reached, continue to randomly select data and input it into the all-optical convolutional neural network until the maximum number of iterations is reached and the training is terminated. After the training is completed, the all-optical convolutional neural network can realize the recognition of the training task category, including untrained images.

[0059] Through steps S1 to S3, the all-optical convolutional neural network is trained, and the image is recognized all-optically based on the trained all-optical convolutional neural network. Compared with the existing digital convolutional neural network, the all-optical convolutional neural network has the advantages of fast task execution speed and low energy consumption.

[0060] During the execution process after the all-optical convolutional neural network training is completed, the image to be detected passes through the lens, multi-core metasurface convolution device and optical fully connected layer in sequence to obtain the total energy of N target areas. The area corresponding to the maximum energy among the N target energies is used as the recognition result of the image to be detected.

[0061] In order to verify that the all-optical convolutional neural network of this embodiment has a high accuracy in image recognition, the existing diffraction neural network is used as a comparison. In this embodiment, the diffraction layer is set to one layer. If the all-optical convolutional neural network of one diffraction layer is higher than the diffraction neural network, the accuracy of the all-optical convolutional neural network in image recognition of this embodiment is higher than that of the existing diffraction neural network.

[0062] Specifically, when the diffraction layer is only one layer, the all-optical convolutional neural network of this embodiment shows better performance than the diffraction neural network. The details are as follows: the all-optical convolutional neural network can achieve nearly 90% accuracy for the MNIST handwriting data set, while the diffraction neural network can only achieve 79% accuracy; in addition, when different numbers are incident, the maximum regional energy in the ten target areas is calculated to predict the input digital category. When the number '6' is incident, it can be seen from the outputs of the all-optical convolutional neural network and the diffraction neural network (such as Figure 5 As shown in Figure 2, the all-optical convolutional neural network has a higher signal-to-noise ratio and lower crosstalk between the ten regions. When other digital images are incident, the output results of the all-optical convolutional neural network are as follows: Figure 5 shown.

[0063] The proposed all-optical convolutional neural network can realize multi-task recognition. Keeping the convolution layer unchanged, re-optimizing the diffraction layer, and using the optical fully connected layer can realize the recognition of other tasks. Figure 6 As shown in the figure, when different gesture images are incident, the trained all-optical convolutional neural network can accurately focus the light on the target area on the output surface, calculate the gesture corresponding to the maximum regional energy in the ten target areas, and achieve correct recognition.

[0064] In this embodiment, the convolution operation extracts image features, making the differentiation of images of different categories greater. Compared with the traditional diffraction neural network, especially when the fully connected layers are all single layers, the image recognition accuracy and output surface signal-to-noise ratio are significantly improved, and the crosstalk of each energy region is reduced. The all-optical convolutional neural network can be used for image recognition, object classification, style recognition, etc. Compared with the existing diffraction neural network, the optical convolution operation in the all-optical convolutional neural network in this embodiment can extract the underlying features of the image. When the diffraction layer is only one layer, the accuracy of the all-optical convolutional neural network is significantly higher than that of the diffraction neural network. In addition, it can also reduce signal crosstalk and improve the signal-to-noise ratio. The optical fully connected layer uses a reconfigurable diffraction unit to achieve multi-task recognition.

[0065] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A method for designing an all-optical convolutional neural network for image recognition, characterized in that: include: Inputting the image to be detected into an all-optical convolutional neural network, which includes a lens, a multi-core metasurface convolution device, and an optical fully connected layer; The target convolution kernel is flattened on a surface and placed in the front focal plane of the lens. A multi-core hypersurface convolver is obtained at the back focal plane of the lens. The multi-core hypersurface convolver is used to perform a convolution operation on the image to be detected to obtain an array image after convolution. The optical fully connected layer is the array image after the diffraction modulation processing convolution, and outputs the recognition result of the image to be detected.

2. The all-optical convolutional neural network design method for image recognition according to claim 1, characterized in that: When there is only one optical fully connected layer, its accuracy, output signal-to-noise ratio, and crosstalk in each energy region are better than those of existing diffraction neural networks.

3. The all-optical convolutional neural network design method for image recognition according to claim 2, characterized in that: The basis for selecting the target convolution kernel is to judge by the difference in the structural similarity index of different categories of images before and after convolution. The judgment basis is to select a convolution kernel with a small structural similarity index and a large differentiation, so as to facilitate the recognition of different categories of images.

4. The all-optical convolutional neural network design method for image recognition according to claim 1, characterized in that: The multi-core metasurface convolver uses standard micro-nano processing technology.

5. The all-optical convolutional neural network design method for image recognition according to claim 1, characterized in that: The training process of the all-optical convolutional neural network is as follows: Construct a training set based on the MNIST handwriting dataset and the label value corresponding to each data; Randomly select N data in the training set, illuminate the image with parallel light, and then sequentially enter the lens, multi-core metasurface convolution device, and optical fully connected layer, and then calculate the total energy of N target areas on the output surface; The energy of each target area and the label value matrix of the corresponding data are calculated to construct a loss function, and the phase of the diffraction layer in the optical fully connected layer is updated using the gradient descent method until the set number of iterations is completed.

6. The all-optical convolutional neural network design method for image recognition according to claim 5, characterized in that: The loss function is a cross entropy loss function.

7. The all-optical convolutional neural network design method for image recognition according to claim 1, characterized in that: The light field distribution E3 of the array image after convolution on the convolution plane is as follows: Among them, (x3, y3) is the position coordinate of the convolution surface, is the result without modulation by the multi-core metasurface convolver. is the target convolution kernel, PSF is the electric field distribution of the image to be detected on the convolution surface, U(x0,y0) is the light field distribution of the input image, is the equivalent aperture on the lens spectrum plane, h is the amplitude and phase distribution of the multi-core metasurface convolver, Represents a convolution operation.

8. An image recognition system based on an all-optical convolutional neural network, characterized in that: The all-optical convolutional neural network includes a lens, a multi-core metasurface convolutional unit, and an optical fully connected layer; The target convolution kernel is flattened on a surface and placed in the front focal plane of the lens. A multi-core hypersurface convolver is obtained at the back focal plane of the lens. The multi-core hypersurface convolver is used to perform a convolution operation on the image to be detected to obtain an array image after convolution. The optical fully connected layer is used to perform diffraction modulation processing on the convolved array image and output the recognition result of the image to be detected.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the all-optical convolutional neural network design method as described in any one of claims 1-7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of classification programs, and the plurality of classification programs are used to be called by the processor and execute the all-optical convolutional neural network design method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • All-optical image processing system based on metasurface device and processing method

    CN113589409A

  • Image recognition method of all-optical nonlinear convolutional neural network

    CN114255387A

  • Associated imaging target identification method and device based on all-optical neural network

    CN118072097A

  • All optical neural network

    US20200327403A1

  • Devices and methods employing optical-based machine learning using diffractive deep neural networks

    US20210142170A1