An image recognition device and recognition method based on a single-chip metasurface
Through the image recognition device based on a monolithic metasurface, the convolution layer is converted to the optical domain, and image convolution is achieved using nanopillars and end-to-end optimization is solved, which solves the problems of slow operation speed and high cost in the prior art, and achieves efficient and low-cost image recognition.
Patent Information
- Application Number
- CN202510661892.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-05-22
AI Technical Summary
Existing image recognition systems have poor real-time performance due to low operating speed and high design and integration costs.
Image recognition devices based on a monolithic metasurface, including a metasurface nanophoton front end, a light intensity detector and a data processing module, are used to convert the convolution layer from the digital domain to the optical domain, image convolution is achieved using the arrangement of nanopillars, and phase and weight are trained using an end-to-end optimization framework.
It significantly improves the portability and flexibility of the system, reduces design and integration costs, improves identification efficiency and accuracy, and is suitable for complex application scenarios.
Smart Images

Figure CN120182557B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image recognition device and method, and particularly to an image recognition device and recognition method based on a single-chip metasurface. Background Art
[0002] The massive data processing tasks in the big data era and the development of artificial intelligence technology have generated an explosive demand for computing power. Due to the constraints of integrated circuit technology and its own architecture, traditional electronic computing systems are difficult to achieve a qualitative improvement in computing power and computing speed. Therefore, the insufficient computing power has become a bottleneck in the development of artificial intelligence technology, and the development of new computing methods has changed from basic research to an urgent practical need.
[0003] Optical computing systems process data through light transmission. Since photons are bosons and have no interaction, they can transmit and process data without interference, with low power consumption and high speed; the multi-wavelength and broadband characteristics of light are conducive to giving full play to the advantages of parallel computing; the propagation characteristics of light determine that light can complete low-precision vector matrix multiplication operations at the speed of light through designed optical paths, with the advantages of real-time and high speed, which has important application value for artificial intelligence models that contain a large number of matrix multiplication operations and have relatively low requirements for precision. It can be seen that the characteristics of high speed, high parallelism and low power consumption of optical computing systems meet the requirements for an ideal computing method. At present, optical computing systems have been proposed as a potential way to alleviate several inherent limitations of digital electronics, and some remarkable progress has been made in optical and photonic processors customized for artificial intelligence, such as spatial differential, integral and convolution calculation methods, whose performance far exceeds that of contemporary electronic processors. Most notably, the optical computing neural network in the optical computing system can perform artificial intelligence inference tasks, such as image recognition.
[0004] Existing optical computing neural networks can be roughly divided into two categories. One is based on integrated photonics, such as Mach-Zehnder interferometers, micro-ring resonators, multi-mode optical fibers, etc., for physically implementing multiply-accumulate floating-point operations; the other is based on free-space optics to implement convolutional layers, and light propagates through diffraction elements. Among them, the main ones for implementing image recognition are 3D printed surfaces and metasurfaces.
[0005] Metasurfaces are composed of arrays of nano-units at the sub-wavelength scale, capable of locally imposing control over the phase, amplitude, and polarization state of incident light. They possess the ability to regulate the optical field and offer a high degree of freedom in the structural design process, and can replace optical devices such as lenses, gratings, and beam splitters. Therefore, metasurfaces exhibit high flexibility and diversity in the design of optical computing devices, and optical computing devices based on different principles and architectures can be realized through metasurface structures with different materials, shapes, and sizes. Different materials such as phase change materials, metals, and all-dielectric materials are usually utilized to achieve various mathematical operations such as integration, differentiation, and filtering, and can be used as optical differentiators, optical integrators, optical filters, and optical computing neural networks, which are widely applied in fields such as image recognition, edge detection, autonomous driving, and machine vision.
[0006] For example, the Chinese invention patent with the publication number CN109344916A discloses a method for microwave imaging and target recognition of field-programmable deep learning. This method controls the switching of different metasurface encodings through an FPGA, and adopts a deep learning network including a DeepNIS network and a CNN network. This deep learning network is sample-driven, captures image features, and conducts target recognition and classification. However, this method mainly focuses on microwave imaging and target recognition. In a complex electromagnetic environment, the signal-to-noise ratio of microwave imaging is relatively low, and it is easily affected by interference sources, resulting in a decline in imaging quality. In addition, this method requires the control of an FPGA to switch different metasurface encodings, and this process restricts the operating speed of the system to a certain extent, thus posing higher requirements for real-time performance.
[0007] The Chinese invention patent with the publication number CN118535970A discloses a method and system for quasi-neural network classification based on metasurfaces. This method programs a neural network model using multiple programmable metasurface layers. Each programmable metasurface layer corresponds to a metasurface board, and the programmable neurons of the programmable metasurface layer are composed of metasurface neurons on the corresponding metasurface board. Multiple metasurface boards are placed vertically in parallel, and the distance between adjacent metasurface boards is the same. The number of metasurface neurons on each metasurface board can be the same or different. Finally, the electromagnetic wave transmittance of the metasurface neurons is controlled through a field-programmable gate array (FPGA) to achieve target classification. However, the training process of the neural network model in this system requires multi-level synchronous iterative updates, the training process is relatively complex, and the construction requirements for the initial network model are relatively high. At the same time, the design of this system requires the collaborative work of multiple metasurface boards, which brings higher process manufacturing and integration costs at the hardware level. In addition, the preparation of metasurface boards requires high-precision processing technology, and the integration of multiple metasurface boards also requires precise alignment and calibration. Therefore, the design of this system undoubtedly increases its implementation difficulty and investment cost.
[0008] The Chinese invention patent with the publication number CN119065142A discloses a metasurface device with polarization-switchable image recognition function and its implementation method. This method simulates the real optical diffraction process through a diffraction neural network, which consists of multiple layers. The phase distribution of each layer is precisely designed to modulate the incident light wave; the parameters of the diffraction layer are optimized by using the gradient descent algorithm and the backpropagation algorithm to achieve the image recognition function; by using the principle of dynamic phase combined with geometric phase, a metasurface structure array is generated according to the optimized phase parameters of the diffraction layer to achieve the corresponding function. However, the diffraction neural network in this metasurface device consists of multiple layers. Since the phase distribution of each layer needs to be precisely designed to modulate the incident light wave, the integration of multiple metasurface plates is required to achieve optical diffraction, thus increasing the integration cost of the system. Summary of the Invention
[0009] The object of the present invention is to solve the technical problems of poor real-time performance caused by low operating speed in the existing image recognition system, as well as relatively high design cost and integration cost, and to provide an image recognition device and recognition method based on a single metasurface.
[0010] To achieve the above object, the technical solution provided by the present invention is as follows:
[0011] An image recognition device based on a single metasurface, which is characterized in that:
[0012] It includes a metasurface nano-optical front end, a light intensity detector and a data processing module;
[0013] The metasurface nano-optical front end has a periodic structure, which includes a substrate layer and a plurality of nano-columns arranged in an array on the substrate layer; the side of the substrate layer away from the nano-columns is connected to the detection surface of the light intensity detector and is adapted to the size of the detection surface; the plurality of nano-columns are arranged on the substrate layer according to the metasurface transmission phase principle, and are used to provide corresponding phases to achieve convolution of the target image;
[0014] The data processing module is connected to the light intensity detector. The light intensity detector is used to receive the convolved target image and transmit it to the data processing module, and the data processing module is used to identify and classify the target image.
[0015] Further, the phase delay of the metasurface nano-optical front end needs to satisfy:
[0016]
[0017] Where, is the working wavelength band of the metasurface nano-optical front end, is the refractive index difference between the ordinary axis and the extraordinary axis of the metasurface nano-optical front end,h is the height of the nanocolumns.
[0018] Furthermore, the working wavelength band of the metasurface nanophotonic front end is a single wavelength, and the error of its phase delay within this working wavelength band is less than 5%.
[0019] Furthermore, the arrangement of the nanocolumns is determined by the following method:
[0020] Step 1: Use CIFAR-10 as the data set and input it into the convolutional neural network for classification training. After the training is completed, extract the first-layer convolutional weights in the digital domain of the convolutional neural network, and then splice these convolutional weights to form the target point spread function in the optical domain;
[0021] Step 2: Randomly generate a phase and perform Fourier transform on this phase to obtain a point spread function;
[0022] Step 3: Compare the point spread function with the target point spread function and determine whether the loss value between the two meets the requirements of the target loss value. If not, execute Step 4; if so, execute Step 5;
[0023] Step 4: Convert the point spread function into the weights in the digital domain and form an end-to-end optimization framework with the data processing module; then optimize the phase in Step 2 through the optimization framework according to the loss value until the loss value between the obtained point spread function and the target point spread function meets the requirements of the target loss value, and execute Step 5;
[0024] Step 5: Use the phase corresponding to the point spread function at this time as the optimal phase to guide the arrangement of the nanocolumns.
[0025] Furthermore, in Step 1, the convolutional neural network is the ResNet50 convolutional neural network;
[0026] In Step 3, the requirement for the target loss value is < 10%.
[0027] Furthermore, the period of the metasurface nanophotonic front end is 0.4 micrometers.
[0028] Furthermore, the substrate layer selects a silica substrate;
[0029] Each nanocolumn is a columnar structure made of silicon nitride, and the size of the area surrounded by all columnar structures is 1 * 1 mm.
[0030] Furthermore, one side of the substrate layer away from the nanocolumns is connected to the detection surface of the light intensity detector by optical bonding or material growth.
[0031] Furthermore, the light intensity detector selects a visible light detector.
[0032] Meanwhile, the present invention also provides an image recognition method based on a single-chip metasurface. Using the above-mentioned image recognition device based on a single-chip metasurface, it includes the following steps:
[0033] Step 1: The target image is convolved by the nanocolumns of the metasurface nanophotonic front end and then received by the light intensity detector, and is transmitted to the data processing module;
[0034] Step 2: The data processing module identifies and classifies the target image information based on the ResNet50 convolutional neural network, thereby realizing the recognition of the target image.
[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0036] 1. The image recognition device based on a single-chip metasurface provided by the present invention includes a metasurface nanophotonic front end, a light intensity detector, and a data processing module. Among them, the metasurface nanophotonic front end is used to replace the bulky Fourier filter based on the 4f system, realizing the high integration and miniaturization of the system, greatly optimizing the space utilization rate, not only reducing the design cost and integration cost, but also significantly improving the portability and flexibility of the system, making it show higher adaptability and practicality in complex application scenarios.
[0037] 2. In the image recognition device based on a single-chip metasurface provided by the present invention, the arrangement of the nanocolumns in the metasurface nanophotonic front end is achieved by converting the convolutional layer from the digital domain to the optical domain, thereby reducing the remaining computational amount in the digital domain and improving the recognition efficiency of the image recognition device. At the same time, it can also be integrated into the camera optical system under natural light conditions, making full use of the design space of optical convolution.
[0038] 3. For the image recognition device and recognition method based on a single-chip metasurface provided by the present invention, the arrangement of the nanocolumns in the metasurface nanophotonic front end adopts an end-to-end optimization framework, and simultaneously optimizes and selects the phase in the optical domain and the weight in the digital domain, thereby maximizing the overall recognition accuracy and recognition efficiency.
[0039] 4. The image recognition device based on a single-chip metasurface provided by the present invention has a simple structure, strong compatibility, can be integrated into the light intensity detector, has functional independence, and can complete image recognition within a single component. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a schematic structural diagram of an embodiment of the present invention;
[0041] Figure 2 is a schematic structural diagram of a nanocolumn in the metasurface nanophotonic front end obtained by simulation of an embodiment of the present invention;
[0042] Figure 3 It is a partial side view of the metasurface nano - photon front - end obtained by simulation in an embodiment of the present invention;
[0043] Figure 4 It is a phase schematic diagram of the metasurface nano - photon front - end in an embodiment of the present invention;
[0044] Figure 5 It is a confusion matrix diagram based on an embodiment of the present invention.
[0045] The reference numerals are as follows:
[0046] 1 - metasurface nano - photon front - end, 11 - substrate layer, 12 - nano - pillars, 2 - light intensity detector, 3 - data processing module. Detailed implementation manners
[0047] To make the objectives, advantages, and features of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Those skilled in the art should understand that these implementation manners are only used to explain the technical principles of the present invention, and the purpose is not to limit the protection scope of the present invention.
[0048] As Figure 1 shown, this embodiment provides an image recognition device based on a monolithic metasurface, including a metasurface nano - photon front - end 1, a light intensity detector 2, and a data processing module 3. Among them, the metasurface nano - photon front - end 1 has a periodic structure, which transforms the first layer in the convolutional neural network from the digital domain to the optical domain to execute a spatially - varying convolutional layer with a large kernel, thereby reducing the remaining computational amount in the digital domain and greatly improving the recognition efficiency of the image recognition device.
[0049] The metasurface nano - photon front - end 1 of this embodiment includes a substrate layer 11 and a number of nano - pillars 12 arranged in an array on the substrate layer 11. One side of the substrate layer 11 away from the nano - pillars 12 is connected to the detection surface of the light intensity detector 2 by optical bonding or material growth, and is adapted to the size of the detection surface.
[0050] A number of nano - pillars 12 are arranged on the substrate layer according to the metasurface transmission phase principle to provide corresponding phases to implement convolution of the target image. The basic idea of the transmission phase is to modulate the refractive index through spatial transformation of the geometric dimensions of the micro - structure. By changing the size of the nano - pillars 12, that is, the size of the micro - structure, the control of the transmission phase can be achieved.
[0051] In this embodiment, the arrangement of the nano - pillars 12 is determined by the following method:
[0052] Step 1: Use CIFAR-10 as the dataset and input it into the ResNet50 convolutional neural network for classification training. After training, extract the first-layer convolutional weights in the digital domain of the convolutional neural network, and then splice these convolutional weights to form the target point spread function in the optical domain.
[0053] Step 2: Randomly generate a phase and perform Fourier transform on this phase to obtain a point spread function.
[0054] Step 3: Compare the point spread function with the target point spread function and determine whether the loss value between the two meets the requirement of the target loss value (<10%). If not, execute Step 4; if so, execute Step 5.
[0055] Step 4: Convert the point spread function into the weights in the digital domain and form an end-to-end optimization framework with the data processing module 3; then optimize the phase in Step 2 through the optimization framework according to the loss value until the loss value between the obtained point spread function and the target point spread function meets the requirement of the target loss value, and then execute Step 5.
[0056] Step 5: Take the phase corresponding to the point spread function at this time as the optimal phase , and use this to guide the arrangement of the nanocolumns 12, and construct the metasurface nanophotonic front-end 1 according to the design principle of the metasurface.
[0057] Combined Figure 2 and Figure 3 As shown, in this embodiment, the finite-difference time-domain (FDTD) method is used, and at the same time, based on the transmission phase principle of the metasurface, the metasurface nanophotonic front-end 1 is modeled and simulated. The period of the metasurface nanophotonic front-end 1 is P, and the height of each nanocolumn 12 is h , and the length and width are L and W respectively, and the scanning ranges of the length and width are determined according to the processing aspect ratio.
[0058] In addition, the phase delay of the metasurface nanophotonic front-end 1 also needs to satisfy:
[0059]
[0060] where is the working wavelength band of the metasurface nanophotonic front-end 1, is the difference in refractive indices between the ordinary axis and the extraordinary axis of the metasurface nanophotonic front-end 1, h is the height of the nanocolumn 12. When When this occurs, the metasurface nanophotonic front-end 1 can perform convolution on the target image. In this embodiment, the working band of the metasurface nanophotonic front-end 1 is a single wavelength, and the error of its phase delay within this working band is less than 5%.
[0061] Figure 4 This is the phase schematic diagram of the metasurface nanophotonic front-end 1 in this embodiment. Taking the case of light beam perpendicular incidence as an example, during FDTD simulation, the light after passing through the metasurface has a phase delay, which is achieved by the propagation of light in the microstructure with a high aspect ratio to accumulate the phase. After sampling and selection, in this embodiment, it is finally determined that the substrate layer 11 in the metasurface nanophotonic front-end 1 selects a silica substrate, each nanorod 12 selects a columnar structure made of silicon nitride material, the period P of the metasurface nanophotonic front-end 1 is 0.4 microns, and the height h of the nanorod 12 is 0.8 microns, and the size of the area where all nanorods 12 are located (that is, the area surrounded by all nanorods 12) is 1 * 1 mm.
[0062] The light intensity detector 2 selects a visible light detector, and its photosensitive area corresponds to the area where a number of nanorods 12 are located. The light intensity detector 2 is connected to the data processing module 3, and is used to receive the convolution target image and transmit it to the data processing module 3, and the data processing module 3 performs recognition and classification on the target image.
[0063] The present invention proposes an image recognition device with a single-chip metasurface, fast operating speed, and real-time imaging. This recognition device can achieve image recognition by using a single-chip metasurface, without the need to switch the integration and alignment with multiple metasurfaces, thereby reducing the integration and manufacturing costs. At the same time, the present invention transforms the convolution layer in the convolutional neural network from the mathematical domain to the optical domain to execute the spatially varying convolution layer of the large kernel, captures the output feature map for the image sensor, and finally processes it in the data processing module 3. By moving the convolution layer into the optics, the present invention significantly reduces the remaining computational amount in the digital domain, thereby improving the recognition efficiency. In addition, the end-to-end optimization framework can jointly train the optical parameters (phase) and digital layer parameters (weights), thereby maximizing the overall recognition accuracy and recognition efficiency of the device, and avoiding the design of the optical neural network architecture being fundamentally restricted by the underlying network design, that is, the challenges of expanding to a large number of neurons and the lack of scalable energy-saving nonlinear optical operators. The present invention can directly extract and process information in the optical domain, providing the possibility for the development of compact and high-performance image recognition devices.
[0064] The image recognition device based on the single-chip metasurface of the present invention has a simple structure, strong compatibility, and functional independence, and can complete image recognition within a single component. Its low-power consumption characteristics and high spatial utilization efficiency are especially suitable for the field of image recognition.
[0065] In addition, this embodiment also provides an image recognition method based on a single-chip metasurface, including the following steps:
[0066] Step 1: The target image is convolved by the nanocolumns 12 of the metasurface nano-optical front end 1 and then received by the light intensity detector 2, and is transmitted to the data processing module 3.
[0067] Step 2: The data processing module 3 identifies and classifies the target image information based on the ResNet50 convolutional neural network, thereby realizing the recognition of the target image.
[0068] As Figure 5 shown, this embodiment uses a confusion matrix to represent the recognition result. This figure can visually display the classification and recognition effect of the device, and the accuracy rate of classification and recognition can reach more than 82%, and the recognition effect is better.
[0069] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the present invention.
Claims
1. An image recognition device based on a monolithic metasurface, characterized by: It includes a metasurface nanophotonic front end (1), a light intensity detector (2) and a data processing module (3); The metasurface nanophotonic front end (1) has a periodic structure, comprising a substrate layer (11) and a plurality of nanopillars (12) arranged on the substrate layer (11) and arranged in an array; a side of the substrate layer (11) away from the nanopillars (12) is connected to a detection surface of a light intensity detector (2) and is adapted to the size of the detection surface; the plurality of nanopillars (12) are arranged on the substrate layer according to the metasurface transmission phase principle and are used to provide corresponding phases to achieve convolution of a target image; the data processing module (3) is connected to the light intensity detector (2), the light intensity detector (2) is used to receive the convolved target image and transmit it to the data processing module (3), and the data processing module (3) is used to identify and classify the target image; Phase delay of the metasurface nanophotonic front end (1) Need to meet: Wherein, λ is the operating wavelength band of the metasurface nanophotonic front end (1), Δn is the difference in refractive index between the ordinary axis and the extraordinary axis of the metasurface nanophotonic front end (1), and h is the height of the nanopillar (12); The working wavelength band of the metasurface nanophotonic front end (1) is a single wavelength, and the error of its phase delay within the working wavelength band is less than 5%; The arrangement of the nanopillars (12) is determined by the following method: Step 1: Use CIFAR-10 as the dataset and input it into a convolutional neural network for classification training. After training, extract the first layer of convolutional weights in the digital domain of the convolutional neural network. Then, concatenate the convolutional weights to form the target point spread function in the optical domain. Step 2: randomly generate a phase and perform Fourier transform on the phase to obtain a point spread function; Step 3: Compare the point spread function with the target point spread function and determine whether the loss value between the two meets the target loss value requirement. If not, proceed to step 4; if yes, proceed to step 5. Step 4: Convert the point spread function into a weight in the digital domain and form an end-to-end optimization framework with the data processing module (3); then optimize the phase of step 2 through the optimization framework according to the loss value until the loss value between the obtained point spread function and the target point spread function meets the target loss value requirement, and then execute step 5; Step 5: The phase corresponding to the point spread function at this time is used as the optimal phase to guide the arrangement of the nanorods (12).
2. The image recognition device based on a single-chip metasurface according to claim 1, characterized in that: In step 1, the convolutional neural network is a ResNet50 convolutional neural network; In step 3, the target loss value is required to be less than 10%.
3. The image recognition device based on a single-chip metasurface according to claim 1 or 2, characterized in that: The period of the supersurface nanophotonic front end (1) is 0.4 micrometers.
4. The image recognition device based on a single-chip metasurface according to claim 3, characterized in that: The substrate layer (11) is a silicon dioxide substrate; Each nanocolumn (12) is a columnar structure made of silicon nitride, and the size of the area enclosed by all the columnar structures is 1*1 mm.
5. The image recognition device based on a single-chip metasurface according to claim 4, characterized in that: The side of the substrate layer (11) away from the nanocolumns (12) is connected to the detection surface of the light intensity detector (2) by means of optical bonding or material growth.
6. The image recognition device based on a single-chip metasurface according to claim 5, characterized in that: The light intensity detector (2) is a visible light detector.
7. An image recognition method based on a single-chip metasurface, characterized in that: The image recognition device based on a single-chip metasurface according to any one of claims 1 to 6 comprises the following steps: Step 1: The target image is convolved with the nanopillars (12) of the metasurface nanophotonic front end (1), received by the light intensity detector (2), and transmitted to the data processing module (3); Step 2: The data processing module (3) recognizes and classifies the target image information based on the ResNet50 convolutional neural network, thereby realizing the recognition of the target image.
Citation Information
Patent Citations
A microwave imaging and target recognition method based on in-situ programmable depth learning
CN109344916A
Quasi-neural network classification method and system based on metasurface
CN118535970A
Polarization-switchable metasurface device with image recognition function and implementation method of polarization-switchable metasurface device
CN119065142A
Device components formed of geometric structures
US20180045953A1
Monocular snapshot four-dimensional imaging method and system
US20230289989A1