A metasurface device for dual-target spatial position and image recognition and its implementation method

By designing a metasurface device based on diffraction neural network and geometric phase, spatial position and image recognition of dual targets is achieved, the problem of low efficiency of multi-objective recognition and classification is solved, the recognition accuracy and efficiency are improved, and it is suitable for intelligent perception and artificial intelligence fields.

CN118484891BActive Publication Date: 2025-08-12UNIV OF SHANGHAI FOR SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410663158.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2025-08-12
Estimated Expiration
2044-05-27

AI Technical Summary

Technical Problem

The identification and classification of existing diffraction neural networks in scenarios where multiple targets have spatial location information are challenged, limiting their application potential in complex environments.

Method used

A metasurface device for spatial position and image recognition of dual targets is designed, using diffraction neural network and geometric phase principles, the diffraction layer parameters are optimized through gradient descent algorithm and backpropagation algorithm to generate a metasurface structural array to achieve spatial positioning and image recognition of dual targets.

Benefits of technology

It improves the efficiency of multi-objective detection tasks, has the advantages of small size, high integration, low computing resource consumption, and short training time, and is suitable for intelligent perception and artificial intelligence fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118484891B_ABST
    Figure CN118484891B_ABST
Patent Text Reader

Abstract

The present invention relates to a metasurface device for dual-target spatial position and image recognition and an implementation method thereof. The metasurface device comprises a metal digital imaging plate, a metasurface layer, and a detection plane; the metasurface layer comprises a dielectric substrate and an array of nanopillar structures placed on the dielectric substrate; the diffraction neural network corresponding to the target recognition of the device is composed of a multi-layer structure, wherein the phase distribution of each diffraction layer is precisely designed to modulate the incident light wave. The diffraction layer parameters are optimized by applying the gradient descent algorithm and the back-propagation algorithm, thereby realizing the functions of spatial positioning and image recognition of different targets. The geometric phase principle is utilized to generate a metasurface structure array based on the optimized diffraction layer phase parameters to realize the corresponding functions. This solves the problem that current metasurface devices can only realize the recognition and classification of a single target, and simultaneously recognizes and classifies the spatial position and image recognition of dual targets, greatly improving the efficiency of processing multi-target detection tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a metasurface device, and in particular to a metasurface device for realizing dual-target spatial position and image recognition based on a diffraction neural network and geometric phase. Background Art

[0002] Traditional image recognition systems rely on electronic computers to process large amounts of data and execute complex algorithms to achieve spatial positioning and recognition of targets, which has limitations in resource consumption and speed. All-optical deep learning is an emerging research field that shows great potential to change the status quo. In 2018, the Ozcan research group first disclosed an artificial neural network technology based on optical devices, namely the diffraction neural network (DNN). 2 NN), using the principle of optical diffraction to complete information processing. 2 In NN, the units on each metasurface layer can be compared to artificial neurons, and the connections between them are realized through the Huygens-Fresnel diffraction law. When light passes through these metasurfaces, it is modulated by amplitude and phase. Using computational optimization methods, the response of each metasurface unit can be adjusted so that D 2 NN can perform functions such as image recognition, classification, and solving eigenvalue problems. The technology operates at the speed of light and has low energy consumption and high parallel processing capabilities, so it has obvious advantages in data processing speed and energy efficiency.

[0003] Metasurface technology—a nanostructure that can control light wavefronts—has been proven to precisely manipulate light. Metasurfaces are composed of arrays of subwavelength-scale nanounits that can locally control the phase, amplitude, and polarization state of incident light. Geometric phase, also known as Pancharatnam-Berry (PB) phase, is achieved by designing and selecting appropriate unit sizes and shapes, arranging them according to the phase distribution calculated in a diffraction neural network, and introducing the special effects of geometric phase. This allows light to obtain a phase delay related to its rotation angle when passing through these structures. This allows for direct information extraction and processing in the optical domain, making it possible to develop compact, high-performance image recognition equipment.

[0004] By combining the diffraction neural network with the geometric phase of the metasurface, the micro-nano unit structure of the metasurface is used to simulate neurons in deep learning and realize parallel processing of optical signals. This optical neural network significantly enhances the data processing capability and brings breakthrough progress to all-optical computing and machine vision. The integrated design enables the device to perform high-efficiency optical calculations while maintaining the miniaturization and integration of the system, which lays the foundation for more complex optical signal processing tasks and intelligent perception applications. Existing diffraction neural networks have been successful in single-target recognition, but the recognition and classification of multiple targets with spatial position information still face challenges, limiting their application potential in complex environments. Summary of the Invention

[0005] In view of the problem that current metasurface devices can only realize the recognition and classification of a single target, in order to enhance the ability of multi-target detection position detection and classification tasks, a metasurface device and implementation method for dual-target spatial position and image recognition is proposed, which can distinguish the positions of dual-target inputs and perform recognition and classification.

[0006] The present invention belongs to the fields of optical neural networks, micro-nano optics, optical information processing and machine vision. The diffraction neural network corresponding to the target identification of the device is composed of a multi-layer structure, in which the phase distribution of each diffraction layer is precisely designed to modulate the incident light wave. The gradient descent algorithm and the back propagation algorithm are applied to optimize the parameters of the diffraction layer, thereby realizing the functions of spatial positioning and image recognition of different targets. The geometric phase principle is utilized to generate a metasurface structure array according to the optimized diffraction layer phase parameters to realize the corresponding functions. The metasurface device of the present invention can realize the simultaneous recognition and classification of the spatial position and image recognition of dual targets, greatly improving the efficiency of processing multi-target detection tasks, and has the advantages of small size, high integration, low computing resource consumption, short training time, etc., and is suitable for wide application in the fields of intelligent perception and artificial intelligence.

[0007] The technical solution of the present invention is:

[0008] A metasurface device for dual-target spatial position and image recognition includes a metal digital imaging plate, a metasurface layer, and a detection plane, which are sequentially arranged. The metal digital imaging plate includes a silicon dioxide substrate and a gold film disposed on the transparent silicon dioxide substrate, and is shaped like the pattern to be recognized. The metasurface layer has a periodic structure, including a dielectric substrate and a nanopillar structure array disposed on the dielectric substrate. The dielectric substrate is made of silicon. The nanopillar structure array can provide a corresponding phase distribution to achieve light diffraction. The phase distribution of the nanopillar structure array is derived from the model training results of a diffraction neural network that matches the application.

[0009] The diffraction neural network structure includes an input layer, a diffraction layer, and an output layer: the input layer encodes the pattern to be identified into the amplitude of the input field of the diffraction neural network, thereby converting the image information into an optical signal, which is then input into the next-level diffraction layer for processing; the diffraction layer is composed of a finely controlled phase distribution, which is equivalent to the weight matrix in the deep learning neural network, and modulates the propagation of the light wave by precisely controlling the phase of the incident light wave; the output layer is used to detect the light wave output by the diffraction layer and perform image recognition and spatial positioning tasks; the output layer is a pre-set detection plane, and the diffraction network is trained to map the input target to the corresponding detection area; the detection plane is divided into several detection areas, and the number of sub-detection areas corresponds to the number of objects to be identified and the number of position classifications;

[0010] Extract the optimized diffraction layer phase distribution text file from the diffraction neural network structure, and construct the metasurface nanopillar array using the geometric phase principle;

[0011] The metasurface device realizes diffraction of incident light of a specific frequency after passing through the metal digital imaging plate, and finally forms a diffraction spot at the preset corresponding position of the detection plane, realizing the dual-target spatial position and image recognition function; first, the incident light is input and passes through the metal digital imaging plate placed on the silicon dioxide substrate. The metal part of the imaging plate is opaque, and the patterned hollow part is transmissive. The transmitted light is then diffracted by the nano-pillar structure array placed on the dielectric substrate, and finally forms a diffraction spot on the detection plane at a certain distance from the metasurface, and the target recognition result is obtained.

[0012] Preferably, the nanocolumn structure is designed according to transmittance requirements and can operate at different light wavelengths.

[0013] Preferably, the metal digital imaging plate is formed by coating a gold film on a silicon dioxide substrate using a vacuum coating technique, wherein the thickness of the silicon dioxide substrate is 500 microns and the thickness of the gold film is 5 microns.

[0014] Preferably, the dielectric substrate is made of silicon material with a thickness of 500 microns and a structural length and width of 12,000 microns.

[0015] Preferably, the incident light is circularly polarized light, which is an electromagnetic wave whose electric field vector has a constant amplitude but rotates in a constant angular velocity.

[0016] A method for realizing dual-target spatial position and image recognition, based on a metasurface device for realizing dual-target spatial position and image recognition, comprises the following steps:

[0017] The diffraction neural network consists of three parts: input layer, diffraction layer, and output layer. The gradient descent algorithm and back-propagation algorithm are used to train and optimize the diffraction layer parameters. The recognition task data set is used to train the diffraction neural network diffraction layer parameters and determine the diffraction layer phase distribution. The optimized diffraction layer phase distribution text file is extracted, and the geometric phase principle is applied to construct a metasurface nanopillar array. The designed silicon nanopillars realize the function of a half-wave plate at a specific frequency point. The phase change of each silicon nanopillar is twice its own rotation angle. The corresponding phase distribution of the design is achieved by changing the rotation angle of each nanopillar in the nanopillar array, thereby generating a metasurface structure file. At the same time, the pattern to be recognized is processed, and the pattern is hollowed out using a gold film. A silicon dioxide substrate is added, and it is placed at a set distance from the metasurface layer and aligned parallel to the metasurface layer. The irradiation light source passes through the metal imaging plate and the metasurface layer in turn to realize the metasurface device design of dual-target spatial position and image recognition.

[0018] The input layer of the diffraction neural network encodes the image information to be recognized to simulate the amplitude of the modulated light beam, converting the image information into an optical signal, which is then input into the next diffraction layer for processing. The optical signal carrying the amplitude distribution of the image shape interacts with the diffraction layer layer by layer. The optical signal from the input layer to each diffraction layer is interconnected through neurons between layers. The output layer detects the light waves after diffraction transmission and performs image recognition and spatial positioning tasks. Finally, the optimal diffraction layer phase parameters obtained through training are applied to the distribution of the nanopillar structure array to form a metasurface layer, forming a metasurface device for dual-target spatial positioning and image recognition to achieve optimal target recognition results.

[0019] The specific training process of the diffraction neural network model is as follows:

[0020] Step 1) In the diffraction neural network model, the phase and amplitude changes of the light wave are calculated according to the Rayleigh-Sommerfeld diffraction integral formula. At the same time, the network model determines the bias term through the transmission coefficient of each layer of the metasurface. The model uses the forward propagation process to guide the incident light wave carrying the target information through the specially configured multi-layer metasurface structure, forming a diffraction pattern reflecting the characteristics of the target object on the detection plane. The Rayleigh-Sommerfeld diffraction theory describes the propagation mode of the light field from the nth layer to the n+1th layer, where represents the transmission coefficient of the nth layer, where A is the amplitude, is the phase, i is the imaginary unit, and a=1 is set when only the pure phase type diffraction neural network is trained; the light field distribution of the propagation form of the Rayleigh-Sommerfeld diffraction theory is expressed as:

[0021] U(r n+1 )=t(r n )∫∫S U(r n )h(r n+1 -r n )dxdy (1)

[0022] Among them, U(r n ) is the transmitted electric field distribution of the upper layer, U(r n+1 ) is the transmission electric field distribution on the n+1th layer, n is the number of diffraction layers, and the formula h(r n+1 -r n ) represents the impulse response during the transmission process, and the expression is:

[0023]

[0024] in, is the point r on the nth layer n To point r on the n+1th layer n+1 The Euclidean distance between them, λ is the incident wavelength, is the imaginary unit, x n+1 、y n +1 、z n+1 are the spatial coordinate components of a point in the n+1 layer, are the spatial coordinate components of the ath neuron in the nth layer; S represents the integral area, which is all points in the entire wavefront area; the integral represents the weighted and superimposed light field distribution of the entire nth layer to calculate the light field distribution of a point on the n+1th layer; this method can be used to predict and calculate the propagation of light waves in media with different layer structures; the number of diffraction layers is represented by n, where a single diffraction layer n=1 is used, and the n+1th layer is used to represent the detection plane, and the energy distribution S of the sub-detection area of the detection plane is calculated. j =|U j | 2 , j is the sub-detector region number, and the recognition result is determined by the category corresponding to the region with the highest energy;

[0025] Step 2) The MNIST dataset is selected as the training dataset, and a combination of images containing two detection targets is randomly generated as the training set. The purpose is to train the diffraction layer phase parameters of the diffraction neural network to maximize the energy distribution of the corresponding category sub-detection area and minimize the energy distribution of other sub-areas. During the optimization process, the mean square error loss function (MSE) is used to evaluate the difference between the energy distribution of different sub-detector areas and the target energy distribution. The number of identified targets is represented by C:

[0026]

[0027] where s crepresents the energy distribution received in the final detection area, g c represents the target energy distribution, Represents the error between the target area and the detection area; Based on the obtained diffraction pattern and its deviation from the expected result, a loss function is defined to evaluate the recognition accuracy of the diffraction neural network under the current parameters;

[0028] Step 3) Apply the back propagation module and combine it with the gradient descent algorithm to calculate the gradient of the loss function with respect to the phase parameter of the diffraction layer to guide the optimization process of the phase distribution; after multiple iterations until Stop training when convergence is achieved and the accuracy cannot be improved; repeat the above steps to adjust the number of diffraction layers to find the number of diffraction layers that can achieve the object recognition accuracy required by the designer;

[0029] Step 4) iteratively executing steps 1) to 3), continuously updating the diffraction layer phase parameters until a predetermined recognition accuracy or other performance index is met, thereby completing the training and optimization of the diffraction neural network;

[0030] Based on the phase distribution of the diffraction layer determined in step 2), rectangular silicon nanopillar arrays of the same size but different rotation angles are selected to form a multilayer metasurface; the unit structural parameters of these metasurfaces include the length L, width W, height H, and period length P of the nanopillars; the performance of nanopillar structures of different sizes under circularly polarized light is calculated to select the optimal parameter combination required at a specific wavelength and determine the size of the silicon nanopillars; according to the geometric phase principle, the desired phase distribution is achieved by rotating the nanopillar structure; finally, the nanopillar structure is arranged according to the phase distribution obtained by training the diffraction layer of the diffraction neural network, and a metasurface device capable of recognizing the spatial position and image of two targets is constructed based on this arrangement;

[0031] The final phase distribution is in is the phase obtained by training the all-optical diffraction neural network;

[0032] This method can calculate the phase distribution of the diffraction layer according to different recognition tasks, so that the metasurface device can adapt to the resolution and recognition of various objects.

[0033] Preferably, the diffraction neural network training process uses a single diffraction layer, sets the distance from the input layer to the diffraction layer to 10,000 microns, the distance from the diffraction layer to the output layer to 10,000 microns, the learning rate to 0.005, the size of the neuron to 120 microns, and in order to save training time, eight samples are taken as a batch and input into the neural network.

[0034] Preferably, the training model software uses the TensorFlow framework, which is based on the Python programming language.

[0035] Preferably, electromagnetic full-wave simulation software can be used to calculate the performance of nanorod structures of different sizes under circularly polarized light irradiation.

[0036] The beneficial effects of the present invention are:

[0037] 1. The metasurface device disclosed in the present invention realizes dual-target spatial position and image recognition based on diffraction neural network and geometric phase. It is the first metasurface to realize dual-target spatial position and image recognition in the terahertz band. This method can not only use diffraction neural network for image recognition, but also simultaneously identify the spatial position and image features of dual-target objects, thereby improving the efficiency and accuracy of the recognition task.

[0038] 2. The metasurface device disclosed in the present invention, which realizes dual-target spatial position and image recognition based on diffraction neural network and geometric phase, is designed with the ability to operate in different bands such as terahertz, near-infrared and microwave, and its scalability and flexibility in multi-band operation taken into consideration.

[0039] 3. Compared with traditional image recognition devices, the proposed metasurface device based on diffraction neural network and geometric phase to realize dual-target spatial position and image recognition has the advantages of low energy consumption, fast response speed, small size and high integration; the system does not involve mechanical components and can operate at the speed of light, which greatly improves the computing efficiency and response speed; the metasurface nanocolumns have a simple structure, a wide range of materials, are easy to process, and have a wide range of applications; the proposed metasurface device based on diffraction neural network and geometric phase to realize dual-target spatial position and image recognition can achieve frequency band conversion, multi-frequency point conversion, and multi-target recognition by adjusting structural size parameters, setting response frequency bands, etc. according to actual application scenarios and requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 A flow chart of the calculation method for designing the metasurface device of the present invention;

[0041] Figure 2 Schematic diagram of the position of diffraction spots of the input light sources of different shapes and their corresponding output detection layers of the metasurface device of the present invention;

[0042] Figure 3 is a top view of a single periodic structure of the metasurface device of the present invention, wherein: Figure 3 (a) and (b) are schematic diagrams of the unit periodic structure without rotation and the unit periodic structure with rotation, respectively. Angle diagram;

[0043] Figure 4 It is a left view of a single periodic structure of the metasurface device of the present invention;

[0044] Figure 5A top view of the thirty-six periodic structures of the metasurface device of the present invention;

[0045] Figure 6 This is a front view of the metasurface device structure of the present invention;

[0046] Figure 7 A top view of a digital optical imaging plate of a metasurface device according to the present invention;

[0047] Figure 8 These are the images and simulation effects of dual-target handwritten digital image recognition in an embodiment of the present invention.

[0048] Figure ID:

[0049] 1. Nanopillar structure array; 2. Dielectric substrate; 3. Patterned gold film; 4. Silicon dioxide substrate. DETAILED DESCRIPTION

[0050] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0051] A diffraction neural network is a computational model that simulates the principles of optical diffraction. It mimics the way light waves transmit information through a medium, interconnecting layers of neurons. In this model, linear computations at each layer are performed based on the Huygens-Fresnel principle. As light waves propagate through each layer, each diffraction layer generates secondary waves, enabling a top-down signal transmission process. These layers can perform transmission or reflection operations, and each point within a layer is equivalent to a neuron. In a diffraction neural network, the diffraction patterns generated by the propagation and interaction of light waves between different media are used to process information. The network is trained using deep learning on a computer, involving multiple rounds of training and parameter adjustments to optimize the network structure and improve its performance.

[0052] This invention discloses a metasurface device that uses a diffraction neural network and geometric phase to achieve dual-target spatial position and image recognition. The metasurface's phase distribution is calculated by a diffraction neural network, which consists of three components: an input layer, a diffraction layer, and an output layer. Each diffraction layer is precisely calculated and designed to simulate the weight matrix function of a deep learning neural network and is optimized based on specific target tasks. Gradient descent and backpropagation algorithms are used to train the diffraction layer parameters, and the optimal diffraction layer phase distribution is determined based on a specific dataset.

[0053] The metasurface of the present invention has a single periodic structure, including a dielectric substrate and a nanocolumn array structure for phase control. The appropriate geometric dimensions of the metasurface nanocolumn structure are obtained by scanning parameters, and the direction of the metasurface nanocolumns is rotated using the geometric phase principle. According to the optimized diffraction layer phase distribution, the structure file of the metasurface is generated to ensure that the diffraction effect best matches the dual-target recognition requirements and achieves strict modulation of the phase of the incident light beam. Planar light is used to illuminate a metal digital imaging plate to obtain a light beam carrying target image information. When the light beam is projected onto the metasurface, it will follow a predefined path to generate a specific diffraction pattern and form a diffraction spot containing the spatial position of the target object and image recognition information on the output plane. The metasurface device developed by the present invention has the ability to simultaneously process the spatial positioning and image feature processing of two independent targets, significantly improving the efficiency and accuracy of target recognition tasks.

[0054] The metasurface nanopillars are made of silicon, with unit dimensions of 99 microns long, 35 microns wide, a period of 120 microns, and a thickness of 500 microns. At a set operating frequency of 0.6 THz, they function as a half-wave plate, with an array of 100 x 100 nanopillars.

[0055] The dielectric substrate is made of silicon material with a thickness of 500 microns and a structure length and width of 12,000 microns.

[0056] The metal digital imaging plate is made of gold with a thickness of 5 microns, and its substrate is made of silicon dioxide with a thickness of 500 microns.

[0057] These metasurfaces are designed based on the Pancharatnam-Berry (PB) phase principle, also known as geometric phase. When exposed to circularly polarized light, the specific geometric structure introduces additional phase delay related to the polarization state and azimuth angle. The geometric phase modulation method causes the light beam to undergo different phase changes at each azimuth position, thereby achieving full phase control from 0 to 2π while maintaining minimal dispersion characteristics. Circularly polarized light, such as left-handed or right-handed circularly polarized light, when passing through a specially designed metasurface unit, will be converted into circularly polarized light with opposite helical properties due to the influence of the geometric phase, and at the same time carry a phase delay determined by the azimuth angle θ of the nanopillar, which is 2θ.

[0058] A metasurface device based on diffractive neural networks and geometric phase control enables simultaneous spatial position and image recognition of two targets. By designing a specific metasurface structure, the device simulates functionality similar to that of traditional deep learning networks, but performs processing in the optical domain, enabling fast and efficient target detection and classification. This technology belongs to the fields of optical neural networks, micro-nano optics, optical information processing, and machine vision applications.

[0059] The diffraction neural network architecture of a metasurface device that achieves dual-target spatial position and image recognition based on a diffraction neural network and geometric phase consists of three parts: an input layer, a diffraction layer, and an output layer. The input layer uses a digitally encoded imaging plate to encode the image information to be recognized. This image information is applied to the beam amplitude channel and then fed into the diffraction layer as information. The diffraction layer, composed of a phase distribution, simulates the weight matrix in a deep learning neural network and is physically implemented by the metasurface micro-nanostructure. The geometric phase control mechanism, applied to the diffraction layer, consists of a computationally designed nanopillar structure that precisely modulates the beam phase distribution. An output layer detects the light waves output by the diffraction layer and performs image recognition and spatial positioning. The output layer is a pre-prepared detection plane. The recognition targets are dual targets, with two objects to be recognized in each target area. Four equally sized sub-detection areas are defined on the detection plane.

[0060] Among them, the diffraction neural network calculates the entire device in an optical computing manner, and its output layer is an optical signal.

[0061] like Figure 1 As shown, the design process method specifically implemented in this embodiment is as follows:

[0062] An all-optical diffraction neural network was trained using a dual-target combination of handwritten datasets as the training and test sets. The diffraction layer parameters of the diffraction neural network were trained using gradient descent and backpropagation algorithms, combined with a specific recognition task dataset, to obtain the diffraction layer phase distribution of the diffraction neural network. Based on the determined diffraction layer phase distribution, a nanopillar array on the metasurface was designed using the concept of geometric phase. The corresponding metasurface structural engineering design file was created, and a simulation model was established. By implementing a metasurface-based diffraction neural network combined with geometric phase, a metasurface device capable of simultaneously identifying the spatial position and images of dual targets was obtained, promoting the miniaturization and functional integration of multi-target image recognition and classification equipment.

[0063] The Diffraction Deep Neural Network is an all-optical deep learning framework whose core consists of a series of multi-layer diffraction surfaces working together. These surfaces physically form a neural network that can optically perform statistical learning of complex functions. The calculation and prediction mechanism of the physical network is all-optical, and the learning part of its design is completed by computer. This embodiment verifies the effectiveness and reasoning ability of D2NN (Diffraction Deep Neural Network) through simulation, realizing a metasurface device for dual-target spatial position and image recognition. The simulation software uses a time-domain finite difference method.

[0064] The example used is 0.6THz circularly polarized light;

[0065] like Figure 2As shown, the input layer uses a metal imaging plate to encode the information of the object image to be identified, encoding the image information into the beam amplitude channel. The detection plane of the dual-target handwritten digits is divided into four discrete areas, representing the handwritten digits "2" and "6" on the left and the handwritten digits "2" and "6" on the right, respectively. The two sub-detection areas on the left correspond to the left input image. The upper left sub-detection area corresponds to the input left handwritten digit "2", and the lower left sub-detection area corresponds to the input left handwritten digit "6". Similarly, the two sub-detection areas on the right correspond to the right input image.

[0066] like Figure 3 、 4 As shown in the figure, the metasurface device is composed of multiple periods and is designed as a 100*100 period array structure. The FDTD simulation software is used to perform parameter scanning at 0.6 THz to find the optimal structural parameters. The thickness of the nanocolumns in a single period is designed to be h1=500nm, l1=99nm, w1=35nm, and the substrate thickness is h2=500nm. The total thickness h 1+ h2, the length of the unit period in the x and y directions is p = 120nm (the length and width of the substrate are both p, and each periodic structure is a square substrate), and the incident wavelength is 500nm. The long axis of the nanopillar rotates around the center of the unit periodic structure, forming an angle θ with the x-axis of the substrate (the angle of rotation of the nanopillar structure). When the incident light is set to circularly polarized light, the phase modulation value obtained by the transmitted cross-circularly polarized light can cover a range from 0 to 2π.

[0067] like Figure 5 As shown, a total of thirty-six periodic structure diagrams of a partial area are displayed (should be 100*100).

[0068] like Figure 6 As shown, the optical digital imaging plate of the metasurface device uses a patterned gold film with a thickness of h3 = 5 microns, which is placed on a silicon dioxide substrate with a thickness of h4 = 500 microns. The number of nanopillar structure arrays used is 100*100, so the length and width of the metal digital imaging plate are both 100p, and the overall position is located 10,000 microns away from the upper surface of the metasurface layer.

[0069] In this embodiment, all-silicon materials were selected as the dielectric substrate and the metasurface nanopillar structure, and the nanopillar structure was obtained using photolithography and magnetron sputtering coating technology. This metasurface device, which uses a diffraction neural network and geometric phase to achieve dual-target spatial position and image recognition, uses left-handed circularly polarized light at a frequency of 0.6 THz, passing through a digital imaging plate and then irradiating the metasurface layer. This allows for position and image recognition of dual-target digital inputs, with the recognition results displayed on a plane 10,000 microns away from the other side of the metasurface layer.

[0070] like Figure 7 As shown, one of the metal imaging plates is selected for demonstration. The side length is 100 times that of the unit periodic structure, that is, 12,000 microns, which realizes the control of the amplitude of the incident circularly polarized light and applies the image information encoding to the beam amplitude channel.

[0071] like Figure 8 As shown, in a classification simulation, the handwritten digit classification component achieved over 95% accuracy after 15 iterations of training. A training dataset of 11,876 dual-target handwritten digit images was constructed using randomly combined images of two detection targets, and 1,990 dual-target handwritten digit images were selected as the test set. The excellent simulation results demonstrate the effectiveness of the design theory and the device's ability to correctly recognize handwritten digits. Four dual-digit combination patterns were selected for demonstration. When the input number was 22, bright diffraction spots appeared in the upper two locations of the four pre-set squares in the detection area behind the metasurface. When the input number changed to 26, two bright diffraction spots appeared in the main diagonal area. When the input number changed to 62, two bright diffraction spots appeared in the secondary diagonal area. Finally, when the input number changed to 66, bright diffraction spots appeared in the lower two locations of the four pre-set squares in the detection area behind the metasurface. The two detection areas on the left correspond to the recognition results for the left input image, and the two detection areas on the right correspond to the recognition results for the right input image.

[0072] The present invention uses a diffraction neural network to design and optimize a metasurface device that uses a diffraction neural network and geometric phase to achieve dual-target spatial position and image recognition. This is an advanced method that applies deep learning algorithms to optical wavefront design. It considers the physical constraints in optical systems and leverages the powerful fitting capabilities of neural networks to find the optimal nanopillar arrangement. The diffraction layer phase distribution is obtained using a metasurface-based diffraction neural network optimization method. Based on this phase distribution, a corresponding phase-modulated nanopillar structure is selected and a single-layer metasurface is prepared. This results in a metasurface device that uses a diffraction neural network and geometric phase to achieve dual-target spatial position and image recognition.

[0073] The above-described embodiment merely represents one embodiment of the present invention. While the description is relatively specific and detailed, it should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, and these modifications and improvements fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.

Claims

1. A metasurface device for dual-target spatial position and image recognition, characterized in that: The metal digital imaging plate comprises a metal digital imaging plate, a super surface layer, and a detection plane which are arranged in sequence; the metal digital imaging plate comprises a silicon dioxide substrate and a gold film placed on the transparent silicon dioxide substrate, and the shape is the shape of the pattern to be recognized; The metasurface layer has a periodic structure, including a dielectric substrate and a nanorod structure array placed on the dielectric substrate; The dielectric substrate is made of silicon material; The nanopillar array can provide a corresponding phase distribution to achieve light diffraction. The phase distribution of the nanopillar array is derived from the model training results of the diffraction neural network that matches the application. The diffraction neural network structure includes an input layer, a diffraction layer, and an output layer: the input layer encodes the pattern to be recognized into the amplitude of the input field of the diffraction neural network, thereby converting the image information into an optical signal, which is then input into the next level of diffraction layer for processing; The diffraction layer, composed of a finely tuned phase distribution, is equivalent to the weight matrix in a deep learning neural network. It modulates the propagation of light waves by precisely controlling the phase of the incident light waves. The output layer is used to detect the light waves output by the diffraction layer and perform image recognition and spatial positioning tasks. The output layer is a pre-set detection plane, and the diffraction network is trained to map the input target to the corresponding detection area. The detection plane is divided into several detection areas, and the number of sub-detection areas corresponds to the number of objects to be identified and the number of position categories; Extract the optimized diffraction layer phase distribution text file from the diffraction neural network structure, and construct the metasurface nanopillar array using the geometric phase principle; The metasurface device diffracts incident light of a specific frequency after passing through a metal digital imaging plate, ultimately forming a diffraction spot at a preset corresponding position on the detection plane, achieving dual-target spatial position and image recognition functions. First, the incident light is input and passes through a metal digital imaging plate placed on a silicon dioxide substrate. The metal portion of the imaging plate is opaque, while the patterned hollow portion is transmissive. The transmitted light is then diffracted by an array of nanopillar structures placed on a dielectric substrate, ultimately forming a diffraction spot on a detection plane at a certain distance from the metasurface, obtaining the target recognition result. The input layer uses a metal imaging plate to encode the information of the object image to be identified, and encodes the image information into the beam amplitude channel. The detection plane of the double-target handwritten digits is divided into 4 discrete areas, representing the handwritten digits 2 and 6 on the left and the handwritten digits 2 and 6 on the right, respectively. The two sub-detection areas on the left correspond to the left input image, the upper sub-detection area on the left corresponds to the 2 of the input left handwritten digit, the lower sub-detection area on the left corresponds to the 6 of the input left handwritten digit, and the two sub-detection areas on the right correspond to the right input image; four double-digit combination patterns to be identified are selected. When the input digit is When the input number becomes 22, there are bright diffraction spots at the two upper positions of the four preset boxes in the detection area behind the metasurface; when the input number becomes 26, there are two bright diffraction spots on the main diagonal; when the input number becomes 62, there are two bright diffraction spots on the secondary diagonal; and when the input number becomes 66, there are bright diffraction spots at the two lower positions of the four preset boxes in the detection area behind the metasurface. The two detection areas on the left correspond to the recognition results of the input image on the left, and the two detection areas on the right correspond to the recognition results of the input image on the right.

2. The metasurface device for dual-target spatial position and image recognition according to claim 1, characterized in that: The nanocolumn structure is designed according to the transmittance requirements and can operate at different light wavelengths.

3. The metasurface device for dual-target spatial position and image recognition according to claim 1, characterized in that: The metal digital imaging plate is made by coating a gold film on a silicon dioxide substrate using vacuum coating technology. The thickness of the silicon dioxide substrate is 500 microns and the thickness of the gold film is 5 microns.

4. The metasurface device for dual-target spatial position and image recognition according to claim 1, characterized in that: The dielectric substrate is made of silicon material with a thickness of 500 microns and a structure length and width of 12,000 microns.

5. The metasurface device for dual-target spatial position and image recognition according to claim 1, characterized in that: The incident light is circularly polarized light, which is an electromagnetic wave whose electric field vector has a constant amplitude but rotates in a constant angular velocity.

6. A method for realizing dual-target spatial position and image recognition, characterized in that: The implementation of a metasurface device for dual-target spatial position and image recognition based on any one of claims 1 to 5 comprises the following steps: The diffraction neural network consists of three parts: input layer, diffraction layer, and output layer. The gradient descent algorithm and back-propagation algorithm are used to train and optimize the diffraction layer parameters. The recognition task data set is used to train the diffraction neural network diffraction layer parameters and determine the diffraction layer phase distribution. The optimized diffraction layer phase distribution text file is extracted, and the geometric phase principle is applied to construct a metasurface nanopillar array. The designed silicon nanopillars realize the function of a half-wave plate at a specific frequency point. The phase change of each silicon nanopillar is twice its own rotation angle. The corresponding phase distribution of the design is achieved by changing the rotation angle of each nanopillar in the nanopillar array, thereby generating a metasurface structure file. At the same time, the pattern to be recognized is processed, and the pattern is hollowed out using a gold film. A silicon dioxide substrate is added, and it is placed at a set distance from the metasurface layer and aligned parallel to the metasurface layer. The irradiation light source passes through the metal imaging plate and the metasurface layer in turn to realize the metasurface device design of dual-target spatial position and image recognition. The input layer of the diffraction neural network encodes the image information to be recognized to simulate the amplitude of the modulated light beam, converting the image information into an optical signal, which is then input into the next diffraction layer for processing. The optical signal carrying the amplitude distribution of the image shape interacts with the diffraction layer layer by layer. The optical signal from the input layer to each diffraction layer is interconnected through neurons between layers. The output layer detects the light waves after diffraction transmission and performs image recognition and spatial positioning tasks. Finally, the optimal diffraction layer phase parameters obtained through training are applied to the distribution of the nanopillar structure array to form a metasurface layer, forming a metasurface device for dual-target spatial positioning and image recognition to achieve optimal target recognition results. The specific training process of the diffraction neural network model is as follows: Step 1) In the diffraction neural network model, the phase and amplitude changes of the light wave are calculated according to the Rayleigh-Sommerfeld diffraction integral formula. At the same time, the network model determines the bias term through the transmission coefficient of each layer of the metasurface. The model uses the forward propagation process to guide the incident light wave carrying the target information through the specially configured multi-layer metasurface structure, forming a diffraction pattern reflecting the characteristics of the target object on the detection plane. The Rayleigh-Sommerfeld diffraction theory describes the propagation mode of the light field from the nth layer to the n+1th layer, where represents the transmission coefficient of the nth layer, where A is the amplitude, is the phase, i is the imaginary unit, and a=1 is set when only the pure phase type diffraction neural network is trained; the light field distribution of the propagation form of the Rayleigh-Sommerfeld diffraction theory is expressed as: U(r n+1 )=t(r n )∫∫ S U(r n )h(r n+1 -r n )dxdy (1) Among them, U(r n ) is the transmitted electric field distribution of the upper layer, U(r n+1 ) is the transmission electric field distribution on the n+1th layer, n is the number of diffraction layers, and the formula h(r n+1 -r n ) represents the impulse response during the transmission process, and the expression is: in, is the point r on the nth layer n To point r on the n+1th layer n+1 The Euclidean distance between them, λ is the incident wavelength, is the imaginary unit, x n+1 、y n+1 、z n+1 are the spatial coordinate components of a point in the n+1 layer, are the spatial coordinate components of the ath neuron in the nth layer; S represents the integral area, which is all points in the entire wavefront area; the integral represents the weighted and superimposed light field distribution of the entire nth layer to calculate the light field distribution of a point on the n+1th layer; this method can be used to predict and calculate the propagation of light waves in media with different layer structures; the number of diffraction layers is represented by n, where a single diffraction layer n=1 is used, and the n+1th layer is used to represent the detection plane, and the energy distribution S of the sub-detection area of the detection plane is calculated. j =|U j | 2 , j is the sub-detector region number, and the recognition result is determined by the category corresponding to the region with the highest energy; Step 2) The MNIST dataset is selected as the training dataset, and a combination of images containing two detection targets is randomly generated as the training set. The purpose is to train the diffraction layer phase parameters of the diffraction neural network to maximize the energy distribution of the corresponding category sub-detection area and minimize the energy distribution of other sub-areas. During the optimization process, the mean square error loss function (MSE) is used to evaluate the difference between the energy distribution of different sub-detector areas and the target energy distribution. The number of identified targets is represented by C: where s c represents the energy distribution received in the final detection area, g c represents the target energy distribution, Represents the error between the target area and the detection area; Based on the obtained diffraction pattern and its deviation from the expected result, a loss function is defined to evaluate the recognition accuracy of the diffraction neural network under the current parameters; Step 3) Apply the back propagation module and combine it with the gradient descent algorithm to calculate the gradient of the loss function with respect to the phase parameter of the diffraction layer to guide the optimization process of the phase distribution; after multiple iterations until Stop training when convergence is achieved and the accuracy cannot be improved; repeat the above steps to adjust the number of diffraction layers to find the number of diffraction layers that can achieve the object recognition accuracy required by the designer; Step 4) iteratively executing steps 1) to 3), continuously updating the diffraction layer phase parameters until a predetermined recognition accuracy or other performance index is met, thereby completing the training and optimization of the diffraction neural network; Based on the phase distribution of the diffraction layer determined in step 2), rectangular silicon nanopillar arrays of the same size but different rotation angles are selected to form a multilayer metasurface; the unit structural parameters of these metasurfaces include the length L, width W, height H, and period length P of the nanopillars; the performance of nanopillar structures of different sizes under circularly polarized light is calculated to select the optimal parameter combination required at a specific wavelength and determine the size of the silicon nanopillars; according to the geometric phase principle, the desired phase distribution is achieved by rotating the nanopillar structure; finally, the nanopillar structure is arranged according to the phase distribution obtained by training the diffraction layer of the diffraction neural network, and a metasurface device capable of recognizing the spatial position and image of two targets is constructed based on this arrangement; The final phase distribution is in is the phase obtained by training the all-optical diffraction neural network; This method can calculate the phase distribution of the diffraction layer according to different recognition tasks, so that the metasurface device can adapt to the resolution and recognition of various objects.

7. The method for realizing dual-target spatial position and image recognition according to claim 6, characterized in that: The diffraction neural network training process uses a single diffraction layer, sets the distance from the input layer to the diffraction layer to 10,000 microns, the distance from the diffraction layer to the output layer to 10,000 microns, the learning rate to 0.005, the neuron size to 120 microns, and in order to save training time, eight samples are taken as a batch and input into the neural network.

8. The method for realizing dual-target spatial position and image recognition according to claim 6, characterized in that: The training model software uses the TensorFlow framework, which is based on the Python programming language.

9. The method for realizing dual-target spatial position and image recognition according to claim 6, characterized in that: Electromagnetic full-wave simulation software can be used to calculate the performance of nanopillar structures of different sizes under circularly polarized light.

Citation Information

Patent Citations

  • Pluggable diffraction neural network optimization method based on metasurface and task identification device

    CN116596050A