A reverse design method for superlens based on selective dual-channel perception fusion network
Through the design method of selective dual-channel perception fusion network, the problem of insufficient accuracy of image and spectrum generation in superlens inverse design is solved, efficient image and spectral response mapping and generation are achieved, the computational cost is reduced and the generation accuracy is improved.
Patent Information
- Application Number
- CN202510072808.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-01-16
AI Technical Summary
Existing superlens inverse design methods have shortcomings in the accuracy of image and spectrum generation, especially in multimodal tasks, where there are problems of information redundancy, information loss, and difficulty in alignment supervision.
A design method based on a selective dual-channel perception fusion network is adopted, including an adaptive difference frequency feature extraction network ADFFNet, a selective non-common feature interaction network SFIFNet, and a selective multi-dimensional common feature extraction network MMFMamba. Through adaptive difference frequency feature extraction, non-common feature interaction, and multi-dimensional common feature fusion, the non-common frequency features between images and spectra are deeply mined and fused.
It greatly reduces the computational and time costs of metalens design, improves the mapping accuracy and generation efficiency between images and spectral responses, alleviates the information redundancy and loss problems in multimodal feature fusion, and improves generation accuracy.
Smart Images

Figure CN119620389B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a superlens inverse design method based on a selective dual-channel perception fusion network, and belongs to the technical field of superlens. Background Art
[0002] A metalens is a microlens array composed of nanoscale structures that utilizes engineered micro-atoms to determine its optical response. With advances in nanofabrication technology, metalens have found widespread application in holography, beam splitting, optical imaging, sensing, and miniaturized lenses. Therefore, efficiently designing metalens that perform operations such as polarization modulation, filtering, and holography is crucial in the design of emerging optical devices. The design of metalens is challenging. Traditionally, the electromagnetic field model of Maxwell's equations is used to describe the underlying physics of light. However, this model involves a set of coupled partial differential equations, making explicit solutions difficult to obtain. Consequently, traditional metalens design employs full-wave numerical simulations using the finite element method (FEM) and finite-difference time-domain method (FDTD). These methods sweep parameters multiple times to find the optimal solution. This iterative solution strategy consumes significant computational resources and time, and fails to yield unique and optimal nano-atomic designs. Furthermore, due to the mismatch in the solution space, multiple solutions exist when mapping the geometry based on the spectral response. This makes inverse design, such as designing the geometry based on the spectral response, generally infeasible.
[0003] In recent years, with the development of neural networks, significant progress has been made in solving the inverse design problem. In 2020, Abhishek Mall et al. proposed a bidirectional encoder (biAE) in the paper Fast Design of Plasmonic Metasurfaces Enabled by DeepLearning. This model uses a Gan network to learn the forward and inverse mapping between the geometric structure and response spectrum of the metasurface, and simultaneously optimizes in multiple metasurface topological spaces, rather than the previous optimization in a single metasurface topological space, thus solving the one-to-many mapping problem in inverse design. However, this model may lead to overfitting in certain situations, causing model collapse and the generation of significant noise, resulting in poor model performance. In 2024, Yuansan Liu et al. proposed a bidirectional adversarial autoencoder (BiAAE) in the paper Bidirectional Adversarial Autoencoders for the design of Plasmonic Metasurfaces. This model combines the advantages of traditional variational autoencoders in representing information through latent space vectors while retaining the training method of generative adversarial networks (GANs), greatly improving the overfitting and noise suppression problems. However, this adversarial encoder still suffers from mode collapse and unstable training, as the internal game between the generator and discriminator can lead to vanishing or exploding gradients. Essentially, the image-based meta-atom design and spectral response mapping and generation tasks are actually multimodal fusion and multimodal generation tasks. In 2024, Wenbing Li et al. first introduced the selective state transfer network (Mamba) to multimodal tasks in their paper "Coupled Mamba: Enhanced Multi-modal Fusion with Coupled State Space Model." Experimental data show that this network performs well in forward and reverse generation tasks and multimodal fusion tasks. However, existing methods for handling multimodal tasks typically only align features between modalities into a unified representation space, but fail to fully utilize complementary cross-modal information exchange. Alternatively, they ignore the internal propagation of individual modalities and aggregate modal features into a single entity. This can lead to serious information redundancy and feature loss during modal interaction, as well as difficulties in alignment supervision.
[0004] All of the above neural networks have issues such as poor quality of meta-atom generation parameters and poor fitting results when the spectral curve has multiple peaks. In summary, existing metalens inverse design methods have great room for improvement and exploration. Summary of the Invention
[0005] The purpose of the present invention is to provide a superlens inverse design method based on a selective dual-channel perception fusion network to solve the problem in the prior art that the accuracy of image and spectrum generation needs to be improved.
[0006] The technical solution of the present invention is:
[0007] A method for reverse design of a superlens based on a selective dual-channel perception fusion network includes the following steps:
[0008] S1. Get K pairs of original images I p The corresponding original phase spectrum I s , get the original image-original phase spectrum pair {I p ,I s} K The constructed image-phase spectrum dataset is divided into training set and test set;
[0009] S2, the original image I p The corresponding original phase spectrum I s Input to the adaptive difference frequency feature extraction network ADFFNet to extract high-frequency features F with inter-modal differences s H , low-frequency feature F s L .
[0010] S3, the original image I p With the original phase spectrum I s , and the high-frequency feature F obtained in step S2 s H , low-frequency feature F s L , input the selective non-common feature interaction network SFIFNet and output image frequency information F pD and spectral frequency information F sD , corresponding to the original image I p and the original phase spectrum I s Add together to get the image frequency feature F pDo and spectral frequency characteristics F sDo ;
[0011] S4, the image frequency feature F pDo and spectral frequency characteristics F sDo Input the selective multi-dimensional common feature extraction network MMFMamba to obtain the image fusion reconstruction features F D p Fusion with the spectrum to reconstruct the feature F D s ;
[0012] S5, reconstruct the image fusion feature F D p Fusion with the spectrum to reconstruct the feature F D s After inputting the modal reconstruction network for image reconstruction and spectral reconstruction, the generated image results and generated spectral results are obtained.
[0013] Furthermore, in step S2, the adaptive differential frequency feature extraction network ADFFNet includes an encoder network, a masked self-attention mechanism module, a dual-channel adaptive high- and low-frequency information decomposition module, and a convolutional neural network CNN.
[0014] Encoder network: input original image I p and the original phase spectrum I S Perform dimensionality transformation to obtain parameters k, q, v;
[0015] Masked self-attention mechanism module: obtains the attention mask feature X from the input parameters k, q, v l ;
[0016] Dual-channel adaptive high- and low-frequency information decomposition module: input adaptive attention mask feature X l After the image self-attention feature extraction channel, discrete wavelet transform DWT is performed to obtain the global high-frequency features F of the image. s pH and image low-frequency features F s pL , while inputting the adaptive attention mask feature X l After the spectral self-attention feature extraction channel, discrete wavelet transform DWT is performed to obtain the spectral global high-frequency features F s sH and the spectral low-frequency feature F s sL ;
[0017] Convolutional neural network CNN: global high-frequency features F of the input image s pH and image low-frequency features F s pL , spectral global high-frequency characteristics F s sH and the spectral low-frequency feature F s sL After learning the frequency difference features respectively, the high-frequency features F with inter-modal differences are output s H , low-frequency feature F s L .
[0018] Furthermore, in step S3, the selective non-common feature interaction network SFIFNet outputs the image frequency information F pD , specifically,
[0019] y plinear =Linear1(LayerNorm(X p ));
[0020] y pm =σ1(CDWConv(y plinear ))
[0021] y pfreq =FS{y pm ,F s H ,F s L}
[0022]
[0023] Among them, X p The original input image I p The encoded image features are obtained by the convolutional neural network encoder CNN Encoder, LayerNorm is the normalization processing layer, Linear1 and Linear2 are linear layer 1 and linear layer 2 respectively, y plinear is the image feature output by the linear layer 1, y pm is the image preprocessing feature, CDWConv(·) is a cascaded depth convolution, which is used to capture depth information and shallow time features; σ1(·) and σ2(·) are both activation function operations, y pfreq is the inter-modal image frequency feature, F s H is a high-frequency feature with inter-modal differences, F s L For low-frequency features with inter-modal differences, FS is an optional inter-modal frequency space state transfer model FFISSM; It is Hadamard.
[0024] Furthermore, in the optional inter-modal frequency space state transfer model FFISSM, the inter-modal image frequency feature y is obtained pfreq , specifically,
[0025] The image preprocessing feature y pm The transformation parameter matrix A', B' and the discrete time matrix △t are obtained through the linear layer Linear; the transformation parameter matrix A' is initialized by an increasing number method, that is, from 1 to N 2Assign values incrementally, then take the log; according to the zero-order hold method formula, the transformation parameter matrix A', B' is discretized to obtain the time discretization transformation parameter matrix A t ,B t ;
[0026] Get the new state transfer matrix A n , with the new state transfer matrix A n As the core, and the time discretization transformation parameter matrix B t Combined, we get the hidden state matrix h t ', and then obtain the inter-modal image frequency feature y pfreq :
[0027] A n =A t +αf low +βf high
[0028] h t '=A n h t +B t y pm
[0029] y pfreq =C`h t '+IDWT(f low ,f high ).
[0030] Among them, α, β are learnable parameters, high frequency enhancement information f high With low frequency enhancement information f low They are the high-frequency features of the image F obtained in step S2 s pH and image low-frequency features F s pL After convolution, feature enhancement is performed, y pm is the image preprocessing feature, h t is the state matrix, and IDWT(·) represents the inverse discrete wavelet transform.
[0031] Furthermore, in step S3, the selective non-common feature interaction network SFIFNet outputs the spectral frequency information F sD , specifically,
[0032] y Slinear =Linear1(LayerNorm(X s ));
[0033] y Sm =σ1(CDWConv(y Slinear ))
[0034] y Sfreq =FS{y Sm ,F s H ,F s L}
[0035]
[0036] Among them, X s The input original phase spectrum I s The encoded spectral features are obtained by the convolutional neural network encoder CNN Encoder, LayerNorm is the normalization processing layer, Linear1 and Linear2 are linear layer 1 and linear layer 2 respectively, y Slinear is the spectral feature of the linear layer output, y Sm is the spectral preprocessing feature, CDWConv(·) is a cascaded deep convolution, which is used to capture the depth information and capture the shallow time features; ρ1(·) and ρ2(·) are both activation function operations, y Sfreq is the inter-modal spectral frequency characteristic, F s H is a high-frequency feature with inter-modal differences, F s L For low-frequency features with inter-modal differences, FS is an optional inter-modal frequency space state transfer model FFISSM; It is Hadamard.
[0037] Furthermore, in step S4, the selective multi-dimensional common feature extraction network MMFMamba includes a first normalization layer LayerNorm1, a multi-branch selective commonality fusion module and a second normalization layer LayerNorm2, and the image frequency feature F pDo and spectral frequency characteristics F sDo After passing through the corresponding normalized LayerNorm layer, the image normalized feature F is obtained LN p and spectral normalization feature F LN s And output to the feature fusion module, the multi-branch selective commonality fusion module performs selective commonality fusion on the features between the image and spectral modalities and obtains the image fusion feature F p I1 , image fusion feature F p I2 , spectral fusion feature F s I3 and spectral fusion feature F s I4 , image fusion feature F p I1and image fusion feature F p I2 After addition, it passes through the second normalization layer LayerNorm2 and then is combined with the image frequency feature F pDo Add up to get the fusion reconstruction feature F of the image D p ; Spectral fusion feature F s I3 and spectral fusion feature F s I4 After addition, it passes through the second normalization layer LayerNorm2 and then is combined with the spectral frequency feature F sDo Add up to get the fusion reconstruction feature F of the spectrum D s .
[0038] Furthermore, in the multi-branch selective commonality fusion module, the image fusion feature F is obtained. p I1 , image fusion feature F p I2 , spectral fusion feature F s I3 and spectral fusion feature F s I4 , specifically:
[0039] F p I1 =Mamba(CDWConv(F LN p ))⊙SiLU(F LN p )
[0040] F p I2 =F2Block(F LN p , F LN s )⊙SiLU(F LN p )
[0041] F s I3 =F2Block(F LN p , F LN s )⊙SiLU(F LN s )
[0042] F s I4 =Mamba(CDWConv(F LN s ))⊙SiLU(FLN s )
[0043] Among them, F LN p is the normalized image feature, F LN s is the spectral normalization feature, CDWConv(·) is the cascaded depthwise convolution, ⊙ is the element-wise multiplication, SiLU(·) is the activation operation, Mamba(·) is the selective state transfer operation, and F2Block(·) is the dual-channel feature selective extraction module.
[0044] Furthermore, the dual-channel feature selective extraction module F2Block includes a dual-channel convolutional neural encoding network, a convolution layer DWConv, an activation function layer SiLU, a selective deep alignment state transfer model AFSSM and a normalization processing layer LayerNorm.
[0045] The dual-channel feature selective extraction module F2Block includes a dual-channel convolutional neural encoding network, the first convolution layer DWConv1, the first activation function layer SiLU1, the selective deep alignment state transfer model AFSSM and the third normalization layer LayerNorm3.
[0046] Dual-channel convolutional neural encoding network: input image normalized feature F LN p After three layers of convolutional neural network i p , where i = 1, 2, 3, convolutional neural coding network Layer i p It includes the second convolutional layer DWConv2, the second activation function layer SiLU2 and the first fully connected layer Dense1, which performs dimensional transformation and image feature alignment, and outputs the new features as Input spectral normalization feature F LN s After four layers of convolutional neural network i s , where i = 1, 2, 3, 4. After dimensional transformation and image feature alignment, the new features are output as The new feature F s I and the new feature F p I After being sent to the second fully connected layer Dense2 for feature fusion, the fusion feature F is obtained. sp I ;
[0047] The first convolutional layer DWConv1: fusion feature F of the input spI After the convolution operation, the output is sent to the activation function layer SiLU;
[0048] The first activation function layer SiLU1: used to obtain the fused forward feature F sp I' And output to the selective deep alignment state transfer model AFSSM;
[0049] Selective Deep Alignment State Transfer Model AFSSM: Input fusion feature F sp I and fusion forward features F sp I' , through the selective state transfer equation, the intermediate state characteristic F is obtained out `And output to the third normalization layer LayerNorm3;
[0050] The third normalization layer LayerNorm3: normalizes and restores the dimension to output the intermediate feature F out .
[0051] Furthermore, in the selective deep alignment state transfer model AFSSM, the forward feature F is fused sp I' After initialization, the original state transfer matrix A, B, C, D is obtained, and then the input fusion feature F sp I The original state transfer matrix B and C are sent to the linear layer Linear3 and the linear layer Linear4 for hybrid alignment to obtain the fusion state transfer matrix M and N respectively. Finally, the intermediate state feature F is obtained through the selective state transfer equation. out `, the specific process formula is as follows:
[0052] M=Linear3(B,F sp I )
[0053] N=Linear4(C,F sp I )
[0054] h'=Ah t +MF sp I'
[0055] F out `=Nh'+DF sp I'
[0056] Among them, F sp I is the input fusion feature, h tis the state matrix, and h' is the aligned hidden state matrix.
[0057] The beneficial effects of the present invention are:
[0058] First, this metalens reverse design method based on a selective dual-channel perception fusion network replaces the traditional metalens design method with a new neural network design method, greatly reducing the calculation and time costs of metalens design. At the same time, the new neural network can deeply explore the non-common frequency characteristics between metalens meta-atoms and corresponding spectra and perform deep fusion, which can increase the accuracy of the mapping between metalens meta-atoms and corresponding spectral responses, and achieve the accuracy and efficiency of the generation of metalens meta-atoms and spectra in metalens reverse design.
[0059] 2. The present invention adopts the adaptive differential frequency feature extraction network ADFFNet, which can preliminarily extract the non-common frequency features between the meta-atom images and the corresponding spectra from the implicit space, deeply retain the non-common features between different modalities, and provide feature input for subsequent feature fusion.
[0060] 3. This superlens inverse design method based on the selective dual-channel perception fusion network deeply fuses non-common features through the selective non-common feature interaction network SFIFNet, which can greatly alleviate the information redundancy and loss problems in feature fusion between multiple modalities, thereby improving the accuracy of subsequent images and spectral responses.
[0061] 4. This superlens inverse design method based on the selective dual-channel perception fusion network not only focuses on the internal propagation between individual modes but also fully integrates the complementary feature information between modes through the selective multi-dimensional common feature extraction network MMFMamba, which can greatly increase the accuracy of the mapping between meta-atoms and spectral responses. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 1 is a flow chart of a method for reverse designing a superlens based on a selective dual-channel perception fusion network according to an embodiment of the present invention;
[0063] Figure 2 2 is a schematic diagram illustrating an adaptive differential frequency feature extraction network ADFFNet in an embodiment;
[0064] Figure 3 1 is a schematic diagram illustrating the selective non-common feature interaction network SFIFNet according to an embodiment;
[0065] Figure 4 1 is a schematic diagram illustrating an optional inter-modal frequency-space state transfer model FFISSM in an embodiment;
[0066] Figure 52 is a schematic diagram illustrating the selective multi-dimensional common feature extraction network MMFMamba in an embodiment;
[0067] Figure 6 2 is a schematic diagram illustrating a dual-channel feature selective extraction module F2Block in an embodiment;
[0068] Figure 7 1 is a schematic diagram illustrating the selective deep alignment state transfer model AFSSM in an embodiment;
[0069] Figure 8 The image result and the original image I are generated in the embodiment p Schematic diagram of the comparison;
[0070] Figure 9 The spectral results and the original phase spectrum I are generated in the embodiment s Schematic comparison diagram. DETAILED DESCRIPTION
[0071] To make the purpose, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them.
[0072] The embodiment provides a method for reverse design of a super lens based on a selective dual-channel perception fusion network. Figure 1 , including the following steps:
[0073] S1. Get K pairs of original images I P The original phase spectrum I corresponding to the original image S , get the original image-original phase spectrum pair {I p ,I s} K The image-phase spectrum dataset is divided into training set and test set.
[0074] In step S1, the original image The original image can be a metasurface cylindrical image, and the corresponding original phase spectrum Where H, W, and C represent the height, width, and dimension of the image, respectively, and L represents the length of the spectrum.
[0075] S2, the original image I p The corresponding original phase spectrum I s Input to the adaptive difference frequency feature extraction network ADFFNet to extract high-frequency features F with inter-modal differences s H , low-frequency feature F s L .
[0076] In step S2, the adaptive difference frequency feature extraction network ADFFNet is as follows Figure 2 :
[0077] 1. First, the original image and the original phase spectrum Perform dimension transformation in the input encoder network Input Embedding to obtain the corresponding parameters as well as Send the parameters k, q, v to the mask attention encoder MaskAttention to obtain the attention mask feature
[0078] 2. Attention mask feature X l The dual-channel adaptive high- and low-frequency information decomposition module includes the picture self-attention feature extraction channel Picture Sel-Atten and the spectrum self-attention feature extraction channel Spec Sel-Atten, and the image self-attention feature X is mined. p l and enhanced spectral self-attention feature X s l ;
[0079] 3. The image self-attention feature X p l and enhanced spectral self-attention feature X s l Perform discrete wavelet transform (DWT) processing to decompose the global high-frequency features F of the image s pH , image low-frequency features F s pL , spectral global high-frequency characteristics F s sH , spectral low-frequency characteristics F s sL ;
[0080] 4. The global high-frequency feature F of the image s pH , image low-frequency features F s pL , spectral global high-frequency characteristics F s sH , spectral low-frequency characteristics F s sL Frequency difference features are learned through a four-layer convolutional neural network (CNN). The convolutional neural network (CNN) includes a normalization layer (LayerNorm), a convolution layer (Conv), and a linear layer (Linear). The scale is then adjusted to ensure that the feature size remains unchanged. The final output is the high-frequency feature F with inter-modal differences.s H , low-frequency feature F s L .
[0081] In step S2, the mask self-attention mechanism module performs preliminary global information extraction on the image and spectrum, and extracts the adaptive attention mask feature X l , which can preliminarily distinguish the non-common features between images and spectra. The dual-channel adaptive high- and low-frequency information decomposition module amplifies the non-common features of images and spectra, so that the global information of the image and the global information of the spectrum are fully extracted and the information between the two modalities is not lost. After that, the global information of the image and spectrum are respectively subjected to discrete wavelet transformation to decompose the global high-frequency feature F of the image. s pH and image low-frequency information F s pL , spectral global high-frequency characteristics F s sH and spectral low-frequency information F s sL , deeply mining frequency features, such as image edges, color, and shape characteristics, as well as spectral inflection points, smoothness, and other deep frequency details. The non-common frequency domain feature fusion module aligns multi-modal frequency information to extract differential frequency information between different modalities, outputting differential high- and low-frequency information that is a fusion of image and spectrum.
[0082] S3, the original image I p With the original phase spectrum I s , and the high-frequency feature F with inter-modality differences extracted in step S2 s H , low-frequency feature F s L , input the selective non-common feature interaction network SFIFNet and output image frequency information F pD and spectral frequency information F sD , corresponding to the original image I p and the original phase spectrum I s Add together to get the image frequency feature F pDo and spectral frequency characteristics F sDo .
[0083] In step S3, the selective non-common feature interaction network SFIFNet includes a pre-processing module, an optional inter-modal frequency space state transfer model FFISSM, and a feature output module, such as Figure 3 :
[0084] Preprocessing module: input original image I pThe original phase spectrum I corresponding to the original image s Perform preliminary feature extraction and alignment encoding CNN Encoder operation to obtain the encoded image feature X p and the encoded spectral signature X s , then pass through the LayerNorm layer, the linear layer Linear1, the CDWConc convolution layer and the activation layer σ to obtain the image preprocessing feature y pm and spectral preprocessing features y sm ; Among them, the linear layer Linear1 simultaneously outputs the image feature y plinear With spectral characteristics y slinear To the feature output module;
[0085] Selectable inter-modal frequency space state transfer model FFISSM: The image preprocessing feature X is transformed into p With high frequency features F s H And the low-frequency feature F s L , spectral preprocessing feature X s With high frequency features F s H And the low-frequency feature F s L , perform feature interaction respectively to obtain the inter-modal image frequency feature y pfreq and the inter-modal spectral frequency characteristics y sfreq ;
[0086] Feature output module: input image feature y plinear With spectral characteristics y slinear After the activation function operation is performed respectively, the corresponding image frequency information F pD and spectral frequency information F sD After performing the Hadamard product and passing through the linear layer Linear2, the image frequency features F are output respectively. pD and spectral frequency characteristics F sD ,
[0087] In step S3, the selective non-common feature interaction network SFIFNet is as follows Figure 3 , specifically:
[0088] 31) First, the original image Enter the CNN Encoder network, perform preliminary special diagnostic information extraction and pre-coding operations, and output the encoded image features
[0089] 32) Encode image feature X pAfter the normalization layer LayerNorm and the linear layer Linear1, the image feature y output by the linear layer is obtained. plinear Afterwards plinear Perform cascaded deep convolution to extract potential time features; the extracted potential time features are activated by the function to obtain the image preprocessing features y pm ; At the same time, high-frequency features F s H And the low-frequency feature F s L Joint image preprocessing feature y pm Enter the optional inter-modal frequency space state transfer model FFISSM to obtain the inter-modal image frequency feature y pfreq At the same time plinear After cascade depth convolution and inter-modal image frequency feature y pfreq Multiply, pass through Linear2 layer to align the dimensions, and align with the original image I p Add together and finally get the image frequency feature F pDo Generate F pDo The specific process is as follows:
[0090] y plinear =Linear1(LayerNorm(X p ));
[0091] y pm =σ1(CDWConv(y plinear ))
[0092] y pfreq =FS{y pm ,F s H ,F s L}
[0093]
[0094] F pDo =F pD +I p
[0095] Among them, X p The original input image I p The encoded image features are obtained by the convolutional neural network encoder CNN Encoder, LayerNorm is the normalization processing layer, Linear1 and Linear2 are linear layer 1 and linear layer 2 respectively, y plinear is the image feature output by the linear layer 1, y pmis the image preprocessing feature, CDWConv(·) is a cascaded depth convolution, which is used to capture depth information and capture shallow time features; σ1(·) and σ2(·) are both activation function operations, F s H is a high-frequency feature with inter-modal differences, F s L For low-frequency features with inter-modal differences, FS is an optional inter-modal frequency space state transfer model FFISSM; is the Hadamard product. Image frequency feature F pDo Carrying spectral difference information.
[0096] 33) Similarly, the original phase spectrum Send it to the CNN Encoder network for preliminary special diagnosis information extraction and feature mapping encoding operation, and output the encoded spectral features
[0097] 34) Repeat the process of 32) and the final spectral frequency characteristic F sDo Generate F sDo The formula is as follows:
[0098] y Slinear =Linear1(LayerNorm(X s ));
[0099] y Sm =σ1(CDWConv(y Slinear ))
[0100] y Sfreq =FS{y Sm ,F s H ,F s L}
[0101]
[0102] F sDo =F sD +I s
[0103] Among them, X s The input original phase spectrum I s The encoded spectral features are obtained by the convolutional neural network encoder CNN Encoder, LayerNorm is the normalization processing layer, Linear1 and Linear2 are the linear layer 1 and linear layer 2 respectively, y Slinear is the spectral feature of the linear layer output, y Smis the spectral preprocessing feature, CDWConv(·) is a cascaded deep convolution, which is used to capture the depth information and capture the shallow time features; σ1(·) and σ2(·) are both activation function operations, y Sfreq is the inter-modal spectral frequency characteristic, F s H is a high-frequency feature with inter-modal differences, F s L For low-frequency features with inter-modal differences, FS is an optional inter-modal frequency space state transfer model FFISSM; is the Hadamard product, spectral frequency characteristic F sDo Carry image difference information.
[0104] In step 32) and step 34), the inter-modal frequency space state transfer model FFISSM can be selected to obtain the inter-modal image frequency feature y pfreq and the inter-modal spectral frequency characteristics y sfreq , the process is similar to obtain the inter-modal image frequency feature y pfreq Take the following as an example: Figure 4 :
[0105] 321) The high frequency feature F obtained in step S2 s H And the low-frequency feature F s L , after convolution, feature enhancement is performed to obtain high-frequency enhanced information f high With low frequency enhancement information f low ;
[0106] 322) Image preprocessing features The transformation parameter matrix A`∈R is obtained through the linear layer Linear N×1 , B`∈R N×1 And the discrete time matrix △t; for the transformation parameter matrix Use the incremental initialization method, that is, from 1 to N 2 Assign values incrementally, then take the log; according to the zero-order hold method formula, the transformation parameter matrix A', B' is discretized to obtain the time discretization transformation parameter matrix A t ,B t .
[0107] 323) By transforming the parameter matrix A based on time discretization t To further optimize, the high frequency enhancement information f high and low-frequency enhancement information f low Use learnable parameters α, β to add to the time discretization transformation parameter matrix A t Among them, the new state transfer matrix A is obtained n. Thus, the high-frequency features of the image with inter-modality differences F are realized s pH and image low-frequency features F s pL Able to dynamically adjust the state transfer matrix A t , guiding the state update, and then enabling the fusion of multi-modal differential high-low frequency features to enter the learning of single-modal feature extraction. The specific transformation formula is as follows:
[0108] A n =A t +αf low +βf high
[0109] 324) With the new state transfer matrix A n As the core, and the time discretization transformation parameter matrix B t Combined, we can get the hidden state matrix h t ′:
[0110] h t ′=A n h t +B t y m
[0111] where h t is the state matrix.
[0112] 325) Finally, the high frequency enhancement signal f high and low-frequency enhancement information f low After the inverse discrete wavelet transform IDWT and the transformation parameter matrix C' through the output layer equation, the output y is obtained pfreq , to achieve feature extraction carrying image / spectral differences;
[0113] y pfreq =C`h t '+ISWT(f low ,f high ).
[0114] Where IDWT(·) represents the inverse discrete wavelet transform.
[0115] S4, the image frequency feature F pDo and spectral frequency characteristics F sDo Input the selective multi-dimensional common feature extraction network MMFMamba to obtain the image fusion reconstruction features F D p Fusion with the spectrum to reconstruct the feature F D s .
[0116] In step S4, the selective multi-dimensional common feature extraction network MMFMamba includes the first normalization layer LayerNorm1, the multi-branch selective commonality fusion module and the second normalization layer LayerNorm2, the image frequency feature F carrying spectral difference information pDo and the spectral frequency feature F that carries image difference information sDo After passing through the corresponding normalized LayerNorm layer, the image normalized feature F is obtained LN p and spectral normalization feature F LN s And output to the multi-branch selective commonality fusion module, which performs selective commonality fusion on the features between the image and spectral modalities and obtains the image fusion feature F p I1 , image fusion feature F p I2 , spectral fusion feature F s I3 and spectral fusion feature F s I4 , image fusion feature F p I1 and image fusion feature F p I2 After addition, it passes through the second normalization layer LayerNorm2 and is combined with the image frequency feature F that carries spectral difference information. pDo Add up to get the fusion reconstruction feature F of the image D p ; Spectral fusion feature F s I3 and spectral fusion feature F s I4 After addition, it passes through the second normalization layer LayerNorm2 and is combined with the spectral frequency feature F that carries the image difference information. sDo Add up to get the fusion reconstruction feature F of the spectrum D s .
[0117] In the multi-branch selective commonality fusion module, the image fusion feature F is obtained p I1 , image fusion feature F p I2 , spectral fusion feature F s I3 and spectral fusion feature F s I4 , specifically:
[0118] F p I1 =Mamba(CDWConv(FLN p ))⊙SiLU(F LN p )
[0119] F p I2 =F2Block(F LN p , F LN s )⊙SiLU(F LN p )
[0120] F s I3 =F2Block(F LN p , F LN s )⊙SiLU(F LN s )
[0121] F s I4 =Mamba(CDWConv(F LN s ))⊙SiLU(F LN s )
[0122] Among them, F LN p is the normalized image feature, F LN s is the spectral normalization feature, ⊙ is the element-wise multiplication, CDWConv(·) is the 1x1 cascaded depthwise convolution, SiLU(·) is the activation operation, Mamba(·) is the selective state transfer operation, and F2Block(·) is a dual-channel feature selective extraction module used to align the features between image and spectral modalities and further integrate the common features of the two.
[0123] The dual-channel feature selective extraction module F2Block includes a dual-channel convolutional neural encoding network, the first convolution layer DWConv1, the first activation function layer SiLU1, the selective deep alignment state transfer model AFSSM and the third normalization layer LayerNorm3.
[0124] Dual-channel convolutional neural encoding network: input image normalized feature F LN p After three layers of convolutional neural network i p , where i = 1, 2, 3, after dimensional transformation and image feature alignment, the new features are output as Input spectral normalization feature F LN s After four layers of convolutional neural network i s , where i = 1, 2, 3, 4. After dimensional transformation and image feature alignment, the new features are output as The new feature F s I and the new feature F p I After being sent to the second fully connected layer Dense2 for feature fusion, the fusion feature F is obtained. sp I ;
[0125] The first convolutional layer DWConv1: fusion feature F of the input sp I After the convolution operation, the output is sent to the activation function layer SiLU;
[0126] The first activation function layer SiLU1: used to obtain the fused forward feature F sp I' And output to the selective deep alignment state transfer model AFSSM;
[0127] Selective Deep Alignment State Transfer Model AFSSM: Input fusion feature F sp I and fusion forward features F sp I' , through the selective state transfer equation, the intermediate state characteristic F is obtained out `And output to the third normalization layer LayerNorm3;
[0128] The third normalization layer LayerNorm3: normalizes and restores the dimension to output the intermediate feature F out .
[0129] Specifically, in the selective deep alignment state transfer model AFSSM, the forward feature F is fused sp I' After initialization, the original state transfer matrix A, B, C, D is obtained, and then the input fusion feature F sp I The original state transfer matrices B and C are respectively sent to the linear layer Linear3 and the linear layer Linear4 for hybrid alignment to obtain the fusion state transfer matrices M and N respectively. Finally, the intermediate state feature F is obtained through the selective state transfer equation. out , the specific process formula is as follows:
[0130] M=Linear3(B,F spI )
[0131] N=Linear4(C,F sp I )
[0132] h'=Ah t +MF sp I'
[0133] F out =Nh'+DF sp I'
[0134] Among them, F sp I is the input fusion feature, h t is the state matrix, and h' is the aligned hidden state matrix.
[0135] In step S4, the selective multi-dimensional common feature extraction network uses specific modal features to guide the generation of features of another module, aiming to combine local detail features of different modalities, fuse common features, and finally obtain the fusion reconstruction features F of the image. D p Fusion with the spectrum to reconstruct the feature F D s .
[0136] S5, reconstruct the image fusion feature F D p Fusion with the spectrum to reconstruct the feature F D s After inputting the modal reconstruction network for image reconstruction and spectral reconstruction, the generated image results and generated spectral results are obtained.
[0137] In step S5, the modal reconstruction network includes a selective state transfer output model Mamba and a CNN Decoder network to achieve dimensional reconstruction of the generated meta-atom image and the corresponding spectral response.
[0138] After training using the training set through steps S2 to S5, a trained model is obtained. After parameter optimization using the test set, a final model is obtained. After inputting the image to be generated or the phase spectrum to be generated into the final model, the corresponding generated spectrum or generated image is obtained.
[0139] Figure 8 The image result and the original image I generated in the embodiment are p A comparative diagram of . Figure 8 It can be seen that the present invention generates a relatively accurate fitting effect when there are multiple peaks in the spectral curve, which verifies the effectiveness of the present invention.
[0140] Figure 9 The spectral results and the original phase spectrum I are generated in the embodiment s A comparative diagram of . Figure 9 It can be seen that the quality of the meta-atom generation parameters is relatively consistent with the real meta-atom parameters, which verifies the effectiveness of the present invention.
[0141] This superlens inverse design method based on a selective dual-channel perception fusion network can deeply explore the non-common frequency characteristics between images and corresponding spectra and perform deep fusion, which can increase the accuracy of the mapping between images and corresponding spectral responses and improve the accuracy and efficiency of image and spectrum generation.
[0142] This metalens inverse design method, based on a dual-channel extraction network and a selective feature fusion mechanism, introduces an adaptive differential frequency feature extraction network. It utilizes a masked self-attention mechanism, a dual-channel adaptive high- and low-frequency information decomposition module, and a non-common frequency domain feature fusion module to preliminarily extract non-common frequency features between meta-atom images and corresponding spectra from an implicit space. The selective non-common feature interaction network introduces an optional inter-modal frequency space state transfer model to amplify and extract non-common features of a certain modality in the image / spectrum, while implicitly incorporating non-common features from another module and achieving channel alignment. A multi-branch structure and a dual-channel feature selective extraction module are introduced into the selective multi-dimensional common feature extraction network. This not only focuses on the internal propagation between individual modalities but also fully integrates complementary feature information between modalities, significantly improving the accuracy of the mapping between meta-atom and spectral responses. A selective state transfer network is then used in the modal reconstruction network to reconstruct and restore the meta-atom image and spectrum.
[0143] This metalens reverse design method, based on a selective dual-channel perception fusion network, addresses the difficulties inherent in traditional metalens reverse design, the significant computational resources and time required, and the inability to effectively capture the unique optimal meta-atom design structure. By combining four key networks—an adaptive differential frequency feature extraction network (ADFFNet), a selective non-common feature interaction network (SFIFNet), a selective multi-dimensional common feature extraction network (MMFMamba), and a modal reconstruction network—it improves the accuracy and efficiency of image and spectrum generation. The adaptive differential frequency feature extraction network (ADFFNet) extracts features from the cylindrical image and corresponding spectral information in the frequency space of the metalens surface, ensuring that deep, hidden features between modules are not lost. The selective non-common feature interaction network (SFIFNet) amplifies and extracts non-common features from one modality while implicitly incorporating non-common features from another, significantly alleviating information redundancy and loss in multimodal feature fusion. The selective multi-dimensional common feature extraction network (MMFMamba) deeply fuses common features between modalities and fully ensures feature alignment between them, addressing issues of feature misalignment and redundancy. The modality reconstruction network is used to generate reconstructions of images and spectra.
[0144] The above description is merely a preferred embodiment of the present invention and does not limit the present invention in any way. Any person skilled in the art who, without departing from the scope of the present invention, makes any equivalent substitution, modification, or other variation to the technical solution and technical contents disclosed in the present invention shall be deemed to be within the scope of the present invention and still fall within the scope of protection of the present invention.
Claims
1. A method for reverse design of a metalens based on a selective dual-channel perception fusion network, characterized by: The following steps are included: S1. Get K pairs of original images I p The corresponding original phase spectrum I s , get the original image-original phase spectrum pair {I p ,I s } K The constructed image-phase spectrum dataset is divided into training set and test set; S2, the original image I p The corresponding original phase spectrum I s Input to the adaptive difference frequency feature extraction network ADFFNet to extract high-frequency features F with inter-modal differences s H , low-frequency feature F s L ; S3, the original image I p With the original phase spectrum I s , and the high-frequency feature F obtained in step S2 s H , low-frequency feature F s L , input the selective non-common feature interaction network SFIFNet and output image frequency information F pD and spectral frequency information F sD , corresponding to the original image I p and the original phase spectrum I s Add together to get the image frequency feature F pDo and spectral frequency characteristics F sDo ; S4, the image frequency feature F pDo and spectral frequency characteristics F sDo Input the selective multi-dimensional common feature extraction network MMFMamba to obtain the image fusion reconstruction features F D p Fusion with the spectrum to reconstruct the feature F D s ; S5, reconstruct the image fusion feature F D p Fusion with the spectrum to reconstruct the feature F D s After inputting the modal reconstruction network for image reconstruction and spectral reconstruction, the generated image results and generated spectral results are obtained.
2. The method for reverse designing a metalens based on a selective dual-channel sensing fusion network according to claim 1, wherein: In step S2, the adaptive differential frequency feature extraction network ADFFNet includes an encoder network, a masked self-attention mechanism module, a dual-channel adaptive high- and low-frequency information decomposition module, and a convolutional neural network CNN. Encoder network: input original image I p and the original phase spectrum I S Perform dimensionality transformation to obtain parameters k, q, v; Masked self-attention mechanism module: obtains the attention mask feature X from the input parameters k, q, v l ; Dual-channel adaptive high- and low-frequency information decomposition module: input adaptive attention mask feature X l After the image self-attention feature extraction channel, discrete wavelet transform DWT is performed to obtain the global high-frequency features F of the image. s pH and image low-frequency features F s pL , while inputting the adaptive attention mask feature X l After the spectral self-attention feature extraction channel, discrete wavelet transform DWT is performed to obtain the spectral global high-frequency features F s sH and the spectral low-frequency feature F s sL ; Convolutional neural network CNN: global high-frequency features F of the input image s pH and image low-frequency features F s pL , spectral global high-frequency characteristics F s sH and the spectral low-frequency feature F s sL After learning the frequency difference features respectively, the high-frequency features F with inter-modal differences are output s H , low-frequency feature F s L .
3. The method for reverse designing a metalens based on a selective dual-channel sensing fusion network according to claim 1, wherein: In step S3, the selective non-common feature interaction network SFIFNet outputs the image frequency information F pD , specifically, y plinear =Linear1(LayerNorm(X p )); y pm =σ1(CDWConv(y plinear )) y pfreq =FS{y pm ,F s H ,F s L } Among them, X p The original input image I p The encoded image features are obtained by the convolutional neural network encoder CNN Encoder, LayerNorm is the normalization processing layer, Linear1 and Linear2 are linear layer 1 and linear layer 2 respectively, y plinear is the image feature output by the linear layer 1, y pm is the image preprocessing feature, CDWConv(·) is a cascaded depth convolution, which is used to capture depth information and shallow time features; σ1(·) and σ2(·) are both activation function operations, y pfreq is the inter-modal image frequency feature, F s H is a high-frequency feature with inter-modal differences, F s L For low-frequency features with inter-modal differences, FS is an optional inter-modal frequency space state transfer model FFISSM; It is Hadamard.
4. The method for reverse designing a metalens based on a selective dual-channel sensing fusion network according to claim 3, wherein: In the optional inter-modal frequency space state transfer model FFISSM, the inter-modal image frequency feature y is obtained pfreq , specifically, The image preprocessing feature y pm The transformation parameter matrix A', B' and the discrete time matrix △t are obtained through the linear layer Linear; the transformation parameter matrix A' is initialized by an increasing number method, that is, from 1 to N 2 Assign values incrementally, then take the log; according to the zero-order hold method formula, the transformation parameter matrix A', B' is discretized to obtain the time discretization transformation parameter matrix A t ,B t ; Get the new state transfer matrix A n , with the new state transfer matrix A n As the core, and the time discretization transformation parameter matrix B t Combined, we get the hidden state matrix h t ', and then obtain the inter-modal image frequency feature y pfreq : A n =A t +αf low +βf high h t ′=A n h t +B t y pm y pfreq =C`h t '+IDWT(f low ,f high ) Among them, α, β are learnable parameters, high frequency enhancement information f high With low frequency enhancement information f low They are the high-frequency features of the image F obtained in step S2 s pH and image low-frequency features F s pL After convolution, feature enhancement is performed, y pm is the image preprocessing feature, h t is the state matrix, and IDWT(·) represents the inverse discrete wavelet transform.
5. The method for reverse design of a metalens based on a selective dual-channel perception fusion network according to claim 3, wherein: In step S3, the selective non-common feature interaction network SFIFNet outputs the spectral frequency information F sD , specifically, y Slinear =Linear1(LayerNorm(X s )); y Sm =σ1(CDWConv(y Slinear )) y Sfreq =FS{y Sm ,F s H ,F s L } Among them, X s The input original phase spectrum I s The encoded spectral features are obtained by the convolutional neural network encoder CNNEncoder. LayerNorm is the normalization processing layer. Linear1 and Linear2 are linear layers 1 and 2 respectively. Slinear is the spectral feature of the linear layer output, y Sm is the spectral preprocessing feature, CDWConv(·) is a cascaded deep convolution, which is used to capture the depth information and capture the shallow time features; σ1(·) and σ2(·) are both activation function operations, y Sfreq is the inter-modal spectral frequency characteristic, F s H is a high-frequency feature with inter-modal differences, F s L For low-frequency features with inter-modal differences, FS is an optional inter-modal frequency space state transfer model FFISSM; It is Hadamard.
6. The method for reverse design of a metalens based on a selective dual-channel perception fusion network according to any one of claims 1 to 5, characterized in that: In step S4, the selective multi-dimensional common feature extraction network MMFMamba includes a first normalization layer LayerNorm1, a multi-branch selective commonality fusion module and a second normalization layer LayerNorm2, and the image frequency feature F pDo and spectral frequency characteristics F sDo After passing through the corresponding normalized LayerNorm layer, the image normalized feature F is obtained LN p and spectral normalization feature F LN s And output to the feature fusion module, the multi-branch selective commonality fusion module performs selective commonality fusion on the features between the image and spectral modalities and obtains the image fusion feature F p I1 , image fusion feature F p I2 , spectral fusion feature F s I3 and spectral fusion feature F s I4 , image fusion feature F p I1 and image fusion feature F p I2 After addition, it passes through the second normalization layer LayerNorm2 and then is combined with the image frequency feature F pDo Add up to get the fusion reconstruction feature F of the image D p ; Spectral fusion feature F s I3 and spectral fusion feature F s I4 After addition, it passes through the second normalization layer LayerNorm2 and then is combined with the spectral frequency feature F sDo Add up to get the fusion reconstruction feature F of the spectrum D s .
7. The method for reverse designing a metalens based on a selective dual-channel sensing fusion network according to claim 6, wherein: In the multi-branch selective commonality fusion module, the image fusion feature F is obtained p I1 , image fusion feature F p I2 , spectral fusion feature F s I3 and spectral fusion feature F s I4 , specifically: F p I1 =Mamba(CDWConv(F LN p ))⊙SiLU(F LN p ) F p I2 =F2Block(F LN p ,F LN s )⊙SiLU(F LN p ) F s I3 =F2Block(F LN p ,F LN s )⊙SiLU(F LN s ) F s I4 =Mamba(CDWConv)F LN s ))⊙SiLU(F LN s ) Among them, F LN p is the image normalization feature, F LN s is the spectral normalization feature, CDWConv·) is the cascaded depth convolution, ⊙ is the element-wise multiplication, SiLU(·) is the activation operation, Mamba(·) is the selective state transfer operation, and F2Block(·) is the dual-channel feature selective extraction module.
8. The method for reverse designing a metalens based on a selective dual-channel sensing fusion network according to claim 6, wherein: The dual-channel feature selective extraction module F2Block includes a dual-channel convolutional neural encoding network, the first convolution layer DWConv1, the first activation function layer SiLU1, the selective deep alignment state transfer model AFSSM and the third normalization layer LayerNorm3. Dual-channel convolutional neural encoding network: input image normalized feature F LN p After three layers of convolutional neural network i p , where i = 1, 2, 3, convolutional neural coding network Layer i p It includes the second convolutional layer DWConv2, the second activation function layer SiLU2 and the first fully connected layer Dense1, which performs dimensional transformation and image feature alignment, and outputs the new features as Input spectral normalization feature F LN s After four layers of convolutional neural network i s , where i = 1, 2, 3, 4. After dimensional transformation and image feature alignment, the new features are output as The new feature F s I and the new feature F p I After being sent to the second fully connected layer Dense2 for feature fusion, the fusion feature F is obtained. sp I ; The first convolutional layer DWConv1: fusion feature F of the input sp I After the convolution operation, the output is sent to the activation function layer SiLU; The first activation function layer SiLU1: used to obtain the fused forward feature F sp I' And output to the selective deep alignment state transfer model AFSSM; Selective Deep Alignment State Transfer Model AFSSM: Input fusion feature F sp I and fusion forward features F sp I' , through the selective state transfer equation, the intermediate state characteristic F is obtained out `And output to the third normalization layer LayerNorm3; The third normalization layer LayerNorm3: normalizes and restores the dimension to output the intermediate feature F out .
9. The method for reverse designing a metalens based on a selective dual-channel sensing fusion network according to claim 8, wherein: In the selective deep alignment state transfer model AFSSM, the forward feature F is fused sp I' After initialization, the original state transfer matrix A, B, C, D is obtained, and then the input fusion feature F sp I The original state transfer matrix B and C are sent to the linear layer Linear3 and the linear layer Linear4 for hybrid alignment to obtain the fusion state transfer matrix M and N respectively. Finally, the intermediate state feature F is obtained through the selective state transfer equation. out `, the specific process formula is as follows: M=Linear3(B,F sp I ) N=Linear4(C,F sp I ) h'=Ah t +MF sp I' F out `=Nh'+DF sp I' Among them, F sp I is the input fusion feature, h t is the state matrix, and h' is the aligned hidden state matrix.
Citation Information
Patent Citations
Medium metasurface reverse design algorithm utilizing cascaded deep neural network
CN112214719A
Neural network light field image deblurring method based on multi-head cross attention mechanism
CN116152103A