Superlens meta-atom structure reverse design method based on multi-view structural image and electronic equipment
By combining a multi-view coding module, a spectral coupling module, a cross-modal fusion network, and a projection-enhanced decoding network, the problems of insufficient single-view modeling and difficulty in modal alignment in superlens design are solved, achieving high-precision and stable reconstruction of three-dimensional meta-atomic structures and supporting rapid iteration of multifunctional, multi-band superlenses.
Patent Information
- Application Number
- CN202610424921.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-02
- Publication Date
- 2026-06-26
- Estimated Expiration
- 2046-04-02
AI Technical Summary
Existing reverse engineering methods for superlenses suffer from problems such as insufficient geometric information from a single viewpoint, difficulty in aligning modal features, lack of spatial awareness in the decoder, and poor consistency in cross-viewpoint reconstruction. These issues result in insufficient design accuracy and interpretability, making it difficult to meet the rapid iteration requirements of multifunctional, multi-band superlenses.
A superlens inverse design method for multi-view structural images is constructed. Through a multi-view coding module, a spectral coupling module, a cross-modal fusion network, and a projection enhancement decoding network, an end-to-end inverse mapping from the target spectral response to the three-dimensional atomic structure is achieved. The correlation between multi-view geometric information and spectral response is explicitly modeled, and a reprojection consistency supervision mechanism is introduced.
It significantly improves the accuracy, stability, and physical interpretability of superlens design, can efficiently fuse multi-view geometric information and spectral response data, is suitable for high-precision recovery of complex asymmetric structures, and supports the rapid development of multifunctional, multi-band superlenses.
Smart Images

Figure CN121960228B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of meta-atom structure design of superlenses, and more particularly to a reverse design method and electronic device for meta-atom structures of superlenses based on multi-view structural images. Background Technology
[0002] A superlens is a planar optical device composed of subwavelength-scale artificial superatomic periodic arrays, capable of precisely controlling parameters such as the phase, amplitude, and polarization of light within an ultrathin structure. Through the design of unit geometry and material parameters, it can achieve functions difficult to accomplish with traditional optical elements, such as achromatic focusing, beam shaping, light field modulation, and polarization control, showing broad application prospects in infrared imaging, laser processing, optical communication, sensing and detection, and integrated optics. Traditional superlens design mainly relies on parameter scanning optimization methods based on electromagnetic simulation, such as the finite element method (FEM) or the finite difference time-domain method (FDTD). These methods typically require ergonomic simulations of a large number of structural parameters to approximate the target spectral response, resulting in high computational costs, low optimization efficiency, and susceptibility to local optima. Furthermore, designs for different operating wavelengths or material systems often require rebuilding the simulation process, leading to poor model transferability, long development cycles, and difficulty in meeting the rapid iterative design requirements of multifunctional, multi-band superlenses.
[0003] In recent years, deep learning technology has been widely introduced into the field of optical inverse design, providing new ideas for the efficient design of superlenses. By constructing a nonlinear mapping relationship between the atomic geometry and the optical response, neural networks can quickly predict or generate structures that meet target performance without the need for complex electromagnetic simulations. For example, convolutional neural networks (CNNs) can be used to learn the direct mapping between structure images and spectral responses; autoencoders and generative adversarial networks (GANs) can achieve forward and inverse design in the latent space; some studies have further introduced the Transformer architecture to enhance cross-modal feature fusion capabilities. However, existing deep learning-based superlens inverse design methods still face several key challenges:
[0004] First, single-view structural modeling is insufficient. Most methods only use a single-view image of the atomic structure as input, failing to fully characterize the changes in its true 3D geometry under different observation angles. This results in a lack of awareness of spatial geometric consistency, affecting reconstruction accuracy and generalization ability. Second, semantic alignment between image and spectral modes is difficult. Since structural images and spectral responses are heterogeneous data modes, traditional fusion strategies (such as feature stitching or weighted averaging) struggle to establish effective semantic relationships, easily leading to feature conflicts or information loss between modes. Furthermore, the decoder lacks spatial awareness. Existing reconstruction modules often rely on standard convolution or deconvolution operations, failing to explicitly model 3D spatial coordinates and geometric distribution patterns. This results in inconsistencies between the generated structure and physical optical behavior, reducing the model's interpretability and practicality. Finally, there is a lack of cross-view consistency constraints. Ideally, the same atomic structure should correspond to a unique and stable 3D structure and spectral response under different viewpoints. However, existing methods typically do not supervise this, leading to deviations in reconstruction results under multi-view inputs and affecting design reliability.
[0005] While existing research has attempted to introduce graph neural networks or multi-view representation learning to enhance spatial correlation and has made progress in areas such as 3D perception, their application in nanoscale metalens design remains in the exploratory stage. Furthermore, although Transformers have advantages in capturing long-range dependencies and cross-modal fusion, in multi-view physical modeling scenarios, they lack explicit spatial projection mechanisms and position-aware capabilities, making it difficult to accurately recover the reversible mapping relationship between 3D geometry and spectral response.
[0006] Therefore, there is an urgent need for a novel superlens reverse design method that can integrate multi-view geometric information, achieve modal semantic alignment, possess spatial position awareness capabilities, and support cross-view consistency constraints, in order to break through the current technical bottlenecks in intelligent optical design. Summary of the Invention
[0007] To overcome key technical bottlenecks in existing superlens inverse design methods, such as insufficient single-view geometric information, difficulties in modal feature alignment, lack of spatial awareness in the decoder, and poor consistency in cross-view reconstruction, this invention proposes a superlens atomic structure inverse design method based on multi-view structural images, including:
[0008] A meta-atom sample library is constructed, which includes multiple meta-atom samples. Each meta-atom sample includes structural images of a meta-atom from multiple perspectives and spectral response data of the meta-atom under preset incident conditions. The meta-atom samples are preprocessed to obtain target meta-atom samples.
[0009] A superlens inverse design network is constructed, which includes a multi-view coding module, a spectral coupling module, a cross-modal fusion network, and a projection enhancement decoding network.
[0010] The superlens inverse design network is trained using individual target atomic samples to obtain the superlens inverse design model, including:
[0011] The structural image corresponding to each viewpoint in each target atomic sample is input into the multi-view encoding module. The module uses graph convolution operation to generate the multi-view structural features corresponding to the target atomic sample. The spectral response data in the target atomic sample is input into the spectral coupling module. The module jointly characterizes the variation law of the spectral response data in the wavelength dimension and its frequency domain characteristics to generate the spectral modal features corresponding to the target atomic sample.
[0012] Multi-view structural features and spectral modal features are embedded, and multi-level cross-modal fusion is performed through a cross-modal fusion network based on the obtained embedded representation to generate cross-modal fusion features; the cross-modal fusion features are input into the projection enhancement decoding network to output the three-dimensional geometric structure reconstruction results of the corresponding meta-atoms and their spectral response prediction data;
[0013] The target spectral response data is acquired and input into the superlens reverse design model to obtain the three-dimensional geometric structure reconstruction result of the corresponding elementary atom.
[0014] Further, the meta-atom sample is preprocessed to obtain the target meta-atom sample, specifically as follows:
[0015] All structural images are uniformly scaled to a preset size to obtain scaled structural images;
[0016] The spectral response data is interpolated to have a fixed number of sampling points within a preset wavelength range, resulting in interpolated spectral response data; the spectral response data is a phase-wavelength curve, or a combination of an amplitude-wavelength curve and a phase-wavelength curve.
[0017] Normalization processing is performed on the scaled structural image and the interpolated spectral response data to obtain normalized structural image and spectral response data;
[0018] A data augmentation operation is applied to the normalized structural image, the data augmentation operation including at least one of rotation, scaling and Gaussian noise perturbation within a preset angle range.
[0019] Further, the step of inputting the structural images corresponding to each viewpoint in each target atomic sample into the multi-view encoding module, and generating the multi-view structural features corresponding to the target atomic sample through graph convolution operation by the module, specifically involves:
[0020] The structural image from each viewpoint is treated as a node, and an adjacency relationship is constructed based on the angle difference between any two viewpoints: if the angle difference is less than a preset threshold, a connection is established between the corresponding two nodes, thereby forming an adjacency relationship matrix;
[0021] The number of connections for each node is calculated based on the adjacency matrix, and the adjacency matrix is normalized accordingly to obtain a normalized adjacency matrix.
[0022] Feature extraction is performed on the structural images from each viewpoint to obtain the corresponding node features. The node features corresponding to each node are input into a multi-layer graph convolutional structure along with the normalized adjacency matrix. The node features of adjacent viewpoints are aggregated layer by layer through the multi-layer graph convolutional structure, so that the node features corresponding to viewpoints with an angle difference less than the preset threshold are aligned with each other, and finally the multi-view structural features corresponding to the target atomic sample are output.
[0023] Further, the spectral response data of the target atomic sample is input into the spectral coupling module. This module then jointly characterizes the variation of the spectral response data in the wavelength dimension and its frequency domain characteristics to generate the spectral modal features corresponding to the target atomic sample. Specifically:
[0024] The spectral response data is constructed into a spectral sequence in wavelength order: when the spectral response data only contains phase-wavelength curves, it is constructed into a one-dimensional spectral sequence; when the spectral response data contains both amplitude-wavelength curves and phase-wavelength curves, it is constructed into a dual-channel spectral sequence, where one channel corresponds to the amplitude-wavelength curve and the other channel corresponds to the phase-wavelength curve.
[0025] When the spectral sequence is one-dimensional, a comprehensive feature vector for each wavelength position is generated by analyzing the numerical changes and frequency domain characteristics of the phase-wavelength curve in the neighborhood of each wavelength position.
[0026] When the spectral sequence is dual-channel, the numerical changes and frequency domain characteristics of the phase-wavelength curve and amplitude-wavelength curve in the neighborhood of each wavelength position are analyzed separately, generating the phase comprehensive feature vector and amplitude comprehensive feature vector corresponding to each wavelength position, and then splicing the two into the comprehensive feature vector of that wavelength position.
[0027] The comprehensive feature vectors of all wavelength positions are combined in wavelength order to form the spectral modal features corresponding to the target atomic sample.
[0028] Furthermore, when the spectral sequence is one-dimensional, by analyzing the numerical changes and frequency domain characteristics of the phase-wavelength curve in the neighborhood of each wavelength position, a comprehensive feature vector is generated for each wavelength position, specifically as follows:
[0029] When the spectral sequence is one-dimensional, based on the numerical changes of the phase-wavelength curve in the neighborhood of each wavelength position, the local phase trend features corresponding to each wavelength position are calculated, and frequency domain analysis is performed on the neighborhood of each wavelength position to obtain the local phase frequency domain distribution features corresponding to each wavelength position. At each wavelength position, the local phase trend features and the local phase frequency domain distribution features of that position are fused to generate a comprehensive feature vector for that wavelength position.
[0030] Furthermore, when the spectral sequence is dual-channel, the numerical changes and frequency domain characteristics of the phase-wavelength curve and amplitude-wavelength curve in the neighborhood of each wavelength position are analyzed to generate a phase comprehensive feature vector and an amplitude comprehensive feature vector corresponding to each wavelength position, and the two are concatenated into a comprehensive feature vector for that wavelength position, specifically:
[0031] When the spectral sequence is dual-channel, the local phase trend features and local amplitude trend features corresponding to each wavelength position are calculated based on the numerical changes of the phase-wavelength curve and amplitude-wavelength curve in the neighborhood of each wavelength position. Frequency domain analysis is then performed on the neighborhood of each wavelength position to obtain the local phase frequency domain distribution features and local amplitude frequency domain distribution features corresponding to each wavelength position. At each wavelength position, the local phase trend features and local phase frequency domain distribution features, as well as the local amplitude trend features and local amplitude frequency domain distribution features, are fused to obtain the phase comprehensive feature vector and the amplitude comprehensive feature vector. These two are then concatenated to form the comprehensive feature vector for that wavelength position.
[0032] Furthermore, the cross-modal fusion network includes a multi-head attention layer, a modal fusion layer, and a global modeling layer connected in sequence;
[0033] Multi-view structural features and spectral modal features are embedded, and based on the obtained embedding representation, multi-level cross-modal fusion is performed through a cross-modal fusion network to generate cross-modal fused features; specifically:
[0034] The multi-view structural features and the spectral modal features are mapped to a feature space of the same dimension through linear transformation to obtain corresponding embedding representations; wherein, the embedding representation of the spectral modal features contains multiple feature vectors, and the feature vectors correspond one-to-one with the wavelength positions in the spectral sequence;
[0035] The embedding representations corresponding to the multi-view structural features and the spectral modal features are respectively input into the cross-modal fusion network:
[0036] In the multi-head attention layer, the embedded representation of multi-view structural features is used as the query, and the embedded representation of spectral modal features is used as the key and value, and cross-modal features are generated through the attention mechanism.
[0037] In the modality fusion layer, the embedded representations of cross-modal features, multi-view structural features, and spectral modality features are fused to generate a joint feature representation;
[0038] In the global modeling layer, multiple cascaded cross-attention modules are used to perform multiple rounds of cross-modal interaction between the embedded representation of multi-view structural features and the embedded representation of spectral modal features to generate enhanced features, which are then fused with the joint feature representation to output cross-modal fused features.
[0039] Furthermore, the projection enhancement decoding network includes a spatial projection module, an upsampling module, a structure reconstruction module, and a spectral reconstruction module;
[0040] The method involves inputting the cross-modal fusion features into the projection enhancement decoding network, and outputting the three-dimensional geometric structure reconstruction results and spectral response prediction data of the corresponding meta-atoms; specifically:
[0041] The spatial projection module uses a learnable spatial projection matrix to map the cross-modal fusion features to a three-dimensional geometric space, and superimposes position codes corresponding to each spatial location to obtain projection features; the position codes are used to represent spatial coordinate information.
[0042] The projection features are upsampled at multiple levels by the upsampling module. After each level of upsampling, a position attention mechanism is applied to the features obtained by the upsampling at that level to model the dependencies between spatial positions. After multi-level processing, a position-aware spatial feature representation is output.
[0043] The structure reconstruction module reconstructs the three-dimensional geometric structure of the corresponding meta-atom based on the position-aware spatial feature representation and uses three-dimensional convolution operations to reconstruct the three-dimensional geometric structure and outputs voxelized data or 3D model representing the three-dimensional geometric structure.
[0044] The spectral reconstruction module maps the spatial features of the location awareness into spectral response prediction data that corresponds one-to-one with each wavelength position in the spectral sequence. The spectral response prediction data includes the phase response at each wavelength position, or the phase response and amplitude response at each wavelength position.
[0045] Furthermore, the step of training the superlens inverse design network through various target atomic samples also includes:
[0046] The reconstructed voxelized data or 3D model is reprojected onto the imaging plane of each input viewpoint to generate the reconstructed image of the corresponding viewpoint.
[0047] The difference between each reconstructed image and its corresponding viewpoint structural image is calculated as the viewpoint consistency error.
[0048] The consistency error between viewpoints and the 3D geometric reconstruction error are weighted and combined to form a total loss function. The model parameters of the superlens inverse design network are then optimized based on the total loss function.
[0049] To address the aforementioned problems, the present invention also provides an electronic device, comprising: a processor, and a memory storing a program, the program including instructions that, when executed by the processor, cause the processor to perform the methods described above.
[0050] Compared with the prior art, the present invention has at least the following beneficial effects:
[0051] (1) The present invention proposes a reverse design method for the atomic structure of a superlens based on multi-view structural images. By constructing a superlens reverse design model including a multi-view encoding module, a spectral coupling module, a cross-modal fusion network, and a projection enhancement decoding network, it realizes end-to-end reverse mapping from the target spectral response to a high-fidelity three-dimensional atomic structure. This method abandons the traditional parameter scanning process that relies on massive electromagnetic simulations. While significantly reducing the computational cost, it effectively integrates multi-view geometric information and spectral response data, systematically solving multiple technical bottlenecks such as insufficient single-view modeling, difficulty in modal alignment, and lack of spatial perception. It greatly improves the accuracy, stability, and physical interpretability of superlens design, and provides a reliable technical path for the rapid development of multifunctional, multi-band planar optical devices.
[0052] (2) To address the problem that existing reverse engineering methods generally rely on single-view image input and are difficult to fully represent the true three-dimensional geometric morphology of meta-atoms, this invention introduces a multi-view graph feature modeling mechanism. This mechanism treats the structural images from each viewpoint as graph nodes and constructs an adjacency matrix based on the angle differences between viewpoints. A multi-layer graph convolutional network is then used to aggregate the node features of adjacent viewpoints layer by layer. This mechanism can explicitly model the geometric correlation between different observation angles, enabling the network to achieve viewpoint alignment and spatial consistency at the feature level. This effectively captures the complete morphological information of meta-atoms in three-dimensional space, significantly improving the geometric integrity and robustness of the reconstruction results, and is particularly suitable for high-precision restoration of complex asymmetric structures.
[0053] (3) In view of the problem that structural images and spectral responses are heterogeneous modes and traditional fusion methods are difficult to establish an effective correlation, this invention achieves high-precision feature interaction through the collaborative operation of a spectral coupling module and a cross-modal fusion network. The spectral coupling module constructs a corresponding one-dimensional or dual-channel spectral sequence based on the input spectral response data. It then extracts local trend features reflected by numerical changes and local frequency domain distribution features obtained from frequency domain analysis within the neighborhood of each wavelength position. These are then fused to generate a comprehensive feature vector for each wavelength position, forming a physically interpretable spectral modal feature. Based on this, the cross-modal fusion network first maps multi-view structural features and spectral modal features to a unified dimension. Then, using the embedded representation of the structural features as the query and the embedded representation of the spectral features as the key and value, a multi-head attention layer calculates the correlation score between the query and the feature vectors corresponding to each wavelength position in the key. After normalization, the values are weighted and summed to generate the cross-modal feature. Finally, the modal fusion layer integrates this aligned feature with the original embedded representation, and a cascaded cross-attention module in the global modeling layer performs multiple rounds of deep interaction between the structural and spectral features, ultimately outputting an enhanced cross-modal fusion feature. This fusion mechanism abandons the coarse-grained fusion approach of directly splicing heterogeneous features or statically weighting them in traditional methods. By introducing a query-key-value interaction structure based on an attention mechanism, it enables multi-view geometric information to adaptively focus on the physically related wavelength positions in the spectral response, thereby achieving point-by-point alignment and deep fusion of structure and spectrum in the wavelength dimension. This significantly improves the accuracy, robustness, and physical interpretability of cross-modal mapping.
[0054] (4) To address the issues of traditional decoders lacking spatial awareness and inconsistent generated structures with physical behavior, this invention proposes a projection-enhanced decoding network. This network first uses a learnable spatial projection matrix to map cross-modal fusion features to a three-dimensional geometric space, and then superimposes positional codes corresponding to each spatial location. Subsequently, through multi-level upsampling combined with a positional attention mechanism, it explicitly models the long-range dependencies between spatial locations, generating position-aware spatial feature representations. Finally, based on these representations, it reconstructs the three-dimensional geometric structure of the corresponding meta-atoms through three-dimensional convolution operations. This mechanism endows the decoding process with explicit spatial coordinate awareness, ensuring that the output structure is highly self-consistent in geometric distribution and optical response, significantly improving the physical interpretability and engineering practicality of the model.
[0055] (5) To address the problem of lack of consistency constraints and easy generation of viewpoint-dependent bias in reconstruction results under multi-view input, this invention introduces a reprojection consistency supervision mechanism during training. Specifically, the reconstructed voxelized data or 3D model is reprojected onto the imaging plane of each input viewpoint to generate the corresponding reconstructed image, and the pixel-level or structural-level difference between it and the original structural image is calculated as the viewpoint consistency error; this error is weighted and combined with the 3D geometric structure reconstruction error to form the total loss function, which is used to jointly optimize the network parameters. This constraint minimizes the difference between the reconstructed 3D structure reprojected onto each input viewpoint and the corresponding original structural image, prompting the model to generate a geometrically self-consistent 3D structure under multi-view, thereby ensuring the intrinsic consistency of multi-view geometric information, effectively suppressing viewpoint-related bias in structural reconstruction, and significantly improving the stability and reliability of the model under multi-view joint input. Attached Figure Description
[0056] Figure 1 This is a flowchart of a reverse design method for meta-atomic structures of a superlens based on multi-view structural images;
[0057] Figure 2 This is a schematic diagram of the electronic device in this embodiment. Detailed Implementation
[0058] The following are specific embodiments of the present invention, which are described in conjunction with the accompanying drawings. However, the present invention is not limited to these embodiments.
[0059] To overcome key technical bottlenecks in existing superlens inverse design methods, such as insufficient single-view geometric information, difficulty in modal feature alignment, lack of spatial awareness in the decoder, and poor consistency in cross-view reconstruction, etc. Figure 1 As shown, this invention proposes a method for reverse design of meta-atomic structures of superlenses based on multi-view structural images, including:
[0060] A meta-atom sample library is constructed, which includes multiple meta-atom samples. Each meta-atom sample includes structural images of a meta-atom from multiple perspectives and spectral response data of the meta-atom under preset incident conditions. The meta-atom samples are preprocessed to obtain target meta-atom samples.
[0061] The meta-atom sample is preprocessed to obtain the target meta-atom sample, specifically as follows:
[0062] All structural images are uniformly scaled to a preset size to obtain scaled structural images;
[0063] The spectral response data is interpolated to have a fixed number of sampling points within a preset wavelength range, resulting in interpolated spectral response data; the spectral response data is a phase-wavelength curve, or a combination of an amplitude-wavelength curve and a phase-wavelength curve.
[0064] Normalization processing is performed on the scaled structural image and the interpolated spectral response data to obtain normalized structural image and spectral response data;
[0065] Normalization processing is performed on the scaled structural image and the interpolated spectral response data, specifically including:
[0066] The scaled structural image is subjected to a minimum-maximum normalization function to linearly map its pixel values to the [0,1] interval;
[0067] The interpolated spectral response data is subjected to zero mean and unit variance standardization. That is, the mean and standard deviation of the spectral response data are calculated at all sampling points, and the response value of each sampling point is subtracted from the mean and then divided by the standard deviation.
[0068] A data augmentation operation is applied to the normalized structural image, the data augmentation operation including at least one of rotation, scaling and Gaussian noise perturbation within a preset angle range.
[0069] A superlens inverse design network is constructed, which includes a multi-view coding module, a spectral coupling module, a cross-modal fusion network, and a projection enhancement decoding network.
[0070] The superlens inverse design network is trained using individual target atomic samples to obtain the superlens inverse design model, including:
[0071] The structural image corresponding to each viewpoint in each target atomic sample is input into the multi-view encoding module. The module uses graph convolution operation to generate the multi-view structural features corresponding to the target atomic sample. The spectral response data in the target atomic sample is input into the spectral coupling module. The module jointly characterizes the variation law of the spectral response data in the wavelength dimension and its frequency domain characteristics to generate the spectral modal features corresponding to the target atomic sample.
[0072] The step of inputting the structural image corresponding to each viewpoint in each target atomic sample into the multi-view encoding module, and generating the multi-view structural features corresponding to the target atomic sample through graph convolution operation by the module, specifically involves:
[0073] The structural image from each viewpoint is treated as a node, and an adjacency relationship is constructed based on the angle difference between any two viewpoints: if the angle difference is less than a preset threshold, a connection is established between the corresponding two nodes, thereby forming an adjacency relationship matrix;
[0074] The number of connections for each node is calculated based on the adjacency matrix, and the adjacency matrix is normalized accordingly to obtain a normalized adjacency matrix.
[0075] Feature extraction is performed on the structural images from each viewpoint to obtain the corresponding node features. The node features corresponding to each node are input into a multi-layer graph convolutional structure along with the normalized adjacency matrix. The node features of adjacent viewpoints are aggregated layer by layer through the multi-layer graph convolutional structure, so that the node features corresponding to viewpoints with an angle difference less than the preset threshold are aligned with each other, and finally the multi-view structural features corresponding to the target atomic sample are output.
[0076] In this embodiment, to construct multi-view structural features, the structural image from each viewpoint is considered as a node in the graph. Then, the observation angle difference between any two views is calculated, and a preset angle threshold is set. If the angle difference between two views is less than this threshold, a connection is established between the corresponding two nodes. Furthermore, an additional connection pointing to itself is added to each node to preserve its original viewpoint structural information. In the resulting adjacency matrix, if a connection (including self-connections) is established between two views, the original weight at the corresponding position in the adjacency matrix is set to 1; otherwise, it is set to 0. The original weight corresponding to a self-connection is 1.
[0077] To ensure a more stable and balanced subsequent graph propagation process, the adjacency matrix is normalized. Specifically, for each node, the total number of connections it establishes with other nodes (including itself) is counted, i.e., the node's connection count (also called "degree," and its value is at least 1 due to self-connections). Then, by dividing the original weight of each connection by the geometric mean of the number of connections between the two nodes associated with that connection, the information transfer weight corresponding to that connection is obtained, thus forming the normalized adjacency matrix. In this way, nodes with more connections will not have an excessive influence in feature aggregation simply because they have more neighbors, thus ensuring a fairer and more reasonable contribution from different perspectives to feature fusion.
[0078] In each layer of the graph convolutional structure, each node, based on the information transfer weights corresponding to its neighboring nodes (including itself) in the normalized adjacency matrix, performs a weighted summation of the node features corresponding to itself and its angularly neighboring viewpoints to obtain the aggregated neighbor feature representation. Subsequently, this aggregation result is subjected to a learnable linear transformation and a non-linear activation function to generate the updated features of the node in the current layer. Since adjacency relationships are established between viewpoints (including themselves) with an angle difference less than the preset threshold, through layer-by-layer aggregation in the multi-layer graph convolutional structure, the features of each node gradually incorporate the geometric information of angularly neighboring viewpoints, making the node features corresponding to these viewpoints tend to be consistent in the feature space, thereby achieving mutual alignment. Finally, all the output node features together constitute the multi-view structural features corresponding to the target meta-atomic sample.
[0079] To address the problem that existing reverse engineering methods generally rely on single-view image input and struggle to fully represent the true 3D geometry of atomic elements, this invention introduces a multi-view graph feature modeling mechanism. This mechanism treats structural images from different viewpoints as graph nodes and constructs an adjacency matrix based on the angle differences between viewpoints. A multi-layer graph convolutional network is then used to aggregate node features from adjacent viewpoints layer by layer. This mechanism explicitly models the geometric relationships between different observation angles, enabling the network to achieve viewpoint alignment and spatial consistency at the feature level. This effectively captures the complete morphological information of atomic elements in 3D space, significantly improving the geometric integrity and robustness of the reconstruction results, and is particularly suitable for high-precision recovery of complex asymmetric structures.
[0080] The spectral response data of the target atomic sample is input into the spectral coupling module. This module then jointly characterizes the variation of the spectral response data in the wavelength dimension and its frequency domain characteristics to generate the spectral modal features corresponding to the target atomic sample. Specifically:
[0081] The spectral response data is constructed into a spectral sequence in wavelength order: when the spectral response data only contains phase-wavelength curves, it is constructed into a one-dimensional spectral sequence; when the spectral response data contains both amplitude-wavelength curves and phase-wavelength curves, it is constructed into a dual-channel spectral sequence, where one channel corresponds to the amplitude-wavelength curve and the other channel corresponds to the phase-wavelength curve.
[0082] When the spectral sequence is one-dimensional, a comprehensive feature vector for each wavelength position is generated by analyzing the numerical changes and frequency domain characteristics of the phase-wavelength curve in the neighborhood of each wavelength position; specifically:
[0083] When the spectral sequence is one-dimensional, based on the numerical changes of the phase-wavelength curve in the neighborhood of each wavelength position, the local phase trend features corresponding to each wavelength position are calculated, and frequency domain analysis is performed on the neighborhood of each wavelength position to obtain the local phase frequency domain distribution features corresponding to each wavelength position. At each wavelength position, the local phase trend features and the local phase frequency domain distribution features at that position are fused to generate a comprehensive feature vector for that wavelength position.
[0084] When the spectral sequence is dual-channel, the numerical changes and frequency domain characteristics of the phase-wavelength curve and amplitude-wavelength curve in the neighborhood of each wavelength position are analyzed separately. This generates a phase composite feature vector and an amplitude composite feature vector corresponding to each wavelength position, and these two are then concatenated to form the composite feature vector for that wavelength position. Specifically:
[0085] When the spectral sequence is dual-channel, the local phase trend features and local amplitude trend features corresponding to each wavelength position are calculated based on the numerical changes of the phase-wavelength curve and amplitude-wavelength curve in the neighborhood of each wavelength position. Frequency domain analysis is then performed on the neighborhood of each wavelength position to obtain the local phase frequency domain distribution features and local amplitude frequency domain distribution features corresponding to each wavelength position. At each wavelength position, the local phase trend features and local phase frequency domain distribution features, as well as the local amplitude trend features and local amplitude frequency domain distribution features, are fused to obtain the phase comprehensive feature vector and the amplitude comprehensive feature vector. These two are then concatenated to form the comprehensive feature vector for that wavelength position.
[0086] The comprehensive feature vectors of all wavelength positions are combined in wavelength order to form the spectral modal features corresponding to the target atomic sample.
[0087] Multi-view structural features and spectral modal features are embedded, and multi-level cross-modal fusion is performed through a cross-modal fusion network based on the obtained embedded representation to generate cross-modal fusion features; the cross-modal fusion features are input into the projection enhancement decoding network to output the three-dimensional geometric structure reconstruction results of the corresponding meta-atoms and their spectral response prediction data;
[0088] The cross-modal fusion network comprises a multi-head attention layer, a modal fusion layer, and a global modeling layer connected in sequence;
[0089] Multi-view structural features and spectral modal features are embedded, and based on the obtained embedding representation, multi-level cross-modal fusion is performed through a cross-modal fusion network to generate cross-modal fused features; specifically:
[0090] The multi-view structural features and the spectral modal features are mapped to a feature space of the same dimension through linear transformation to obtain corresponding embedding representations; wherein, the embedding representation of the spectral modal features contains multiple feature vectors, and the feature vectors correspond one-to-one with the wavelength positions in the spectral sequence;
[0091] The embedding representations corresponding to the multi-view structural features and the spectral modal features are respectively input into the cross-modal fusion network:
[0092] In the multi-head attention layer, the embedded representation of multi-view structural features is used as the query, and the embedded representation of spectral modality features is used as the key and value. The correlation score between the query and the feature vector corresponding to each wavelength position in the key is calculated. After normalization, a set of attention weights corresponding one-to-one with each wavelength position in the key is obtained. Based on the attention weights, the feature vectors corresponding one-to-one with each wavelength position in the value are weighted and summed to generate cross-modal features.
[0093] Through the aforementioned multi-head attention mechanism, the model can dynamically establish fine-grained correspondences between multi-view structural features and spectral modal features. Specifically, the multi-view structural features (as queries) are correlated with the feature vectors corresponding to each wavelength position in the key to calculate relevance scores. The resulting attention weights reflect the importance of spectral information at different wavelength positions to the multi-view structural features. Subsequently, based on these weights, the feature vectors corresponding to each wavelength position in the values are weighted and summed, so that the final generated cross-modal features retain both structural context information and semantically relevant spectral response content. This process enables the model to adaptively associate geometric morphology and optical behavior at the feature level, thereby modeling the semantic relationship between structural information and spectral response.
[0094] In the modality fusion layer, the embedded representations of cross-modal features, multi-view structural features, and spectral modality features are fused to generate a joint feature representation;
[0095] In the global modeling layer, the embedded representation of multi-view structural features is used as the query, and the embedded representation of spectral modality features is used as the key and value. The relevance score between the query and the key is calculated layer by layer through multiple cascaded cross-attention modules, and the attention weight is obtained after normalization. The value is weighted and summed using the attention weight to obtain the enhanced feature. The enhanced feature is added to the joint feature representation to output the cross-modal fusion feature.
[0096] To address the challenge of establishing an effective correlation between structural images and spectral responses, which are heterogeneous modalities, traditional fusion methods struggle. This invention achieves high-precision feature interaction through the collaborative operation of a spectral coupling module and a cross-modal fusion network. The spectral coupling module constructs a corresponding one-dimensional or dual-channel spectral sequence based on the input spectral response data. It then extracts local trend features reflected by numerical changes and local frequency domain distribution features obtained from frequency domain analysis within the neighborhood of each wavelength position, fusing them to generate a comprehensive feature vector for each wavelength position, forming physically interpretable spectral modal features. Building upon this, the cross-modal fusion network first maps multi-view structural features and spectral modal features to a unified dimension. Then, using the embedded representation of structural features as the query and the embedded representation of spectral features as the key and value, a multi-head attention layer calculates the correlation score between the query and the feature vectors corresponding to each wavelength position in the key. After normalization, the values are weighted and summed to generate cross-modal features. Finally, the modal fusion layer integrates this aligned feature with the original embedded representation, and a cascaded cross-attention module in the global modeling layer performs multiple rounds of deep interaction between the structural and spectral features, ultimately outputting enhanced cross-modal fusion features. This fusion mechanism abandons the coarse-grained fusion approach of directly splicing heterogeneous features or statically weighting them in traditional methods. By introducing a query-key-value interaction structure based on an attention mechanism, it enables multi-view geometric information to adaptively focus on the physically related wavelength positions in the spectral response, thereby achieving point-by-point alignment and deep fusion of structure and spectrum in the wavelength dimension. This significantly improves the accuracy, robustness, and physical interpretability of cross-modal mapping.
[0097] The projection enhancement decoding network includes a spatial projection module, an upsampling module, a structure reconstruction module, and a spectral reconstruction module.
[0098] The method involves inputting the cross-modal fusion features into the projection enhancement decoding network, and outputting the three-dimensional geometric structure reconstruction results and spectral response prediction data of the corresponding meta-atoms; specifically:
[0099] The spatial projection module uses a learnable spatial projection matrix to map the cross-modal fusion features to a three-dimensional geometric space, and superimposes position codes corresponding to each spatial location to obtain projection features; the position codes are used to represent spatial coordinate information.
[0100] The projection features are upsampled at multiple levels by the upsampling module. After each level of upsampling, a position attention mechanism is applied to the features obtained by the upsampling at that level to model the dependencies between spatial positions. After multi-level processing, a position-aware spatial feature representation is output.
[0101] The structure reconstruction module reconstructs the three-dimensional geometric structure of the corresponding meta-atom based on the position-aware spatial feature representation and uses three-dimensional convolution operations to reconstruct the three-dimensional geometric structure and outputs voxelized data or 3D model representing the three-dimensional geometric structure.
[0102] The spectral reconstruction module maps the spatial features of the location awareness into spectral response prediction data that corresponds one-to-one with each wavelength position in the spectral sequence. The spectral response prediction data includes the phase response at each wavelength position, or the phase response and amplitude response at each wavelength position.
[0103] To address the issues of traditional decoders lacking spatial awareness and exhibiting inconsistencies between generated structures and physical behavior, this invention proposes a projection-enhanced decoding network. This network first utilizes a learnable spatial projection matrix to map cross-modal fused features onto a 3D geometric space, and then superimposes positional codes corresponding to each spatial location. Subsequently, through multi-level upsampling combined with a positional attention mechanism, it explicitly models the long-range dependencies between spatial locations, generating position-aware spatial feature representations. Finally, based on these representations, it reconstructs the 3D geometric structure of the corresponding meta-atoms through 3D convolution operations. This mechanism endows the decoding process with explicit spatial coordinate awareness, ensuring high self-consistency in geometric distribution and optical response of the output structure, significantly improving the model's physical interpretability and engineering practicality.
[0104] The method of training the superlens inverse design network through each target element atomic sample also includes:
[0105] The reconstructed voxelized data or 3D model is reprojected onto the imaging plane of each input viewpoint to generate the reconstructed image of the corresponding viewpoint.
[0106] The difference between each reconstructed image and its corresponding viewpoint structural image is calculated as the viewpoint consistency error.
[0107] The consistency error between viewpoints and the 3D geometric reconstruction error are weighted and combined to form a total loss function. The model parameters of the superlens inverse design network are then optimized based on the total loss function.
[0108] To address the issues of lack of consistency constraints and viewpoint-dependent bias in reconstruction results under multi-view input, this invention introduces a reprojection consistency supervision mechanism during training. Specifically, the reconstructed voxelized data or 3D model is reprojected onto the imaging plane of each input viewpoint to generate a reconstructed image for that viewpoint. The pixel-level or structural-level difference between this reconstructed image and the original structural image is calculated as the viewpoint consistency error. This error is weighted and combined with the 3D geometric reconstruction error to form a total loss function, which is used to jointly optimize the network parameters. This constraint minimizes the difference between the reconstructed 3D structure reprojected onto each input viewpoint and the corresponding original structural image, prompting the model to generate a geometrically self-consistent 3D structure under multi-view input. This ensures the intrinsic consistency of multi-view geometric information, effectively suppresses viewpoint-related bias in structural reconstruction, and significantly improves the stability and reliability of the model under multi-view joint input.
[0109] In this embodiment, the meta-atom sample library is constructed using a full-wave electromagnetic simulation method. Each meta-atom sample corresponds to a preset superlens meta-atom three-dimensional geometric structure (as the true value of the three-dimensional reconstruction) and includes: (1) two-dimensional structural images of the structure from multiple perspectives; (2) spectral response data of the structure under preset incident conditions (such as normal incidence, specific polarization state). The multiple perspectives are uniformly distributed within a ±30° range relative to the surface normal direction of the meta-atom, for example, with 5° intervals, acquiring structural images from 13 perspectives (including –30°, –25°, …, 0°, …, +25°, +30°). During model training, the multi-view structural images and spectral response data are used as input, with the goal of reconstructing the original three-dimensional geometric structure. By minimizing the three-dimensional geometric structure reconstruction error between the reconstruction result and the true value, the network parameters are continuously optimized. After training, the effectiveness and accuracy of the proposed method can be evaluated by calculating this reconstruction error.
[0110] It is important to note that the 3D geometric reconstruction error is used to quantify the geometric difference between the reconstruction result output by the model and the preset superlens atomic 3D geometric structure (i.e., the true value) of the corresponding atomic sample. Specifically, after aligning the reconstructed 3D structure (voxelated data or 3D model) with the true structure in the same spatial coordinate system, the error is calculated by comparing the deviations of the two in key geometric features. In one embodiment, the 3D geometric reconstruction error is measured by surface distance: first, the reconstructed 3D structure is aligned with the corresponding true 3D structure in a unified spatial coordinate system; then, uniform sampling is performed on the surfaces of the two structures to obtain the reconstructed point set and the true point set. The average of the shortest distances from the reconstructed point set to the true surface and the average of the shortest distances from the true point set to the reconstructed surface are calculated, and the sum or mean of the two is taken as the final 3D geometric reconstruction error. This error directly reflects the local deviation of the two structures in terms of geometric shape; the smaller the value, the closer the reconstruction result is to the true value.
[0111] In this embodiment, the training of the superlens reverse design network is mainly based on the total loss function formed by the weighted combination of the inter-viewpoint consistency error and the 3D geometric reconstruction error. In a preferred embodiment, to further improve the physical consistency of the reconstruction results, the spectral reconstruction loss can also be introduced into the training process: Specifically, the model maps the position-aware spatial feature representation into spectral response prediction data corresponding one-to-one with each wavelength position in the spectral sequence through the spectral reconstruction module. This prediction data includes the phase response at each wavelength, or simultaneously includes the phase response and amplitude response. Subsequently, the prediction data is compared with the spectral response data provided in the atomic sample (i.e., the real spectral response data corresponding to the same atomic atom under preset incident conditions), the spectral reconstruction loss is calculated, and it is combined with the aforementioned geometric correlation loss to form an extended total loss function, which is used to optimize the network parameters end-to-end.
[0112] The specific calculation method for the spectral reconstruction loss is as follows: For each wavelength position in the spectral sequence, the square of the difference between the predicted phase response and the original phase response is calculated; if the predicted amplitude response is also included, the square of the difference between the predicted amplitude response and the original amplitude response is also calculated. Subsequently, the squared phase errors at all wavelength positions are averaged to obtain the phase mean square error; the squared amplitude errors at all wavelength positions are averaged to obtain the amplitude mean square error. Finally, the phase mean square error is used as the spectral reconstruction loss, or the phase mean square error and the amplitude mean square error are added together according to a preset weight to obtain the spectral reconstruction loss.
[0113] The target spectral response data is acquired and input into the superlens reverse design model to obtain the three-dimensional geometric structure reconstruction result of the corresponding elementary atom.
[0114] This invention proposes a reverse design method for the atomic structure of a superlens based on multi-view structural images. By constructing a superlens reverse design model including a multi-view encoding module, a spectral coupling module, a cross-modal fusion network, and a projection enhancement decoding network, it achieves end-to-end reverse mapping from the target spectral response to a high-fidelity three-dimensional atomic structure. This method abandons the traditional parameter scanning process that relies on massive electromagnetic simulations. While significantly reducing computational costs, it effectively integrates multi-view geometric information and spectral response data, systematically solving multiple technical bottlenecks such as insufficient single-view modeling, difficulties in modal alignment, and lack of spatial awareness. It significantly improves the accuracy, stability, and physical interpretability of superlens design, providing a reliable technical path for the rapid development of multifunctional, multi-band planar optical devices.
[0115] This invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, which, when executed by the at least one processor, causes the electronic device to perform the method of this invention.
[0116] The present invention also provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform the method of the present invention.
[0117] This invention also provides a computer program product, including a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform the method of this invention.
[0118] refer to Figure 2 The present invention will now describe a structural block diagram of an electronic device that can serve as a server or client in embodiments of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0119] like Figure 2As shown, the electronic device includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. The RAM 403 may also store various programs and data required for the operation of the electronic device. The computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0120] Multiple components in the electronic device are connected to I / O interface 405, including: input unit 406, output unit 407, storage unit 408, and communication unit 409. Input unit 406 can be any type of device capable of inputting information into the electronic device. Input unit 406 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of the electronic device. Output unit 407 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 408 may include, but is not limited to, disks and optical discs. Communication unit 409 allows the electronic device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0121] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, CPUs, graphics processing units (GPUs), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the methods and processes described above. For example, in some embodiments, the method embodiments of the present invention may be implemented as a computer program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on an electronic device via ROM 402 and / or communication unit 409. In some embodiments, the computing unit 401 may be configured to perform the methods described above by any other suitable means (e.g., by means of firmware).
[0122] Computer programs for implementing the methods of embodiments of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0123] In the context of embodiments of the present invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0124] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.
[0125] Furthermore, in this invention, descriptions involving terms such as "first," "second," and "a" are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0126] In this invention, unless otherwise explicitly specified and limited, the terms "connection," "fixed," etc., should be interpreted broadly. For example, "fixed" can mean a fixed connection, a detachable connection, or an integral part; it can mean a mechanical connection or an electrical connection; it can mean a direct connection or an indirect connection through an intermediate medium; it can mean the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0127] Furthermore, the technical solutions of the various embodiments of the present invention can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.
Claims
1. A method for reverse design of meta-atomic structures of superlenses based on multi-view structural images, characterized in that, include: A meta-atom sample library is constructed, which includes multiple meta-atom samples. Each meta-atom sample includes structural images of a meta-atom from multiple perspectives and spectral response data of the meta-atom under preset incident conditions. The meta-atom samples are preprocessed to obtain target meta-atom samples. A superlens inverse design network is constructed, which includes a multi-view coding module, a spectral coupling module, a cross-modal fusion network, and a projection enhancement decoding network. The superlens inverse design network is trained using individual target atomic samples to obtain the superlens inverse design model, including: The structural image corresponding to each viewpoint in each target atomic sample is input into the multi-view encoding module. The module uses graph convolution operation to generate the multi-view structural features corresponding to the target atomic sample. The spectral response data in the target atomic sample is input into the spectral coupling module. The module jointly characterizes the variation law of the spectral response data in the wavelength dimension and its frequency domain characteristics to generate the spectral modal features corresponding to the target atomic sample. The step of inputting the structural image corresponding to each viewpoint in each target atomic sample into the multi-view encoding module, and generating the multi-view structural features corresponding to the target atomic sample through graph convolution operation by the module, specifically involves: The structural image from each viewpoint is treated as a node, and an adjacency relationship is constructed based on the angle difference between any two viewpoints: if the angle difference is less than a preset threshold, a connection is established between the corresponding two nodes, thereby forming an adjacency relationship matrix; The number of connections for each node is calculated based on the adjacency matrix, and the adjacency matrix is normalized accordingly to obtain a normalized adjacency matrix. Feature extraction is performed on the structural images from each viewpoint to obtain the corresponding node features. The node features corresponding to each node are input into a multi-layer graph convolutional structure along with the normalized adjacency matrix. The node features of adjacent viewpoints are aggregated layer by layer through the multi-layer graph convolutional structure, so that the node features corresponding to viewpoints with an angle difference less than the preset threshold are aligned with each other, and finally the multi-view structural features corresponding to the target element atom sample are output. Multi-view structural features and spectral modal features are embedded, and multi-level cross-modal fusion is performed through a cross-modal fusion network based on the obtained embedded representation to generate cross-modal fusion features; the cross-modal fusion features are input into the projection enhancement decoding network to output the three-dimensional geometric structure reconstruction results of the corresponding meta-atoms and their spectral response prediction data; The target spectral response data is acquired and input into the superlens reverse design model to obtain the three-dimensional geometric structure reconstruction result of the corresponding elementary atom.
2. The method for reverse design of meta-atom structures of superlenses based on multi-view structural images according to claim 1, characterized in that, The meta-atom sample is preprocessed to obtain the target meta-atom sample, specifically as follows: All structural images are uniformly scaled to a preset size to obtain scaled structural images; The spectral response data is interpolated to have a fixed number of sampling points within a preset wavelength range, resulting in interpolated spectral response data; the spectral response data is a phase-wavelength curve, or a combination of an amplitude-wavelength curve and a phase-wavelength curve. Normalization processing is performed on the scaled structural image and the interpolated spectral response data to obtain normalized structural image and spectral response data; A data augmentation operation is applied to the normalized structural image, the data augmentation operation including at least one of rotation, scaling and Gaussian noise perturbation within a preset angle range.
3. The method for reverse design of meta-atom structures of superlenses based on multi-view structural images according to claim 1, characterized in that, The spectral response data of the target atomic sample is input into the spectral coupling module. This module then jointly characterizes the variation of the spectral response data in the wavelength dimension and its frequency domain characteristics to generate the spectral modal features corresponding to the target atomic sample. Specifically: The spectral response data is constructed into a spectral sequence in wavelength order: when the spectral response data only contains phase-wavelength curves, it is constructed into a one-dimensional spectral sequence; when the spectral response data contains both amplitude-wavelength curves and phase-wavelength curves, it is constructed into a dual-channel spectral sequence, where one channel corresponds to the amplitude-wavelength curve and the other channel corresponds to the phase-wavelength curve. When the spectral sequence is one-dimensional, a comprehensive feature vector for each wavelength position is generated by analyzing the numerical changes and frequency domain characteristics of the phase-wavelength curve in the neighborhood of each wavelength position. When the spectral sequence is dual-channel, the numerical changes and frequency domain characteristics of the phase-wavelength curve and amplitude-wavelength curve in the neighborhood of each wavelength position are analyzed separately, generating the phase comprehensive feature vector and amplitude comprehensive feature vector corresponding to each wavelength position, and then splicing the two into the comprehensive feature vector of that wavelength position. The comprehensive feature vectors of all wavelength positions are combined in wavelength order to form the spectral modal features corresponding to the target atomic sample.
4. The method for reverse design of meta-atom structures of superlenses based on multi-view structural images according to claim 3, characterized in that, When the spectral sequence is one-dimensional, a comprehensive feature vector for each wavelength position is generated by analyzing the numerical changes and frequency domain characteristics of the phase-wavelength curve in the neighborhood of each wavelength position. Specifically: When the spectral sequence is one-dimensional, based on the numerical changes of the phase-wavelength curve in the neighborhood of each wavelength position, the local trend characteristics of the phase corresponding to each wavelength position are calculated, and frequency domain analysis is performed on the neighborhood of each wavelength position to obtain the local frequency domain distribution characteristics of the phase corresponding to each wavelength position. At each wavelength position, the local phase trend characteristics and local phase frequency domain distribution characteristics at that position are fused to generate a comprehensive feature vector for that wavelength position.
5. The method for reverse design of meta-atom structures of superlenses based on multi-view structural images according to claim 4, characterized in that, When the spectral sequence is dual-channel, the numerical changes and frequency domain characteristics of the phase-wavelength curve and amplitude-wavelength curve in the neighborhood of each wavelength position are analyzed respectively, generating a phase comprehensive feature vector and an amplitude comprehensive feature vector corresponding to each wavelength position, and concatenating the two to form a comprehensive feature vector for that wavelength position, specifically: When the spectral sequence is dual-channel, the local phase trend features and local amplitude trend features corresponding to each wavelength position are calculated based on the numerical changes of the phase-wavelength curve and amplitude-wavelength curve in the neighborhood of each wavelength position. Frequency domain analysis is then performed on the neighborhood of each wavelength position to obtain the local phase frequency domain distribution features and local amplitude frequency domain distribution features corresponding to each wavelength position. At each wavelength position, the local phase trend features and local phase frequency domain distribution features, as well as the local amplitude trend features and local amplitude frequency domain distribution features, are fused to obtain the phase comprehensive feature vector and the amplitude comprehensive feature vector. These two are then concatenated to form the comprehensive feature vector for that wavelength position.
6. The method for reverse design of meta-atom structures of superlenses based on multi-view structural images according to claim 5, characterized in that, The cross-modal fusion network comprises a multi-head attention layer, a modal fusion layer, and a global modeling layer connected in sequence; Multi-view structural features and spectral modal features are embedded, and multi-level cross-modal fusion is performed through a cross-modal fusion network based on the obtained embedded representation to generate cross-modal fusion features; Specifically: The multi-view structural features and the spectral modal features are mapped to a feature space of the same dimension through linear transformation to obtain corresponding embedding representations; wherein, the embedding representation of the spectral modal features contains multiple feature vectors, and the feature vectors correspond one-to-one with the wavelength positions in the spectral sequence; The embedding representations corresponding to the multi-view structural features and the spectral modal features are respectively input into the cross-modal fusion network: In the multi-head attention layer, the embedded representation of multi-view structural features is used as the query, and the embedded representation of spectral modal features is used as the key and value, and cross-modal features are generated through the attention mechanism. In the modality fusion layer, the embedded representations of cross-modal features, multi-view structural features, and spectral modality features are fused to generate a joint feature representation; In the global modeling layer, multiple cascaded cross-attention modules are used to perform multiple rounds of cross-modal interaction between the embedded representation of multi-view structural features and the embedded representation of spectral modal features to generate enhanced features, which are then fused with the joint feature representation to output cross-modal fused features.
7. The method for reverse design of meta-atom structures of superlenses based on multi-view structural images according to claim 6, characterized in that, The projection enhancement decoding network includes a spatial projection module, an upsampling module, a structure reconstruction module, and a spectral reconstruction module. The method involves inputting the cross-modal fusion features into the projection enhancement decoding network, and outputting the three-dimensional geometric structure reconstruction results and spectral response prediction data of the corresponding meta-atoms; specifically: The spatial projection module uses a learnable spatial projection matrix to map the cross-modal fusion features to a three-dimensional geometric space, and superimposes the position codes corresponding to each spatial position to obtain the projection features. The location code is used to represent spatial coordinate information; The projection features are upsampled at multiple levels by the upsampling module. After each level of upsampling, a position attention mechanism is applied to the features obtained by the upsampling at that level to model the dependencies between spatial positions. After multi-level processing, a position-aware spatial feature representation is output. The structure reconstruction module reconstructs the three-dimensional geometric structure of the corresponding meta-atom based on the position-aware spatial feature representation and uses three-dimensional convolution operations to reconstruct the three-dimensional geometric structure and outputs voxelized data or 3D model representing the three-dimensional geometric structure. The spectral reconstruction module maps the spatial features of the location awareness into spectral response prediction data that corresponds one-to-one with each wavelength position in the spectral sequence. The spectral response prediction data includes the phase response at each wavelength position, or the phase response and amplitude response at each wavelength position.
8. The method for reverse design of meta-atomic structures of superlenses based on multi-view structural images according to claim 7, characterized in that, The method of training the superlens inverse design network through each target element atomic sample also includes: The reconstructed voxelized data or 3D model is reprojected onto the imaging plane of each input viewpoint to generate the reconstructed image of the corresponding viewpoint. The difference between each reconstructed image and its corresponding viewpoint structural image is calculated as the viewpoint consistency error. The consistency error between viewpoints and the 3D geometric reconstruction error are weighted and combined to form a total loss function. The model parameters of the superlens inverse design network are then optimized based on the total loss function.
9. An electronic device, comprising: A processor and a memory storing a program, characterized in that the program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Superatom structure on-demand design method based on diffusion model and contrast learning
CN119049612A
All-dielectric metasurface target spectral response reverse design method based on deep learning
CN120745407A