A cross-modal driven superlens inverse design method, device and electronic medium
By employing a cross-modal driven design approach, multimodal features of metasurface structures are extracted and fused, addressing the issues of insufficient model accuracy and information redundancy in existing technologies. This enables high-precision mapping and improved generalization performance in the reverse design of metalenses.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 浙江优众新材料科技有限公司
- Filing Date
- 2026-05-12
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies in the reverse design of superlenses suffer from problems such as insufficient accuracy of model parameters, poor fitting of spectral curve peaks, and information redundancy and feature loss during multimodal feature fusion, making it impossible to achieve accurate control and prediction of multiple types of superstructures.
A cross-modal driven design approach is adopted to extract semantic, visual, and spectral features of metasurface structures through a feature extraction network, calculate feature similarity in the CSCLNet network, obtain local feature information in the LSS2Net network, and calculate feature point difference using the GSPNet network to achieve image and spectrum reconstruction.
It significantly improves the mapping accuracy between geometric features and spectral responses and the generalization performance of the model, alleviates the problems of information redundancy and semantic offset, and enhances cross-modal feature alignment and generalization capabilities.
Smart Images

Figure CN122172449B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of superlens technology, and in particular to a method, apparatus and electronic medium for reverse engineering of superlenses driven by cross-modal operation. Background Technology
[0002] In recent years, the reverse design of metalens has received widespread attention. Numerous research efforts have been conducted using methods such as deep learning. Unlike existing forward design, reverse design defines optical targets and uses optimization algorithms to search for structural parameters. The fundamental problem is improving the optimization process for finding the optimal atomic design parameters of the metastructure, which consumes significant computational resources and time. With the development of deep learning, some researchers have used generative adversarial neural networks to learn the mapping relationship between metasurface structures and spectral responses. Others have proposed using autoencoders to achieve image-based meta-atomic design and spectral response mapping. However, in the generation of meta-atomic structural parameters, these methods suffer from insufficient model parameter accuracy and poor peak fitting of spectral curves.
[0003] Furthermore, these works primarily focus on predicting specific structures, merely understanding the relationship between input and output, and cannot achieve accurate control and prediction of multiple types of superstructures. Achieving multimodal fusion of metaatomic images and spectral responses requires not only considering information decoupling but also mining additional supervisory information. Relying solely on pixel information has historically lacked accurate descriptive and discriminative capabilities. Simultaneously, existing methods for multimodal tasks typically focus on aligning different modal features to a unified representation space. However, in the task of metasurface inverse design, image and spectral information largely does not require alignment due to poor correlation. Existing research has neglected the complementary information interaction between modalities and the propagation of features within modalities. What is needed is how to utilize complementary model information to adjust the degree of fusion to address the problems of information redundancy, feature loss, and difficulty in alignment supervision that arise during feature fusion, thereby affecting the model's expressive and generative capabilities. Summary of the Invention
[0004] This invention provides a cross-modal driven superlens reverse design method, device, and electronic medium, which aims to effectively alleviate the problems of information redundancy and semantic offset in the multimodal feature fusion process, and significantly improve the mapping accuracy between geometric features and spectral response and the generalization performance of the reverse design model.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A cross-modal driven superlens reverse design method, wherein the method includes the following steps: Feature extraction networks are used to extract features from text, images, and spectra of metasurface structures to obtain semantic features, visual features, and spectral features. The three types of features are then fused and aligned to achieve a three-modal mapping. In the CSCLNet network, the feature similarity among three types of features—semantic features, visual features, and spectral features—is compared to obtain a similarity score. In LSS 2 In the Net network, the three types of features are used as input to the short-range state-aware SS module to obtain local feature information of the three types of features, and the local feature information and similarity score are used as input to the long-range state-aware LS module to obtain reconstructed features. The reconstructed features are input into a feature-aware mapping network to calculate the feature point difference of the reconstructed features, and the feature point difference is used for image reconstruction and spectral reconstruction.
[0006] Furthermore, the metasurface structure text representation input to the feature extraction network is a geometric semantic description of the corresponding metasurface structure.
[0007] Furthermore, based on the description of the shape of the real atomic image and the description of the geometric statistical values of the shape of the real atomic image, various types of geometric semantic descriptions of the metasurface structure are obtained.
[0008] Furthermore, the metasurface structure spectrum input to the feature extraction network is represented as the phase spectral response of the corresponding metasurface structure within a preset range of wavelengths.
[0009] Furthermore, the feature extraction network includes a CLIP model and a spectral response distribution encoder. The CLIP model includes an image feature encoder and a text feature encoder. The image feature encoder and the text feature encoder are used to extract the visual and semantic features of the metasurface structure, respectively, while the spectral response distribution encoder is used to extract the feature representation of the phase spectrum.
[0010] Furthermore, the image feature encoder extracts features in a dual-channel manner, where one channel uses prior knowledge from the CLIP model to obtain features, and the other channel directly uses the CLIP model to obtain features.
[0011] Furthermore, in the CSCLNet network, the similarity between visual features and semantic features is represented as a similarity score α, and the similarity between visual features and spectral features is represented as a similarity score β.
[0012] Further steps include: passing the reconstructed features through a CNN decoding network to restore the image and spectrum to obtain the generated image and the generated spectrum.
[0013] The present invention also provides a cross-modal driven superlens reverse design apparatus, the apparatus including a processor, which executes a computer program stored in a memory to implement the steps of the above-described cross-modal driven superlens reverse design method.
[0014] The present invention also provides a computer-readable storage medium that, when the instructions in the storage medium are executed by a processor within the device, enables the device to perform the above-described cross-modal driven superlens reverse engineering method.
[0015] Compared with the prior art, the present invention has at least the following beneficial effects: (1) Based on the cross-modal multi-type feature extraction network and selective feature fusion mechanism, this scheme realizes multi-modal joint representation of meta-atomic images, text semantics and spectral response features from the latent space through multi-type superstructure text semantic generation, dual-channel pre-trained feature extraction and spectral response feature extraction; (2) In this scheme, text semantic generation is based on CLIP prior knowledge to design text prompts. By semantically describing the geometric shape and structural parameters of meta-atoms, the cross-modal feature alignment and generalization ability are significantly enhanced. In the dual-channel feature extraction stage, the meta-atom image and text prompts are calculated through a contrastive learning mechanism to dynamically adjust the knowledge transfer intensity and achieve adaptive cross-modal selective fusion. (3) This scheme introduces a two-level long / short selection space module (LSS) 2 Net captures multi-scale structural relationships between modalities from both long and short dependency perspectives, significantly improving the sufficiency and relevance of feature representation; (4) The cross-modal geometric space semantic feature perception mapping network (GSPNet) designed in this scheme realizes the consistency modeling and difference perception control of geometric-spectral space by calculating the feature point difference between geometric features and optical response distribution, which provides a precise weight allocation basis for subsequent feature fusion. Attached Figure Description
[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a flowchart of the steps of the cross-modal driven superlens reverse design method provided in this embodiment; Figure 2 This is a schematic diagram of the network architecture for cross-modal fusion in the method provided in this embodiment. Detailed Implementation
[0018] The following are specific embodiments of the present invention, and the technical solutions of the present invention will be further described in conjunction with the accompanying drawings. However, the present invention is not limited to these embodiments.
[0019] like Figure 1 and Figure 2 As shown, this embodiment provides a superlens inverse design method based on a cross-modal multi-type feature extraction network and a selective feature fusion mechanism. It mainly includes a three-branch feature extraction network TFENNet, a cross-modal similarity comparison learning network CSCLNet, and a two-level long / short selection space network LSS. 2 Networks such as Net and GSPNet, a modal geometry spatial semantic feature perception mapping network for modal reconstruction.
[0020] The three-branch feature extraction network TFENNet consists of three parts: a multi-type superstructured text semantic generation module, a dual-channel pre-trained fine-tuning feature extraction module, and a spectral response feature extraction module. It initially extracts features between multimodal meta-atomic images, text, and corresponding spectra from the latent space.
[0021] The design method provided in this embodiment includes the following steps: S1. Based on the feature extraction network, feature extraction is performed on the text, image and spectrum of the metasurface structure to obtain semantic features, visual features and spectral features, and the three types of features are fused and aligned to achieve the corresponding mapping of the three modes; S2. In the CSCLNet network, compare the feature similarity among the three types of features: semantic features, visual features, and spectral features to obtain similarity scores; S3, in LSS 2 In the Net network, the three types of features are used as input to the short-range state-aware SS module to obtain local feature information of the three types of features, and the local feature information and similarity score are used as input to the long-range state-aware LS module to obtain reconstructed features. S4. Input the reconstructed features into the GSPNet network to calculate the feature point difference of the reconstructed features, and use the feature point difference for image reconstruction and spectral reconstruction.
[0022] Specifically, before proceeding with this method, it is necessary to first acquire images, texts, and corresponding phase spectrum datasets of various types of metasurface structures, and then divide the data into training and testing sets.
[0023] The dataset contains samples in triplet form {geometric text semantics, metasurface structure image, phase spectrum}: { , , }, , as well as These correspond to the geometric semantic description of the metasurface structure, its structural image, and its phase spectral response in a specific wavelength band. , , H, W, and C represent the height, width, and dimension of the image, respectively; L represents the length of the spectrum; and N represents the text length.
[0024] In geometric semantic description, various types of geometric semantic descriptions of real atomic metasurface structures are obtained based on the descriptions of the shapes of real atomic images and the descriptions of the geometric statistical values of those shapes. These include: the description of the shape of the real atomic metasurface as a circle, square, cross, butterfly, H-shape, V-shape, etc. Correspondingly, the geometric statistical values of these shapes are represented as the physical parameters of the shapes, including dimensions, height, radius, etc. This results in geometric textual descriptions of various types of metasurface structures describing real atomic metasurface structures. Its geometrical statistical values .
[0025] By aligning the image embedding of the metasurface structure, the phase spectral dataset embedding of the metasurface structure, and the average text embedding of the image geometry using this method, a geometric semantic description of the metasurface structure can be obtained. .
[0026] Furthermore, in the feature extraction network, a Contrastive Language-Image Pre-training (CLIP) model is used to extract features from textual cues and images of the metasurface structure. This model consists of two sub-networks: an image feature encoder f(·) and a text feature encoder g(·). The image feature encoder f(·) and various types of geometric textual descriptions of the metasurface shape are then processed. Its geometrical statistical values Embedding, fusion, and comparison are performed to extract visual features. and semantic features Alignment is achieved in the shared feature space. Simultaneously, the feature representation of the phase spectrum is extracted using the spectral response distribution encoder s(·). This enables the mapping between spectral modes and visual and semantic modes.
[0027] Among them, metasurface images are preferred. Feature extraction is performed using a dual-channel approach, with one channel fine-tuned using prior knowledge from the CLIP model to obtain features. The other channel directly uses the CLIP model to obtain features. It can efficiently extract and learn the geometric shape features of multiple types of superstructures.
[0028] In this embodiment, the feature extraction network TFENNet utilizes a multi-type superstructured text semantic generation module, a dual-channel pre-trained fine-tuned feature extraction module, and a spectral response feature extraction module to initially extract features between multimodal meta-atomic images, text, and corresponding spectra from the latent space. The multi-type superstructured text semantic generation module leverages prior knowledge from the CLIP model to improve the generalization performance of feature detection across various modalities.
[0029] Furthermore, this embodiment designs a cross-modal similarity comparison learning network CSCLNet, which compares the similarity between meta-atom images and meta-atom text descriptions, as well as the similarity between meta-atom images and meta-atom spectral response distribution information, based on a comparison learning approach.
[0030] The Cross-Modal Similarity Contrastive Learning Network (CSCLNet) is based on a contrastive learning mechanism. It calculates the feature similarity between meta-atomic images and geometric semantic descriptions, and between meta-atomic images and phase spectral responses, to measure the consistency of different modalities in semantic and physical spaces. In the contrastive learning network, the similarity between visual and semantic features is represented as a similarity score α, and the similarity between visual and spectral features is represented as a similarity score β. By comparing the similarity scores α and β obtained from the loss function, a similarity probability is obtained to dynamically characterize the matching degree between modalities. This score is used as a control factor in the subsequent knowledge fusion module to adaptively adjust the weights of cross-modal feature transfer, thereby improving the accuracy and stability of multimodal feature fusion.
[0031] Specifically, let , , , First, perform normalization: , , , .
[0032] Define cosine similarity: , .
[0033] Comparative losses: ; ; in, N is a constant. a It is represented as {a,b,c,d}.
[0034] Symmetry is usually performed: ; .
[0035] Using probabilistic similarity as a criterion for quantifying similarity, we define it as a similarity score, α and β: ; .
[0036] Furthermore, this embodiment introduces a two-level long-short selection space net (LSS). 2 This module (Net) combines the global modeling capabilities of the Mamba model with the local feature correlations based on CNNs. It consists of cascaded LSS... 2 The Net global modeling unit is composed of parallel multi-kernel deep convolutional units, forming a multi-scale feature extraction structure that is "global-local" collaborative.
[0037] The obtained image, text, and spectral features As input to the short-range state-aware SS1 and SS2 modules, similarity scores α and β are used as input to the long-range and short-range state-aware LS modules to adjust the degree of fusion of various mixed knowledge from images, text, and spectra. The similarity score α mainly adjusts the shallow detail features from the text encoder. (Such as words, sentences, etc.) and the deep abstract features of the SS2 module Meanwhile, the similarity score β adjustment originates from the shallow detail features of the spectral encoder. (Such as intensity, amplitude, etc.) and the deep abstract features of the SS2 module .
[0038] Then, based on the obtained shallow detail features and Deep abstract features and The input is fed into a long- and short-range state-aware LS module, aiming to combine local detail features from different modalities with common features. Ultimately, fused reconstructed features from the image and spectrum are obtained respectively. and .
[0039] Specifically, at the local level, the Short-Space (SS) submodule models fine-grained feature associations between meta-atomic images and text prompts, and between meta-atomic images and spectral responses, respectively, through multi-kernel convolution operations, capturing local feature information such as geometric edges, textures, and semantic terms; at the global level, the Long-Space (LS) submodule achieves globally consistent expression of cross-modal features by performing contextual modeling of feature dependencies between the two-level SS submodules.
[0040] Furthermore, the LS module utilizes the similarity scores α and β obtained during the contrastive learning phase as dynamic adjustment factors to establish skip connections between shallow features and deep abstract features, achieving adaptive fusion of cross-level information. By calculating the similarity between modalities at each feature point, LSS... 2 Net can automatically adjust the intensity of knowledge transfer, thereby significantly improving the model's ability to express complex cross-modal features and overall task accuracy.
[0041] Local Short Spatial Enhancement (SS) Module: To capture the local feature dependencies between image-text and image-spectrum, two independent short spatial enhancement submodules are defined: ; .
[0042] in, , All of them use the same Local Short Spatial Augmentation (SS) module, which mainly includes Layer Normalization (LN), linear layers, depthwise convolution, and the SiLU activation function. Short Spatial Augmentation Encoder It consists of 4 convolutional layers, a state-space model (SSM), and a short-space augmentation encoder. It connects with residuals to extract local geometric edges and semantic dependencies.
[0043] ; ; .
[0044] , The components are combined using a splicing method, and after normalization, two branches emerge from the normalization layer. One branch enters the linear layer (Linear) and DWConv to obtain... Cascaded depthwise convolutions are performed to extract latent temporal features. The extracted spatially enhanced features are then processed through an activation function to obtain preprocessed feature results. .
[0045] DWConv is a concatenated depthwise convolution that uses a kernel size of 3 and an increasing dilation factor. To enlarge the convolution kernel, capture depth information and capture shallow, short-space features. Then, the features... Input to short space enhancement encoder In the process, short-space enhancement features between modes are obtained through the feature fusion SSM module. .
[0046] Another branch directly enters the linear layer, passes through an activation function, multiplies the result of the previous branch, and then performs dimensional alignment through the linear layer, ultimately yielding a local short-space enhancement feature carrying spectral information. This enhances the difference between text and images. .generate The formula is as follows: ; Similarly, for , Local short-space enhancement features are combined by stitching together information based on the differences between the image and spectral response obtained from the above process. .
[0047] To capture global dependencies between modalities, a global long spatial (LS) module is further constructed by cascading the aforementioned Local Short Space Enhancement (SS) modules and combining them with a CNN module. This achieves local short space enhancement features for the difference information between the image and the spectral response. Local short spatial enhancement features that provide difference information between text and images Information exchange between them. Simultaneously, a similarity-guided adaptive fusion mechanism is introduced based on modal similarity scores. Adaptive adjustment of feature fusion weights.
[0048] Among them, the Local Short Space Enhancement (SS) module Local short-space enhancement features that obtain differential information between text and images through cascading nesting. Further construct global long-space augmentation features Local short spatial enhancement features that provide information about the difference between the image and the spectral response. Short local information is obtained through cascaded deep convolution and convolution modules. The degree of fusion is controlled by the similarity score α, and finally the global long-space features of image and text are obtained. .
[0049] ; ; .
[0050] Similarly The operation is as follows. ; ; .
[0051] in, The similarity score obtained through contrastive learning dynamic learning is defined as the adaptive weight of semantic feature fusion.
[0052] The two-level long-short selection space network (LSS) introduced in this embodiment 2 NET captures multi-scale structural relationships between modalities from both long and short dependency perspectives, significantly improving the sufficiency and relevance of feature representation.
[0053] Furthermore, the features are reconstructed by fusing the image and spectrum output by the LS module. and The inputs are fed together into a cross-modal geometric space semantic feature-aware mapping network (GSPNet) to compute... and Feature point dissimilarity is dynamically evaluated, and its dissimilarity distribution is dynamically assessed to adaptively adjust the intensity of cross-modal knowledge transfer, thereby enabling image reconstruction and spectral reconstruction.
[0054] The degree of difference between different data modalities is a major factor affecting the effectiveness of knowledge transfer. The greater the difference in data features, the worse the effect of knowledge transfer, and directly learning cross-modal data distribution from highly different modalities may adversely affect segmentation prediction. Therefore, we propose an adaptive feature selection scheme—the Cross-Modal Geometric Space Semantic Feature Aware Mapping Network (GSPNet)—to ensure LSS (Local Semantic Segmentation). 2 The effectiveness of cross-modal knowledge transfer in Net output.
[0055] The Cross-Modal Geometric Space Semantic Feature Aware Mapping Network (GSPNet) comprises a cosine similarity layer, an adaptive factor layer, a weighted fusion layer, and an MSE loss layer. When the feature distributions of various modal paths differ significantly, the degree of knowledge transfer decreases. Therefore, we designed an adaptive factor λ to harmonize the LSS. 2 The impact of Net's cross-modal output features on metasurface geometry images and spectral reconstruction.
[0056] LSS 2 Cross-modal features of Net output , As input to the cosine similarity layer of the GSPNet cross-modal geometric space semantic feature perception mapping network, and For example, for any corresponding point on the feature image w×h and (Where w and h represent the height and width of the feature map, respectively), we calculate the vector and To prevent the curse of dimensionality, we calculate the cosine similarity R between the vectors. and Cosine similarity between It is calculated and used as an indicator to assess the magnitude of feature differences between different modalities. .
[0057] Subsequently, the cosine similarity R is normalized and smoothed to fall within the [0, 1] interval to obtain... To harmonize the automatic adjustment of LSS 2 The "strength" of knowledge transfer between the metasurface geometry image and the spectrum output by Net. If the feature differences between the two modes are too large, less is transferred to avoid "misleading"; if the differences are small, more is transferred to make full use of complementary information.
[0058] .
[0059] Cosine similarity plays a role in progressive fusion. When the feature differences are too large, it can prevent the features from changing rapidly and significantly, thereby preserving the utilization rate of the inherent features of the modality.
[0060] Then, an adaptive factor is used to weight the cross-modal features point by point in the spatial dimension, thereby achieving fine-grained modal feature fusion.
[0061] The cross-modal geometric space semantic feature perception mapping network (GSPNet) designed in this embodiment realizes consistent modeling and difference perception control of geometric-spectral space by calculating the feature point difference between geometric features and optical response distribution, providing a precise basis for subsequent feature fusion.
[0062] In addition, the reconstructed features are processed by a CNN decoding network to restore the image and spectral dimensions, ultimately yielding the generated image and the generated spectral results.
[0063] Through the synergistic effect of the aforementioned networks, this embodiment can effectively alleviate the problems of information redundancy and semantic offset in the multimodal feature fusion process, and significantly improve the mapping accuracy between geometric features and spectral responses as well as the generalization performance of the inverse design model.
[0064] This embodiment also provides a cross-modal driven superlens reverse design apparatus, the apparatus including a processor, which executes a computer program stored in a memory to implement the steps of the above-described cross-modal driven superlens reverse design method.
[0065] This embodiment also provides a computer-readable storage medium that, when the instructions in the storage medium are executed by a processor within the device, enables the device to perform the aforementioned cross-modal driven superlens reverse engineering method.
[0066] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.
[0067] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a specific posture. If the specific posture changes, the directional indication will also change accordingly.
[0068] Furthermore, in this invention, descriptions involving terms such as "first," "second," and "a" are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0069] In this invention, unless otherwise explicitly specified and limited, the terms "connection," "fixed," etc., should be interpreted broadly. For example, "fixed" can mean a fixed connection, a detachable connection, or an integral part; it can mean a mechanical connection or an electrical connection; it can mean a direct connection or an indirect connection through an intermediate medium; it can mean the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0070] Furthermore, the technical solutions of the various embodiments of the present invention can be combined with each other, but only if they are feasible for those skilled in the art. If the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.
Claims
1. A method for inverse design of a cross-modality driven metalens, characterized in that, The method includes the following steps: Feature extraction networks are used to extract features from text, images, and spectra of metasurface structures to obtain semantic features, visual features, and spectral features. The three types of features are then fused and aligned to achieve a three-modal mapping. In the CSCLNet network, the feature similarity among the three types of features—semantic features, visual features, and spectral features—is compared to obtain a similarity score. In LSS 2 In the Net network, the three types of features are used as input to the short-range state-aware (SS) module to obtain local feature information of the three types of features, and the local feature information and the similarity score are used as input to the long-range state-aware (LS) module to obtain reconstructed features. The reconstructed features are input into the GSPNet network to calculate the feature point difference of the reconstructed features, and the feature point difference is used for image reconstruction and spectral reconstruction. The textual representation of the metasurface structure input to the feature extraction network is used as a geometric semantic description of the corresponding metasurface structure; multiple types of geometric semantic descriptions of the metasurface structure are obtained based on the description of the shape of the real meta-atomic image and the description of the geometric statistical values of the shape of the real meta-atomic image.
2. The method of claim 1, wherein, The metasurface structure spectrum input to the feature extraction network is represented as the phase spectral response of the corresponding metasurface structure in a preset range of wavelengths.
3. The method of claim 1, wherein, The feature extraction network includes a CLIP model and a spectral response distribution encoder. The CLIP model includes an image feature encoder and a text feature encoder. The image feature encoder and the text feature encoder are used to extract the visual and semantic features of the metasurface structure, respectively. The spectral response distribution encoder is used to extract the feature representation of the phase spectrum.
4. The method of claim 3, wherein, The image feature encoder extracts features in a dual-channel manner, where one channel uses prior knowledge from the CLIP model to obtain features, and the other channel directly uses the CLIP model to obtain features.
5. The method of claim 1, wherein, In the CSCLNet network, the similarity between the visual features and the semantic features is represented as a similarity score α, and the similarity between the visual features and the spectral features is represented as a similarity score β.
6. The method of claim 1, wherein, The steps include: passing the reconstructed features through a CNN decoding network to restore the image and spectrum, thereby obtaining the generated image and the generated spectrum.
7. An apparatus for inverse design of a cross-modally driven metalens, comprising: The device includes a processor that executes a computer program stored in a memory to implement the steps of the cross-modal driven superlens reverse engineering method as described in any one of claims 1-6.
8. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by a processor within the device, the device is able to perform the cross-modal driven superlens reverse engineering method according to any one of claims 1-6.