Metasurface sensor design method and device based on weak supervision and electronic equipment

By using a weakly supervised bi-branch encoder network and a global context injection module, combined with a progressive resolution decoder, the problems of insufficient training samples and insufficient prediction accuracy in metasurface sensor design are solved. This enables accurate prediction of structural parameters and spectral response, simplifies the design process, and reduces costs.

CN121997776AActive Publication Date: 2026-05-08浙江优众新材料科技有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
浙江优众新材料科技有限公司
Filing Date
2026-04-08
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing metasurface sensor design methods suffer from problems such as difficulty in obtaining training samples, insufficient high-precision spectral feature capture, and lack of model interpretability, resulting in inaccurate sensor performance evaluation. In particular, peak position shifts and inaccurate linewidth predictions often occur when dealing with sharp resonance peak optical responses.

Method used

We adopt a weakly supervised metasurface sensor design method, which extracts local and global features through a dual-branch encoder network, and improves prediction accuracy step by step by combining a global context injection module and a progressive resolution decoder. Furthermore, we simplify the design process and reduce the dependence on large-scale labeled datasets by iteratively optimizing the total task loss.

Benefits of technology

It achieves accurate prediction of structural parameters and spectral response, simplifies the design process, improves design efficiency, reduces design costs, solves the problems of disconnect between local details and global laws and insufficient prediction accuracy, and enhances the accuracy of sensor performance evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997776A_ABST
    Figure CN121997776A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of optical device design, and relates to a metasurface sensor design method and device based on weak supervision and electronic equipment. The method comprises the following steps: acquiring an attribute data set of a metasurface sensor and carrying out data processing, inputting the attribute data set into a double-branch encoder network, extracting local features and global features, inputting the local features and the global features into a global context injection module, and outputting the local features and the global features into a global context injection module; enhancing local features, performing feature splicing and fusion with global features to obtain comprehensive features, inputting the comprehensive features into a progressive resolution decoder, outputting an attribute prediction result according to an intermediate prediction result of each resolution decoding stage, calculating the total loss of tasks according to the attribute prediction result and an attribute data set, and calculating the total loss of the tasks according to the attribute prediction result and the attribute data set. And iteratively optimizing and updating the attribute prediction result according to the total loss of the task. And meanwhile, depending on weak supervision core logic, the dependence on a large-scale annotation data set is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of optical device design technology, and relates to a design method, device and electronic device for metasurface sensors based on weak supervision. Background Technology

[0002] In recent years, metasurface sensors based on continuous-domain bound states have shown significant potential in molecular fingerprint detection and refractive index sensing. Continuous-domain bound states can modulate radiation loss and enhance local optical field intensity. Quasi-continuous-domain bound state modes, after symmetry protection is broken, are expected to enhance the near-field localization capability and sensing performance of sensors due to their excellent resonant lifetime and high-quality factor. Traditional numerical optimization design methods are time-consuming and labor-intensive. While deep learning-driven design methods have accelerated the process, they suffer from limitations such as difficulty in obtaining training samples, insufficient high-precision spectral feature capture, and a lack of model interpretability.

[0003] Weakly supervised learning methods offer a novel approach to alleviating the problem of insufficient training samples. Their core principle is to leverage the model's own predictive capabilities to progressively improve the quality of training samples, forming a self-reinforcing learning loop. In computer vision, this approach has achieved good performance through iterative optimization strategies for segmentation models. However, applying it to metasurface sensor design still faces challenges such as integrating physical constraints, representing multi-scale features, and accurately modeling complex optical responses. Currently, intelligent design of metasurface sensors also faces technical bottlenecks such as insufficient model generalization ability, loss of detail in high-precision spectral feature prediction, and opaque model decision-making processes. Existing methods often encounter problems such as peak position shifts and inaccurate linewidth predictions when dealing with sharp resonance peak optical responses, severely impacting the accuracy of sensor performance evaluation. Summary of the Invention

[0004] The purpose of this invention is to address the aforementioned problems in existing technologies by proposing a metasurface sensor design method based on weak supervision.

[0005] The objective of this invention can be achieved through the following technical solution: a design method for metasurface sensors based on weak supervision, comprising: Acquire and process the property dataset of the metasurface sensor, wherein the property dataset includes structural parameter vectors and spectral response vectors; The attribute dataset is input into a dual-branch encoder network to extract local and global features; The local and global features are input into the global context injection module to enhance the local features, which are then concatenated and fused with the global features to obtain comprehensive features. The integrated features are input into the progressive resolution decoder, and the attribute prediction results are output based on the intermediate prediction results of each resolution decoding stage. The total task loss is calculated based on the attribute prediction results and the attribute dataset, and the attribute prediction results are iteratively optimized and updated based on the total task loss.

[0006] As an optional embodiment of the present invention, data processing includes: The spectral response vector is downsampled to obtain spectral data at different resolutions; Standardize the structural parameter vector and the spectral response vector respectively; The standardized spectral response vector and the spectral data are concatenated along the channel dimension to obtain sequence data of the corresponding dimension; Data augmentation strategies are used to augment the structural parameter vector, the spectral response vector, and the sequence data.

[0007] As an optional embodiment of the present invention, the attribute dataset is input into a dual-branch encoder network to extract local and global features, including: By using different receptive field feature extractors, local features of different ranges are extracted from the sequence data; The encoder is used to segment the sequence data, and positional encoding is added to the segmented segments. Based on the multi-head self-attention mechanism of the encoder, the global features corresponding to each positional encoding segment are extracted.

[0008] As an optional embodiment of the present invention, the local features and global features are input into the global context injection module to enhance the local features, including: The global features are globally pooled to transform them into intermediate vectors of the corresponding dimensions. The intermediate vector is input into a multi-layer neural network for nonlinear transformation to obtain the global context vector. Calculate the association weights between the global context vector and the local features; Based on the association weights, the global context vector is injected into the local features through element-wise multiplication.

[0009] As an optional embodiment of the present invention, the integrated features are input into a progressive resolution decoder, and attribute prediction results are output based on the intermediate prediction results of each resolution decoding stage, including: The resolution decoding stage includes a first resolution stage, a second resolution stage, and a third resolution stage. The first resolution stage has the lowest resolution, and the third resolution stage has the highest resolution. The first resolution stage analyzes the comprehensive features to generate intermediate prediction results of structural parameters within a first preset accuracy range; The second resolution stage analyzes the comprehensive features and the intermediate prediction results of the structural parameters to generate intermediate prediction results of the spectral response within a second preset accuracy range; The third resolution stage performs forward design tasks and reverse design tasks respectively based on the comprehensive features, the intermediate prediction results of the structural parameters, the intermediate prediction results of the spectral response, and the enhanced local features, and generates structural parameter prediction results and spectral response prediction results within a third preset accuracy range. Among them, the first preset accuracy range has the lowest accuracy requirement, and the third preset accuracy range has the highest accuracy requirement.

[0010] As an optional embodiment of the present invention, the total task loss is calculated based on the attribute prediction results and the attribute dataset, and the attribute prediction results are iteratively optimized and updated based on the total task loss, including: Calculate the difference loss between the predicted structural parameters and the structural parameter vector to obtain the structural parameter loss result; Calculate the difference loss between the spectral response prediction result and the spectral response vector to obtain the spectral response loss result; Based on the spectral response prediction results, the reverse design task is executed again to generate secondary structure parameter prediction results. The difference loss between the secondary structure parameter prediction results and the structure parameter vector is calculated to obtain the consistency loss results. The total task loss is obtained by weighted summing of the structural parameter loss results, the spectral response loss results, and the consistency loss results. Based on the total task loss and the physical constraint rules corresponding to the attribute dataset, the loss results are corrected, and the structural parameter prediction results and the spectral response prediction results are iteratively optimized and updated.

[0011] As an optional embodiment of the present invention, after data processing, it further includes: The initial training of the model is completed based on the processed attribute dataset. The trained model is then used to predict attributes on the unlabeled dataset and generate pseudo-labels for the unlabeled data. Calculate the confidence score of the pseudo-label. If the confidence score is greater than a preset confidence threshold, add the corresponding unlabeled data and pseudo-label to the attribute dataset to complete sample augmentation. Based on the expanded attribute dataset, the steps of model training, unlabeled data attribute prediction, and pseudo-label filtering and addition are repeatedly executed to iteratively optimize the attribute dataset until a preset number of iterations or a preset early stopping condition is reached.

[0012] As an optional embodiment of the present invention, after outputting the attribute prediction result, the following further includes: Obtain the attention weights and association weights corresponding to the branches used for extracting global features; The attention weights and the association weights are mapped using the gradient-weighted class activation mapping method, and the prediction contribution of each spectral region to the attribute prediction result is calculated. Based on the attention weights and their corresponding predicted contributions, and the association weights and their corresponding predicted contributions, corresponding attention heatmaps are generated respectively.

[0013] This invention also proposes a metasurface sensor design device based on weak supervision, comprising: The data acquisition and processing module is used to acquire and process the attribute dataset of the metasurface sensor, wherein the attribute dataset includes structural parameter vectors and spectral response vectors. The feature extraction module is used to input the attribute dataset into the dual-branch encoder network and extract local and global features; The feature fusion module is used to input the local features and global features into the global context injection module, enhance the local features, and perform feature concatenation and fusion with the global features to obtain comprehensive features; The attribute prediction module is used to input the comprehensive features into the progressive resolution decoder and output the attribute prediction results based on the intermediate prediction results of each resolution decoding stage. The iterative optimization module is used to calculate the total task loss based on the attribute prediction results and the attribute dataset, and to iteratively optimize and update the attribute prediction results based on the total task loss.

[0014] The present invention also provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to implement the aforementioned weakly supervised metasurface sensor design method when executing executable instructions.

[0015] Compared with existing technologies, this invention uses a dual-branch encoder network to extract local and global features respectively, combines a global context injection module to achieve feature enhancement and fusion, and gradually improves prediction accuracy with a progressive resolution decoder. Finally, through iterative optimization of the total task loss, it effectively solves the problems of "separation between local details and global laws" and "insufficient prediction accuracy" in metasurface sensor design, achieves accurate prediction of structural parameters and spectral response, simplifies the design process, improves design efficiency, and reduces the dependence on large-scale labeled datasets by relying on weakly supervised core logic, thereby reducing the design cost and threshold of metasurface sensors. Attached Figure Description

[0016] Figure 1This is a flowchart of a metasurface sensor design method based on weak supervision according to an embodiment of the present invention; Figure 2 This is a block diagram of a metasurface sensor design device based on weak supervision according to an embodiment of the present invention. Detailed Implementation

[0017] The following are specific embodiments of the present invention, which are described in conjunction with the accompanying drawings. However, the present invention is not limited to these embodiments.

[0018] Example 1

[0019] Based on the technical problems highlighted in the background, this embodiment proposes a metasurface sensor design method based on weak supervision, such as... Figure 1 As shown, it includes: S1, acquire the attribute dataset of the metasurface sensor and perform data processing. The attribute dataset includes structural parameter vectors and spectral response vectors. S2, input the attribute dataset into the dual-branch encoder network to extract local and global features; S3, input the local features and global features into the global context injection module to enhance the local features, and perform feature concatenation and fusion with the global features to obtain comprehensive features; S4, input the comprehensive features into the progressive resolution decoder, and output the attribute prediction results based on the intermediate prediction results of each resolution decoding stage; S5. Calculate the total task loss based on the attribute prediction results and the attribute dataset, and iteratively optimize and update the attribute prediction results based on the total task loss.

[0020] First, the attribute dataset of the metasurface sensor is obtained. Each sample in the attribute dataset exists as a pair of structural parameter vectors and spectral response vectors. Then, the attribute dataset is processed to expand the sample diversity and eliminate dimensional differences, so that the model can learn features at different scales and alleviate the problem of insufficient training data.

[0021] The model employs a dual-branch encoder network, consisting of a local feature branch and a global feature branch. The local feature branch captures fine structures such as the shape of formant peaks, while the global feature branch captures the overall spectral contour, forming a complementary feature extraction architecture.

[0022] A further design incorporates a global context injection module. First, global features are aggregated to generate a global context vector. Then, an attention mechanism is used to calculate the association weights between local features and the global context vector, injecting global information into the local features. Finally, the enhanced local features are concatenated with the original global features to obtain comprehensive features. This allows the model to dynamically adjust the weights of local features based on the global spectrum, focusing on key regions such as high-quality factor resonance peaks. Furthermore, a progressive resolution decoder is constructed, with low, medium, and high resolution levels for progressively analyzing comprehensive features, ultimately yielding high-precision attribute prediction results.

[0023] Preferably, data processing includes: The spectral response vector is downsampled to obtain spectral data at different resolutions; Standardize the structural parameter vector and the spectral response vector respectively; The standardized spectral response vector and the spectral data are concatenated along the channel dimension to obtain sequence data of the corresponding dimension; Data augmentation strategies are used to augment the structural parameter vector, the spectral response vector, and the sequence data.

[0024] Structure parameter vector It includes K geometric parameters such as period, patch length, height, and asymmetric silver, and a spectral response vector. This represents the optical response at N wavelengths. To alleviate the problem of insufficient samples, the spectral response vector is downsampled to obtain spectral data at various resolutions:

[0025] in, This indicates a downsampling operation, where r is the downsampling ratio. Through multi-resolution representation, the model can learn spectral features at different scales, gradually improving prediction accuracy.

[0026] Data processing also includes a preprocessing stage, which involves standardizing the structural parameter vector and the spectral response vector. , .

[0027] in, The result is the normalized spectral response vector. The result of standardization of the structural parameter vector. , These are the mean and standard deviation of the spectral response vector, respectively. , The mean and standard deviation of the structural parameter vector are given. The preprocessing stage also includes data augmentation, employing strategies such as adding Gaussian noise, spectral shifting, and scaling to the standardized spectral response vector and structural parameter vector to increase sample diversity.

[0028] Furthermore, the standardized spectral response vector and spectral data are concatenated along the channel dimension. Channel-dimensional concatenation refers to the concatenation of downsampled spectral data at different resolutions and the standardized spectral response vector as two independent feature channels, forming a multi-channel one-dimensional sequence data. The number of channels in this sequence data is the sum of the number of channels in each input data set, and its length is consistent with the number of spectral wavelength points. Data augmentation is performed on the sequence data, maintaining consistency in the augmentation process, and it is used as input for subsequent feature extraction branches.

[0029] The final attribute dataset includes standardized and data-augmented structural parameter vectors and spectral response vectors, as well as data-augmented sequence data. Multi-resolution spectral data is obtained by downsampling the spectral response vectors, preserving the global spectral trend while reducing redundant computation. Standardizing the structural parameter vectors and spectral response vectors eliminates prediction bias caused by dimensional differences, improving model training stability. Channel splicing integrates multi-dimensional spectral information, making subsequent feature extraction more comprehensive. Simultaneous data augmentation of the three core data types enriches the diversity of training samples, effectively alleviating model overfitting and enhancing the model's generalization ability to different structures and spectral response scenarios of metasurface sensors, providing high-quality, highly adaptable training data support for subsequent feature extraction and prediction tasks.

[0030] Preferably, the attribute dataset is input into a dual-branch encoder network to extract local and global features, including: By using different receptive field feature extractors, local features of different ranges are extracted from the sequence data; The encoder is used to segment the sequence data, and positional encoding is added to the segmented segments. Based on the multi-head self-attention mechanism of the encoder, the global features corresponding to each positional encoding segment are extracted.

[0031] The processed attribute dataset is input into a dual-branch encoder network. The local feature branch employs different receptive field feature extractors, specifically a multi-layer convolutional neural network structure, focusing on extracting local detail features and high-frequency components in the spectral response. This branch first inputs the sequence data into a one-dimensional convolutional layer to extract spectral-related local features. ; ;

[0032] in, This represents a one-dimensional convolutional layer. express The local feature map output by the first one-dimensional convolutional layer, This represents the local feature map output by the second one-dimensional convolutional layer. This represents the local feature map output by the third one-dimensional convolutional layer. `padding` represents the zero-padding operation during convolution and is one of the core hyperparameters of the one-dimensional convolutional layer. `k` is the kernel size, and `s` is the stride. One-dimensional convolutional layers can accurately capture the local correlation features of adjacent wavelengths in a spectral sequence. By using multi-scale convolutional kernels, they capture local features of different ranges, reflecting different local dependencies, and focusing on capturing the local shape and fine structure of resonance peaks.

[0033] The global feature branch employs a Transformer encoder structure, a commonly used feature encoding structure in existing technologies. It is primarily used for feature extraction and encoding of the input feature sequence, outputting an encoded feature vector. Focusing on modeling the long-range dependencies and global context information of the spectral response, this branch first segments the sequence data into multiple fragments. And perform position encoding : ;

[0034] Indicates the first The input feature sequence is obtained by projecting and positionally encoding spectral segments, where Let be the projection weight matrix. The location is encoded, and then the feature is transformed through multiple layers of Transformer blocks. This represents the global feature sequence output after feature transformation through multiple Transformer Blocks. This represents the core computational unit (Transformer block) of the Transformer encoder. Each Transformer block contains a multi-head self-attention mechanism and a feedforward network.

[0035] in, This represents Scaled Dot-Product Attention, the core computational function of the Transformer's self-attention mechanism. The normalized exponential function (QF) is the core activation function that transforms attention similarity scores into a valid probability distribution. Q, K, and V are the query matrix, key matrix, and value matrix, respectively, derived from the input sequence. composition, This represents the scaling factor, where It is the dimension of the query matrix Q divided by the dimension of the key matrix K (i.e., the feature dimension of each attention head). The global feature branch can effectively model the relationship between different resonance peaks in the spectrum, capturing the overall spectral profile and mode distribution.

[0036] By utilizing different receptive field feature extractors, we can accurately capture local spectral details (such as formant peak values, linewidths, and other high-frequency components) in different ranges of sequence data, avoiding the problem of incomplete local feature extraction. By segmenting the sequence data with an encoder and adding position encoding, combined with a multi-head self-attention mechanism, we can effectively mine the global correlation patterns of different wavelength segments in the spectral sequence, accurately extract global features, and achieve separate and accurate extraction of local details and global patterns. This provides a high-quality feature foundation for subsequent feature fusion and further improves the accuracy of attribute prediction.

[0037] Preferably, the local features and global features are input into the global context injection module to enhance the local features, including: The global features are globally pooled to transform them into intermediate vectors of the corresponding dimensions. The intermediate vector is input into a multi-layer neural network for nonlinear transformation to obtain the global context vector. Calculate the association weights between the global context vector and the local features; Based on the association weights, the global context vector is injected into the local features through element-wise multiplication.

[0038] A global context injection module is introduced into the model to effectively fuse local and global features. This injects global information from global features into local features, enhancing their representational power. Specifically, global pooling is performed on global features to compress high-dimensional global features into low-dimensional ones while retaining core global information.

[0039] in, This is a global pooling operation. Represents global features.

[0040] Intermediate vector of corresponding dimension This is the result obtained after the global pooling operation. The intermediate vector of the corresponding dimension is then input into a multi-layer neural network for non-linear transformation to obtain the global context vector. :

[0041] Among them, the multilayer neural network uses the common multilayer perceptron, namely the MLP in the formula.

[0042] Next, the association weights between local features and the global context vector are calculated using an attention mechanism. :

[0043] in, The normalized exponential function, denoted as , is the core activation function that transforms attention similarity scores into a valid probability distribution. This represents the scaling factor, where It is the dimension of the query matrix Q divided by the key matrix K (i.e., the feature dimension of each attention head).

[0044] Context injection of local features based on association weights:

[0045] This represents the local features after context injection, where ⊙ denotes element-wise multiplication. Finally, the enhanced local features are concatenated and fused with the global features. ;

[0046] express , Features after feature concatenation This indicates that the features are concatenated according to dimension 1, where k=1 represents dimension 1. This refers to the fused comprehensive features, which simultaneously include enhanced local details and global context, enabling precise targeting of key regions such as high-quality factor resonance peaks. The global context injection mechanism allows the model to dynamically adjust the weights of local features based on the global spectral features, focusing on key spectral regions that significantly impact overall performance, such as high-quality factor resonance peaks.

[0047] By performing global pooling on global features and applying nonlinear transformations through multi-layer neural networks, global context information can be accurately extracted, generating global context vectors that adapt to local features. By calculating correlation weights, key matching points between global context and local features can be accurately located. Element-wise multiplication enables the precise injection of global context information into local features, effectively strengthening key local features, suppressing irrelevant and redundant features, and solving the problem of local features deviating from global patterns. This allows enhanced local features to retain their own detailed advantages while conforming to global spectral patterns, improving the quality and adaptability of subsequent comprehensive features, and thus enhancing the accuracy of attribute prediction.

[0048] Preferably, the integrated features are input into the progressive resolution decoder, and attribute prediction results are output based on the intermediate prediction results of each resolution decoding stage, including: The resolution decoding stage includes a first resolution stage, a second resolution stage, and a third resolution stage. The first resolution stage has the lowest resolution, and the third resolution stage has the highest resolution. The first resolution stage analyzes the comprehensive features to generate intermediate prediction results of structural parameters within a first preset accuracy range; The second resolution stage analyzes the comprehensive features and the intermediate prediction results of the structural parameters to generate intermediate prediction results of the spectral response within a second preset accuracy range; The third resolution stage performs forward design tasks and reverse design tasks respectively based on the comprehensive features, the intermediate prediction results of the structural parameters, the intermediate prediction results of the spectral response, and the enhanced local features, and generates structural parameter prediction results and spectral response prediction results within a third preset accuracy range. Among them, the first preset accuracy range has the lowest accuracy requirement, and the third preset accuracy range has the highest accuracy requirement.

[0049] This embodiment employs a progressive resolution enhancement strategy, using a progressive resolution decoder to gradually reconstruct the target output from low to high resolution. This decoder contains multiple resolution levels, each responsible for the prediction task at its corresponding resolution. Specifically, this embodiment includes a first resolution stage, a second resolution stage, and a third resolution stage. The first resolution stage is a low-resolution decoding stage, the second resolution stage is a medium-resolution decoding stage, and the third resolution stage is a high-resolution decoding stage. The low-resolution decoding stage has the lowest resolution, and the high-resolution decoding stage has the highest resolution.

[0050] The low-resolution decoding stage first analyzes and synthesizes features to generate coarse-grained intermediate prediction results of structural parameters. :

[0051] in The Global Pooling Layer is a key operation layer that aggregates global information from the feature map. The result reflects the approximate range of structural parameters, with relatively low precision but providing a basic framework for structural parameters.

[0052] The medium-resolution decoding stage combines comprehensive features and intermediate prediction results of coarse-grained structural parameters for analysis, generating a medium-precision spectral response. :

[0053] in, The repeat operation represents a tensor operation that performs dimensional expansion and element copying on a tensor. This represents the medium-resolution dimension. The accuracy of this result is higher than that of intermediate predictions for coarse-grained structural parameters.

[0054] The high-resolution decoding stage utilizes the outputs from the low-resolution and medium-resolution decoding stages, combining synthesized features and enhanced local features to perform forward and inverse design tasks, respectively. The forward design task predicts the spectral response from the structural parameters, while the inverse design task predicts the structural parameters from the spectral response, ultimately generating high-precision structural parameter prediction results. and spectral response prediction results : ;

[0055] in This is a fully connected layer, its function is to achieve nonlinear mapping and dimensional transformation of the feature space. This is a global pooling layer, whose function is to aggregate global statistical information from local features to generate a global feature vector that represents the whole. These are the enhanced local features, which are spectral detail features. Spectral detail features refer to features that can reflect the fine structure of the local spectral response, such as the peak intensity, linewidth, steepness, and minute fluctuations of resonance peaks.

[0056] The progressive resolution decoder enables the model to progressively optimize structural parameter predictions and spectral response predictions, achieving high-precision predictions with limited training data. Residual connections are used in each decoding stage. and layer normalization Ensure training stability:

[0057] in, This represents the input feature tensor for the decoding stage, which is the original input for the current resolution decoding step. This represents the core sub-layer operations (sub-network units) in the decoding stage, which are the core computational modules for implementing feature transformation. The representation layer normalization layer normalizes the feature tensors, accelerating model training convergence and improving training stability. This represents the prediction result output at the current resolution decoding stage. A progressive resolution decoding design is employed, gradually increasing prediction accuracy from low to high precision, consistent with the logical rules of metasurface sensor property prediction and avoiding error accumulation caused by direct high-precision prediction. The three resolution stages are sequentially linked and progressively advance, utilizing intermediate prediction results from previous stages to assist predictions in subsequent stages. Simultaneously, forward and reverse design tasks are executed in the highest accuracy stage to achieve bidirectional accurate prediction of structural parameters and spectral response, meeting the dual requirements of forward design (knowing the structure to obtain the spectrum) and reverse design (knowing the spectrum to obtain the structure) of metasurface sensors. Furthermore, the third stage incorporates enhanced local features to further improve the prediction reliability of the highest accuracy stage and adapt to design scenarios with different accuracy requirements.

[0058] Preferably, the process of calculating the total task loss based on the attribute prediction results and the attribute dataset, and iteratively optimizing and updating the attribute prediction results based on the total task loss, includes: Calculate the difference loss between the predicted structural parameters and the structural parameter vector to obtain the structural parameter loss result; Calculate the difference loss between the spectral response prediction result and the spectral response vector to obtain the spectral response loss result; Based on the spectral response prediction results, the reverse design task is executed again to generate secondary structure parameter prediction results. The difference loss between the secondary structure parameter prediction results and the structure parameter vector is calculated to obtain the consistency loss results. The total task loss is obtained by weighted summing of the structural parameter loss results, the spectral response loss results, and the consistency loss results. Based on the total task loss and the physical constraint rules corresponding to the attribute dataset, the loss results are corrected, and the structural parameter prediction results and the spectral response prediction results are iteratively optimized and updated.

[0059] A multi-task joint optimization strategy is adopted to simultaneously optimize the reverse design task and the forward design task, and to improve the physical rationality of the model by combining physical constraints.

[0060] The multi-task loss function includes reverse design loss, forward design loss, and consistency loss:

[0061] Among them, through , , For each weight of the loss function, the total task loss is obtained by weighted summation. .

[0062] Reverse design loss To calculate the predicted results of structural parameters Compared with the structure parameter vector before preprocessing The structural parameter loss results obtained from the difference loss:

[0063] in The decoded prediction result, This represents the loss function for mean squared error. Forward design loss. To calculate the spectral response prediction results Compared with the spectral response vector before preprocessing The spectral response loss results obtained from the difference loss are as follows:

[0064] Consistency loss To ensure consistency between the results of the two tasks, specifically, the spectral response prediction results are used as input for the reverse design task. The reverse design task is executed again to obtain the secondary structure parameter prediction results. The difference loss between the secondary structure parameter prediction results and the structure parameter vector before preprocessing is calculated to obtain the consistency loss result.

[0065] in, This indicates that the forward design network predicts structural parameters. The generated spectral response prediction results This indicates the reverse-design network's response to the predicted spectrum. The generated structural parameter prediction results.

[0066] In this embodiment, the physical constraint rules corresponding to the attribute dataset are prior knowledge from electromagnetic simulation. A physical constraint module is designed in the model, which integrates the prior knowledge from electromagnetic simulation into the model training process, including the approximate relationship between resonant wavelength and structural parameters, the dependence of quality factor and asymmetry factor, etc.

[0067] in and These are weighting factors. For the predicted quality factor, The quality factor is calculated based on the physical formula. For the predicted resonance wavelength, For the physically constrained resonant wavelength, the formula is... The L2 norm (mean squared error) loss term represents the difference between the predicted quality factor and the physical prior calculated quality factor. The formula is... This represents the L2 norm (mean squared error) loss term between the predicted resonant wavelength and the physically constrained resonant wavelength. The physical constraints and total loss are backpropagated to optimize each loss term, making the model's attribute predictions more consistent with physical laws and improving practicality and reliability.

[0068] Preferably, after data processing, the method further includes: The initial training of the model is completed based on the processed attribute dataset. The trained model is then used to predict attributes on the unlabeled dataset and generate pseudo-labels for the unlabeled data. Calculate the confidence score of the pseudo-label. If the confidence score is greater than a preset confidence threshold, add the corresponding unlabeled data and pseudo-label to the attribute dataset to complete sample augmentation. Based on the expanded attribute dataset, the steps of model training, unlabeled data attribute prediction, and pseudo-label filtering and addition are repeatedly executed to iteratively optimize the attribute dataset until a preset number of iterations or a preset early stopping condition is reached.

[0069] To address the issue of limited training samples, this embodiment employs a weakly supervised iterative learning mechanism. After processing the attribute dataset, the encoder and decoder architecture of the model is trained using the attribute dataset, which is a labeled dataset. After training, attribute prediction is performed on unlabeled samples in the unlabeled dataset to generate corresponding pseudo-labels. For example, structural parameter vectors, spectral response vectors, etc.

[0070] in, This represents the encoder-decoder neural network model trained using this method. This indicates unlabeled samples. Next, a confidence assessment is performed. The confidence score of the pseudo-labels is evaluated by calculating the entropy of the predicted distribution. A confidence threshold is set to filter out high-quality pseudo-labels. :

[0071] in Calculate the entropy of the predicted distribution. The confidence score (confidence rating) represents the confidence level of the pseudo-labels. When the confidence score is greater than the confidence threshold, the pseudo-labels with high confidence and the unlabeled data are added as a group to the attribute dataset for the next round of model training.

[0072] in, Let represent the training dataset at iteration t. This represents the pseudo-label data generated for the m-th unlabeled sample. This represents the filter mask for the m-th unlabeled sample, where t is the iteration number and m is the number of unlabeled samples. A maximum number of iterations and an early stopping mechanism are set to prevent error accumulation, forming a self-reinforcing learning loop that gradually expands the effective training sample size, continuously improves model performance, and thus improves the prediction accuracy of the structural parameter prediction results and the spectral response prediction results.

[0073] This step, based on the weakly supervised approach, trains the model using an initial labeled attribute dataset. Then, pseudo-labels are generated for the unlabeled dataset and filtered using confidence scores. This enables the effective use of high-quality unlabeled samples, addressing the core pain point of the difficulty and high cost of acquiring labeled datasets for metasurface sensors. Through iterative expansion of samples and repeated model training, the quality and scale of the attribute dataset are continuously optimized, allowing the model to learn more diverse metasurface structures and spectral response features. This gradually improves the model's generalization ability and prediction accuracy. At the same time, by setting the number of iterations and early stopping conditions, resource waste caused by excessive iteration is avoided, balancing optimization effect and training efficiency.

[0074] Preferably, after outputting the attribute prediction results, the method further includes: Obtain the attention weights and association weights corresponding to the branches used for extracting global features; The attention weights and the association weights are mapped using the gradient-weighted class activation mapping method, and the prediction contribution of each spectral region to the attribute prediction result is calculated. Based on the attention weights and their corresponding predicted contributions, and the association weights and their corresponding predicted contributions, corresponding attention heatmaps are generated respectively.

[0075] To enhance model interpretability, an attention mechanism visualization module was designed to reveal the key spectral regions the model focuses on during decision-making. Specifically, it obtains the attention weights and association weights when the global feature branch extracts global features: ;

[0076] in, Represents the original attention weights. The feature fusion module is the core component in this method used to fuse the outputs of local and global feature branches. Attention weights are a learnable weight matrix output by the feature fusion module, used to characterize the importance of different features.

[0077] Feature importance analysis calculates the contribution of different spectral regions to the final attribute prediction results by applying gradient-weighted class activation mapping to attention weights and association weights. :

[0078] in Let be the gradient weights of the k-th feature map in relation to the target output. This is the corresponding feature map.

[0079] The attention weights and their predictive contributions, as well as the association weights and their predictive contributions, are visualized and displayed in an attention heatmap. This visually demonstrates the model's attention to different spectral segments, such as resonance regions and non-resonant regions, helping to understand the model's learning mechanism and decision-making basis.

[0080] This step extracts two types of attention weights and combines them with a gradient-weighted activation mapping method to accurately quantify the contribution of different spectral regions to the attribute prediction results, clarifying the core spectral basis of the model's predictions. It generates corresponding attention heatmaps, transforming the model's "black box" feature extraction and decision-making process into intuitive visualizations. This clearly demonstrates the model's attention to different spectral segments such as resonance peak regions and non-resonant regions, helping researchers accurately understand the model's learning mechanism and decision-making logic, identify the core reasons for prediction biases, further improve the model's interpretability and credibility, and provide clear directional support for subsequent model optimization and metasurface sensor design improvements.

[0081] This embodiment employs a dual-branch encoder network to extract local and global features separately, and combines a global context injection module to achieve feature enhancement and fusion. It is paired with a progressive resolution decoder to gradually improve prediction accuracy. Finally, through iterative optimization of the total task loss, it effectively solves the problems of "separation between local details and global laws" and "insufficient prediction accuracy" in metasurface sensor design, achieving accurate prediction of structural parameters and spectral response, simplifying the design process, improving design efficiency, and relying on weakly supervised core logic to reduce dependence on large-scale labeled datasets, thereby lowering the design cost and threshold of metasurface sensors.

[0082] Example 2

[0083] Based on the principles described in Example 1, a metasurface sensor design device 100 based on weak supervision is proposed, such as... Figure 2 As shown, it includes: The data acquisition and processing module 110 is used to acquire and process the attribute dataset of the metasurface sensor, wherein the attribute dataset includes a structural parameter vector and a spectral response vector. The feature extraction module 120 is used to input the attribute dataset into the dual-branch encoder network and extract local and global features. The feature fusion module 130 is used to input the local features and global features into the global context injection module, enhance the local features, and perform feature concatenation and fusion with the global features to obtain comprehensive features; The attribute prediction module 140 is used to input the comprehensive features into the progressive resolution decoder and output the attribute prediction results based on the intermediate prediction results of each resolution decoding stage. The iterative optimization module 150 is used to calculate the total task loss based on the attribute prediction results and the attribute dataset, and to iteratively optimize and update the attribute prediction results based on the total task loss.

[0084] Example 3

[0085] Furthermore, an electronic device is proposed, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to implement a weakly supervised metasurface sensor design method of Embodiment 1 when executing executable instructions.

[0086] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.

[0087] Furthermore, it should be noted that the use of terms such as "first," "second," and "a" in this invention is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified. The terms "connection," "fixed," etc., should be interpreted broadly. For example, "fixed" can mean a fixed connection, a detachable connection, or an integral part; it can mean a mechanical connection or an electrical connection; it can mean a direct connection or an indirect connection through an intermediate medium; it can mean the internal communication of two elements or the interaction between two elements, unless otherwise explicitly specified. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0088] Furthermore, the technical solutions of the various embodiments of the present invention can be combined with each other, but only if they are feasible for those skilled in the art. If the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0089] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.

Claims

1. A method for designing metasurface sensors based on weak supervision, characterized in that, include: Acquire and process the attribute dataset of the metasurface sensor, wherein the attribute dataset includes structural parameter vectors and spectral response vectors; The attribute dataset is input into a dual-branch encoder network to extract local and global features; The local and global features are input into the global context injection module to enhance the local features, which are then concatenated and fused with the global features to obtain comprehensive features. The integrated features are input into the progressive resolution decoder, and the attribute prediction results are output based on the intermediate prediction results of each resolution decoding stage. The total task loss is calculated based on the attribute prediction results and the attribute dataset, and the attribute prediction results are iteratively optimized and updated based on the total task loss.

2. The method according to claim 1, characterized in that, Data processing includes: The spectral response vector is downsampled to obtain spectral data at different resolutions; Standardize the structural parameter vector and the spectral response vector respectively; The standardized spectral response vector and the spectral data are concatenated along the channel dimension to obtain sequence data of the corresponding dimension; Data augmentation strategies are used to augment the structural parameter vector, the spectral response vector, and the sequence data.

3. The method according to claim 2, characterized in that, The attribute dataset is input into a dual-branch encoder network to extract local and global features, including: By using different receptive field feature extractors, local features of different ranges are extracted from the sequence data; The encoder is used to segment the sequence data, and positional encoding is added to the segmented segments. Based on the multi-head self-attention mechanism of the encoder, the global features corresponding to each positional encoding segment are extracted.

4. The method according to claim 2, characterized in that, The local and global features are input into the global context injection module to enhance the local features, including: The global features are globally pooled to transform them into intermediate vectors of the corresponding dimensions. The intermediate vector is input into a multi-layer neural network for nonlinear transformation to obtain the global context vector. Calculate the association weights between the global context vector and the local features; Based on the association weights, the global context vector is injected into the local features through element-wise multiplication.

5. The method according to claim 2, characterized in that, The integrated features are input into the progressive resolution decoder, and attribute prediction results are output based on the intermediate prediction results of each resolution decoding stage, including: The resolution decoding stage includes a first resolution stage, a second resolution stage, and a third resolution stage. The first resolution stage has the lowest resolution, and the third resolution stage has the highest resolution. The first resolution stage analyzes the comprehensive features to generate intermediate prediction results of structural parameters within a first preset accuracy range; The second resolution stage analyzes the comprehensive features and the intermediate prediction results of the structural parameters to generate intermediate prediction results of the spectral response within a second preset accuracy range; The third resolution stage performs forward design tasks and reverse design tasks respectively based on the comprehensive features, the intermediate prediction results of the structural parameters, the intermediate prediction results of the spectral response, and the enhanced local features, and generates structural parameter prediction results and spectral response prediction results within a third preset accuracy range. Among them, the first preset accuracy range has the lowest accuracy requirement, and the third preset accuracy range has the highest accuracy requirement.

6. The method according to claim 5, characterized in that, Calculate the total task loss based on the attribute prediction results and the attribute dataset, and iteratively optimize and update the attribute prediction results based on the total task loss, including: Calculate the difference loss between the predicted structural parameters and the structural parameter vector to obtain the structural parameter loss result; Calculate the difference loss between the spectral response prediction result and the spectral response vector to obtain the spectral response loss result; Based on the spectral response prediction results, the reverse design task is executed again to generate secondary structure parameter prediction results. The difference loss between the secondary structure parameter prediction results and the structure parameter vector is calculated to obtain the consistency loss results. The total task loss is obtained by weighted summing of the structural parameter loss results, the spectral response loss results, and the consistency loss results. Based on the total task loss and the physical constraint rules corresponding to the attribute dataset, the loss results are corrected, and the structural parameter prediction results and the spectral response prediction results are iteratively optimized and updated.

7. The method according to claim 1, characterized in that, After data processing, it also includes: The initial training of the model is completed based on the processed attribute dataset. The trained model is then used to predict attributes on the unlabeled dataset and generate pseudo-labels for the unlabeled data. Calculate the confidence score of the pseudo-label. If the confidence score is greater than a preset confidence threshold, add the corresponding unlabeled data and pseudo-label to the attribute dataset to complete sample augmentation. Based on the expanded attribute dataset, the steps of model training, unlabeled data attribute prediction, and pseudo-label filtering and addition are repeatedly executed to iteratively optimize the attribute dataset until a preset number of iterations or a preset early stopping condition is reached.

8. The method according to claim 4, characterized in that, The output attribute prediction results also include: Obtain the attention weights and association weights corresponding to the branches used for extracting global features; The attention weights and the association weights are mapped using the gradient-weighted class activation mapping method, and the prediction contribution of each spectral region to the attribute prediction result is calculated. Based on the attention weights and their corresponding predicted contributions, and the association weights and their corresponding predicted contributions, corresponding attention heatmaps are generated respectively.

9. A design device for a metasurface sensor based on weak supervision, characterized in that, include: The data acquisition and processing module is used to acquire and process the attribute dataset of the metasurface sensor, wherein the attribute dataset includes structural parameter vectors and spectral response vectors. The feature extraction module is used to input the attribute dataset into the dual-branch encoder network and extract local and global features; The feature fusion module is used to input the local features and global features into the global context injection module, enhance the local features, and perform feature concatenation and fusion with the global features to obtain comprehensive features; The attribute prediction module is used to input the comprehensive features into the progressive resolution decoder and output the attribute prediction results based on the intermediate prediction results of each resolution decoding stage. The iterative optimization module is used to calculate the total task loss based on the attribute prediction results and the attribute dataset, and to iteratively optimize and update the attribute prediction results based on the total task loss.

10. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the weakly supervised metasurface sensor design method according to any one of claims 1-8 when executing executable instructions.

Citation Information

Patent Citations

  • Resonance metasurface design method and device based on spectrum sensing improvement

    CN120030900A

  • Cross-scale space-time fusion ground feature classification method based on double-branch architecture

    CN121121529A

  • Weak supervision image semantic segmentation system and method based on attention mechanism

    CN121330297A

  • Method for co-design of hardware and neural network architectures using coarse-to-fine search, two-phased block distillation and neural hardware predictor

    US20220147680A1