Micro-nano structure optical filter reverse design method based on deep learning

Through the deep learning micro-nano structure filter reverse design method, the problems of narrow application scope and low design accuracy in the existing technology are solved, and the automated design of multi-category micro-nano structures is realized, and design efficiency and accuracy are improved.

CN120294973APending Publication Date: 2025-07-11HANGZHOU INST FOR ADVANCED STUDY UCAS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510343561.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing reverse design methods for micro-nano structures have a narrow range of applications, making it difficult to choose multiple structural types, and are inexpensive in design accuracy and efficiency.

Method used

The reverse design method of micro-nano structure filters based on deep learning is adopted, and the structure type classification network and structural parameter regression network are combined with data initialization, cosine annealing strategy and multi-stage gradient cutting to realize the automated design of multi-category micro-nano structures.

Benefits of technology

It improves the versatility and scope of application of the design, reduces design time, improves design efficiency and accuracy, and realizes rapid prediction from transmission spectrum to structural parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120294973A_ABST
    Figure CN120294973A_ABST
Patent Text Reader

Abstract

The invention discloses a micro-nano structure optical filter reverse design method based on deep learning. The method comprises the following steps: data preprocessing and data set construction; building a reverse design network model based on deep learning; jointly training the reverse design network to obtain a high-performance model; structural information of the optical filter is derived from the target transmission spectrum, and reverse design of the micro-nano structure optical filter is achieved. According to the micro-nano structure optical filter reverse design method based on deep learning, structure type classification and structure parameter regression are combined, batch processing of multiple micro-nano structures is supported, the limitation that an existing method only aims at a single structure type (such as a metasurface grating or a hole digging structure) is solved, and the method is suitable for large-scale popularization and application. And the universality and the application range of the reverse design are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of micro-nano optical structure design, and particularly relates to a reverse design method of a micro-nano structure filter based on deep learning. Background Art

[0002] Since the 20th century, the expansion of the boundaries of physics and the rapid development of micro-nano processing technology have promoted the research in the field of optics to enter the sub-wavelength range. A large number of micro-nano optical structures, including metasurfaces, metamaterials, Fabry-Perot cavities, quantum dots, nanowires, etc., have been designed and used. These micro-nano optical structures usually have characteristics that exceed natural materials, such as being artificially prepared, having feature sizes much smaller than the working wavelength, showing negative refractive index in the optical field, and having near-zero dielectric constant.

[0003] The design of micro-nano optical structures usually uses methods such as the finite-difference time-domain method or the finite element method. These methods generally require setting parameters such as the structure, materials, light source, and boundary conditions to simulate the optical properties of the micro-nano structure. Such methods not only have high requirements for time and computing power but are also relatively cumbersome for the design of specific requirements with clear goals. Taking the filter structure as an example, it is difficult to achieve precise regulation for specific transmission spectrum requirements and requires repeated modification of the structure to gradually approach. Due to the non-linear relationship between a large number of parameters and the transmission spectrum, the design of high-precision micro-nano structures is very difficult.

[0004] Reverse design thus emerges. Different from the huge computational workload of solving mathematical equations such as Maxwell's equations and partial differential equations in forward design, reverse design usually uses deep learning. After model training in the early stage, the structural parameters can be quickly obtained from the spectral response, improving the design complexity of micro-nano structures and increasing the feasibility of customized requirement design.

[0005] However, the current micro-nano structure reverse design methods still have certain problems. They usually design for a specific structure type, such as a metasurface grating structure or a metasurface hole structure, with a narrow scope of application. For multiple structures, multiple neural networks need to be designed for prediction, and they do not have the ability to select the micro-nano structure type.

[0006] Therefore, how to solve the problem of selecting the micro-nano structure type and provide a reverse design method of a micro-nano structure filter based on deep learning with a stronger scope of application and higher accuracy is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention

[0007] The purpose of the present invention is to provide a reverse design method of a micro-nano structure filter based on deep learning to solve the problem that the selection of micro-nano structure types and high-quality design cannot be achieved in the prior art.

[0008] To achieve the above object of the present invention, the following technical solutions are adopted:

[0009] A reverse design method for micro-nano structure filters based on deep learning, comprising the following steps:

[0010] S1: Data preprocessing and dataset construction:

[0011] Collect the filter structure parameters and transmission spectrum data;

[0012] Represent the filter structure parameters as a vector, with the first bit of the vector being the structure identification bit, encoding the material and structure type, and the remaining bits of the vector being the structure parameter bits, encoding the resonant unit size, coating thickness, and cycle period;

[0013] Convert the transmission spectrum data into a vector, with the elements arranged in ascending order of wavelength band;

[0014] Establish a one-to-one correspondence between each filter structure vector and the corresponding transmission spectrum vector to construct a dataset;

[0015] S2: Build a reverse design network model based on deep learning:

[0016] The reverse design network model is divided into a structure type classification network and a structure parameter regression network;

[0017] The structure type classification network extracts spectral features and outputs classification probabilities;

[0018] The structure parameter regression network weights and adjusts the spectral features according to the classification probabilities and outputs the predicted values of the structure parameters;

[0019] S3: Train the reverse design network model to obtain a high-performance model:

[0020] Strip the structure identification bit of the dataset and convert it into a label type for classification prediction;

[0021] Initialize the dataset to improve the model convergence speed and reduce the influence of parameter value ranges on the weights;

[0022] Divide the initialized dataset into a training set, a test set, and a validation set;

[0023] Use the cross-entropy loss function for the classification task and the mean squared error loss function for the regression task, and weight them to obtain the total loss function;

[0024] Use the AdamW optimizer based on the cosine annealing strategy and combine it with multi-stage gradient clipping to optimize the loss function for training;

[0025] Evaluate the trained model on the validation set, calculate the total loss function of the validation set, select the model with the minimum total loss function as the best model, and save the network model parameters and the initialization parameters of the dataset;

[0026] S4: Input the target transmission spectrum into the inverse design network model to derive the structural information of the filter, and realize the inverse design of the micro-nano structure filter:

[0027] Preprocess the target transmission spectrum, and perform normalization operation on the target transmission spectrum using the initialization parameters of the dataset;

[0028] Input the normalized target transmission spectrum vector into the stored inverse design network model. The category probability output by the structure type classification network is taken as the prediction category by taking the maximum value, and the structure parameter regression network outputs the initialized structure parameters;

[0029] Use the initialization parameters of the dataset to perform an inverse transformation on the output of the regression network to obtain the actual physical quantities of the structure parameters,

[0030] Output the filter structure information.

[0031] While adopting the above technical solutions, the present invention can also adopt or combine the following technical solutions:

[0032] As a preferred technical solution of the present invention: The filter structures included in the dataset include micro-nano grating structures, Fabry-Perot cavity structures, and mosaic coding structures, or only any two of them;

[0033] Step S1 specifically includes the following steps:

[0034] S101, randomly generate filter structure parameters:

[0035] According to the minimum structure size that the processing equipment can process and the limitation of the filter size in the actual use scenario, set the range of the size of each filter structure, randomly generate a set of parameters within this range, and define the specific structure of the filter;

[0036] S102, simulate the transmission spectrum:

[0037] Use the finite-difference time-domain method or the finite element method to perform optical simulation on the generated filter structure and calculate its transmission spectrum; The transmission spectra of different structures are generated within the same wavelength range, sampling points, and intervals to ensure data consistency; The transmission spectrum data is represented in vector form,

[0038] S103, splice the structure identification bit and parameters:

[0039] Assign a unique structure identification bit number to each filter structure to distinguish the structure type;

[0040] Concatenate the structure identification digit with the randomly generated structure parameters and convert them into a vector;

[0041] S104, construct a data set: Combine the vector of each filter structure with the corresponding transmission spectrum vector to form a complete data sample.

[0042] As a preferred technical solution of the present invention: The structure type classification network uses a fully connected feedforward neural network to extract deep features of spectral data, and the features are input into the classification head and output the classification probability of the structure type through Softmax normalization.

[0043] As a preferred technical solution of the present invention: The structure type classification network includes a fully connected layer feedforward neural network and a classification head, which are respectively used to extract deep features of spectral data and predict the structure type corresponding to the input transmission spectrum;

[0044] The fully connected feedforward neural network in the classification network consists of 3 fully connected layers. Each layer performs BatchNorm1d batch normalization, LeakyReLU activation, and Dropout regularization operations on the data. First, expand the input transmission spectrum data to 1024 dimensions, allowing the network to extract more complex features in the high-dimensional space, and then compress it to 512 dimensions to obtain deep features through information refinement;

[0045] The classification head in the classification network contains two fully connected layers, gradually reducing the 512-dimensional deep features to the same number of dimensions as the number of filter structure types. The values of the data output through the softmax operation on each dimension represent the probability that the transmission spectrum is a filter of this type.

[0046] As a preferred technical solution of the present invention: The structure parameter regression network includes an attention mechanism network and regression heads equal to the number of structure types. The deep features and classification probabilities are fused through the attention mechanism network to adjust the feature weights. The fused attention-weighted features enter the corresponding regression heads according to the classification probabilities through the masking mechanism, and the regression heads output the predicted values of the filter structure sizes corresponding to the transmission spectrum of this category.

[0047] As a preferred technical solution of the present invention: The attention mechanism network in the regression network contains two fully connected layers, and LeakyReLU activation and Sigmoid activation are respectively used in the two layers of neurons. First, concatenate the 512-dimensional deep features with the classification probabilities, and achieve feature fusion and secondary refinement through two mappings to convert them into 512-dimensional attention-weighted features; the attention-weighted features enter the regression heads corresponding to their structure types through the masking mechanism;

[0048] The number of regression heads in the regression network is equal to the number of structural type. Each regression head corresponds to only one structural type and contains two fully connected layers. BatchNorm1d batch normalization and LeakyReLU activation are performed on the first layer, and Tanh activation is performed on the second layer. The 512-dimensional fused features input gradually reduce in dimension, and the dimension of the final output layer is the same as the number of physical parameter of the structural size parameters of this category. The output data represents each physical parameter after the initialization of the filter structure of this category.

[0049] As a preferred technical solution of the present invention: Step S3 includes the following stages

[0050] S301, constructing a training set:

[0051] Data processing: Stripping the identification bits in the filter structure vector and converting them into label types for classification prediction; initializing the data set, when the data distribution is uniform, using the maximum-minimum normalization method, and when there are extreme values in the distribution, using robust standardization;

[0052] Data set division: Dividing the data set after data processing into a training set, a test set, and a validation set according to the ratio of 80%, 10%, and 10%;

[0053] S302, constructing a loss function:

[0054] Using cross-entropy loss as the loss function for the classification task, the calculation formula of the cross-entropy loss function is:

[0055]

[0056] Among them, B is the number of samples calculated in the same batch, y i is the true class label of the sample, i is the class index, indicating which class of structure, the class label uses one-hot encoding, with 1 only at the index of the correct class and 0 at other positions; p i is the predicted probability of the model for class i;

[0057] Using the mean square error MSE as the loss function for the regression task, the calculation formula of the mean square error loss function is:

[0058]

[0059] Among them, B is the number of samples calculated in the same batch, K represents the feature dimension of each sample, i is the sample index, indicating the current i-th sample, and j is the feature index, indicating the j-th feature dimension of the current sample, is the predicted value of the j-th dimension of the i-th sample, θ i.j is the true value of the j-th dimension of the i-th sample;

[0060] The total loss function of the inverse design network is obtained by weighting the loss function of the classification task and the loss function of the regression task, and the weight ratio is adjusted according to the training accuracy of the two parts; the calculation formula is:

[0061] L mse = w1L cls + w2L mse

[0062] Where w1 is the weight of the loss function of the classification task, and w2 is the weight of the loss function of the regression task. Considering that in the inverse design process, the output data dimension of the regression task is relatively high and it is relatively more difficult to obtain excellent network model parameters, the default value of w1 is 0.4 and the default value of w2 is 0.6;

[0063] S303, training method:

[0064] Use the AdamW optimizer based on the cosine annealing strategy, with an initial learning rate of 1×10 -4 , a weight decay of 1×10 -5 , a batch size of 64, train the model for 200 epochs, adopt the method of multi-stage gradient clipping, loosely clip in the first 50% of the training epochs, and strictly clip in the last 50% of the training epochs, and save the model parameters with the smallest total loss on the validation set.

[0065] As a preferred technical solution of the present invention: in step S4, the transmission spectrum of the filter to be designed is represented as a one-dimensional vector, and the vector dimension and band information are consistent with the data set. If the number of spectral bands or the band range is inconsistent with the data set, use spline interpolation to fit the data to align with the data set;

[0066] Process the target transmission spectrum using the initialization parameters of the data set to obtain a transmission spectrum vector, input the transmission spectrum vector into the inverse design network to obtain a structure category label and its corresponding initialization structure parameters;

[0067] Perform an inverse transformation on the predicted initialization structure parameters based on the initialization parameters of the data set to restore the actual structure parameters, and then convert the structure category label and structure parameters output by the network into actual physical quantities to finally obtain the filter structure information.

[0068] Compared with the prior art, a method for inverse design of a micro-nano structure filter based on deep learning of the present invention has the following beneficial effects:

[0069] The present invention proposes dual tasks of structural type classification and structural parameter regression, uses a conditional multi-branch regression network, and combines a mask mechanism to realize the automatic operation of the entire process. It uses a unified encoding of structural identifiers and structural parameters to support batch processing of various micro-nano structures such as micro-nano gratings, Fabry-Perot cavities, and mosaic coding, and solves the problem that the existing reverse design methods are usually targeted at a single structural type such as a metasurface grating or a hole-digging structure, have a limited scope of application, and are difficult to fit all transmission spectrum requirements. The joint application of a unified encoding strategy and a conditional multi-branch regression network realizes the fully automatic prediction of multi-category micro-nano structure filters, meets the design requirements of complex transmission spectra, and improves the versatility and scope of application of reverse design.

[0070] The present invention utilizes a deep learning neural network to establish a mapping relationship between transmission spectra and structural parameters; adopts data initialization, cosine annealing strategy and multi-stage gradient clipping to improve model convergence efficiency, reduce training time, introduces an attention mechanism, enhances network feature extraction capabilities, and improves design accuracy. While ensuring design accuracy, the network solves the problems of traditional simulation software design processes requiring manual repeated parameter adjustment, large amount of calculation, and low efficiency, greatly reducing design time, improving design efficiency, and realizing rapid prediction from transmission spectra to structural parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 The overall flow chart of the method of the present invention is as follows;

[0072] Figure 2 A schematic diagram of a method for converting a micro-nano structure into a structural vector encoding method applicable to the present invention;

[0073] Figure 3 This is a schematic diagram of the reverse-engineered neural network structure of the present invention;

[0074] Figure 4 The predicted structure values ​​of the reverse-designed neural network for two transmission spectra are compared with the actual structure values. DETAILED DESCRIPTION

[0075] The present invention will be described in further detail with reference to the accompanying drawings and specific embodiments.

[0076] Example 1

[0077] Combination Figures 1-4 As shown, a reverse design method of a micro-nano structure filter based on deep learning of the present invention comprises the following steps:

[0078] S1: Data preprocessing and dataset construction:

[0079] Collect filter structure parameters and transmission spectrum data;

[0080] Represent the structural parameters of the filter in a vector. The first element of the vector is the structure identification bit, encoding the material and structure type. The remaining bits of the vector are the structural parameter bits, encoding the resonant unit size, coating thickness, and cycle period;

[0081] Convert the transmission spectrum data into a vector, with the elements arranged in ascending order of wavelength bands;

[0082] Establish a one-to-one correspondence between each filter structure vector and the corresponding transmission spectrum vector to construct a dataset;

[0083] S2: Build an inverse design network model based on deep learning:

[0084] The inverse design network model is divided into a structure type classification network and a structural parameter regression network;

[0085] The structure type classification network extracts spectral features and outputs classification probabilities;

[0086] The structural parameter regression network adjusts the spectral features based on the classification probabilities and outputs the predicted values of the structural parameters;

[0087] S3: Train the inverse design network model to obtain a high-performance model:

[0088] Strip the structure identification bit of the dataset and convert it into a label type for classification prediction;

[0089] Initialize the dataset to improve the model convergence speed and reduce the impact of parameter value ranges on the weights;

[0090] Divide the initialized dataset into a training set, a test set, and a validation set;

[0091] Use the cross-entropy loss function for classification tasks and the mean squared error loss function for regression tasks, and obtain the total loss function through weighting;

[0092] Use the AdamW optimizer based on the cosine annealing strategy and combine it with multi-stage gradient clipping to optimize the loss function for training;

[0093] Evaluate the trained model on the validation set, calculate the total loss function of the validation set, select the model with the minimum total loss function as the best model, and save the network model parameters and the initialization parameters of the dataset;

[0094] S4: Input the target transmission spectrum into the inverse design network model to deduce the structural information of the filter and realize the inverse design of the micro-nano structure filter:

[0095] Preprocess the target transmission spectrum and perform normalization operations on the target transmission spectrum using the initialization parameters of the dataset;

[0096] The target transmission spectrum vector after normalization operation is input into the stored reverse design network model, the category probability output by the structure type classification network is taken as the maximum value as the predicted category, and the structure parameter regression network outputs the initialized structure parameters;

[0097] Use the initialization parameters of the data set to inversely transform the output of the regression network to obtain the actual physical quantity of the structural parameters.

[0098] Output filter structure information.

[0099] S1: Data preprocessing and dataset construction:

[0100] Construct a data set, where the data set samples include filter structures and their corresponding transmission spectra;

[0101] The transmission spectrum is represented as a vector. Each element in the vector represents the transmittance of a band. The order of the elements corresponds to the change from small to large spectral bands. The spectral bands are equally spaced, and the vector dimensions of different transmission spectra are the same.

[0102] The filter structure is represented as a vector. The first bit of the vector is the structure identification bit, which is determined by the material and structure type of the filter. The remaining bits are the structure parameter bits, representing specific structural dimensions such as coating thickness and cycle period.

[0103] The filter structure vectors in the data set correspond one-to-one to the transmission spectrum vectors;

[0104] In step S1, the filter structure selected in this example is a micro-nano grating-shaped metasurface structure and a Fabry-Perot cavity structure.

[0105] Furthermore, the substrate of the micro-nano grating-shaped metasurface structure is SiO2, a layer of Si3N4 is coated on the substrate, and a square periodic array composed of Si is made on the Si3N4 layer by a patterning process. The structure is determined by four parameters, specifically the side length of the square Si, the height of the square Si, the period interval of the square Si, and the thickness of the Si3N4 layer. The ranges of these four parameters are 100-200nm, 50-200nm, 300-400nm, and 200-400nm, respectively.

[0106] Furthermore, the substrate of the Fabry-Perot cavity structure is SiO2, on which two high and low refractive index materials TiO2 and SiO2 are alternately evaporated to form a total of ten layers. The structure is represented by the thickness sequence of the ten coating layers, and the thickness of each layer is between 150-300nm.

[0107] Further, for each of the two types of filters, 6000 sets of different parameters are generated within a defined range. The specific structure of the filter is defined, and the finite-difference time-domain method is used to perform optical simulation on the generated filter structure to obtain the transmission spectrum corresponding to the filter structure. The wavelength range of the transmission spectrum is from 400 nm to 700 nm, and a total of 151 sampling points are selected at intervals of 2 nm.

[0108] Further, as Figure 2 shown, the filter structure is encoded. Each micro-nano grating-shaped metasurface structure is represented by a five-dimensional vector. The first digit is the structure identification bit, all represented as 1, and the second to fifth digits are arranged in the order of structure parameters. Each Fabry-Perot cavity structure is represented by an eleven-dimensional vector. The first digit is the structure identification bit, all represented as 2, and the second to eleventh digits are arranged in the order of structure parameters. There are 12000 sets of structure vectors and the corresponding 12000 sets of transmission spectra for the two types of filters, which together form a dataset.

[0109] S2: Build a reverse design network model based on deep learning:

[0110] The reverse design network model includes a structure type classification network and a structure parameter regression network;

[0111] The structure type classification network uses a fully connected feedforward neural network to extract the deep features of the spectral data. This feature is input into the classification head and the classification probability of the structure type is output after Softmax normalization;

[0112] The structure parameter regression network includes an attention mechanism network and regression heads equal in number to the number of structure types. The deep features and classification probabilities are fused through the attention mechanism network to adjust the feature weights. The fused attention-weighted features enter the corresponding regression heads through the masking mechanism according to the classification probabilities, and the regression heads output the predicted values of the structure size parameters;

[0113] In step S2, the reverse design network model based on deep learning is as Figure 3 shown, including two parts: a structure type classification network and a structure parameter regression network.

[0114] Further, the structure type classification network includes a fully connected layer feedforward network and a classification head. It takes the 151-dimensional transmission spectrum as the input and passes through three fully connected layers (151→512→1024→512) in sequence. Each layer contains BatchNorm1d batch normalization, LeakyReLU(0.1) activation, and Dropout(0.2) regularization to extract 512-dimensional deep features. This feature is input into the classification head (512→256→2) and the two-class classification probability is output after Softmax normalization.

[0115] Furthermore, the structural parameter regression network includes an attention mechanism network and regression heads equal in number to the number of structural types. It concatenates the two-class classification probabilities and the 512-dimensional deep features into 514-dimensional joint features, generates a Sigmoid weight vector through a regression attention module (514→512→512), multiplies the weight vector element-wise with the deep features to obtain 512-dimensional attention-weighted features, inputs the attention-weighted features into the corresponding regression heads according to the classification probabilities through a masking mechanism. The attention-weighted features pass through two fully connected layers in sequence in the regression head, perform BatchNorm1d batch normalization and LeakyReLU(0.1) activation on the first layer, perform Tanh activation on the second layer, and finally reduce the dimension to the number of structural size parameters of this type (4 for structural type 1 and 10 for structural type 2), and output the initialized structural parameters of this type of filter structure.

[0116] S3: Train the inverse design network:

[0117] Strip the structure identification bit of the data set and convert it into a label type for classification prediction;

[0118] Initialize the data set to improve the model convergence speed and reduce the influence of parameter value ranges on the weights;

[0119] Divide the initialized data set into a training set, a test set, and a validation set;

[0120] The classification network uses a cross-entropy loss function, the regression network uses a mean squared error loss function, and the loss function of the inverse design network is obtained by weighting the two;

[0121] Use an AdamW optimizer based on the cosine annealing strategy and combine multi-stage gradient clipping to optimize the loss function for training;

[0122] Store the network model parameters with the minimum loss function on the validation set and the data set initialization parameters;

[0123] In step S3, strip the first and last bits of the 12,000 structural vectors of the two types of optical sheets as the label categories, perform max-min normalization on the remaining structural parameter bits and the transmission spectrum vectors, and divide the initialized data set into a training set, a test set, and a validation set according to the ratio of 80%, 10%, and 10%.

[0124] Furthermore, the classification task uses a cross-entropy loss function, the regression task uses a mean squared error MSE loss function, and the total loss function of the inverse design network is obtained by weighting the two, where the weight of the classification task loss function is 0.4 and the weight of the regression task loss function is 0.6.

[0125] Furthermore, use an AdamW optimizer based on the cosine annealing strategy to dynamically adjust the learning rate, and the initial learning rate is 1×10-4 with a weight decay of 1×10 -5 , a batch size of 64, the model is trained for 200 epochs, with a loose clipping of maximum norm 2 in the first 100 training epochs and a strict clipping of maximum norm 1 in the last 100 training epochs.

[0126] Furthermore, during the 200 epochs of training, the minimum loss of the inverse design network on the validation set is 0.0047, and at this time the classification accuracy rate is 99.92%. The model parameters at the epoch with the minimum loss on the validation set and the initialization parameters during the initialization process of this dataset are saved.

[0127] S4: Inverse design of the micro-nano structure filter:

[0128] The transmission spectrum of the required filter is represented as a one-dimensional vector, and the vector dimension and band information are consistent with the dataset. If the number of spectral bands or the band range is inconsistent with the dataset, spline interpolation is used to fit the data to align with the dataset;

[0129] The initialization parameters of the dataset are used to process the target transmission spectrum to obtain a transmission spectrum vector. The transmission spectrum vector is input into the inverse design network to obtain a structure category label and initialization structure parameters;

[0130] Based on the initialization parameters of the dataset, an inverse transformation is performed on the predicted initialization structure parameters to restore the actual structure parameters. Then, the structure category label and structure parameters output by the network are converted into actual physical quantities, and finally the filter structure information is obtained.

[0131] In step S4, a transmission spectrum of a filter of structure type 1 that is not included in the training set is randomly selected. Between 400 nm and 700 nm, a total of 151 sampling points are selected at intervals of 2 nm. After being processed by the saved initialization parameters for maximum-minimum normalization, it is input into the inverse design network to obtain structure type 1 and four initialized structure parameters. After inverse transformation based on the initialization parameters of the dataset and comparison with the actual filter structure size, the average error of the four parameters is 5.58 nm.

[0132] Furthermore, a transmission spectrum of a filter of structure type 2 that is not included in the training set is randomly selected. Between 400 nm and 700 nm, a total of 151 sampling points are selected at intervals of 2 nm. After being processed by the saved initialization parameters for maximum-minimum normalization, it is input into the inverse design network to obtain structure type 2 and ten corresponding initialized structure parameters. After inverse transformation based on the initialization parameters of the dataset and comparison with the actual filter structure size, the average error of the ten parameters is 12.00 nm.

[0133] Such as Figure 4As shown, the comparison between the filter structure sizes predicted by two transmission spectra and the actual sizes.

[0134] Compared with the prior art, a reverse design method for micro-nano structure filters based on deep learning of the present invention has the following beneficial effects:

[0135] (1) The present invention proposes a dual task of structure type classification and structure parameter regression, uses unified encoding of structure identifiers and structure parameters, supports batch processing of various micro-nano structures such as micro-nano gratings, Fabry-Perot cavities, and mosaic coding, solves the problem that existing reverse design methods usually target a single structure type such as metasurface gratings or hole-drilling structures, have limited applicability, and are difficult to fit all transmission spectrum requirements, realizes the full-automatic prediction of multi-category micro-nano structure filters, and improves the generality and applicability of reverse design;

[0136] (2) The present invention uses a deep learning neural network to establish a mapping relationship between the transmission spectrum and the structure parameters; with the efficient feature extraction ability of the neural network, redundant calculations are reduced, and the problems of the traditional simulation software design process that requires manual repeated parameter adjustment, large computational amount, and low efficiency are solved, the design time is greatly reduced, the design efficiency is improved, and the rapid prediction from the transmission spectrum to the structure parameters is realized;

[0137] (3) The present invention optimizes the model training process through data initialization, cosine annealing strategy, and multi-stage gradient clipping, improves the model convergence speed, reduces the training time, and enhances the training speed; by introducing an attention mechanism, the attention to important spectral information is enhanced, and a conditional multi-branch regression network is used, combined with a masking mechanism, to solve the structure parameters of different categories respectively, solves the problem that the reverse design of a single structure type is difficult to meet the requirements of complex transmission spectra, and the switching of multiple models requires manual judgment, increasing the error risk, improves the accuracy of structure type classification and parameter regression, reduces the error of the regression task, and enhances the accuracy of design.

[0138] The above specific embodiments are used to explain the present invention, and are only the preferred embodiments of the present invention, rather than limiting the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.

Claims

1. A reverse design method for micro-nano structure filters based on deep learning, comprising the following steps: S1: Data preprocessing and dataset construction: Collect the structural parameters and transmission spectrum data of the filter; Represent the filter structural parameters as a vector, with the first bit of the vector being the structure identification bit, encoding the material and structure type, and the remaining bits of the vector being the structural parameter bits, encoding the resonant unit size, coating thickness, and cycle period; Convert the transmission spectrum data into a vector, with the elements arranged in ascending order of wavelength band; Establish a one-to-one correspondence between each filter structure vector and the corresponding transmission spectrum vector to construct a dataset; S2: Build a reverse design network model based on deep learning: The reverse design network model is divided into a structure type classification network and a structural parameter regression network; The structure type classification network extracts spectral features and outputs classification probabilities; The structural parameter regression network weights and adjusts the spectral features according to the classification probabilities and outputs the predicted values of the structural parameters; S3: Train the reverse design network model to obtain a high-performance model: Strip the structure identification bit of the dataset and convert it into a label type for classification prediction; Initialize the dataset to improve the model convergence speed and reduce the influence of parameter value ranges on the weights; Divide the initialized dataset into a training set, a test set, and a validation set; Use the cross-entropy loss function for the classification task and the mean squared error loss function for the regression task, and weight them to obtain the total loss function; Use the AdamW optimizer based on the cosine annealing strategy and combine multi-stage gradient clipping to optimize the loss function for training; Evaluate the trained model on the validation set, calculate the total loss function of the validation set, select the model with the smallest total loss function as the best model, and save the network model parameters and the initialization parameters of the dataset; S4: Input the target transmission spectrum into the reverse design network model to deduce the structural information of the filter, realizing the reverse design of the micro-nano structure filter: Preprocess the target transmission spectrum and perform a normalization operation on the target transmission spectrum using the initialization parameters of the dataset; Input the normalized target transmission spectrum vector into the stored reverse design network model. The category probability output by the structure type classification network is taken as the maximum value as the predicted category, and the structural parameter regression network outputs the initialized structural parameters; Perform an inverse transformation on the output of the regression network using the initialization parameters of the dataset to obtain the actual physical quantities of the structural parameters; Output the filter structure information.

2. The reverse design method of the micro-nano structure optical filter based on deep learning according to claim 1, characterized in that: The filter structures included in the dataset include micro-nano grating structures, Fabry-Perot cavity structures, and mosaic coding structures, or only any two of them; Step S1 specifically includes the following steps: S101, randomly generate filter structural parameters: According to the minimum structure size that the processing equipment can process and the limitation of the filter size in the actual use scenario, set the range of each filter structure size, randomly generate a set of parameters within this range, and define the specific structure of the filter; S102, simulate the transmission spectrum: Optical simulation is performed on the generated filter structure using the finite-difference time-domain method or the finite element method to calculate its transmission spectrum; the transmission spectra of different structures are generated under the same wavelength range, sampling points, and intervals to ensure data consistency; the transmission spectrum data is represented in vector form, S103, Structure identification bit and parameter splicing: Assign a unique structure identification bit number to each filter structure to distinguish the structure type; Splice the structure identification bit number with the randomly generated structure parameters and convert them into a vector; S104, Construct a dataset: Combine the vector of each filter structure with the corresponding transmission spectrum vector to form a complete data sample.

3. The reverse design method of the micro-nano structure filter based on deep learning according to claim 1, characterized in that: The structure type classification network uses a fully connected feedforward neural network to extract the deep features of the spectral data, and the features are input into the classification head and output the classification probability of the structure type through the Softmax normalization operation.

4. The reverse design method of the micro-nano structure filter based on deep learning according to claim 3, characterized in that: The structure type classification network includes a fully connected feedforward neural network and a classification head, which are used to extract the deep features of the spectral data and predict the structure type corresponding to the input transmission spectrum respectively; The fully connected feedforward neural network in the classification network consists of 3 fully connected layers. Each layer performs BatchNorm1d batch normalization, LeakyReLU non-linear activation, and Dropout regularization operations on the data. First, the input transmission spectrum data is expanded to 1024 dimensions to allow the network to extract more complex features in the high-dimensional space, and then compressed to 512 dimensions to obtain deep features through information refinement; The classification head in the classification network contains two fully connected layers, which gradually reduce the 512-dimensional deep features to the same number of dimensions as the number of filter structure types. The values of the data output through the softmax normalization operation on each dimension represent the probability that the transmission spectrum is a filter of this type.

5. The reverse design method of the micro-nano structure filter based on deep learning according to claim 1, characterized in that: The structure parameter regression network includes an attention mechanism network and regression heads equal in number to the number of structure types. The deep features and classification probabilities are fused through the attention mechanism network to adjust the feature weights. The fused attention-weighted features enter the corresponding regression heads according to the classification probabilities through the masking mechanism, and the regression heads output the predicted values of the filter structure dimensions corresponding to the transmission spectrum.

6. The reverse design method of the micro-nano structure filter based on deep learning according to claim 5, characterized in that: The attention mechanism network in the regression network contains two fully connected layers, and LeakyReLU activation and Sigmoid activation are used in the two layers of neurons respectively. First, the 512-dimensional deep features are spliced with the classification probabilities, and feature fusion and secondary refinement are achieved through two mappings, so that they are transformed into 512-dimensional attention-weighted features; The attention-weighted features enter the regression head corresponding to its structure type through the masking mechanism; The number of regression heads in the regression network is equal to the number of structure types. Each regression head corresponds to only one structure type. It contains 2 fully connected layers, BatchNorm1d batch normalization and LeakyReLU activation are performed on the first layer, and Tanh activation is performed on the second layer. The input 512-dimensional fusion features are gradually reduced in dimension, and the dimension of the final output layer is the same as the number of physical parameters of this type of structure size parameter. The output data represents the initial physical parameters of each filter structure of this type.

7. The reverse design method of the micro-nano structure filter based on deep learning according to claim 1, wherein: Step S3 includes the following stages, S301, Construct a training set: Data processing: Strip the identification bits in the filter structure vector and convert them into label types for classification prediction; Initialize the data set. When the data is evenly distributed, use the maximum-minimum normalization method. When there are extreme values in the distribution, use robust standardization; Data set division: Divide the initialized data set into a training set, a test set, and a validation set according to the ratio of 80%, 10%, and 10%; S302, Construct a loss function: Use cross-entropy loss as the loss function for the classification task. The calculation formula of the cross-entropy loss function is: Among them, B is the number of samples calculated in the same batch, y i is the true class label of the sample, i is the class index, indicating which type of structure, and the class label uses one-hot encoding, where only the index of the correct class is 1 and the other positions are 0; p i is the predicted probability of the model for class i; Use mean squared error MSE as the loss function for the regression task. The calculation formula of the mean squared error loss function is: Among them, B is the number of samples calculated in the same batch, K represents the feature dimension of each sample, i is the sample index, indicating the current i-th sample, and j is the feature index, indicating the j-th feature dimension of the current sample. is the predicted value of the j-th dimension of the i-th sample, θ i.j is the true value of the j-th dimension of the i-th sample; The total loss function of the inverse design network is obtained by weighting the loss function of the classification task and the loss function of the regression task. The weight ratio is adjusted according to the training accuracy of the two parts; the calculation formula is: L mse = w1L cls + w2L mse Where w1 is the weight of the loss function of the classification task, and w2 is the weight of the loss function of the regression task. Considering that in the inverse design process, the output data dimension of the regression task is relatively high and it is relatively more difficult to obtain excellent network model parameters, the default value of w1 is 0.4 and the default value of w2 is 0.6; S303, Training method: The learning rate is dynamically adjusted using a weight decay adaptive moment estimation optimizer based on the cosine annealing strategy, with an initial learning rate of 1×10 -4 , a weight decay of 1×10 -5 , a batch size of 64, the model is trained for 200 epochs, and a multi-stage gradient clipping method is adopted. Loose clipping is used for the first 50% of the training epochs, and strict clipping is used for the last 50% of the training epochs. The model parameters with the minimum total loss on the validation set are saved.

8. The inverse design method of the micro-nano structure filter based on deep learning according to claim 1, characterized in that: In step S4, represent the transmission spectrum of the filter to be designed as a one-dimensional vector. The vector dimension and band information are the same as those of the data set. If the number of spectral bands or the band range is inconsistent with the data set, use spline interpolation to fit the data to align with the data set; Process the target transmission spectrum using the initialization parameters of the data set to obtain a transmission spectrum vector. The transmission spectrum vector is input into the inverse design network to obtain the structure category label and its corresponding initialization structure parameters; Perform an inverse transformation on the predicted initialization structure parameters based on the initialization parameters of the data set to restore the actual structure parameters, and then convert the structure category label and structure parameters output by the network into actual physical quantities to finally obtain the filter structure information.

Citation Information

Cited By

  • Intelligent reverse design method and system for vacuum coating

    CN121457143A