Metasurface structure prediction and model training method, system, equipment and medium
Through layer-by-layer greed training and small sample learning to optimize the deep belief network, the problem of insufficient generalization ability of the metasurface structure prediction model is solved, and high-precision metasurface structure parameter prediction is achieved, which improves the reliability and efficiency of the design.
Patent Information
- Application Number
- CN202510438385.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the generalization ability of the metasurface structure prediction model is poor, the training data acquisition cost is high, the samples are limited, and the data annotation requires professional knowledge, which leads to the difficulty of complex metasurface modeling process.
The layer by layer greed training method is used to perform unsupervised pre-training on each layer of the constrained Boltzmann machine of the deep belief network. Combined with the few-sample learning method and the meta-learning method, the metasurface structure prediction model is optimized, and the model parameters are updated by calculating the difference between the predicted structural parameters and the real structural parameters.
The feature extraction and generalization capabilities of the metasurface structure prediction model are improved, the prediction accuracy is improved, the deviation between theoretical calculation results and actual measurement results is reduced, and the reliability and engineering feasibility of metasurface material design are enhanced.
Smart Images

Figure CN120337760A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of electromagnetic metamaterials, and particularly relates to a method, system, device and medium for predicting a metasurface structure and training a model. Background Art
[0002] A metasurface is an artificial electromagnetic material with a sub-wavelength periodic structure, and has electromagnetic properties that do not exist in natural materials, such as negative permittivity, negative permeability, negative refractive index, and inverse Doppler effect. In recent years, metamaterials have been widely used in the fields of stealth technology, communication, acoustic devices, medical examinations, etc. in the microwave, terahertz, and optical frequency bands. Their unique physical properties and flexible design make them have great application potential. At present, the traditional design of metasurfaces relies on numerical simulations such as the finite element method or the finite-difference time-domain method, which are computationally complex and time-consuming, and require a large amount of human participation. The application of artificial neural networks provides an efficient prediction method for metasurface design, avoiding repeated modeling and optimization based on experience. However, deep learning still faces challenges in metasurface design. The main problems are: high cost of obtaining training data, limited samples, limited generalization ability, and the need for professional knowledge for data annotation, which makes the modeling process of complex metasurfaces more difficult. Therefore, it is necessary to provide a method, system, device and medium for predicting a metasurface structure and training a model. Summary of the Invention
[0003] In view of the above-mentioned disadvantages of the prior art, the purpose of the present application is to provide a method, system, device and medium for predicting a metasurface structure and training a model, which improves the problem of poor generalization ability of the existing metasurface structure prediction model.
[0004] To achieve the above-mentioned purpose and other related purposes, the present application provides a method for training a metasurface structure prediction model, including: obtaining an electromagnetic property parameter set of a metasurface material and a corresponding metasurface structure parameter set; inputting the electromagnetic property parameter set into a deep belief network, and using the layer-by-layer greedy training method to pre-train the restricted Boltzmann machine of each layer of the deep belief network in turn. After all layers are trained, a pre-trained deep belief network is obtained; wherein, the deep belief network includes multiple cascaded restricted Boltzmann machines; inputting the electromagnetic property parameter set into the metasurface structure prediction model, and correspondingly obtaining a predicted value of the metasurface structure parameter; wherein, the metasurface structure prediction model includes a pre-trained deep belief network and a prediction layer cascaded therewith; calculating the difference degree between the predicted value of the metasurface structure parameter and the corresponding metasurface structure parameter, and updating the parameters of the metasurface structure prediction model based on the difference degree to obtain a trained metasurface structure prediction model.
[0005] In an embodiment of the present application, for the restricted Boltzmann machine in the first layer of the deep belief network, its training process includes: inputting the electromagnetic performance parameter set into the restricted Boltzmann machine for feature extraction to obtain the metasurface deep feature set; performing a reconstruction process on the metasurface deep feature set to obtain the metasurface reconstruction feature set; calculating the contrast difference degree between the electromagnetic performance parameter set and the metasurface reconstruction feature set based on the contrast divergence algorithm, and updating the restricted Boltzmann machine based on the contrast difference degree to obtain a trained restricted Boltzmann machine, and using the metasurface deep feature set generated by the restricted Boltzmann machine in the last time during the training phase as the metasurface feature set input to the restricted Boltzmann machine in the next layer.
[0006] In an embodiment of the present application, for the restricted Boltzmann machine in each of the remaining layers of the deep belief network, its training process includes: inputting the metasurface feature set generated by the restricted Boltzmann machine in the previous layer into the restricted Boltzmann machine in the current layer for feature extraction to obtain the metasurface deep feature set generated by the restricted Boltzmann machine in the current layer; performing a reconstruction process on the metasurface deep feature set to obtain the metasurface reconstruction feature set; calculating the contrast difference degree between the metasurface deep feature set and the metasurface reconstruction feature set based on the contrast divergence algorithm, and updating the restricted Boltzmann machine in the current layer based on the contrast difference degree to obtain a trained restricted Boltzmann machine in the current layer.
[0007] In an embodiment of the present application, after obtaining the trained metasurface structure prediction model, it further includes: optimizing the obtained metasurface structure prediction model based on the few-shot learning method to obtain the finally trained metasurface structure prediction model.
[0008] In an embodiment of the present application, the few-shot learning is the MAML meta-learning method. Optimizing the obtained metasurface structure prediction model based on the MAML meta-learning method includes: dividing all the electromagnetic performance parameter sets and their corresponding metasurface structure parameter sets used for training the metasurface structure prediction model according to preset different prediction targets to form a plurality of electromagnetic performance parameter combinations and their corresponding metasurface structure parameter combinations; wherein, each prediction target corresponds to a mapping relationship between a set of electromagnetic performance parameters and metasurface structure parameters; performing a preset number of retrainings on the metasurface structure prediction model for each prediction target. For each retraining for the prediction target: inputting the electromagnetic performance parameter combination corresponding to the prediction target into the initially trained metasurface structure prediction model to obtain the predicted value of the metasurface structure parameters; calculating the difference degree between the predicted value of the metasurface structure parameters and the corresponding metasurface structure parameters, and updating the parameters of the metasurface structure prediction model based on the difference degree; calculating the total loss based on the difference degrees of the last retraining for each prediction target, and updating the initially trained metasurface structure prediction model based on the total loss to obtain the finally trained metasurface structure prediction model.
[0009] In an embodiment of the present application, for each retraining of the prediction target, the parameters of the metasurface structure prediction model are updated by dynamically adjusting the learning rate based on the degree of difference.
[0010] In an embodiment of the present application, a metasurface structure prediction method is further provided. The prediction method includes: obtaining the electromagnetic performance parameters of the metasurface material to be predicted; inputting the electromagnetic performance parameters into the metasurface structure prediction model to generate the metasurface structure parameters of the metasurface material; wherein, the metasurface structure prediction model is a model trained by the metasurface structure prediction model method of any one of the above.
[0011] In an embodiment of the present application, a metasurface structure prediction model system is further provided. The system includes: a data acquisition module, configured to acquire a set of electromagnetic performance parameters of the metasurface material and a corresponding set of metasurface structure parameters; an initial training module, configured to input the set of electromagnetic performance parameters into a deep belief network, and adopt a layer-by-layer greedy training method to pre-train the restricted Boltzmann machine of each layer of the deep belief network in turn. After all layers are trained, a pre-trained deep belief network is obtained; wherein, the deep belief network includes multiple cascaded restricted Boltzmann machines; a parameter prediction module, configured to input the set of electromagnetic performance parameters into the metasurface structure prediction model to correspondingly obtain the predicted values of the metasurface structure parameters; wherein, the metasurface structure prediction model includes the pre-trained deep belief network and a prediction layer cascaded therewith; a parameter update module, configured to calculate the degree of difference between the predicted values of the metasurface structure parameters and the corresponding metasurface structure parameters, and update the parameters of the metasurface structure prediction model based on the degree of difference to obtain a trained metasurface structure prediction model.
[0012] In an embodiment of the present application, an electronic device is further provided, including: one or more processors; a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the training method of the metasurface structure prediction model of any one of the above.
[0013] In an embodiment of the present application, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by the processor of the computer, the computer executes the training method of the metasurface structure prediction model of any one of the above.
[0014] As mentioned above, a method, system, device and medium for training a super surface structure prediction and model of the present application have the following beneficial effects: a layer-by-layer greedy training method is used to independently and unsupervisedly pre-train each layer of the restricted Boltzmann machine of the deep belief network, so that the model gradually learns the deep characteristic relationship between electromagnetic properties and structural parameters, thereby improving the model's feature extraction ability and generalization ability. On this basis, a super surface structure prediction model is constructed according to the pre-trained deep belief network and its cascaded prediction layer, and the model is trained and optimized so that it can accurately predict the super surface structure parameters. The prediction accuracy is improved by calculating the difference between the predicted structural parameters and the actual structural parameters and continuously optimizing the model parameters. In the end, the super surface structure parameters can be accurately predicted. The present application effectively improves the problem of insufficient accuracy of super surface structure prediction in the prior art, and improves the reliability and engineering feasibility of super surface material design. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A schematic flow chart of a super surface structure prediction model method provided in an embodiment of the present application;
[0016] Figure 2 A schematic diagram of the structure of a deep belief network provided in an embodiment of the present application;
[0017] Figure 3 A schematic flow chart of a method for predicting a supersurface structure provided in an embodiment of the present application;
[0018] Figure 4 Shown is a structural block diagram of a super surface structure prediction model system provided by an embodiment of the present application;
[0019] Figure 5 Shown is a structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0020] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0021] It should be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present application. Therefore, only the components related to the present application are shown in the illustrations, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0022] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.
[0023] Currently, traditional electromagnetic metasurface design relies on complex and time-consuming numerical simulations, such as the finite element method or the finite difference time domain method, which require a large amount of human participation and supervision. The development of artificial neural networks has provided a simple and efficient solution for the design of metamaterials. Without the need for experience-based predictions and iterative modeling optimization design methods, it is possible to quickly predict the complex optical properties of metasurfaces. However, applying deep learning to the design of metasurfaces has some limitations and challenges. The main problem is that metasurface design requires a large amount of electromagnetic property data, and obtaining this data is costly and time-consuming. Therefore, the number of available training samples is limited, which restricts the performance and generalization ability of the deep learning model. In addition, annotating data in deep learning requires professional knowledge, including information such as electromagnetic properties, material attributes, and geometric structures. For more complex metasurface properties, the annotation process may be very complex. And because the data acquisition of electromagnetic metasurfaces is very limited and time-consuming, the final generalization performance of the model is poor.
[0024] To address the above problems, the present application provides a method for a metasurface structure prediction model. The method uses a layer-by-layer greedy training method to independently perform unsupervised pre-training on each layer of restricted Boltzmann machines of a deep belief network, so that the model gradually learns the deep feature relationship between electromagnetic performance and structural parameters, thereby enhancing the model's feature extraction ability and generalization ability. On this basis, a metasurface structure prediction model is constructed according to the pre-trained deep belief network and its cascaded prediction layer, and the model is trained and optimized to enable it to accurately predict the metasurface structure parameters. By calculating the difference between the predicted structure parameters and the true structure parameters and continuously optimizing the model parameters, the prediction accuracy is improved. Finally, the metasurface structure parameters can be accurately predicted. The present application effectively improves the problem of insufficient accuracy in the prediction of metasurface structures in the prior art. Using the method of the present application reduces the deviation between the theoretical calculation results and the actual measurement results, thereby enhancing the reliability and engineering feasibility of metasurface material design.
[0025] Please refer to Figure 1 , the hyper-surface structure prediction model method proposed in this application includes the following steps:
[0026] S11. Obtain the electromagnetic performance parameter set of the hyper-surface material and the corresponding hyper-surface structure parameter set.
[0027] The hyper-surface structure parameters are used to describe the geometric shape, material properties, periodic arrangement and other characteristics of the hyper-surface, and are the key factors determining the electromagnetic response of the hyper-surface. The structure parameters include but are not limited to the thickness of each layer of the hyper-surface, the length and width of the hyper-surface unit, material properties, the periodic arrangement method of the unit, etc. The electromagnetic performance parameters are used to describe the action mode of the hyper-surface on the incident electromagnetic wave, and determine its transmission, reflection, absorption and phase modulation capabilities. The performance parameters include but are not limited to phase, amplitude, transmittance, reflectance, absorptance and phase change, etc. These hyper-surface structure parameters constitute the training set for training the deep belief network and the hyper-surface structure prediction model. Similarly, the obtained electromagnetic performance parameters constitute the label set for training the deep belief network and the hyper-surface structure prediction model. It should be noted that the hyper-surface structure parameter set in this application refers to a part of the data selected from the training set for training the deep belief network and the hyper-surface structure prediction model of the current batch. It can be understood that different hyper-surface structure parameters correspond to different electromagnetic performance parameters. Exemplarily, the thickness of the dielectric layer can correspond to electromagnetic performance parameters such as transmittance and absorptance, because when the thickness of the hyper-surface dielectric layer increases, its working frequency point will shift to a certain extent, and the transmittance and absorptance will also change.
[0028] In order to improve the stability of training and enhance the data quality, after obtaining the hyper-surface structure parameters and electromagnetic performance parameters, this application will also preprocess the obtained hyper-surface structure parameters and electromagnetic performance parameters. Among them, the preprocessing methods include but are not limited to normalization processing and denoising processing. Through denoising processing, data with measurement errors greater than the preset error tolerance threshold can be removed. Through normalization processing, data of different physical quantities can be adjusted to the same numerical range. Optionally, the normalization processing can adopt the following process: Based on the mean μ k and standard deviation σ k of the original parameters, parameter conversion is performed so that the converted parameters follow the standard normal distribution. Specifically, the original parameter x k is converted into the standardized form shown in formula (1):
[0029]
[0030] where is the k-th normalized parameter, x k is the k-th original parameter, μ k is the mean of the k-th original parameter, and σk is the standard deviation of the k-th original parameter, and the transformed parameter follows a standard normal distribution with a mean of 0 and a standard deviation of 1, so as to make the data distributions of different physical quantities consistent.
[0031] S12. Input the electromagnetic performance parameter set into the deep belief network, and adopt the layer-by-layer greedy training method to pre-train the restricted Boltzmann machine of each layer of the deep belief network in turn. After all layers are trained, a pre-trained deep belief network is obtained; among them, the deep belief network includes multiple cascaded restricted Boltzmann machines.
[0032] According to the complexity of the problem, to construct a deep belief network model, it is necessary to determine the number of layers of the network, the number of neurons in each layer, and other hyperparameters, such as the learning rate, the number of iterations, etc. The Deep Belief Network (DBN) consists of multiple cascaded Restricted Boltzmann Machines (RBMs). Each layer of the restricted Boltzmann machine includes a visible layer and a hidden layer cascaded with it. The visible layer is used to receive input data, and the hidden layer is used to learn and extract the deep features of the input data. During training, the restricted Boltzmann machines of each layer are pre-trained in sequence according to the layer order. When training each layer, the output of the current layer's restricted Boltzmann machine is used as the input of the next layer's restricted Boltzmann machine, so that the model can gradually extract the high-level features of the data. After all layers of the deep belief network are trained, a pre-trained deep belief network can be obtained, thus providing a robust initial weight for subsequent supervised fine-tuning and parameter optimization to improve the learning efficiency and generalization ability of the subsequent metasurface structure prediction model.
[0033] As Figure 2 shown, each layer of the restricted Boltzmann machine consists of a visible layer and a hidden layer. The two layers are fully connected through the weight matrix W, and the units within the same layer have no direct connection. During the processing of the restricted Boltzmann machine, the visible layer units (v1, v2... v m ) are used to receive input data, such as electromagnetic performance parameters. The hidden layer units (h1, h2... h m ) are used to extract deep features. The weight matrix W i j is used to connect the visible units and the hidden units, determining the way of data transfer between layers.
[0034] Specifically, for the restricted Boltzmann machine of the first layer of the deep belief network, its training process includes the following steps:
[0035] First, input the electromagnetic performance parameter set into the restricted Boltzmann machine for feature extraction to obtain the deep metasurface feature set.
[0036] Input the electromagnetic performance parameter set into the first-layer restricted Boltzmann machine. For each electromagnetic performance parameter input into this layer of the restricted Boltzmann machine, use the electromagnetic performance parameters (such as transmittance, reflectance, absorption rate, etc.) as the input vector v of the visible layer input into the restricted Boltzmann machine. Through the weight matrix W and bias term b of the hidden layer, calculate the activation probability of the hidden layer according to formula (2):
[0037] p(h|v) = σ(Wh + b) (2)
[0038] Among them, p(h|v) represents the probability that the neuron vector h of the hidden layer is activated when the visible layer input v is given. v is the input vector of the visible layer, σ(*) is the Sigmoid activation function, b is the bias term of the hidden layer, and W is the weight matrix of the hidden layer. The activation probability p(h|v) reflects the probability that the neurons in the hidden layer are activated. When the activation probability of the hidden layer is relatively high, it indicates that the neuron has a strong response under the current input data situation, thus forming corresponding metasurface deep features in the hidden layer. After these deep features are sampled by probability, they are used as the input of the next layer of the restricted Boltzmann machine to further extract higher-level feature information. For a single training, a corresponding metasurface deep feature set is obtained for the input electromagnetic performance parameter set.
[0039] Secondly, perform reconstruction processing on the metasurface deep feature set to obtain the metasurface reconstruction feature set.
[0040] Using the metasurface deep feature set as the input, through the reconstruction mechanism, calculate the metasurface reconstruction features of each metasurface deep feature from the hidden layer to the visible layer, as shown in formula (3):
[0041] p(v'|h) = σ(W T h + a) (3)
[0042] Among them, v' is the reconstructed visible layer feature, that is, the metasurface reconstruction feature. a is the bias term of the visible layer. p(v'|h) represents the probability distribution of the neuron vector v' of the visible layer when the neuron vector h of the hidden layer is given. Select the neuron vector of the visible layer with a higher probability distribution as the generated metasurface reconstruction feature. For a single training, a corresponding metasurface reconstruction feature set is obtained for the input metasurface deep feature set.
[0043] Finally, calculate the contrast difference degree between the electromagnetic performance parameter set and the metasurface reconstruction feature set based on the contrast divergence algorithm, update the restricted Boltzmann machine based on the contrast difference degree to obtain the trained restricted Boltzmann machine, and use the metasurface deep feature set generated by the restricted Boltzmann machine at the last time in the training stage as the metasurface feature set input into the next layer of the restricted Boltzmann machine.
[0044] For the input electromagnetic performance parameter set and the reconstructed metasurface reconstruction feature set, use the formula (4) to calculate the comparison difference degree Loss between the two:
[0045] Loss = ∑(v - v') 2 (4)
[0046] Among them, v is the electromagnetic performance parameter input in the current training batch, which is used as the input of the visible layer neurons during the training of the restricted Boltzmann machine. v' is the corresponding reconstructed metasurface reconstruction feature, that is, the estimated value of the electromagnetic performance parameter reconstructed in reverse after learning the deep features through the hidden layer. Calculate the total error of all samples in the current training batch, and optimize the weight matrix and bias term of the restricted Boltzmann machine in the current layer through gradient descent, so that the metasurface reconstruction feature v' is closer to the original input v, thereby improving the learning ability and generalization performance of the restricted Boltzmann machine in this layer. Iteratively train the restricted Boltzmann machine in this layer several times according to the above process until the comparison difference degree converges, indicating that the training of the restricted Boltzmann machine in the current layer is completed. At this time, take the output h of the last hidden layer of the restricted Boltzmann machine during the training stage as the metasurface feature set of the next layer of the restricted Boltzmann machine, which is used to further extract higher-level deep features to gradually optimize the mapping relationship between the metasurface structure parameters and the electromagnetic performance.
[0047] For each remaining restricted Boltzmann machine layer of the deep belief network, its training process includes the following steps:
[0048] First, input the metasurface feature set generated by the previous layer of the restricted Boltzmann machine into the restricted Boltzmann machine of the current layer for feature extraction, and obtain the metasurface deep feature set generated by the restricted Boltzmann machine of the current layer.
[0049] In this application, the restricted Boltzmann machine of the current layer does not directly receive the original electromagnetic performance parameters, but uses the deep features extracted by the previous layer as the input, enabling the deep belief network to learn deeper representation information layer by layer. Specifically, take the output h of the last hidden layer of the restricted Boltzmann machine trained in the previous layer as the input data and input it into the visible layer of the restricted Boltzmann machine of the current layer. Use the above formula (2) to calculate the activation probability of the hidden layer neurons of the restricted Boltzmann machine of this layer, and thus obtain the metasurface deep features.
[0050] Then, perform reconstruction processing on the metasurface deep feature set to obtain the metasurface reconstruction feature set. The reconstruction processing method is the same as the processing method of the first-layer metasurface deep features above, and will not be elaborated here.
[0051] Finally, the contrast divergence algorithm is used to calculate the contrast difference between the deep feature set of the metasurface and the reconstructed feature set of the metasurface, and the restricted Boltzmann machine of the current layer is updated based on the contrast difference to obtain the trained restricted Boltzmann machine of the current layer. The processing method is the same as that of the first-layer restricted Boltzmann machine mentioned above and will not be elaborated here.
[0052] S13. Input the electromagnetic performance parameter set into the metasurface structure prediction model to correspondingly obtain the predicted values of the metasurface structure parameters. Among them, the metasurface structure prediction model includes a pre-trained deep belief network and a prediction layer cascaded with it.
[0053] After completing the layer-by-layer greedy pre-training, the parameters of the deep belief network have been well initialized. To further improve the prediction ability of the model, the fine-tuning stage is entered next. In this stage, the model is further optimized through supervised learning to make it more accurately adapt to specific task requirements. Specifically, according to specific task requirements, a Softmax classifier or a linear regression layer can be added on top of the deep belief network as the prediction layer, and the electromagnetic performance parameter set is input into the constructed metasurface structure prediction model again to generate the corresponding predicted values of the metasurface structure parameters.
[0054] S14. Calculate the difference between the predicted values of the metasurface structure parameters and the corresponding metasurface structure parameters, and update the parameters of the metasurface structure prediction model based on the difference to obtain the trained metasurface structure prediction model.
[0055] After completing the pre-training of the deep belief network, the fine-tuning stage is entered, aiming to optimize the entire model through supervised learning to make it more accurately adapt to the metasurface design task. The core goal of fine-tuning is to minimize the error between the prediction result and the true label, thereby improving the overall performance and generalization ability of the model. During the fine-tuning process, a labeled data set is used to train the entire metasurface structure prediction model to further optimize the performance of metasurface structure prediction. Specifically, the backpropagation algorithm is used to optimize the entire metasurface structure prediction model to continuously update its weights, and finally obtain the optimized weight parameters, and finally obtain the updated weight parameters W R =(w l , w l-1 ,…, w1). During the fine-tuning process, the weights and biases of each layer of the restricted Boltzmann machine in the deep belief network will be updated, and the weights and biases of the Softmax layer will also be optimized to ensure that the entire model can better learn the mapping relationship between the electromagnetic performance parameters and the metasurface structure parameters. In addition, the fine-tuning process can effectively avoid problems such as the deep belief network falling into local optimal solutions, insufficient model generalization ability, or excessive training time due to randomly initialized weight values, making the mapping relationship of the feature vectors tend to be optimal to further improve the accuracy and reliability of metasurface structure prediction.
[0056] Optionally, after obtaining the trained metasurface structure prediction model in step S14, it further includes: optimizing the obtained metasurface structure prediction model based on the few-shot learning method to obtain the finally trained metasurface structure prediction model. To further improve the adaptability and generalization performance of the metasurface structure prediction model, in view of the difficulties in obtaining metasurface material performance data and the limitation of the limited number of samples, a few-shot learning strategy is adopted to optimize the model. Specifically, the optimization methods include but are not limited to model optimization based on meta-learning and model optimization based on transfer learning, so that the model can still have strong learning ability and prediction accuracy under limited data, thereby improving the efficiency of metasurface structure design.
[0057] Optionally, the few-shot learning is the meta-learning method, which is used to optimize the adaptability and generalization performance of the metasurface structure prediction model under different tasks. The meta-learning algorithm enables the model to still have a high prediction accuracy under limited data by learning how to quickly adapt to new tasks, thereby significantly improving the training efficiency and practical application ability of the metasurface structure prediction model. The core goal of meta-learning is to endow the model with the ability to quickly learn new tasks, thereby improving the generalization of the metasurface structure prediction model in metasurface prediction. Specifically, meta-learning can optimize the performance of the metasurface structure prediction model in the following ways: on the one hand, meta-learning can fine-tune the parameters of the metasurface structure prediction model so that it can quickly adapt to new tasks while avoiding the model overfitting the features of a specific training set, thereby improving the generalization ability of the model. On the other hand, meta-learning can identify the features crucial for the generalization of the metasurface structure prediction model and preferentially select the data points of these features for training, thereby improving the applicability of the model in different metasurface design tasks. In addition, meta-learning can learn the optimal metasurface structure prediction model structure in different tasks, enabling it to make adaptive adjustments according to specific prediction goals, thereby enhancing its adaptability to different metasurface tasks. The meta-learning method can adopt MAML, Reptile, Meta-SGD, etc. Optionally, the MAML (Model-Agnostic Meta-Learning) meta-learning algorithm is adopted in this application to optimize the metasurface structure prediction model. The core idea of MAML is to train an initial model so that it can quickly adapt to new tasks with only a small number of gradient updates. This strategy enables the metasurface structure prediction model to efficiently adjust parameters using existing experience when facing new metasurface materials or new electromagnetic performance targets, without having to train from scratch, thereby significantly shortening the training time and improving the prediction accuracy.
[0058] Optimizing the obtained metasurface structure prediction model based on the MAML meta-learning method includes the following process:
[0059] First, according to different preset prediction targets, all electromagnetic performance parameter sets used to train the metasurface structure prediction model and their corresponding metasurface structure parameter sets are divided to form multiple electromagnetic performance parameter combinations and their corresponding metasurface structure parameter combinations. Among them, each prediction target corresponds to a mapping relationship between a set of electromagnetic performance parameters and metasurface structure parameters.
[0060] Adopt a meta-learning strategy to further optimize the trained metasurface structure prediction model to improve its adaptability and generalization ability under different tasks. Specifically, according to different preset prediction targets, the electromagnetic performance parameter set used for training and its corresponding metasurface structure parameter set are divided. Among them, each prediction target represents a specific task requirement, such as electromagnetic responses in different frequency bands, different material conditions, or design targets under different application scenarios, etc. It can be understood that when dividing, the original training data is divided into multiple sub-task data sets, and each data set is composed of a set of electromagnetic performance parameter combinations and their corresponding metasurface structure parameter combinations. Each data set represents a task sample under a prediction target, and its essence is a mapping relationship from performance parameters to structure parameters. Exemplarily, when the prediction target is to achieve a high transmittance design in the terahertz band, when dividing the training set, the electromagnetic performance parameters that meet the high transmittance requirements and their corresponding metasurface structure parameter combinations are used as the sub-task data set for this prediction target.
[0061] Then, for each prediction target, the metasurface structure prediction model is retrained a preset number of times, and the following process is executed for each retraining for this prediction target:
[0062] First, the electromagnetic performance parameter combination corresponding to this prediction target is input into the initially trained metasurface structure prediction model to obtain the predicted value of the metasurface structure parameters.
[0063] In the process of optimizing the metasurface structure prediction model based on the meta-learning strategy, for the i-th prediction target T i , K samples are sampled from the electromagnetic performance parameter combination corresponding to this prediction target. Each sample consists of a pair of input-output pairs (x (j) , y (j) ), where x (j) is the electromagnetic performance parameter of the j-th sample, and y (j) is the true metasurface structure parameter of the j-th sample. These electromagnetic performance parameter samples x (j) are input into the previously trained metasurface structure prediction model f θ , and the corresponding predicted value of the structure parameters f θ (x (j) ) is obtained.
[0064] Then, calculate the difference between the predicted values of the metasurface structure parameters and the corresponding metasurface structure parameters, and update the parameters of the metasurface structure prediction model based on the difference.
[0065] For the i-th prediction target T i , calculate the mean squared error between the predicted value and the true structure parameters, as shown in Equation (5):
[0066]
[0067] where is the mean squared error of the initially trained metasurface structure prediction model f θ on the i-th prediction target T i . Based on this error, perform gradient update on the model parameters θ to obtain the adaptive parameter θ' i for the current i-th prediction target T i . It can be understood that in the K-shot regression task, K input / output pairs are provided for each prediction target for learning. The updated parameter vector θ' i is calculated using one or more gradient descent updates on the prediction target T i . Specifically, the parameter update method is as shown in Equation (6):
[0068]
[0069] where α is the step size, which can be set as a fixed value or a meta-learning hyperparameter for optimization, is the gradient of the loss function with respect to the parameter θ, θ is the model parameter before update, and θ' i is the updated model parameter on the i-th prediction target T i . Use the adapted new parameter θ' i to calculate the loss on the task T i Repeat the above iterative process until the preset number of retraining times is reached. In addition, considering that each prediction target corresponds to different types of metasurface structures and electromagnetic response characteristics, optionally, for each retraining of the prediction target, the parameters of the metasurface structure prediction model are updated based on the difference degree by dynamically adjusting the learning rate. Through this adjustment of the dynamic learning rate, the metasurface structure prediction model learns more actively when the error is large and converges more stably when the error is small, thereby enhancing the adaptability of the metasurface structure prediction model under various tasks. After all prediction targets have been retrained for the preset number of times, finally, calculate the total loss based on the difference degrees of the last retraining of each prediction target, and update the initially trained metasurface structure prediction model based on the total loss to obtain the finally trained metasurface structure prediction model.
[0070]
[0071] After all prediction targets have completed retraining for the preset number of times, to further optimize the initially trained metasurface structure prediction model, based on the difference degrees calculated after the last retraining of each prediction target, a cross-task global update is performed to achieve meta-optimization of the model parameters. Specifically, the difference degrees of all n prediction targets are averaged to obtain the total loss L total , as shown in Equation (7):
[0072]
[0073] Based on this total loss, the backpropagation algorithm is used to update the parameters θ of the initially trained metasurface structure prediction model to obtain a new round of metasurface structure prediction model parameters where β is the meta-learning rate, which is used to control the amplitude of the update across prediction targets, is the gradient value of the original parameter. Repeat the above process for multiple rounds until the average loss of the model converges on all tasks, and finally obtain a metasurface structure prediction model with good generalization ability and fast adaptation ability.
[0074] It should be noted that the trained metasurface structure prediction model can either input electromagnetic performance parameters to predict the corresponding metasurface structure parameters, or input metasurface structure parameters to predict the corresponding electromagnetic performance parameters. Optionally, using the above-trained metasurface structure prediction model, by inputting the target electromagnetic performance (phase, amplitude, reflection, transmission, and absorption characteristics, etc.), the model outputs the corresponding metasurface structure parameters. This process may need to combine optimization algorithms, such as Genetic Algorithm (GA) or Particle Swarm Optimization (PSO), to search for the optimal metasurface structure. The performance of the model is evaluated using the held-out test set or through cross-validation, and the finite element method (FEM) or other numerical simulation methods are used to verify whether the generated metasurface structure meets the expected electromagnetic performance. Using the mean squared error (MSE), mean absolute error (MAE), or coefficient of determination (R 2)Calculate the metrics for prediction performance. Based on the validation results, iteratively optimize the metasurface structure prediction model or the metasurface design until satisfactory performance is achieved. Finally, fabricate the metasurface structure generated by the model and measure its electromagnetic properties through experiments. Compare the experimental results with the simulation results and the design goals to verify the accuracy and reliability of the design. Finally, the trained model can be integrated into a user-friendly system, allowing users to quickly obtain performance prediction results by simply inputting the parameters of the metasurface material without the need for high professional knowledge and repeated numerical simulations. As the amount of data increases, the performance of the model can be further improved through continuous learning. Through such a method, researchers and engineers can quickly predict the performance of metasurface materials, thus accelerating the design and development process of new materials. This method is particularly suitable for application scenarios where data is scarce but rapid iteration and decision-making are required.
[0075] As Figure 3 shown, this application also provides a method for predicting the metasurface structure. The prediction method includes:
[0076] S31. Obtain the electromagnetic performance parameters of the metasurface material to be predicted.
[0077] S32. Input the electromagnetic performance parameters into the metasurface structure prediction model to generate the metasurface structure parameters of the metasurface material.
[0078] Obtain the electromagnetic performance parameters (such as transmittance, reflectance, absorptance, etc.) of the metasurface material to be predicted, input the obtained electromagnetic performance parameters into the metasurface structure prediction model, extract electromagnetic features and generate corresponding metasurface structure parameters. Thus, an efficient mapping from electromagnetic performance to structure parameters is achieved, providing accurate prediction support for the design and optimization of metasurface materials.
[0079] As Figure 4As shown in the figure, the training system 100 of the metasurface structure prediction model includes: a data acquisition module 110, a primary training module 120, a parameter prediction module 130, and a parameter update module 140. The above-mentioned data acquisition module 110 is used to acquire the electromagnetic performance parameter set of the metasurface material and the corresponding metasurface structure parameter set. The above-mentioned primary training module 120 is used to input the electromagnetic performance parameter set into the deep belief network, and adopt the layer-by-layer greedy training method to pre-train the restricted Boltzmann machine of each layer of the deep belief network in turn. After all layers are trained, a pre-trained deep belief network is obtained; among them, the deep belief network includes multiple cascaded restricted Boltzmann machines. The above-mentioned parameter prediction module 130 is used to input the electromagnetic performance parameter set into the metasurface structure prediction model, and correspondingly obtain the predicted value of the metasurface structure parameter; among them, the metasurface structure prediction model includes a pre-trained deep belief network and a prediction layer cascaded with it. The above-mentioned parameter update module 140 is used to calculate the difference degree between the predicted value of the metasurface structure parameter and the corresponding metasurface structure parameter, and update the parameters of the metasurface structure prediction model based on the difference degree to obtain a trained metasurface structure prediction model.
[0080] For the specific limitations of the training system of the metasurface structure prediction model, reference can be made to the limitations of the training method of the metasurface structure prediction model in the above text, which will not be elaborated here. Each module in the above-mentioned training system of the metasurface structure prediction model can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in a hardware format, or stored in the memory in the computer device in a software format, so that the processor can call the corresponding operations of the above-mentioned modules.
[0081] It should be noted that in order to highlight the innovative part of this application, modules that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other modules in this embodiment.
[0082] As Figure 5 shown, the electronic device 1 may include a memory 11, a processor 12 and a bus, and may also include a computer program stored in the memory 11 and executable on the processor 12, such as a training of the metasurface structure prediction model or a metasurface structure prediction program.
[0083] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In some other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device 1. The memory 11 can be used not only to store application software installed on the electronic device 1 and various types of data, such as the training of the metasurface structure prediction model or the code for metasurface structure prediction, etc., but also to temporarily store data that has been output or will be output.
[0084] In some embodiments, the processor 12 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 12 is the control core (Control Unit) of the electronic device 1, connecting various components of the entire electronic device 1 through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as the training of the metasurface structure prediction model or the metasurface structure prediction program, etc.), and by calling the data stored in the memory 11, to execute various functions of the electronic device 1 and process data.
[0085] The processor 12 executes the operating system of the electronic device 1 and various installed application programs. The processor 12 executes the application program to implement the steps in the above metasurface structure prediction model method.
[0086] Exemplarily, the computer program can be divided into one or more modules, and one or more modules are stored in the memory 11 and executed by the processor 12 to complete this application. One or more modules may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program can be divided into a data acquisition module 110, a primary training module 120, a parameter prediction module 130, and a parameter update module 140.
[0087] The above-mentioned integrated unit implemented in the form of a software function module can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The above-mentioned software function module is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to perform the training method of the hypersurface structure prediction model or part of the function of the hypersurface structure prediction method of each embodiment of the present application.
[0088] In summary, the present application discloses a method, system, device and medium for training a hypersurface structure prediction and model, and adopts a layer-by-layer greedy training method to independently and unsupervisedly pre-train each layer of the restricted Boltzmann machine of the deep belief network, so that the model gradually learns the deep characteristic relationship between electromagnetic performance and structural parameters, thereby improving the feature extraction ability and generalization ability of the model. On this basis, a hypersurface structure prediction model is constructed according to the pre-trained deep belief network and its cascaded prediction layer, and the model is trained and optimized so that it can accurately predict the hypersurface structure parameters. By calculating the difference between the predicted structural parameters and the real structural parameters and continuously optimizing the model parameters, the prediction accuracy is improved. Finally, the hypersurface structure parameters can be accurately predicted. The present application effectively improves the problem of insufficient accuracy of the hypersurface structure prediction of the prior art, and the method of the present application is used to reduce the deviation between the theoretical calculation results and the actual measurement results, thereby improving the reliability and engineering feasibility of the design of the hypersurface material. Therefore, the present application effectively overcomes the various shortcomings in the prior art and has a high industrial utilization value.
[0089] The above embodiments are merely illustrative of the principles and effects of the present application and are not intended to limit the present application. Anyone familiar with the technology may modify or change the above embodiments without violating the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed in the present application shall still be covered by the claims of the present application.
Claims
1. A training method for a metasurface structure prediction model, characterized in that The training method includes: Obtaining an electromagnetic property parameter set of a metasurface material and a corresponding metasurface structure parameter set; Inputting the electromagnetic property parameter set into a deep belief network, and using a layer-by-layer greedy training method to pre-train the restricted Boltzmann machine of each layer of the deep belief network in turn. After all layers are trained, a pre-trained deep belief network is obtained; wherein, the deep belief network includes multiple cascaded restricted Boltzmann machines; Inputting the electromagnetic property parameter set into a metasurface structure prediction model to correspondingly obtain a predicted value of the metasurface structure parameter; wherein, the metasurface structure prediction model includes the pre-trained deep belief network and a prediction layer cascaded therewith; Calculating the difference degree between the predicted value of the metasurface structure parameter and the corresponding metasurface structure parameter, and updating the parameters of the metasurface structure prediction model based on the difference degree to obtain a trained metasurface structure prediction model.
2. The training method of the metasurface structure prediction model according to claim 1, wherein For the restricted Boltzmann machine of the first layer of the deep belief network, its training process includes: Inputting the electromagnetic property parameter set into the restricted Boltzmann machine for feature extraction to obtain a metasurface deep feature set; Performing a reconstruction process on the metasurface deep feature set to obtain a metasurface reconstruction feature set; Calculating the contrast difference degree between the electromagnetic property parameter set and the metasurface reconstruction feature set based on the contrast divergence algorithm, and updating the restricted Boltzmann machine based on the contrast difference degree to obtain a trained restricted Boltzmann machine, and using the metasurface deep feature set generated by the restricted Boltzmann machine in the last time of the training stage as the metasurface feature set input to the next layer of the restricted Boltzmann machine.
3. The training method of the metasurface structure prediction model according to claim 2, wherein, For each remaining layer of the restricted Boltzmann machine of the deep belief network, its training process includes: Inputting the metasurface feature set generated by the previous layer of the restricted Boltzmann machine into the restricted Boltzmann machine of the current layer for feature extraction to obtain a metasurface deep feature set generated by the restricted Boltzmann machine of the current layer; Performing a reconstruction process on the metasurface deep feature set to obtain a metasurface reconstruction feature set; Calculating the contrast difference degree between the metasurface deep feature set and the metasurface reconstruction feature set based on the contrast divergence algorithm, and updating the restricted Boltzmann machine of the current layer based on the contrast difference degree to obtain a trained restricted Boltzmann machine of the current layer.
4. The training method of the metasurface structure prediction model according to claim 1, characterized in that After obtaining the trained metasurface structure prediction model, it further includes: optimizing the obtained metasurface structure prediction model based on the few-shot learning method to obtain a finally trained metasurface structure prediction model.
5. The training method of the metasurface structure prediction model according to claim 4, characterized in that, The few-shot learning is the MAML meta-learning method. Optimizing the obtained metasurface structure prediction model based on the MAML meta-learning method includes: Dividing all the electromagnetic property parameter sets and their corresponding metasurface structure parameter sets used for training the metasurface structure prediction model according to preset different prediction targets to form multiple electromagnetic property parameter combinations and their corresponding metasurface structure parameter combinations; wherein, each prediction target corresponds to a mapping relationship between a group of electromagnetic property parameters and metasurface structure parameters; Performing a preset number of retrainings on the metasurface structure prediction model for each prediction target. For each retraining for the prediction target: Input the electromagnetic performance parameter combination corresponding to the prediction target into the initially trained metasurface structure prediction model to obtain the predicted value of the metasurface structure parameters; Calculate the difference degree between the predicted value of the metasurface structure parameters and the corresponding metasurface structure parameters, and update the parameters of the metasurface structure prediction model based on the difference degree; Calculate the total loss based on the difference degrees of the last retraining for each prediction target, and update the initially trained metasurface structure prediction model based on the total loss to obtain the finally trained metasurface structure prediction model.
6. The training method of the metasurface structure prediction model according to claim 5, characterized in that For each retraining of the prediction target, the parameters of the metasurface structure prediction model are updated by dynamically adjusting the learning rate based on the difference degree.
7. A method for predicting a metasurface structure, characterized in that The prediction method includes: Obtain the electromagnetic performance parameters of the metasurface material to be predicted; Input the electromagnetic performance parameters into the metasurface structure prediction model to generate the metasurface structure parameters of the metasurface material; wherein, the metasurface structure prediction model is a model trained by the metasurface structure prediction model method according to any one of claims 1-6.
8. A training system for a metasurface structure prediction model, characterized in that, The system includes: A data acquisition module for acquiring the electromagnetic performance parameter set of the metasurface material and the corresponding metasurface structure parameter set; An initial training module for inputting the electromagnetic performance parameter set into a deep belief network, and using the layer-by-layer greedy training method to pre-train the restricted Boltzmann machines of each layer of the deep belief network in sequence. After all layers are trained, obtain the pre-trained deep belief network; wherein, the deep belief network includes multiple cascaded restricted Boltzmann machines; A parameter prediction module for inputting the electromagnetic performance parameter set into the metasurface structure prediction model to correspondingly obtain the predicted value of the metasurface structure parameters; wherein, the metasurface structure prediction model includes the pre-trained deep belief network and a prediction layer cascaded with it; A parameter update module for calculating the difference degree between the predicted value of the metasurface structure parameters and the corresponding metasurface structure parameters, and updating the parameters of the metasurface structure prediction model based on the difference degree to obtain the trained metasurface structure prediction model.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the training method of the metasurface structure prediction model according to any one of claims 1 to 6 or the metasurface structure prediction method according to claim 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which, when executed by the processor of the computer, causes the computer to execute the training method of the metasurface structure prediction model according to any one of claims 1 to 6 or the metasurface structure prediction method according to claim 7.