Asymmetric Optical Polarization Device Structure Based on Asynchronous Reinforcement Learning and Its Design Method
By applying asynchronous reinforcement learning and deep neural networks in polarization conversion metasurface structure design, the structure of asymmetric light polarization conversion devices is automatically optimized, which solves the problem of long simulation time, professional intervention and low manual adjustment efficiency in the traditional design process, and achieves a fast and efficient global optimal design.
Patent Information
- Application Number
- CN202110656127.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-11
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-06-11
AI Technical Summary
In the design process of traditional polarization conversion metasurface structures, there are problems such as long simulation time, requiring professional intervention and low manual adjustment efficiency, resulting in low design efficiency and easy to fall into local optimization.
The design method based on asynchronous reinforcement learning is adopted, and the transmission prediction network is built using deep neural networks, and the asynchronous reinforcement learning algorithm is adaptively optimized the structure of the asymmetric optical polarization conversion device, and the structural attributes and material attributes are automatically selected to reduce the dependence on expert judgments.
It realizes the rapid and efficient design of a metasurface device structure that meets the expected efficiency, saves a lot of simulation time and computing resources, avoids local optimal problems, and achieves global optimal transmission performance.
Smart Images

Figure CN113378388B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of polarization device design, and particularly relates to an asymmetric optical polarization device structure based on asynchronous reinforcement learning and a design method thereof. Background Art
[0002] Currently, polarization is an important property of electromagnetic waves, which has broad research value in imaging, military, navigation, satellite communication, etc. However, traditional polarization control mainly uses half-wave plates and dichroic crystals. The principle is that when electromagnetic waves propagate inside, the phase difference of light with mutually perpendicular polarization directions accumulates with the increase of the propagation distance, resulting in polarization conversion. However, the effects brought by traditional methods are not ideal, such as low conversion efficiency and narrow bandwidth.
[0003] In recent decades, with the rapid development of artificial electromagnetic structures (metamaterials), the unique properties they increasingly exhibit provide a new method for solving the problems brought by traditional materials. Metamaterials are reshaped in space for one or more sub-wavelength units in a certain combination manner, so that arbitrary control of electromagnetic waves can be achieved within the sub-wavelength scale. At the same time, properties that natural media do not possess can be realized, such as negative refractive index property, optical activity, and inverse Doppler effect.
[0004] Generally speaking, all optical systems composed of artificial sub-wavelength structures can become metamaterials. Usually, in metamaterials with polarization control performance, there are two major popular branches, namely anisotropic metamaterials (metasurfaces) and chiral metamaterials. Among them, anisotropic metamaterials (metasurfaces) mainly introduce different phase differences in two orthogonal directions to independently control the responses of different polarization directions. For chiral metamaterials, due to the lack of any symmetry in the structure, electromagnetic coupling occurs when electromagnetic waves propagate in the structure, so that effective regulation of electromagnetic waves can be achieved.
[0005] In traditional polarization conversion metasurface structure design schemes, mainly through such as Figure 6The design is carried out through the two major cycles of the numerical simulation tool and manual adjustment shown. For the designed polarization conversion metasurface, it is first necessary to determine the structural type of the polarization conversion metasurface (whether it is an anisotropic structure, a chiral structure, or a combination of a chiral structure and anisotropy), the period size of the structure (which determines the working band of the polarization device), and the materials of each layer. After determining the structural type, materials, and period size, randomly initialize each metasurface structure parameter (including parameters such as the thickness of each layer and the refractive index of the dielectric layer) to reasonable parameter values. Then, according to the requirements of the simulation tool (such as the FDTD software of Lumerical), use a computer language to establish the mathematical model required for numerical simulation. Then, select one of the structure parameters in the order of each structure parameter or other logical order as the parameter layer to be adjusted, and keep the rest of the structure parameters unchanged. Use the numerical simulation tool to simulate the current structure, so as to obtain the cross-polarization transmittance and co-polarization transmittance of the polarization conversion metasurface. Furthermore, judge whether the performance of the current polarization conversion metasurface is the best. If not, then use the prior knowledge of experts to guide how to adjust the value of the current structure parameter to be optimized; if it is the best, then continue to select the next structure parameter to be adjusted. When all structure parameters have been adjusted to the optimal values, the design process is completed.
[0006] The disadvantages of the traditional design of the polarization conversion metasurface structure are mainly reflected in the following aspects:
[0007] Selecting the structural type, period size of the structure, and materials of each layer of the polarization conversion metasurface requires relying on a large amount of engineering experience from previous experiments. At the same time, the intervention of experts is needed to judge the rationality of the selection of each parameter, which will waste a lot of manpower and material resources.
[0008] After each adjustment of a structure parameter in the polarization conversion metasurface structure, it is necessary to build different simulation models, which requires the intervention of professionals and will also consume unnecessary time in the establishment and fine-tuning of the theoretical model.
[0009] Each time the structure of the polarization conversion metasurface is adjusted, not only does it need to rebuild the simulation model, but also it needs to spend a lot of time on simulation calculation and model solution again to obtain the cross-polarization transmittance T yx and co-polarization transmittance T yy .
[0010] The evaluation and adjustment of the simulation results of the polarization conversion metasurface both require manual intervention. Moreover, since the structural data of the metasurface is not one-dimensional, it is difficult for humans to jointly adjust its structure simultaneously. After optimizing by adjusting one variable at a time and then fixing that parameter to adjust other parameters, the performance of the polarization conversion metasurface obtained by the adjustment is prone to falling into a local optimum, and it is difficult to find the global optimum structure.
[0011] Through the above analysis, the problems and defects existing in the prior art are as follows: In the traditional design process, the simulation time is too long, professional personnel intervention is required, and the efficiency of manual adjustment is low.
[0012] The difficulty in solving the above problems and defects is:
[0013] The present invention can use reinforcement learning to intelligently and automatically select the structural attributes and material attributes of the polarization conversion metasurface, solving the difficulty of relying on experts. Furthermore, it is not necessary to judge the quality of the selection every time an adjustment or selection is made.
[0014] Relying on the deep neural network, the transmittance attribute can be directly obtained from the structural attributes of the polarization conversion metasurface, which can solve the drawback that FDTD needs to rebuild and simulate every time a new structure needs to be simulated.
[0015] Adopting reinforcement learning and deep neural network for optimization can solve the drawback of falling into local optimum when optimizing one variable by one in the traditional optimization method. At the same time, there is no need to manually adjust the structure or material parameters, saving the adjustment time and making the final result reach the global optimum.
[0016] The significance of solving the above problems and defects is:
[0017] Using reinforcement learning to automatically optimize the structural or material attribute parameters of the polarization conversion metasurface can save a large amount of time for expert judgment compared with the traditional optimization process. At the same time, the reinforcement learning algorithm has a unified evaluation standard for the optimization results, eliminating human subjectivity and being completely objective.
[0018] The deep neural network can directly obtain the transmittance attribute from the structural attributes of the polarization conversion metasurface. Then, there is no need to simulate the device with the help of an expensive server, which can save a large amount of computing resources and computing time, providing a faster way to quickly verify the performance of the device.
[0019] Adopting reinforcement learning and deep neural network to optimize the structural and material parameters of each part of the device can achieve the purpose of global optimization, avoid the problem of local optimization, and make the polarization efficiency of the polarization conversion metasurface better than that obtained by the traditional optimization method. Summary of the Invention
[0020] Aiming at the problems existing in the prior art, the present invention provides an asymmetric optical polarization device structure based on asynchronous reinforcement learning and a design method thereof. The present invention is a design method for optimizing the structural parameters of an asymmetric optical polarization conversion device to maximize the transmittance of the asymmetric optical polarization conversion device, and can conveniently and efficiently design a metasurface device structure meeting the expected efficiency.
[0021] The present invention is implemented as follows. A design method for an asymmetric optical polarization device structure based on asynchronous reinforcement learning, the design method for the asymmetric optical polarization device structure based on asynchronous reinforcement learning includes:
[0022] Step 1, preprocessing of the simulation data set;
[0023] Step 2, building and initializing a transmittance prediction network;
[0024] Step 3, training the transmittance prediction network;
[0025] Step 4, optimizing the structure of the asymmetric polarization conversion device by an asynchronous reinforcement learning algorithm.
[0026] Further, in the above Step 1, the specific process of preprocessing the simulation data set is as follows:
[0027] For the data set preprocessing part, the obtained device structure data is normalized to between 0 and 1, and at the same time, the obtained 83-dimensional transmittance spectrum data is downsampled to 27-dimensional data through a ratio of 3:1.
[0028] Further, in the above Step 2, the specific process of building and optimizing the transmittance prediction network is as follows:
[0029] The transmittance prediction network based on the deep neural network with a residual structure takes the structural parameters of the asymmetric optical polarization conversion device as input data and predicts the corresponding transmittance value in a very short time;
[0030] Among all modules of the transmittance prediction network, the basic units adopted are fully connected layers and Relu activation layers; followed by a batch normalization layer, and at the same time, a residual structure is adopted as one of the paths for error backpropagation.
[0031] Further, the basic structure of the transmittance prediction network is divided into three parts: serial input SIN, predicted output PYN of transmittance T yy and predicted output PXN of transmittance T yx ;
[0032] The input is the structural parameters of the asymmetric optical polarization conversion device, and the output is the transmittance attribute T yy of the asymmetric optical polarization conversion device and T yx, in the SIN, four fully connected layers Dense are first adopted, and the Relu function (f(x) = max(0, x)) is used to activate all the fully connected layers, and then the batch normalization layer is used; in the predicted output PYN of the transmittance T yy and the predicted output PXN of the transmittance T yx , the fully connected layer FC and the activation function Relu are also composed, and a residual structure is added therein, and finally the LeakyRelu function is used:
[0033] for activation; using the residual structure is beneficial to the backpropagation of errors and avoids gradient disappearance and gradient explosion.
[0034] Furthermore, after the series input part SIN receives the structure of the asymmetric optical polarization conversion device as input data, it connects 4 fully connected layers and ends with batch normalization, extracts the features of the structure data, uses a relatively large training rate parameter, and uses more batch training data at one time during training; the extracted feature data is used as the input of both the PYN and PXN parts after the batch normalization layer;
[0035] PYN consists of 6 basic units and an output part. Each basic unit includes a fully connected layer Dense and a Relu activation layer. The number of hidden layer neurons included in these basic units is 200, 500, 500, 500, 200, 28, and the activation layer function uniformly uses the Relu function; the final output part is first a fully connected layer of 28 neurons followed by a LeakyRelu activation output;
[0036] PXN consists of 5 basic units and an output part. Each basic unit includes a fully connected layer Dense and a Relu activation layer. The number of hidden layer neurons included in these basic units is 200, 500, 500, 200, 28, and the activation layer function uniformly uses the Relu function; the final output part is first a fully connected layer of 28 neurons followed by a LeakyRelu activation output.
[0037] Furthermore, in the third step, the specific process of training the transmittance prediction network is as follows:
[0038] For the training of the network, 80% of the obtained simulation data set is used as the training set, and the remaining 20% is used as the test set, and the 5-fold cross-validation method is adopted to group and train and test all the data in turn.
[0039] Furthermore, in the fourth step, the asynchronous reinforcement learning algorithm is used to optimize the structure of the asymmetric polarization conversion device, specifically:
[0040] An asynchronous reinforcement learning algorithm is used to adaptively optimize the structure of an asymmetric optical polarization conversion device. The asynchronous reinforcement learning algorithm uses the transmittance prediction network as an agent function and the structure of the asymmetric optical polarization conversion device as an independent variable. At the same time, the mean value of the transmittance T yy ,T yx is used as the target value to be optimized, and the structure of the asymmetric optical polarization conversion device is continuously adjusted by maximizing the optimization target value, and the optimal adjustment strategy is continuously explored during the iteration process; until the asynchronous reinforcement learning algorithm converges finally, at this time the structure of the asymmetric optical polarization conversion device is the optimal.
[0041] Furthermore, the asynchronous reinforcement learning algorithm is specifically as follows:
[0042] After initializing the maximum number of iterations, boundary conditions, and additional constraint conditions; randomly initialize the structure of the asymmetric optical polarization device; at the time of initialization, make the structure parameters within the preset range;
[0043] In the asynchronous reinforcement learning algorithm, the transmittance T is obtained from the structure of the asymmetric optical polarization device through the transmittance prediction network yy and T yx , and then the return value inside the asynchronous reinforcement learning algorithm is calculated through calculation Furthermore, the return value inside the asynchronous reinforcement learning algorithm is calculated.
[0044] The specific calculation method is that when the numerical value of the structure of the asymmetric optical polarization device exceeds the limit, the return value r = -0.5 at this time; when the maximum value of T yy and T yx is less than 0.3, r = exp(max(T yy ) + max(T yx ) - 1) - 1 at this time. When the maximum values of T yy and T yx are both between 0.3 and 0.5, r = 0. When the maximum value of Tyy is between 0.3 and 0.5 and at the same time T yx is greater than or equal to 0.5 or the maximum value of T yx is between 0.3 and 0.5 and at the same time T yy is greater than or equal to 0.5, When the maximum values of T yy and T yx are both greater than 0.5, r = exp(max(T yy ) + max(T yx ) - 1).
[0045] Furthermore, the adaptive termination condition of the asynchronous reinforcement learning algorithm is designed as:
[0046] After finding the optimal value in this iteration, it is compared with the optimal values found in previous iterations. If the optimal value found in this iteration is better, then continue the iteration;
[0047] If it is not better than the previous one, then check the number of iterations. If the number of iterations is greater than 2500 and no better target value is found in the last 300 iterations, then end the optimization.
[0048] Another object of the present invention is to provide an asymmetric optical polarization device structure using the asymmetric optical polarization device structure design method based on asynchronous reinforcement learning. The asymmetric optical polarization device structure is provided with a double-layer anisotropic metasurface, and the double-layer anisotropic metasurface includes: a pair of metal resonant rods and a sub-wavelength metal grating, and a dielectric layer is filled between the two layers of structures.
[0049] Combining all the above technical solutions, the advantages and positive effects of the present invention are:
[0050] During the construction of the transmittance prediction network, a residual structure is used, which can solve the problem of vanishing gradient during the backpropagation of the network, and at the same time solve the problem of network degradation. Furthermore, it speeds up the training speed of the transmittance prediction network. At the same time, since all transmittances are positive, the activation functions output by the network all adopt LeakyRelu, and a batch normalization layer is used, thereby accelerating the training and convergence speed of the network, controlling gradient explosion to prevent gradient disappearance, and preventing overfitting during network training.
[0051] (3) The effect of dependent claim 3.
[0052] Dividing the network into three parts: SIN, PYN, and PXN is beneficial to more accurately predict the transmittance T yy and T yx . The main function of SIN is to extract information from the structural parameters and transfer the extracted information to the next layer. Since the transmittance T yy and T yx are relatively independent, two output prediction networks, PYN and PXN, are used for separate prediction. At the same time, it can reduce the parameters that need to be trained in the network and speed up the later training speed.
[0053] (4) The comparative technical effects or experimental effects.
[0054] By adopting the methods described in claim 1, claim 2, and claim 3, the prediction accuracy of the transmittance prediction network for T yy and T yx can reach 96.6% and 95.5%. And it can complete the prediction of the transmittance attribute of the asymmetric optical polarization device within milliseconds. And it can complete the optimization process in much less time than the traditional optimization method.
[0055] The present invention proposes a design method based on a transmittance prediction network with a residual structure combined with an asynchronous reinforcement learning algorithm. By utilizing the characteristics of a deep neural network, which can approximate non-linear functions with arbitrary precision, the transmittance attribute can be accurately predicted from structural data through the transmittance prediction network, and the structure of an asymmetric optical polarization conversion device can be optimally designed in a reverse manner efficiently and time-savingly by the asynchronous reinforcement learning algorithm. Based on the transmittance prediction network of the deep neural network with a residual structure and the asynchronous reinforcement learning algorithm of the present invention, through effective downsampling of data and reasonable partitioning, the transmittance prediction network is effectively trained and the structure of the asymmetric optical polarization conversion device is optimally designed using the asynchronous reinforcement learning algorithm, improving the design efficiency and making the efficiency of the final asymmetric optical polarization conversion device superior to that obtained by traditional design methods.
[0056] The present invention also has the following advantages:
[0057] First, the device structure is novel. The polarization conversion metasurface proposed by the present invention is composed of a pair of metal resonant rods and sub-wavelength metal gratings, which can achieve asymmetric polarization conversion in the blue light band, concentrating the electromagnetic wave energy distributed in two orthogonal linear polarization states into one polarization state, and providing a feasible technical implementation approach for many low-loss optoelectronic applications.
[0058] Second, the transmittance prediction network is used to replace the traditional numerical simulation tool to predict the transmittance, thereby saving a large amount of simulation time and computing resources. Taking the simulation of 12,500 groups of data of the asymmetric optical polarization conversion device as an example, the transmittance prediction network only requires 0.59 seconds, while the traditional FDTD simulation requires 762,514 seconds. The former can save about 1.3 million times the time. Moreover, the latter needs to complete the computational simulation on a high-performance server or computer cluster, while the design scheme of the present invention can complete the design process on an ordinary home computer.
[0059] Third, the present invention provides an asynchronous reinforcement learning algorithm. Combined with the above-mentioned transmittance prediction network, it can adaptively optimize the structure of the asymmetric optical polarization conversion device in a reverse manner, making the efficiency of the asymmetric optical polarization conversion device reach the expected efficiency. Using the asynchronous reinforcement learning algorithm to replace manual adjustment can save adjustment time, and at the same time, joint adjustment can be performed to find the optimal discrete structure of the asymmetric optical polarization conversion device, solving the disadvantage that the transmittance of the asymmetric optical polarization conversion device designed by the traditional method is prone to fall into a local optimum, and making the efficiency of the asymmetric optical polarization conversion device reach the global optimum. Taking the design of the structure of an asymmetric optical polarization conversion device with four-dimensional adjustable structure data as an example, the optimal average transmittance of the asymmetric polarization conversion device designed by combining the deep neural network with a residual structure and the asynchronous reinforcement learning algorithm is 21% higher than the optimal average transmittance obtained by traditional manual adjustment combined with FDTD simulation.
[0060] IV. The structural design method of the asymmetric optical polarization conversion device proposed by the present invention has strong universality. It can be used for the structural optimization of asymmetric polarization devices with different material selections and different structural types. Moreover, the optimal structure can be quickly found without the intervention of professionals during the design process. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is a flowchart of the method for designing the structure of an asymmetric optical polarization device based on asynchronous reinforcement learning provided by an embodiment of the present invention.
[0062] Figure 2 is a flowchart of the structural design of the asymmetric optical polarization conversion device provided by an embodiment of the present invention.
[0063] Figure 3 is a schematic diagram of the transmittance prediction network provided by an embodiment of the present invention.
[0064] Figure 4 is a flowchart of the design of the structure of the asymmetric optical polarization conversion device by combining the asynchronous reinforcement learning algorithm and the transmittance prediction network provided by an embodiment of the present invention.
[0065] Figure 5 is a flowchart of the asynchronous reinforcement learning algorithm provided by an embodiment of the present invention.
[0066] Figure 6 is a flowchart of the design of the traditional asymmetric optical polarization conversion device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention.
[0068] In view of the problems existing in the prior art, the present invention provides an asymmetric optical polarization device structure and its design method based on asynchronous reinforcement learning, which will be described in detail below with reference to the accompanying drawings.
[0069] Those of ordinary skill in the art in the industry can also implement the method for designing the structure of the asymmetric optical polarization device based on asynchronous reinforcement learning provided by the present invention using other steps. Figure 1 The method for designing the structure of the asymmetric optical polarization device based on asynchronous reinforcement learning provided by the present invention is only a specific embodiment.
[0070] As Figure 1 shown, the method for designing the structure of the asymmetric optical polarization device based on asynchronous reinforcement learning provided by an embodiment of the present invention
[0071] S101: Preprocessing of the simulation data set;
[0072] S102: Build and initialize the transmittance prediction network;
[0073] S103: Train the transmittance prediction network;
[0074] S104: Optimize the structure of the asymmetric polarization conversion device using the asynchronous reinforcement learning algorithm.
[0075] In S101 provided by the embodiment of the present invention, the specific process of preprocessing the simulation data set is as follows:
[0076] For the data set preprocessing part, the obtained device structure data is normalized to be between 0 and 1, and at the same time, the obtained 83-dimensional transmittance spectrum data is downsampled to 27-dimensional data through a 3:1 ratio. The normalization process is as follows: where X norm is the result after normalization, X min and X max represent the minimum and maximum values of the same dimension in the same batch of structure or material property data. After normalization, perform downsampling operations on T yy and T yx the transmittance spectrum data. For example, if the indexes of the data are 0, 1, 2…, 81, 82, the data in the spectral lines with indexes 0, 3, 6…, 81 are retained after downsampling.
[0077] In S102 provided by the embodiment of the present invention, the specific process of building and optimizing the transmittance prediction network is as follows:
[0078] The transmittance prediction network based on the deep neural network with a residual structure takes the structural parameters of the asymmetric optical polarization conversion device as input data and predicts the corresponding transmittance value in a very short time;
[0079] Among all the modules of the transmittance prediction network, the basic units adopted are the fully connected layer and the Relu activation layer; followed by the batch normalization layer, and at the same time, the residual structure is used as one of the paths for error backpropagation.
[0080] The basic structure of the transmittance prediction network is divided into three parts: the series input SIN, the predicted output PYN of the transmittance T yy , and the predicted output PXN of the transmittance T yx .
[0081] The input is the structural parameters of the asymmetric optical polarization conversion device, and the output is the transmittance attributes T yy and T yx . In SIN, first use four fully connected layers Dense, and activate all the fully connected layers using the Relu function (f(x) = max(0, x)), and then use the batch normalization layer. In the transmittance Tyy The predicted outputs PYN and transmittance T yx Among the predicted outputs PXN, the fully connected layer FC and the activation function Relu are also composed, and a residual structure is added therein. Finally, the LeakyRelu function is used:
[0082] For activation. Using the residual structure is beneficial to the backpropagation of errors and avoids gradient vanishing and gradient explosion.
[0083] Among them, the serial input part SIN is connected to 4 fully connected layers after receiving the structure of the asymmetric optical polarization conversion device as input data and ends with batch normalization to extract the features of the structure data. A relatively large training rate parameter is used, and more batch training data is used at one time during training. The extracted feature data is used as the input for both the PYN and PXN parts after the batch normalization layer. PYN consists of 6 basic units and an output part. Each basic unit includes a Dense fully connected layer and a Relu activation layer. The number of hidden layer neurons included in these basic units is 200, 500, 500, 500, 200, 28, and the activation layer function uniformly uses the Relu function. The final output part is first a fully connected layer of 28 neurons followed by a LeakyRelu activation output. PXN consists of 5 basic units and an output part. Each basic unit includes a Dense fully connected layer and a Relu activation layer. The number of hidden layer neurons included in these basic units is 200, 500, 500, 200, 28, and the activation layer function uniformly uses the Relu function. The final output part is first a fully connected layer of 28 neurons followed by a LeakyRelu activation output. All the weight parameters in the network are initialized using the Glorot uniform distribution, and the bias parameters are initialized using the all-0 uniform distribution.
[0084] In S103 provided by the embodiment of the present invention, the specific process of training the transmittance prediction network is as follows:
[0085] For the training of the network, 80% of the obtained simulation data set is used as the training set, and the remaining 20% is used as the test set, and 5-fold cross-validation is used to group-train and test all the data in turn. For 10,000 groups of data, 2,000 groups of data are selected as the validation set each time, and the other data is used as the training set.
[0086] In S104 provided by the embodiment of the present invention, the asynchronous reinforcement learning algorithm is used to optimize the structure of the asymmetric polarization conversion device, specifically as follows:
[0087] An asynchronous reinforcement learning algorithm is used to adaptively optimize the structure of an asymmetric optical polarization conversion device. The asynchronous reinforcement learning algorithm takes the transmittance prediction network as an agent function and the structure of the asymmetric optical polarization conversion device as an independent variable. At the same time, the mean value of the transmittance T yy ,T yx is used as the target value to be optimized. The structure of the asymmetric optical polarization conversion device is continuously adjusted by maximizing the optimization target value, and the optimal adjustment strategy is continuously explored during the iteration process. Until the asynchronous reinforcement learning algorithm converges finally, the structure of the asymmetric optical polarization conversion device is optimal at this time.
[0088] Among them, the asynchronous reinforcement learning algorithm is specifically as follows:
[0089] After initializing the maximum number of iterations, boundary conditions, and additional constraint conditions, randomly initialize the structure of the asymmetric optical polarization device. When initializing, it is necessary to ensure that the structural parameters are within the preset range.
[0090] In the asynchronous reinforcement learning algorithm, the transmittance T is obtained from the structure of the asymmetric optical polarization device through the transmittance prediction network yy and T yx . After that, the return value inside the asynchronous reinforcement learning algorithm is further calculated through . The specific calculation method is as follows: when the structural value of the asymmetric optical polarization device exceeds the limit, the return value r = -0.5 at this time. When T yy and T yx 's maximum value is less than 0.3, r = exp(max(T yy ) + max(T yx ) - 1) - 1 at this time. When the maximum values of T yy and T yx are both between 0.3 and 0.5, r = 0. When the maximum value of T yy is between 0.3 and 0.5 and T yx is greater than or equal to 0.5, or when the maximum value of Tyx is between 0.3 and 0.5 and Tyy is greater than or equal to 0.5, When the maximum values of T yy and T yx are both greater than 0.5, r = exp(max(T yy ) + max(T yx ) - 1). Taking max(T yy ) = 0.15 and max(T yx ) = 0 as an example, r = exp(0.15 + 0.4 - 1) - 1 = -0.36. Then max(T yy ) = 0.55 and max(T yx) Taking 0.67 as an example, r = exp(0.55 + 0.67 - 1) = 1.25. At this time, the return value obtained by asynchronous reinforcement learning is larger, making the algorithm conducive to moving towards a better adjustment strategy.
[0091] The adaptive termination condition of the asynchronous reinforcement learning algorithm is designed as follows: After the algorithm finds the optimal value in this iteration, it compares it with the optimal value found in the previous iteration. If the optimal value found this time is better, then continue the iteration; if it is not better than the previous one, then check the number of iterations. If the number of iterations is greater than 2500 and no better target value is found in the last 300 iterations, then end the optimization.
[0092] The technical solution of the present invention will be described in detail below in conjunction with specific embodiments.
[0093] 1. The method for designing the structure of an asymmetric optical polarization device based on asynchronous reinforcement learning provided by the embodiment of the present invention includes: preprocessing of the simulation data set, construction and initialization of the transmittance prediction network, training of the transmittance prediction network, and optimization of the structure of the asymmetric polarization conversion device by the asynchronous reinforcement learning algorithm.
[0094] Preprocessing of the simulation data set
[0095] For the data set preprocessing part, the obtained device structure data is normalized between 0 and 1, and at the same time, the obtained 83-dimensional transmittance spectrum data is downsampled to 27-dimensional data at a ratio of 3:1.
[0096] The simulation data set preprocessing module collects a total of 83 transmittance data in the T yy , T yx two directions from 430nm to 550nm, and then performs a 3:1 downsampling operation on the 83-dimensional data. Finally, each group of data retains 27 valid data, which is convenient for subsequent feature extraction and observation calculations of the data. At the same time, the data is evenly divided into 5 parts (if it cannot be evenly divided, the number of the last part will be reduced or increased as appropriate) for subsequent 5-fold verification.
[0097] Construction and optimization of the transmittance prediction network
[0098] The transmittance prediction network among them is as Figure 3 shown, where BatchNormalization - batch normalization layer, Dense - fully connected layer, 144 - 144 hidden neurons, Relu - Relu activation function, LeakyRelu - LeakyRelu activation function.
[0099] The transmittance prediction network based on the residual structure of the deep neural network can accept the structural parameters of the asymmetric light polarization conversion device as input data and predict the corresponding transmittance value in a very short time. Among all the modules of the transmittance prediction network, the basic units adopted are the fully connected layer and the Relu activation layer; followed by the batch normalization layer, and the residual structure is used as one of the paths for error backpropagation. At the same time, the basic structure of the network is as Figure 3 shown and divided into three parts: serial input SIN, predicted output PYN of transmittance T yy and predicted output PXN of transmittance T yx . Figure 3 As shown, the input is the structural parameters of the asymmetric light polarization conversion device, and the output is the transmittance attributes T yy and T yx of the asymmetric light polarization conversion device. In SIN, first, four fully connected layers Dense are adopted, and the Relu function (f(x) = max(0, x)) is used to activate all the fully connected layers, and then the batch normalization layer is used. In the predicted output PYN of transmittance T yy and the predicted output PXN of transmittance T yx , similarly, it is composed of the fully connected layer FC and the activation function Relu, and the residual structure is added in it. Finally, the LeakyRelu function:
[0100] is used for activation. Using the residual structure is beneficial to the error backpropagation and avoids gradient disappearance and gradient explosion.
[0101] Among them, after the serial input part SIN accepts the structure of the asymmetric light polarization conversion device as input data, it connects 4 fully connected layers and ends with batch normalization. The purpose is to better extract the features of the structure data, and a larger training rate parameter can be used, and more batch training data can be used at one time during training. After the batch normalization layer, the extracted feature data is used as the input of the PYN and PXN parts. PYN consists of 6 basic units and an output part. Each basic unit includes the fully connected layer Dense and the Relu activation layer. The number of hidden layer neurons contained in these basic units is 200, 500, 500, 500, 200, 28, and the activation layer function uniformly uses the Relu function. The final output part is first a fully connected layer of 28 neurons followed by a LeakyRelu activation output. PXN consists of 5 basic units and an output part. Each basic unit includes the fully connected layer Dense and the Relu activation layer. The number of hidden layer neurons contained in these basic units is 200, 500, 500, 200, 28, and the activation layer function uniformly uses the Relu function. The final output part is first a fully connected layer of 28 neurons followed by a LeakyRelu activation output.
[0102] Transmittance prediction network training
[0103] In the present invention, other parameters are initialized randomly. For the training of the network, 80% of the obtained simulation data set is used as the training set, and the remaining 20% is used as the test set. And a 5-fold cross-validation method is adopted to group and train and test all the data in turn.
[0104] The transmittance prediction network based on the residual structure of the deep neural network can accurately predict the corresponding T from the structure of the asymmetric optical polarization conversion device after training yy ,T yx The transmittance in two directions. The transmittance prediction network can be used to replace professional numerical simulation tools to simulate the transmittance of the asymmetric optical polarization conversion device.
[0105] Asynchronous reinforcement learning algorithm to optimize the structure of the asymmetric polarization conversion device
[0106] In the embodiment of the present invention, an asynchronous reinforcement learning algorithm is used to adaptively optimize the design of the structure of the asymmetric optical polarization conversion device. The process of designing the structure of the asymmetric optical polarization conversion device with adjustable four-dimensional parameters by the asynchronous reinforcement learning algorithm is as follows Figure 4 shown. The asynchronous reinforcement learning algorithm takes the transmittance prediction network as the proxy function, takes the structure of the asymmetric optical polarization conversion device as the independent variable, and at the same time the transmittance T yy ,T yx The mean value of is used as the target value to be optimized, and the structure of the asymmetric optical polarization conversion device is continuously adjusted by maximizing the optimization target value, and the optimal adjustment strategy is continuously explored during the iteration process. Until the asynchronous reinforcement learning algorithm converges finally, at this time the structure of the asymmetric optical polarization conversion device is the optimal.
[0107] Among them, the detailed flow chart of the asynchronous reinforcement learning algorithm is as follows Figure 5 shown. Initialize the maximum number of iterations, boundary conditions, and additional limiting conditions. Randomly initialize the structure of the asymmetric optical polarization device. When initializing, it is necessary to ensure that the structure parameters are within the preset range.
[0108] In the asynchronous reinforcement learning algorithm as shown in Figure 5 The transmittance T can be obtained from the structure of the asymmetric optical polarization device through the transmittance prediction network yy and T yx , and then through calculation The return value inside the asynchronous reinforcement learning algorithm can be further calculated. The specific calculation method is that when the numerical value of the structure of the asymmetric optical polarization device exceeds the limit, the return value r = -0.5 at this time. When T yy and T yxWhen the maximum value is less than 0.3, then r = exp(max(T)(T yy ) + max(T yx ) - 1) - 1. When the maximum values of T yy and T yx are both between 0.3 and 0.5, r = 0. When the maximum value of T yy is between 0.3 and 0.5 and T yx is greater than or equal to 0.5, or when the maximum value of T yx is between 0.3 and 0.5 and T yy is greater than or equal to 0.5, When the maximum values of T yy and T yx are both greater than 0.5, r = exp(max(T yy ) + max(T yx ) - 1). Taking max(T yy ) = 0.15 and max(T yx ) = 0.4 as an example, r = exp(0.15 + 0.4 - 1) - 1 = -0.36. Taking max(T yy ) = 0.55 and max(T yx ) = 0.67 as an example, r = exp(0.55 + 0.67 - 1) = 1.25. At this time, the return value obtained by asynchronous reinforcement learning is larger, making the algorithm conducive to moving towards a better adjustment strategy. At the same time, the adaptive termination condition of the asynchronous reinforcement learning algorithm designed in the present invention is: after the algorithm finds the optimal value in this iteration, it is compared with the optimal value found in the previous iteration. If the optimal value found this time is better, then continue the iteration; if it is not better than the previous one, then check the number of iterations. If the number of iterations is greater than 2500 times and no better target value is found in the last 300 iterations, then end the optimization.
[0109] II. Double - layer anisotropic metasurface
[0110] The novel device structure provided by the embodiment of the present invention - double - layer anisotropic metasurface. The structural unit of this metasurface includes: a pair of metal resonant rods and a sub - wavelength metal grating, and a dielectric layer is filled between the two layers.
[0111] For the x - polarized light incident perpendicularly, there is strong polarization conversion in this structure, which can effectively convert it into y - polarized light. For the y - polarized light incident perpendicularly, there is only a very weak polarization conversion function, and most of the incident light will pass through the metamaterial structure maintaining the original polarization direction, and asymmetric polarization conversion can be achieved.
[0112] The effects of the present invention will be further described below in combination with specific experimental data.
[0113] In this invention, an asymmetric polarization structure is selected as an example. It consists of a double-rod resonator (aluminum) and a tangent structure (aluminum) converter structure. The two-layer metasurface structure is separated by a dielectric material. The structures to be optimized are the length and width of the double rods, the thickness of the dielectric layer, and the refractive index of the dielectric layer material. After optimizing the above-mentioned structure and material parameters through a reinforcement learning algorithm combined with a deep neural network, the optimal structural material property parameters can finally be obtained as: {218nm, 53nm, 348nm, 1.314}. The optimal average refractive index can reach 60.5%. Compared with the previous baseline of 50%, the optimization method of this invention can increase it by 21%. At the same time, the entire optimization time can be controlled within about 45 minutes, which can greatly reduce the time compared with traditional manual optimization. At the same time, the transmittance prediction network has a prediction accuracy of 96.6% and 95.5% for T yy and T yx , and can complete the prediction of the transmittance property of the asymmetric optical polarization device within milliseconds on a home-grade computer, without the need for simulation on an expensive server, thus replacing the cumbersome FDTD simulation.
[0114] In the description of this invention, unless otherwise specified, the meaning of "a plurality of" is two or more; the orientation or positional relationship indicated by terms such as "upper", "lower", "left", "right", "inner", "outer", "front end", "back end", "head", "tail", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to this invention. In addition, terms such as "first", "second", "third", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0115] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, such as provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and their modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software such as firmware.
[0116] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.
Claims
1. A method for designing the structure of an asymmetric optical polarization device based on asynchronous reinforcement learning, characterized in that The method for designing the structure of an asymmetric optical polarization device based on asynchronous reinforcement learning includes: Step 1, preprocessing the simulation data set of the device structure; Step 2, building and initializing the transmittance prediction network; Step 3, training the transmittance prediction network; Step 4, optimizing the structure of the asymmetric polarization conversion device by the asynchronous reinforcement learning algorithm; In the above Step 1, the specific process of preprocessing the simulation data set is as follows: For the data set preprocessing part, the obtained device structure data is normalized to be between 0 and 1, and at the same time, the obtained 83-dimensional transmittance spectrum data is downsampled to 27-dimensional data through a ratio of 3:
1. In the above Step 2, the specific process of building and optimizing the transmittance prediction network is as follows: The transmittance prediction network based on the deep neural network with a residual structure takes the structural parameters of the asymmetric optical polarization conversion device as input data and predicts the corresponding transmittance value in a very short time. Among all the modules of the transmittance prediction network, the basic units adopted are fully connected layers and Relu activation layers; followed by a batch normalization layer, and at the same time, a residual structure is used as one of the paths for error backpropagation. The basic structure of the transmittance prediction network is divided into three parts: series input SIN, predicted output PYN of transmittance T yy and predicted output PXN of transmittance T yx ; The input is the structural parameters of the asymmetric optical polarization conversion device, and the output is the transmittance property T of the asymmetric optical polarization conversion device yy and T yx , in SIN, four fully connected layers Dense are first adopted, and the Relu function f(x) = max(0, x) is used to activate all the fully connected layers, and then the batch normalization layer is used; in the predicted output PYN of the transmittance T yy and the predicted output PXN of the transmittance T yx , it is also composed of the fully connected layer FC and the activation function Relu, and the residual structure is added therein, and finally the LeakyRelu function is used: Activate; using a residual structure is beneficial to the backpropagation of errors and avoids gradient vanishing and gradient explosion.
2. The method for designing the structure of an asymmetric optical polarization device based on asynchronous reinforcement learning according to claim 1, wherein, The series input part SIN connects 4 fully connected layers after receiving the structure of the asymmetric optical polarization conversion device as input data, and ends with batch normalization, extracts the features of the structure data, uses a relatively large training rate parameter, and uses more batch training data at one time during training. After the batch normalization layer, the extracted feature data is used as the input for both the PYN and PXN parts. PYN consists of 6 basic units and an output part. Each basic unit includes a fully connected layer Dense and a Relu activation layer. The number of hidden layer neurons contained in these basic units is 200, 500, 500, 500, 200, 28, and the activation layer function uniformly uses the Relu function; finally, the output part is first a fully connected layer of 28 neurons followed by a LeakyRelu activation output. PXN consists of 5 basic units and an output part. Each basic unit includes a fully connected layer Dense and a Relu activation layer. The number of hidden layer neurons contained in these basic units is 200, 500, 500, 200, 28, and the activation layer function uniformly uses the Relu function; finally, the output part is first a fully connected layer of 28 neurons followed by a LeakyRelu activation output.
3. The method for designing the structure of an asymmetric optical polarization device based on asynchronous reinforcement learning according to claim 1, wherein In the above Step 3, the specific process of training the transmittance prediction network is as follows: For the training of the network, 80% of the obtained simulation data set is used as the training set, and the remaining 20% is used as the test set, and a 5-fold cross-validation method is adopted to group and train and test all the data in turn.
4. The method for designing the structure of an asymmetric optical polarization device based on asynchronous reinforcement learning according to claim 1, wherein In the above Step 4, optimizing the structure of the asymmetric polarization conversion device by the asynchronous reinforcement learning algorithm is specifically as follows: An asynchronous reinforcement learning algorithm is used to adaptively optimize the structure of an asymmetric optical polarization conversion device. The asynchronous reinforcement learning algorithm takes the transmittance prediction network as an agent function, and the structure of the asymmetric optical polarization conversion device as an independent variable. At the same time, the mean value of the transmittance T yy ,T yx is used as the target value to be optimized. The structure of the asymmetric optical polarization conversion device is continuously adjusted by maximizing the optimization target value, and the optimal adjustment strategy is continuously explored during the iteration process; until the asynchronous reinforcement learning algorithm converges finally, at this time the structure of the asymmetric optical polarization conversion device is the optimal one.
5. The method for designing the structure of an asymmetric optical polarization device based on asynchronous reinforcement learning according to claim 3, wherein The specific asynchronous reinforcement learning algorithm is as follows: After initializing the maximum number of iterations, boundary conditions, and additional constraint conditions; randomly initialize the structure of the asymmetric optical polarization device; when initializing, make the structure parameters within the preset range; In the asynchronous reinforcement learning algorithm, the transmittance T is obtained from the structure of the asymmetric optical polarization device through the transmittance prediction network yy and T yx , and then through calculation the return value inside the asynchronous reinforcement learning algorithm is further calculated; The specific calculation method is as follows: when the numerical value of the asymmetric optical polarization device structure exceeds the limit, the return value r = -0.5 at this time; when the maximum value of T yy and T yx is less than 0.3, r = exp(max(T yy ) + max(T yx ) - 1) - 1 at this time. When the maximum values of both T yy and T yx are between 0.3 and 0.5, r = 0. When the maximum value of T yy is between 0.3 and 0.5 and at the same time T yx is greater than or equal to 0.5, or when the maximum value of T yx is between 0.3 and 0.5 and at the same time T yy is greater than or equal to 0.5, when the maximum values of both T yy and T yx are greater than 0.5, r = exp(max(T yy ) + max(T yx ) - 1).
6. The method for designing the structure of an asymmetric optical polarization device based on asynchronous reinforcement learning according to claim 5, characterized in that, The adaptive termination condition of the asynchronous reinforcement learning algorithm is designed as: After finding the optimal value in this iteration, it is compared with the optimal values found in previous iterations. If the optimal value found in this iteration is better, then continue the iteration; If the number of iterations is greater than 2500 and no better objective value is found in the last 300 iterations, end the optimization.
7. An asymmetric optical polarization device structure using the asymmetric optical polarization device structure design method based on asynchronous reinforcement learning according to any one of claims 1 to 6, characterized in that, The asymmetric optical polarization device structure is provided with a double-layer anisotropic metasurface, and the double-layer anisotropic metasurface includes: a pair of metal resonant rods and a sub-wavelength metal grating, and a dielectric layer is filled between the two layers of structures.
Citation Information
Patent Citations
RESURF power device structure automatic optimization method based on device performance
CN111428422A