An infrared stealth film design method based on deep reinforcement learning and multilayer perceptron
Patent Information
- Application Number
- CN202311464886.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-11-06
AI Technical Summary
[0030]第一、本发明采用了逆向设计微纳光学薄膜的方法,由目标光谱结果来设计薄膜结构,与传统正向设计相比,能得到更多与先验结构不同的优良薄膜结构。
Smart Images

Figure CN117369125B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of micro-nano structure design technology, specifically relating to an infrared stealth optical thin film structure design method, which can be used for the design and fabrication of infrared stealth materials. Background Technology
[0002] With the rapid development of science and technology, infrared detection technology has reached a fairly high level. Optoelectronic imaging satellites can obtain visible light and infrared images with a resolution of 0.1m and can photograph ground targets under completely dark conditions. Traditional infrared stealth design faces severe challenges in multi-path target detection and multi-functional compatibility. Therefore, it is of great significance to study optical micro-nano infrared stealth materials that can achieve infrared radiation characteristic selection and fine structure.
[0003] The infrared radiation manipulation of micro / nano optical structures can be achieved through the interference principle of multilayer films. The maximum and minimum values of the reflectivity of the multilayer film structure can be calculated using the transmission matrix. The material radiation characteristics required for infrared stealth are mainly reflected in three bands: 3µm–5µm, 5µm–8µm, and 8µm–14µm. Among them, the low emissivity of the infrared atmospheric window in the 3µm–5µm and 8µm–14µm bands is used to achieve thermal camouflage; the high emissivity in the 5µm–8µm band is used to achieve heat dissipation (according to Kirchhoff's thermal radiation law, this is equivalent to the low absorptivity in the 3µm–5µm and 8µm–14µm bands and the high absorptivity in the 5µm–8µm band).
[0004] In 2017, Peng and Liu proposed a multilayer selective thermal emitter composed of an ultrathin silver / germanium layer. Employing a forward design, it simultaneously achieves good stealth performance and high heat dissipation. The film exhibits low emissivity (0.18 in the 3µm–5µm range and 0.31 in the 8µm–14µm range) and high emissivity (5µm–8µm) for radiative cooling. The multilayer structure is scalable, enabling large-area and flexible fabrication, providing significant advantages for the proposed selective emitter and showing great promise in the field of infrared stealth. However, this method also has significant drawbacks: First, as a forward design method, it requires manual selection of the initial structure and incurs substantial computational costs to obtain the target spectral line, making the design process quite challenging. Second, forward design may yield a locally optimal solution. Given the same materials, a structure with superior stealth and heat dissipation performance may exist but remains undiscovered. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of the existing technologies by proposing an infrared stealth film design method based on deep reinforcement learning and multilayer perceptron. This method fully utilizes the learning capabilities of neural networks and the exploratory capabilities of reinforcement learning to reverse-engineer micro-nano optical structures based on the target ideal spectral results, thereby achieving automated design of micro-nano structures and avoiding the problems of difficulty in forward design and the tendency to get trapped in local optima.
[0006] The technical solution of this invention is implemented as follows:
[0007] I. Technical Principles
[0008] This invention first constructs the multilayer membrane structure to be designed and selects the membrane material. Since membrane structures of different thicknesses have different absorptivity curves, a large amount of absorptivity data of the membrane structure within the target wavelength range is collected for training a multilayer perceptron (MLP) to obtain an accurate absorption spectrum prediction model. Then, a reinforcement learning environment is constructed based on the requirements of infrared stealth for band absorptivity. The reward and penalty values are set by the prediction structure of the MLP, and reinforcement learning is used to adjust the thickness of the multilayer membrane structure, updating the network based on the feedback. Finally, the reinforcement learning algorithm is used to train and iterate in the environment, and the convergence result of the output membrane structure is observed, realizing the design of an infrared stealth optical thin film.
[0009] II. Technical Solution:
[0010] Based on the above principles, the implementation steps of this invention include the following:
[0011] (1) Design n layers of micro-nano optical thin films composed of selected materials. Ports p1 and p2 are set at the top and bottom of the film, respectively, to obtain the reflection spectrum R(λ) and absorption spectrum A(λ) of light, where λ is the wavelength and n≥2;
[0012] (2) Set the parameterized scan step size for each thin film layer to d. i According to step size d i A parameterized scan of the n-layer optical thin film in the mid-infrared wavelength domain is performed to obtain N sets of data. These data are then preprocessed into training sets to obtain a training dataset, where i = 1, 2, ..., n.
[0013] (3) Establish a model with input parameters of n film thicknesses, and output the corresponding absorption spectra A(λ2), A(λ3), ..., A(λ2). i ),...,A(λ 14 A multilayer perceptron, where i = 2, 3, ..., 14 correspond to wavelengths from 2µm to 14µm respectively;
[0014] (4) Based on the training dataset, the backpropagation method is used to train the neural network to obtain a model that can accurately predict absorption spectral lines.
[0015] (5) Constructing a reinforcement learning environment:
[0016] (5a) Initialize a space of size 2 n The discrete space is used to represent the action space, and the n-dimensional action space is mapped to one dimension;
[0017] (5b) At the beginning of each iteration, initialize the thickness of each layer of the film to a random number in the parameter domain;
[0018] (5c) Receive the input action a from the current environment and return the current state s, the next state s', and the reward / penalty value r after the action is executed.
[0019] (6) Thickness adjustment of the thin film is performed using a deep Q-network algorithm in a configured reinforcement learning environment:
[0020] (6a) Set the number of iterations e and the maximum number of execution steps in each iteration, and establish a memory bank with a defined capacity;
[0021] (6b) At each step, the evaluation network of the deep Q network algorithm performs the operation a to increase or decrease the thickness of each film layer; and obtains the current state s, the next state s', and the reward / penalty value r after the action is performed from the environment, and stores the quadruple (s,a,r,s′) in the memory bank.
[0022] (6c) Determine if the memory bank is full:
[0023] If the memory bank is not full, proceed to step (6f);
[0024] If the memory bank is full, the quadruple in (6b) is used to overwrite the quadruple already stored in the memory bank, and (6d) is executed.
[0025] (6d) Sample the data stored in the memory bank to obtain batch sample data, and use the batch sample data to calculate the evaluation value q_eval and the target value q_target using the deep Q network algorithm;
[0026] (6e) Substitute q_eval and q_target into the error function, update the network parameters of the evaluation network using the backpropagation method, and then update the target network using the updated evaluation network.
[0027] (6f) Repeat (6b) to (6e) a total of step times.
[0028] (6g) Repeat (6f) until the set number of iterations e is reached, and output the thickness of each film layer.
[0029] Compared with existing technologies that use traditional methods for designing infrared stealth micro / nano optical thin films, this invention has the following significant advantages:
[0030] First, this invention employs a reverse design method for micro / nano optical thin films, which designs the thin film structure based on the target spectral results. Compared with traditional forward design, this method can obtain more superior thin film structures that differ from the prior structure.
[0031] Secondly, this invention employs reinforcement learning and neural networks for the automated design of target micro- and nano-structures. Specifically, it achieves automatic modulation of thin film structures by iterating through neural networks using input parameters. Compared to traditional methods, this approach involves less computation and has a simpler modulation method.
[0032] Third, this invention fully utilizes the exploration advantages of reinforcement learning in unknown environments and the learning ability of neural networks to design optical thin films, solving the problem that traditional methods are prone to getting trapped in local optima when designing optical thin film structures. By using reinforcement learning, multiple solutions that satisfy the target conditions can be found in the convergence results.
[0033] Fourth, this invention can set different parameters according to different target spectral characteristics to design selective emissivity materials, and has a wide range of applications. Attached Figure Description
[0034] Figure 1 This is a flowchart illustrating the implementation of the present invention;
[0035] Figure 2 This is a schematic diagram of the micro / nano optical thin film structure used in this invention;
[0036] Figure 3 This is a schematic diagram of the multilayer perceptron constructed in this invention;
[0037] Figure 4 This is a graph showing the change in the loss function during training of the multilayer perceptron in this invention.
[0038] Figure 5 This is a network framework diagram of the deep Q-network algorithm used in this invention;
[0039] Figure 6 This is a graph showing the variation of the cumulative reward value R in a single round with the number of iterations when using the deep Q-network algorithm to modulate a thin film in this invention.
[0040] Figure 7 The image shows the absorption spectrum of the micro / nano structure designed using this invention. Detailed Implementation
[0041] The embodiments and effects of the present invention will be further described in detail below with reference to the accompanying drawings.
[0042] Reference Figure 1 The steps in this example are as follows:
[0043] Step 1: Construct micro / nano optical thin films.
[0044] Reference Figure 2 Design an n-layer micro / nano optical thin film composed of a selected material. The film consists of n infinitely extended layers of a defined thickness, with thicknesses l1, l2, ..., ln from bottom to top. i ,...,l n , i = 1, 2, ..., n;
[0045] Ports p1 and p2 are set at the top and bottom of the film, respectively, to obtain the absorption spectrum A(λ) of the material at the mid-infrared wavelength. The step size of λ is selected as 1 μm or 0.5 μm according to the required accuracy of the target absorption spectrum.
[0046] In this embodiment, n=4 is set, but not limited to, Ag material is used for the thin film substrate and the third layer, Ge material is used for the second layer and the top layer, and the step size of λ is set to 1µm.
[0047] Step 2: Perform parametric scanning on the optical thin film and group preprocessing to obtain the training set.
[0048] 2.1) Determine the parameter domain for the thickness of each layer [p] i ,q i ] and parameterized scan step size d i Parameterized scan step size d i The settings should meet For the thicknesses l1, l2, l of the optical thin film i ,...,l n Parametric scanning was performed to obtain the absorption spectrum A(λ) in the wavelength range of 2–14 μm, where λ = 2, 3, 4, ..., 14.
[0049] In this embodiment, n=4, the parameter scan range of parameters l1 and l3 is 5~40nm, the step size d1=d3=5nm, and the number of parameters... The parameter ranges for l2 and l4 are 300–700 nm, with a step size of d2 = d4 = 50 nm and a parameter quantity of [missing information].
[0050] 2.2) The N sets of data obtained after parameterized scanning are grouped into one group, that is, the 13 absorbance data corresponding to 2um to 14um under the same thin film structure are grouped into one group to obtain the training dataset.
[0051] In this embodiment, N = n1 * n2 * n3 * n4 * 13 = 67392. After grouping preprocessing, 67392 / 13 = 5184 training sets are obtained, and each training set represents the absorption spectrum A under a thin film structure. i(λ),i=1,2,...,5184,λ=2,3,...,14.
[0052] Step 3: Establish a multilayer perceptron.
[0053] A multilayer perceptron is constructed, consisting of an input layer, intermediate hidden layers, and an output layer, wherein the input layer has a thin film thickness of l1, l2, l... i ,...,l n The output layer of the sensor is the absorption spectrum A(2), A(3), ..., A(14) of the thin film of this thickness;
[0054] The thicknesses of each thin film layer are l1, l2, l i ,...,l n The input to the neurons in the input layer undergoes linear processing and is then transmitted to the neurons in the intermediate hidden layers. These intermediate hidden layers then sequentially perform linear and nonlinear processing on each layer thickness before transmitting the data to the output layer. Finally, the output layer outputs the absorption spectrum of the thin film, such as... Figure 3 As shown;
[0055] In this embodiment, n=4, the perceptron has five hidden layers, and the number of neurons in the hidden layers is set to 400; the ReLU activation function is used between layers for nonlinear processing in the hidden layers.
[0056] Step 4: Train the multilayer perceptron.
[0057] 4.1) Configure the iterator through the dataloader function, input the training set obtained in step 2 into the iterator, and set the batch_size of the iterator to 128;
[0058] 4.2) The mean square error function (MSE) is selected as the error function of the neural network;
[0059] 4.3) Obtain batch data of a single training round from the training set through an iterator, input these batch data into a multilayer perceptron to obtain the estimated value of the absorption spectrum, and substitute the estimated value and the actual value into the error function MSE to calculate the loss value estimated by the neural network. Use the Adam optimizer with a learning rate of 0.001 to minimize the loss value and update the weights and biases of the neural network.
[0060] 4.4) Repeat step 4.3) to continuously decrease the loss function with the number of iterations until the loss value is less than 0.0005, such as... Figure 4 As shown, a trained multilayer perceptron is obtained.
[0061] Step 5: Build a reinforcement learning environment.
[0062] 5.1) Initialize the action space and perform action mapping within the action space:
[0063] Initialize an n-dimensional array [a0, a1, a2, a...] i ,...,a n-1 ] is used to indicate that the number of actions is 2. n an n-dimensional action space, where a i This indicates that an operation to increase or decrease the thickness of the i-th thin film layer is performed, a i =±d i / 5, i = 0, 1, 2, ..., n-1;
[0064] 2 of the n-dimensional action space n Each action is mapped to a one-dimensional number of elements: 0, 1, 2, ..., 2. n -1 is a discrete number;
[0065] In this embodiment, n = 4, a1 = a3 = ±1, a2 = a4 = ±10;
[0066] 5.2) Before the start of each iteration, initialize the structure of the thin film, and initialize the thickness of each layer to a random number within the parameter domain;
[0067] In this embodiment, l1 is a random number between 5 and 40, l2 is a random number between 300 and 700, l3 is a random number between 5 and 40, and l4 is a random number between 300 and 700.
[0068] 5.3) Receive the input action 'a' from the current environment and return three parameters: the current state 's', the next state 's', and the reward / penalty value 'r' after the action is executed.
[0069] 5.3.1) Based on the current state s=[x0,x1,x2,...,x i ,...,x n-1 ], calculate the state at the next time step s′=s+a=[x′0,x′1,x′2,...,x′ i ,...,x′ n-1 ], where x i Let x′ represent the thickness of the i-th thin film at the current moment. i This represents the thickness of the i-th thin film at the next moment; in this embodiment, n = 4;
[0070] 5.3.2) Let x′ i With the thickness range of each thin film [p i ,q i Compare,
[0071] If x′ i <p i Let x′i =q i ;
[0072] If x′ i >q i Let x′ i =p i ;
[0073] In this embodiment, p1 = p3 = 5, q1 = q3 = 40, p2 = p4 = 300, and q2 = q4 = 700.
[0074] 5.3.3) Based on the absorption spectrum characteristics of infrared stealth, the reward / penalty function r is set as:
[0075] r=200(A(6)+A(7))-100(A(2)+A(3)+A(4)+A(9)+A(10)+A(11)+A(12)+A(13)+A(14))
[0076] In the formula, A(t) represents the absorption rate of the current thin film structure at wavelength λ in the λ = t (um) band. It is obtained by a trained multilayer perceptron, that is, the state s' of the next moment is input into the trained multilayer perceptron and the output value is A(t)t = 2, 3, 4, ..., 14; t ≠ 5; t ≠ 8.
[0077] Step 6: Use the deep Q-network algorithm to adjust the thickness of the thin film in the configured reinforcement learning environment.
[0078] like Figure 5 As shown, the deep Q-network algorithm consists of two networks with the same structure but different purposes: an evaluation network and a target network. The input parameter of both networks is the thickness of each membrane layer, and the output parameter is the Q value corresponding to each action in the action space. Each network has a hidden layer in the middle, and its activation function is the ReLU function.
[0079] The steps for adjusting the film thickness using this deep Q-network algorithm include the following:
[0080] 6.1) Set the batch size for selecting samples in a single run, the learning rate of the Adam optimizer, the greed coefficient ε, the reward decay value γ, the target network update frequency iter, the number of updates of the evaluation network c, the number of exploration steps in a single round step, the number of iteration learning rounds episode, and set the cumulative reward value in a single round R.
[0081] Establish a memory with a capacity of [capacity].
[0082] In this embodiment, batch_size = 128, learning_rate = 0.005, ε = 0.9, γ = 0.9, iter = 100, step = 200, episode = 600, capacity = 2000, R = 0;
[0083] 6.2) At each step, the evaluation network of the deep Q-network algorithm performs the operation of increasing or decreasing the thickness of each film layer: a:
[0084] 6.2.1) Input the current thickness of each film layer into the evaluation network of the depth Q-network algorithm, and output a total of 2... n The Q value of each action;
[0085] 6.2.2) Generate a random number i between 0 and 1, and compare it with the greedy coefficient ε to determine action a:
[0086] If i < ε, then action a is 2. N The action corresponding to the largest Q value among all Q values;
[0087] If i > ε, then action a is 2. n One action is randomly selected from a set of actions;
[0088] 6.3) Obtain the current state s, the next state s', and the reward / penalty value r after performing the action from the environment, and store these data in the memory bank in the form of a quadruple (s, a, r, s'); Let R = R + r;
[0089] 6.4) Determine if the memory bank is full:
[0090] If the memory bank is not full, proceed to step (6.7);
[0091] If the memory bank is full, the quadruple in (6.3) is used to overwrite the quadruple already stored in the memory bank, and (6.5) is executed.
[0092] 6.5) Sample the data already stored in the memory bank to obtain batch sample data, and substitute this batch sample data into the deep Q-network algorithm to calculate the evaluation value q_eval and the target value q_target:
[0093] 6.5.1) Randomly select batch_size quadruplets (s, a, r, s′) from the memory bank, and form the corresponding batch sample data (bs, ba, br, bs′) from these quadruplets, where bs is an array of batch_size s, ba is an array of batch_size a, br is an array of batch_size r, and bs′ is an array of batch_size s′;
[0094] 6.5.2) Input bs into the evaluation network of the deep Q-network algorithm to obtain the evaluation value q_eval;
[0095] 6.5.3) Input bs' into the target network of the deep Q-network algorithm to obtain the intermediate value q_next;
[0096] 6.5.4) Substitute the intermediate value q_next into the following formula to obtain the target value q_taret:
[0097] q_target = br + γ * q_next;
[0098] 6.6) Substitute q_eval and q_target into the error function MSE, update the network parameters of the evaluation network using backpropagation, and then use the updated evaluation network to update the target network:
[0099] 6.6.1) Calculate the error function MSE of the evaluation value q_val and the target value q_target, minimize the error function MSE using the Adam optimizer, and update the weights and biases of the evaluation network;
[0100] 6.6.2) Let c = c + 1. If iter is divisible by c, execute 6.6.3);
[0101] 6.6.3) Assign the network parameters of the evaluation network to the target network, and update the network parameters of the target network;
[0102] 6.7) Update the current state to the state s′ of the next time step, i.e., s = s′;
[0103] 6.8) Repeat steps 6.2) to 6.7) a total of steps, and reset R to 0;
[0104] 6.9) Repeat 6.8) for a total of epochs, output the current state s, and obtain the thickness of each optical thin film, i.e. l1 = 40nm, l2 = 610nm, l3 = 10nm, l4 = 340nm.
[0105] The cumulative reward value R per epoch during the above training process varies with the number of epochs, as follows: Figure 6 As shown, Figure 6 The horizontal axis represents the number of iterations (epoch), and the vertical axis represents the cumulative reward value R in a single round.
[0106] The effects of this invention can be further illustrated by the following simulation experiments:
[0107] The film thicknesses obtained in this example, l1 = 40 nm, l2 = 610 nm, l3 = 10 nm, l4 = 340 nm, and the selected Ag materials for the first and third layers, and Ge materials for the second and fourth layers, were substituted into COMSOL 6.0 software for electromagnetic wave wavelength domain simulation using the wave optics module. This yielded the mid-infrared wavelength domain absorption spectrum of the designed optical thin film, as shown below. Figure 7 As shown, Figure 7 The horizontal axis represents wavelength, and the vertical axis represents absorptivity;
[0108] from Figure 7 It can be seen that the absorption spectrum of this optical film meets the requirements for infrared stealth, which has low absorption rates in the range of 3–5 μm and 8–14 μm and high absorption rates in the range of 5–8 μm.
[0109] The numbering of the steps described above is merely for the purpose of more clearly illustrating the implementation scheme of this example, and the order of the numbers is not limited. Furthermore, the above description is only a specific example of the present invention and does not constitute any limitation on the present invention. Obviously, those skilled in the art, after understanding the content and principles of the present invention, may make various modifications and changes in form and detail without departing from the principles and structure of the present invention. However, these modifications and changes based on the concept of the present invention are still within the scope of protection of the claims of the present invention.
Claims
1. A method for designing infrared stealth films based on deep reinforcement learning and multilayer perceptrons, characterized in that, The steps include the following: (1) Design n layers of micro-nano optical thin films composed of selected materials. Ports p1 and p2 are set at the top and bottom of the thin film, respectively, to obtain the absorption spectrum A(λ) of light, where λ is the wavelength and n≥2; (2) Set the parameterized scan step size for each thin film layer to d. i According to step size d i A parameterized scan of the n-layer optical thin film in the mid-infrared wavelength domain was performed to obtain N sets of data; The data are preprocessed by grouping them into training sets to obtain the training dataset, where i = 1, 2, ..., n (3) Establish a model with input parameters of n film thicknesses, and output the corresponding absorption spectra A(λ2), A(λ3), ..., A(λ2). i ),...,A(λ 14 A multilayer perceptron, where i = 2, 3, ..., 14 correspond to wavelengths from 2µm to 14µm respectively; (4) Based on the training dataset, the backpropagation method is used to train the neural network to obtain a model that can accurately predict absorption spectral lines. (5) Constructing a reinforcement learning environment: (5a) Initialize a space of size 2 n The discrete space is used to represent the action space, and the n-dimensional action space is mapped to one dimension; (5b) At the beginning of each iteration, initialize the thickness of each layer of the film to a random number in the parameter domain; (5c) Receive the input action a from the current environment and return the current state s, the next state s', and the reward / penalty value r after the action is executed. (6) Thickness adjustment of the thin film is performed using a deep Q-network algorithm in a configured reinforcement learning environment: (6a) Set the number of iterations e and the maximum number of execution steps in each iteration, and establish a memory bank with a defined capacity; (6b) At each step, the evaluation network of the deep Q network algorithm performs the operation a to increase or decrease the thickness of each film layer; and obtains the current state s, the next state s', and the reward / penalty value r after the action is performed from the environment, and stores the quadruple (s,a,r,s′) in the memory bank. (6c) Determine if the memory bank is full: If the memory bank is not full, proceed to step (6f); If the memory bank is full, the quadruple in (6b) is used to overwrite the quadruple already stored in the memory bank, and (6d) is executed. (6d) Sample the data stored in the memory bank to obtain batch sample data, and use the batch sample data to calculate the evaluation value q_eval and the target value q_target using the deep Q network algorithm; (6e) Substitute q_eval and q_target into the error function, update the network parameters of the evaluation network using the backpropagation method, and then update the target network using the updated evaluation network. (6f) Repeat (6b) to (6e) a total of step times. (6g) Repeat (6f) until the set number of iterations e is reached, and output the thickness of each film layer.
2. The method according to claim 1, characterized in that, The absorption spectrum A(λ) obtained in step (1) is selected with a step size of 1 μm or 0.5 μm based on the accuracy requirements of the target absorption spectral line.
3. The method according to claim 1, characterized in that, In step (2), the parameterized scan step size for each thin film is set to d. i This involves treating the thickness of each thin film layer as a parameter to determine the parameter domain of the thickness of each thin film layer in the target structure to be designed. i ,q i ], parameterized scan step size d i The settings should meet 4. The method according to claim 1, characterized in that, In step (2), the training set preprocessing of the N sets of data is to divide the 13 absorption rate data corresponding to 2um to 14um under the same thin film structure into one group, that is, to merge the original N sets of data into N / 13 groups to obtain the training dataset.
5. The method according to claim 1, characterized in that, The multilayer perceptron established in step (3) has an input dimension of the number of membrane layers and an output dimension of the number of nodes in the absorption spectrum. The ReLU activation function is used between layers.
6. The method according to claim 1, characterized in that, In step (4), the neural network is trained using backpropagation based on the training dataset. The steps include the following: (4a) Configure the iterator through the dataloader function, and set its batch_size parameter to 128; (4b) The mean square error function (MSE) is selected as the error function of the neural network; (4c) Obtain batch data of a single training round from the training set through an iterator, calculate the error function MSE between the actual and estimated values of this batch of data, use the Adam optimizer to minimize the error function MSE, and update the weights and biases of the neural network. (4d) Repeat (4c) until the error function is less than 0.0005, and you will get a model that can accurately predict absorption lines.
7. The method according to claim 1, characterized in that, Step (5a) initializes a space of size 2 n The discrete space is used to represent the action space, which is initialized with an n-dimensional array [a0, a1, a2, ..., a...]. i ,...,a n-1 ], where a i This indicates that an operation to increase or decrease the thickness of the i-th thin film layer is performed, a i =±d i / 5, i = 0, 1, 2, ..., n-1.
8. The method according to claim 1, characterized in that, Step (5c) receives the input action 'a' from the current environment and returns three parameters: the current state 's', the next state 's', and the reward / penalty value 'r' after the action is executed. The implementation steps include: (5c1) Based on the current state s = [x0, x1, x2, ..., x i ,...,x n-1 ], calculate the state at the next time step s′=s+a=[x′0,x′1,x′2,...,x′ i ,...,x′ n-1 ], where x i Let x′ represent the thickness of the i-th thin film in the current state. i This indicates the thickness of the i-th thin film at the next moment; (5c2) x′ i With the thickness range of each thin film [p i ,q i Compare, If x′ i <p i Let x′ i =q i ; If x′ i >q i Let x′ i =p i ; (5c3) Based on the absorption spectrum characteristics of infrared stealth, the reward / penalty function r is set as: r=200(A(6)+A(7))-100(A(2)+A(3)+A(4)+A(9)+A(10)+A(11)+A(12)+A(13)+A(14)) where A(t) represents the absorption rate of the current thin film structure at wavelength λ in the λ=t(um) band. It is obtained by inputting s' into a trained model that can accurately predict absorption spectral lines and outputting the value A(t), t=2,3,4,...,14;t≠5;t≠8.
9. The method according to claim 1, characterized in that, The deep Q-network algorithm used in step (6) consists of two networks with the same structure but different purposes: an evaluation network and a target network. The input parameter of both networks is the thickness of each membrane layer, and the output parameter is the Q value corresponding to each action in the action space. Each network has a hidden layer in the middle, and its activation function is the ReLU function.
10. The method according to claim 1, characterized in that, Each step in step (6b) involves the evaluation network of the deep Q-network algorithm performing an operation a to increase or decrease the thickness of each film layer. The steps include: (6b1) Set the greed coefficient ε, 0 < ε < 1 (6b2) Evaluation of the Deep Q-Network Algorithm: The network input is the current thickness of each film layer, and the output is a total of 2... n The Q value of each action; (6b3) The system generates a random number i between 0 and 1 and compares it with the greedy coefficient ε: If i < ε, then action a is 2. N The action corresponding to the largest Q value among all Q values; If i > ε, then action a is 2. n One action is randomly selected from a set of actions.
11. The method according to claim 1, characterized in that, Step (6d) samples the data already stored in the memory bank and calculates the evaluation value q_eval and the target value q_target using a deep Q-network algorithm. The steps include: (6d1) Set the batch size to batch_size = 64 and the reward decay value to γ = 0.9; (6d2) Randomly select batch_size quadruplets (s,a,r,s′) from the memory bank, and form the corresponding batch sample data (bs,ba,br,bs′) from these quadruplets, where bs is an array composed of batch_size s, ba is an array composed of batch_size a, br is an array composed of batch_size r, and bs′ is an array composed of batch_size s′; (6d3) Input bs into the evaluation network of the deep Q-network algorithm to obtain the evaluation value q_eval; (6d4) Input bs' into the target network of the deep Q-network algorithm to obtain the intermediate value q_next; (6d5) Substituting the intermediate value q_next into the following formula, we obtain the target value q_taret: q_target = br + γ * q_next.
12. The method according to claim 1, characterized in that, Step (6e) uses backpropagation to update the network parameters of the evaluation network, and then uses the updated evaluation network to update the target network. The steps include: (6e1) Select the mean squared error function (MSE) as the error function of the neural network, set the update frequency of the target network to iter, and set the update number of the evaluation network to c. (6e2) Calculate the error function MSE of the evaluation value q_val and the target value q_target, minimize the error function MSE using the Adam optimizer, and update the weights and biases of the evaluation network. (6e3) Let c = c + 1. When iter can divide c, execute (6e4); (6e4) Assign the network parameters of the evaluation network to the target network and update the network parameters of the target network.
Citation Information
Patent Citations
Method for designing infrared phase change material based on deep learning
CN112182965A
Deep neural network hyperparameter optimization method, electronic device and storage medium
WO2021007812A1