Super-lens inverse design optimization method based on reinforcement learning
Through collaborative exploration in a three-dimensional matrix maze within a multi-agent reinforcement learning framework, the problems of structural irregularity and three-dimensional step-by-step optimization in the inverse design of superlenses were solved, enabling efficient and precise superlens design to meet semiconductor manufacturing requirements.
Patent Information
- Application Number
- CN202511106513.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-08
AI Technical Summary
The structure generated by the generative neural network in the existing superlens inverse design technology is irregular, which affects manufacturability. In addition, the three-dimensional design requires step-by-step optimization, resulting in insufficient coordination, low efficiency and high computing power consumption.
A multi-agent reinforcement learning framework is adopted to achieve synchronous inverse design of the three geometric parameters of the superlens through collaborative exploration of a three-dimensional matrix maze. Combined with physical information constraints and reward function optimization, a structure that meets manufacturing requirements is generated.
It improves design accuracy and global optimization efficiency, reduces manufacturing difficulty and computing power consumption, and ensures that the designed superlens structure meets semiconductor process requirements.
Smart Images

Figure CN120597582A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a superlens inverse design optimization method, and in particular to a superlens inverse design optimization method based on reinforcement learning, belonging to the technical field of superlens inverse design. Background Art
[0002] As a new type of optical device with great application potential, superlens relies on micro-nanoscale structures to achieve efficient control of light fields, showing important value in cutting-edge fields such as imaging, sensing, and quantum technology. The current development of superlens inverse design technology faces two types of technical challenges that need to be overcome: (1) Optimization space of generative neural network methods: Existing solutions often use generative neural networks such as generative adversarial networks to construct the mapping relationship between superlens structure and physical properties. In practical applications, there are two points that can be optimized: first, the generated structure often has irregular geometric shapes, which may affect the manufacturability of the device; second, it is necessary to pre-construct a large-scale random data set. Such data is prone to insufficient correlation, which leads to limited model convergence efficiency and high requirements for computing resources. (2) Dimensional design constraints of the adjoint method: The adjoint method can achieve efficient convergence through the gradient optimization mechanism, but it is limited by the theoretical framework and is currently mainly applicable to the design of two-dimensional top-view superlenses. When expanded to three-dimensional design scenarios, it is necessary to optimize the parameters of different dimensions step by step. This model may lead to insufficient synergy in the optimization of each dimension, and there is room for improvement in balancing device manufacturability and multi-dimensional performance indicators. Summary of the Invention
[0003] The purpose of the present invention is to provide a reinforcement learning-based inverse design optimization method for a metalens. By integrating a multi-agent reinforcement learning collaborative exploration mechanism with physical information constraints, this method addresses the uncontrollable nature of generative neural network structures, poor data correlation, and the step-by-step design required for three-dimensional optimization in adjoint methods. This method enables the simultaneous inverse design and optimization of the three geometric parameters of a metalens.
[0004] The technical solution of the present invention is a method for inverse design optimization of a metalens based on reinforcement learning, which specifically includes the following steps: Step 1: constructing a nanostructure library consisting of nanostructures with different geometric parameters, and arranging the nanoring structures in the nanostructure library in a three-dimensional matrix maze according to the values of the geometric parameters; each coordinate in the three-dimensional matrix maze corresponds to a unique nanostructure; Step 2: Establishing a physical information multi-agent reinforcement learning framework: Using the three-dimensional matrix maze as the agent exploration environment, the matrix maze is divided into multiple two-dimensional maze slices along the dimension representing the radius of the ring nanostructure. Multiple agents are distributed on the multiple maze slices, and the agents achieve synchronous exploration of geometric parameters by interacting with the reinforcement learning environment. Step 3: Set up a reward function that includes the reward value of a single agent and the global reward value. Use angular spectrum transmission to calculate the axial light intensity distribution of the metalens as the basis for performance evaluation. When the reward value of the reward function reaches the preset target or converges, output the design result of the metalens composed of the nanostructure selected by the agent.
[0005] In the aforementioned reinforcement learning-based superlens inverse design optimization method, in step one, the nanostructure is a nanoring; the geometric parameters include three dimensions: radius, coarse width, and fine width; the radius range is 250nm-49750nm; the coarse width range is 100nm-500nm; the fine width is 0nm-16nm, and the fine width is used to widen the coarse width; the parameter dimensions of the three-dimensional matrix maze are arranged as 100×11×5; the coordinates also include phase modulation information and transmittance information corresponding to the nanoring structure, wherein the phase modulation information is used as a label to reduce data processing complexity.
[0006] In the aforementioned reinforcement learning-based superlens inverse design optimization method, in step 2, the physical information includes phase modulation information and transmittance information of the nanostructure; the physical information multi-agent reinforcement learning framework uses a three-dimensional matrix maze as the exploration environment and divides it into 100 two-dimensional 11×5 slices along the radius dimension. 100 agents are distributed on these 100 two-dimensional slices, and the movement and observation of geometric parameters are achieved by interacting with the reinforcement learning environment; the physical information multi-agent reinforcement learning framework adopts a multi-agent proximal strategy optimization algorithm.
[0007] In the aforementioned reinforcement learning-based superlens inverse design optimization method, the actions of the agent during movement include movement operations and phase gradient control operations; the observation values during the agent observation process include the current coordinates of the current agent, the previous coordinates of the current agent, the coordinates of the previous agent, and the phase label corresponding to the coordinates of the current agent.
[0008] In the aforementioned reinforcement learning-based superlens inverse design optimization method, in step 3, the reward function of the single agent reward value is: ; Where, is the local reward value; is the distance between the agent’s current exploration coordinate and the initialization coordinate, is the coefficient used to control the exploration range.
[0009] In the aforementioned reinforcement learning-based metalens inverse design optimization method, the reward function of the global reward value is: ; Where, is the global reward value, is the final peak reward, is the non-peak interval flattening penalty term, is the non-peak penalty term in the non-peak interval.
[0010] In the aforementioned reinforcement learning-based metalens inverse design optimization method, the non-peak interval flattening penalty term The calculation formula is: ; Where, represents the sum of all numbers in the non-target peak area; The non-peak interval non-peak penalty item The calculation formula is: ; Where, Indicates the total number of peaks that appear in the peak interval; The final peak reward The calculation formula is: ; Where, Represents the total average light intensity in the focal area, Indicates the focus intensity ratio compliance, It is a single focus confirmation reward.
[0011] In the aforementioned reinforcement learning-based metalens inverse design optimization method, the sum of the average light intensity in the focal area is The calculation formula is: ; Where, and Represent two peak interval elements respectively; The focus intensity ratio conformity The calculation formula is: ; Where, represents the proportional error; The single focus confirmation reward The calculation formula is: .
[0012] Compared with the prior art, the present invention has the following beneficial effects: This invention overcomes the efficiency bottleneck of traditional adjoint methods, which require step-by-step optimization of three-dimensional parameters (fixing two dimensions while designing the third). Through collaborative exploration by multiple agents within a three-dimensional matrix maze, it achieves simultaneous inverse design of the three geometric dimensions of the metalens, avoiding the local optimality issues caused by step-by-step optimization and improving global optimization efficiency and design accuracy. This invention uses a three-dimensional matrix maze to strictly constrain the parameters of the nanorings within a process-feasible range, overcoming the drawback of generative neural networks in generating irregular structures. This ensures that the designed metalens structure meets semiconductor manufacturing process requirements and reduces manufacturing difficulty and cost. This invention eliminates the generative method's reliance on large, low-correlation random datasets and instead utilizes reinforcement learning agents to actively explore valuable parameter spaces. By memorizing and learning from historical exploration results, it reduces inefficient computations, significantly reducing computing power consumption and accelerating optimization convergence compared to traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a schematic diagram of a three-dimensional matrix maze of the present invention; Figure 2 A schematic diagram of phase distribution scanning corresponding to the nanostructure of the present invention; Figure 3 A schematic diagram of a transmittance distribution scan corresponding to the nanostructure of the present invention; Figure 4 This is the evaluation function training graph for trifocal lenses; Figure 5 Schematic diagram of the phase and transmittance distribution of the metalens designed for the present invention; Figure 6 Schematic diagram of the cross-section of the lens radius designed for the present invention; Figure 7 This is the simulated light intensity axial distribution diagram of the present invention. DETAILED DESCRIPTION
[0014] The present invention will be further described below with reference to the accompanying drawings and examples, but they are not intended to limit the present invention.
[0015] Embodiment: A method for inverse design optimization of a metalens based on reinforcement learning, specifically comprising the following steps: Step 1: constructing a nanostructure library consisting of nanostructures with different geometric parameters, and arranging the nanoring structures in the nanostructure library in a three-dimensional matrix maze according to the values of the geometric parameters; each coordinate in the three-dimensional matrix maze corresponds to a unique nanostructure; In this step, a structure library consisting of manufacturable nanoring structures was first generated and obtained using FDTD software, such as Figure 1As shown, each nanoring has three arbitrary geometric dimensions: radius, width, and height, or radius, coarse width, and fine width. These values can be selected based on the target device function and processing technology. For example, in our trifocal lens design, we planned to use a high-precision EBL process to fabricate the metalens. Therefore, the three dimensions of the matrix maze were radius, coarse width, and fine width. The radius ranged from 250 nm to 49,750 nm; the coarse width ranged from 100 nm to 500 nm; and the fine width, which was used to increase the coarse width, ranged from 0 nm to 16 nm. Based on the differences in these three geometric design dimensions between nanoring structures, the data for a total of 5,500 manufacturable nanoring structures were arranged in a three-dimensional matrix maze measuring 100 (radius) × 11 (width) × 5 (height). Since the position coordinates of each point in the three-dimensional matrix maze contain three values, a position coordinate in the matrix maze can represent the three geometric dimensions of the corresponding nanoring structure. At the same time, the coordinate also contains the phase modulation information and transmittance information of the corresponding nanoring structure. The phase modulation information is used as a label to reduce the complexity of data processing. By accessing this coordinate, the phase modulation information and transmittance information of the incident circularly polarized light obtained by simulation of each nanoring structure stored in the matrix maze can be obtained, such as Figure 2 and Figure 3 As shown in the figure, since the precise phase information of these nanostructures is obtained through simulation, after confirming that the phase response corresponding to each nanoring is unique, the corresponding nanoring geometric parameters can be reversely searched through the phase value. It is worth noting that the phase here only represents the label of the data. In the subsequent inverse design method, the proposed physical information multi-agent reinforcement learning framework uses a variety of explicit or implicit exploration information. Using phase as a label can reduce the complexity of data processing to a certain extent. Figure 1 This is a schematic diagram of a three-dimensional matrix maze. The grayscale of the squares represents the mapping of phase labels according to phase values. Translucent squares represent unselected structures, and opaque squares represent structures selected for constructing metalenses. Corresponding to the 100 radii of the designed metalens, 100 small squares will be selected.
[0016] Step 2: Establishing a physical information multi-agent reinforcement learning framework: Using the three-dimensional matrix maze as the agent exploration environment, the matrix maze is divided into multiple two-dimensional maze slices along the dimension representing the radius of the ring nanostructure. Multiple agents are distributed on the multiple maze slices, and the agents achieve synchronous exploration of geometric parameters by interacting with the reinforcement learning environment. In this step, the physical information includes phase modulation information and transmittance information of the nanostructure; the coordinates of the three-dimensional matrix maze are associated with the phase modulation information, and the phase modulation information is used as a label to reduce the complexity of data processing.
[0017] The physical information multi-agent reinforcement learning framework includes an environment module and an algorithm module. The basic idea of the physical information multi-agent reinforcement learning framework is to map the information required to solve the target problem into a program called an environment. Then the algorithm part of the reinforcement learning project will create and manage the program organism - the agent. The agent will actively interact with the environment according to the set goals to maximize the rewards. If the agent's behavior is conducive to achieving the target task, it will be rewarded, otherwise it will be punished. When the agent completes the number of actions set in each round, or performs an unallowed action, the state of the environment and the agent will be reset to the initial state, and the experience of the previous training round will be saved and analyzed through the neural network.
[0018] The environment model uses the 100x11x5 matrix maze mapped from the previous nanostructure library as the coordinate range of the locations that the agents need to explore the environment. Based on the dimension of the matrix maze being 100 in length, the matrix maze can be divided into 100 11x5 maze slices, and 100 agents are set up to move and observe on these 100 slices. agents, whose observations contain agents The current coordinates, the coordinates before this movement, the agent The coordinates of the agent The phase label corresponding to the current coordinate; its actions are moving one unit up, one unit down, one unit left, one unit right, no movement, and two particularly important actions that imply changing the phase gradient between adjacent structures, one length unit closer to the previous agent, and one length unit away from the previous agent.
[0019] The algorithm module uses the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm. This method sets up a discrete exploration environment in the algorithm part and uses a four-layer recurrent neural network to learn the global observation variables of the agents.
[0020] Through the configuration of algorithm and environment modules, each agent in the physical information multi-agent reinforcement learning framework can observe the state of all agents at each action and take actions based on the state of the previous agent. The global critic network and the actor network equipped with each agent respectively learn the state of all agents and the state of individual agents, and provide guidance for subsequent actions or actions to be taken during the next training session.
[0021] Among them, the physical properties of the superlens composed of the selected 100 nanostructures, that is, its radial light intensity distribution, are calculated through angular spectrum transmission. The calculation of angular spectrum transmission requires the phase distribution data and transmittance distribution data of the designed superlens.
[0022] Step 3: Set up a reward function that includes the reward value of a single agent and the global reward value. Use angular spectrum transmission to calculate the axial light intensity distribution of the metalens as the basis for performance evaluation. When the reward value of the reward function reaches the preset target or converges, output the design result of the metalens composed of the nanostructure selected by the agent.
[0023] In this step, adjustments need to be made based on the actual situation. The ultimate goal is to maximize the result of adding the reward value of a single agent to the global reward value. The larger the value, the more consistent it is with the design goal. Through continuous exploration and training through the program, when the return value stops increasing after multiple iterations of training or reaches the set target value, the result is considered to have converged and the program is stopped: The reward function includes a reward function for a single agent reward value and a reward function for a global reward value. The reward function for a single agent reward value is: ; Where, is the local reward value; is the distance between the agent’s current exploration coordinate and the initialization coordinate, is the coefficient used to control the exploration range, by controlling The size of can control the exploration range of each agent in the two-dimensional slice of the three-dimensional matrix maze to which it belongs.
[0024] The output of the metalens design during the reinforcement learning iteration process can calculate its corresponding radial light intensity distribution through angular spectrum transmission. This distribution is a sequence of 300 numbers. The position of the number in the sequence represents the distance between the metalens and the metalens in the axial direction (the 1st to 300th numbers represent the distance between 0 and 100 microns), and the size of the number represents the light intensity.
[0025] Assume that we need to design an axially distributed equal-intensity trifocal metalens with focal lengths of 30, 60, and 90 μm, respectively, and a target light intensity distribution ratio of 1:1:1. The reward function setting principle is the same.
[0026] Therefore, the 10 points near 90 and 180 in the 300-digit sequence output by the angular spectrum transmission (these numbers represent light intensity, all greater than or equal to 0) are set as the target peak interval.
[0027] The reward function of the global reward value is: ; Where, is the global reward value, is the final peak reward, is the non-peak interval flattening penalty term, is the non-peak penalty term in the non-peak interval.
[0028] The non-peak interval flattening penalty term The calculation formula is: ; Where, represents the sum of all numbers in the non-target peak area; The non-peak interval non-peak penalty item The calculation formula is: ; Where, Indicates the total number of peaks that appear in the peak interval; The final peak reward The calculation formula is: ; Where, Represents the total average light intensity in the focal area, Indicates the focus intensity ratio compliance, It is a single focus confirmation reward.
[0029] The sum of the average light intensity of the focal area The calculation formula is: ; Where, and Represent two peak interval elements respectively; The focus intensity ratio conformity The calculation formula is: ; Where, represents the proportional error; The single focus confirmation reward The calculation formula is: .
[0030] Figure 4 This graph shows the evaluation function training for a trifocal lens. The horizontal axis represents the number of time steps in reinforcement learning. The reward function is averaged over every 200 time steps in the original data and plotted on the graph. The vertical axis represents the global reward function value. The fluctuations in the curve are due to reinforcement learning's continuous exploration to avoid local optimal solutions.
[0031] The phase and transmittance distribution of the metalens designed by the present invention are as follows: Figure 5As shown, the radius cross section of the lens designed by the present invention and the simulated light intensity axial distribution are respectively as shown in FIG. Figure 6 and Figure 7 As shown. The present invention solves the efficiency bottleneck of traditional adjoint methods that require step-by-step optimization of three-dimensional parameters (fixing two dimensions and designing the third dimension). Through the collaborative exploration of multiple agents in a three-dimensional matrix maze, the synchronous reverse design of the three geometric dimensions of the metalens is achieved, avoiding the local optimal problem caused by step-by-step optimization, and improving global optimization efficiency and design accuracy. The present invention strictly constrains the parameters of the nanoring within the process feasible range through a three-dimensional matrix maze, overcoming the defects of generative neural networks in generating irregular structures, ensuring that the designed metalens structure meets the requirements of semiconductor manufacturing processes, and reducing manufacturing difficulty and cost. The present invention abandons the generative method's reliance on large-scale low-correlation random data sets, and uses reinforcement learning agents to actively explore valuable parameter spaces. By memorizing and learning historical exploration results, it reduces invalid calculations, significantly reduces computing power consumption compared to traditional methods, and improves the optimization convergence speed.
[0032] In summary, the present invention solves the problems of uncontrollable generative neural network structure, poor data correlation, and step-by-step design of three-dimensional optimization of the accompanying method by integrating the collaborative exploration mechanism of multi-agent reinforcement learning with physical information constraints, thereby realizing the synchronous inverse design and optimization of the three geometric dimensions of the metalens.
Claims
1. A method for inverse design optimization of a metalens based on reinforcement learning, characterized in that: The specific steps include: Step 1: constructing a nanostructure library consisting of nanostructures with different geometric parameters, and arranging the nanoring structures in the nanostructure library in a three-dimensional matrix maze according to the values of the geometric parameters; each coordinate in the three-dimensional matrix maze corresponds to a unique nanostructure; Step 2: Establishing a physical information multi-agent reinforcement learning framework: Using the three-dimensional matrix maze as the agent exploration environment, the matrix maze is divided into multiple two-dimensional maze slices along the dimension representing the radius of the ring nanostructure. Multiple agents are distributed on the multiple maze slices, and the agents achieve synchronous exploration of geometric parameters by interacting with the reinforcement learning environment. Step 3: Set up a reward function that includes the reward value of a single agent and the global reward value. Use angular spectrum transmission to calculate the axial light intensity distribution of the metalens as the basis for performance evaluation. When the reward value of the reward function reaches the preset target or converges, output the design result of the metalens composed of the nanostructure selected by the agent.
2. The method for inverse design optimization of a metalens based on reinforcement learning according to claim 1, wherein: In step one, the nanostructure is a nanoring; the geometric parameters include three dimensions: radius, rough width and fine width; the radius range is 250nm-49750nm; the rough width range is 100nm-500nm; the fine width is 0nm-16nm, and the fine width is used to widen the rough width; the parameter dimensions of the three-dimensional matrix maze are arranged as 100×11×5; the coordinates also include phase modulation information and transmittance information corresponding to the nanoring structure, wherein the phase modulation information is used as a label to reduce the complexity of data processing.
3. The method for inverse design optimization of a metalens based on reinforcement learning according to claim 1, wherein: In step 2, the physical information includes phase modulation information and transmittance information of the nanostructure; the physical information multi-agent reinforcement learning framework uses a three-dimensional matrix maze as the exploration environment and divides it into 100 two-dimensional 11×5 slices along the radius dimension. 100 agents are distributed on these 100 two-dimensional slices, and the movement and observation of geometric parameters are realized by interacting with the reinforcement learning environment; the physical information multi-agent reinforcement learning framework adopts a multi-agent proximal strategy optimization algorithm.
4. The method for inverse design optimization of a metalens based on reinforcement learning according to claim 3, wherein: The actions of the intelligent agent during movement include movement operations and phase gradient control operations; the observation values during the intelligent agent observation process include the current coordinates of the current intelligent agent, the previous coordinates of the current intelligent agent, the coordinates of the previous intelligent agent, and the phase label corresponding to the coordinates of the current intelligent agent.
5. The method for inverse design optimization of a metalens based on reinforcement learning according to claim 1, wherein: In step 3, the reward function of the single agent reward value is: ; Where, is the local reward value; is the distance between the agent’s current exploration coordinate and the initialization coordinate, is the coefficient used to control the exploration range.
6. The method for inverse design optimization of a metalens based on reinforcement learning according to claim 1, wherein: The reward function of the global reward value is: ; Where, is the global reward value, is the final peak reward, is the non-peak interval flattening penalty term, is the non-peak penalty term in the non-peak interval.
7. The method for inverse design optimization of a metalens based on reinforcement learning according to claim 6, wherein: The non-peak interval flattening penalty term The calculation formula is: ; Where, Represents the sum of all numbers in the non-target peak area; The non-peak interval non-peak penalty item The calculation formula is: ; Where, Indicates the total number of peaks that appear in the peak interval; The final peak reward The calculation formula is: ; Where, Represents the total average light intensity in the focal area, Indicates the focus intensity ratio compliance, It is a single focus confirmation reward.
8. The method for inverse design optimization of a metalens based on reinforcement learning according to claim 7, wherein: The sum of the average light intensity of the focal area The calculation formula is: ; Where, and Represent two peak interval elements respectively; The focus intensity ratio conformity The calculation formula is: ; Where, represents the proportional error; The single focus confirmation reward The calculation formula is: 。
Citation Information
Patent Citations
A Reinforcement Learning-Based Method for Reverse Design and Optimization of Optical Resonators
CN114676635B
Wargame multi-entity asynchronous collaborative decision-making method and device based on reinforcement learning
CN114880955A
Planar super-oscillation lens optimization design method based on deep learning
CN117148568A
Design method of super lens for light splitting, super lens and image sensor
CN118348679A
Intelligent super-lens light field camera circuit board detection method based on deep reinforcement learning
CN119147549A