Method and device for joint control of quadrupole lens in radio frequency section of high energy ion implanter
The quadrupole lens in the RF section of the high-energy ion implanter is jointly controlled through the reinforcement learning method of the Actor-Critic architecture, which solves the problem of time-consuming and poor control in the existing technology, achieves efficient and rapid improvement of beam transmission efficiency, and adapts to changes in equipment status.
Patent Information
- Application Number
- CN202411728246.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-11-28
AI Technical Summary
In the existing technology, the quadrupole lens control method of the RF section of the high-energy ion implanter is time-consuming and ineffective, and cannot achieve effective beam stability and transmission efficiency. In addition, the neural network model needs to be frequently trained to adapt to changes in equipment status, resulting in low control efficiency.
A reinforcement learning method based on the Actor-Critic architecture is adopted. By acquiring and classifying operating parameters to train the reinforcement learning model, joint control of each quadrupole lens is achieved. The reinforcement learning agent is used to adjust the set voltage in real time to improve transmission efficiency. The reward mechanism is designed by combining environmental parameters and action parameters to optimize the control strategy.
It achieves efficient and fast quadrupole lens control, improves transmission efficiency, shortens debugging time, reduces costs, does not rely on additional equipment, has strong migration capabilities, and can adapt to changes in equipment status.
Smart Images

Figure CN119885821B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of high-energy ion implanters, and in particular to a method and device for jointly controlling a quadrupole lens in a radio frequency section of a high-energy ion implanter. Background Art
[0002] High-energy ion implanters are high-precision, industrial-grade equipment that accelerates ions using electric fields, injecting them into the surface of solid materials at a specific energy, thereby modifying the material's physical properties. Compared to electrostatic ion implanters, such as medium- and high-beam ion implanters, high-energy ion implanters utilize radio frequency (RF) acceleration, making particle control more challenging. The RF section of a high-energy RF ion implanter is a unique optical structure, characterized by its complex structure, including multiple RF barrels and quadrupole lenses. This complexity is significantly greater than that of electrostatic particle implanters. In particular, the transverse dynamics are difficult to adjust and match, requiring consideration of both the acceleration effect of the electric field and the guiding and focusing effect of the magnetic field. This makes dynamic calculations challenging, requiring accurate understanding of the evolutionary mechanisms of particle energy differences and positional dispersion within the beam and the development of appropriate control measures to ensure efficient and stable beam propagation and focusing along the central trajectory. In addition to uniformity and stability, precise control of the RF and magnetic field intensity and distribution is required to ensure beam phase space stability. The vast solution space during the adjustment process makes it difficult to quickly achieve good transmission efficiency results through debugging.
[0003] For the regulation of the RF section of a high-energy ion implanter, the existing technology usually adopts a method of adjusting the quadrupole lens individually, that is, debugging each quadrupole lens independently. However, this method is not only time-consuming, but also cannot use the coordinated adjustment of all quadrupole lenses to adjust the beam as a whole. Structures such as the ion source RF barrel will affect the adjustment effect. The coupling effect between each quadrupole lens and the mutual influence with other systems will also make data analysis cumbersome and complicated. In addition, the influence of the adjustment of a single quadrupole lens on the beam may be transmitted to the back-end quadrupole lens or even the entire downstream components of the RF section, causing collective failures. Therefore, the traditional method of regulating a single quadrupole lens cannot achieve the optimal target debugging effect.
[0004] Some practitioners have proposed using neural network methods to control the beam state of high-energy ion implanters. However, the model trained based on the neural network method corresponds to the machine state at the time of data acquisition, and the equipment state of the ion implanter changes slightly over time. When the state changes, the neural network model needs to be retrained, which also requires a lot of time and the actual control effect is not good. Summary of the Invention
[0005] The technical problem to be solved by the present invention is: in response to the technical problems existing in the prior art, the present invention provides a method and device for jointly controlling the quadrupole lens in the RF section of a high-energy ion implanter, which has a simple implementation method, low cost, high control efficiency and precision, and strong portability.
[0006] In order to solve the above technical problems, the technical solution proposed by the present invention is:
[0007] A method for jointly controlling a quadrupole lens in a radio frequency section of a high-energy ion implanter, comprising the following steps:
[0008] Acquire various operating parameters of a controlled high-energy ion implanter during its motion under different operating conditions. The operating parameters include the set voltage, readback current, and beam transmission efficiency of each quadrupole lens. The readback current is generated by beam loss on the quadrupole lens electrode plate. The acquired operating parameters are classified into environmental parameters, action parameters, state parameters, and constraints required for reinforcement learning and constructed into a data set. The environmental parameters are parameters that characterize the state of the ion implanter equipment. The action parameters include the set voltage of the quadrupole lens and the change in the set voltage relative to the previous moment. The state parameters are the beam transmission efficiency and the readback current of the quadrupole lens.
[0009] The data set is used to train a reinforcement learning model using an Actor-Critic architecture. During the training process, the setting voltage of each quadrupole lens is adjusted according to the change in the feedback transmission efficiency. During the adjustment process, the quadrupole lens setting voltage at the current moment and the change in the quadrupole lens setting voltage relative to the previous moment are input into the Actor network as action parameters. After the Actor network, the transmission efficiency and the quadrupole lens readback current are output as state parameters. The Critic network further evaluates the Actor network results based on the state parameters to find the optimal strategy, so that the intelligent agent learns the mapping relationship in the transmission efficiency adjustment process. After the training is completed, a reinforcement learning intelligent agent with the highest transmission efficiency is obtained;
[0010] The readback current of each quadrupole lens and the beam transmission efficiency of the controlled high-energy ion implanter are obtained in real time and read into the trained reinforcement learning agent. The reinforcement learning agent controls and adjusts the setting voltage of each quadrupole lens until the beam transmission efficiency is maximized.
[0011] Furthermore, when obtaining various operating parameters during the movement of the controlled high-energy ion implanter under different working conditions, the RF segment inlet current intensity Iin and the RF segment outlet current intensity Iout are obtained by reading the RF segment inlet Faraday cup current intensity Iinj and the RF segment outlet Faraday cup current intensity Ifem, and the transmission efficiency is calculated according to TransportRatio(TR)=Iin / Iout.
[0012] Furthermore, when obtaining various operating parameters during the movement of the controlled high-energy ion implanter under different working conditions, it also includes reconstructing the cavity pressure Uc and phase P based on the set voltage V and the readback current Ir, and the operating parameters also include any number of particle type, valence beam parameter intensity I, beam energy E, and ion source vacuum, beam line vacuum, target chamber vacuum, filament usage time, and cathode cap usage time.
[0013] Furthermore, the environmental parameters include any one or more of water temperature, water resistance, air temperature, air humidity, ion source vacuum, beam line vacuum, target chamber vacuum, filament usage time, and cathode cap usage time; the action parameters include the setting voltage of the quadrupole lens and the change in the setting voltage relative to the previous moment; the state parameters include the quadrupole lens readback current and transmission efficiency; and the constraints include any one or more of ion energy, beam size, and beam uniformity.
[0014] Furthermore, in the process of using the data set to train the reinforcement learning model using the Actor-Critic architecture, the total reward is composed of the distance reward, the trend reward and the value reward. The distance reward is the Euclidean distance between the readback current and 0, the trend reward is the reward set for the change trend of the transmission efficiency, and the value reward is the reward set for the value of the readback current. The total reward for a single time step is r total Expressed as:
[0015] r total =r distance +r trend +r value
[0016] Among them, r distance is the distance reward, r trend For trend rewards, r value Numerical rewards.
[0017] Furthermore, the trend reward r trend The calculation steps are:
[0018] At any moment, the beam transmission efficiency and the readback current value of the quadrupole lens are read to form a vector B = [Ir1, Ir2, ..., Iri ... IrN, TR], where Iri represents the readback current value of the i-th quadrupole lens, N is the number of quadrupole lenses, and TR is the transmission efficiency;
[0019] Traverse each element b in vector B to obtain the change Δb of each element b, that is, Δb=bt i -bt i-1 , bt i Indicates the value of element b at the current moment, bt i-1Represents the value of element b at the previous moment. The final change Δb is obtained by combining the changes Δb of all elements b. The trend reward r is set according to the value of the change Δb. trend , where when the change Δb is less than 0, the trend reward r trend-single is a value greater than 0, otherwise the trend reward r trend-single is a value less than 0;
[0020] The trend reward r of each element in vector B trend-single The sum of the trend reward r trend .
[0021] Furthermore, the trend reward r of each element is calculated according to the following formula: trend-single :
[0022]
[0023] Where N is the number of quadrupole lenses.
[0024] The trend reward r of each element in vector B trend-single The sum of the trend reward r trend .
[0025] Furthermore, the numerical reward r of each element b in vector B is calculated according to the following formula: value-ssingle :
[0026]
[0027] The numerical reward r of all elements b in vector B value-single The sum of the total numerical reward r value .
[0028] A high-energy ion implanter radio frequency section quadrupole lens joint control device comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to perform the above method.
[0029] A computer-readable storage medium storing a computer program, wherein the computer program implements the above method when executed by a processor.
[0030] Compared with the prior art, the advantages of the present invention are: by combining the reinforcement learning method in machine learning, the present invention learns the strategy of joint control of each quadrupole lens through advance training, and obtains a reinforcement learning model with the highest transmission efficiency. The model can be quickly migrated and applied to the control of the quadrupole lens in the RF section of a real-time high-energy ion implanter. Since it is a joint control strategy for each quadrupole lens, compared with the control strategy for a single quadrupole lens, it can fully consider the coupling effect and influence between different quadrupole lenses, which can not only speed up the calculation speed and shorten the actual debugging time, but also greatly improve the control effect, quickly obtain better transmission efficiency, and effectively solve the ion implanter's demand for fast, stable and automatic beam adjustment. At the same time, it can also have extremely strong migration capabilities, without relying on complex calculations, and without installing additional equipment such as beam diagnostic equipment, and can achieve effective improvement of control performance at a lower cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a schematic diagram of the principle of the high-energy ion implanter layout.
[0032] Figure 2 It is a schematic diagram of the structural principle of the radio frequency section of a high-energy ion implanter in a specific application embodiment.
[0033] Figure 3 It is a schematic diagram of the implementation process of the method for jointly controlling the quadrupole lens in the radio frequency section of the high-energy ion implanter in this embodiment.
[0034] Figure 4 This is a schematic diagram of the principle of implementing the verification of reinforcement learning model training language in this embodiment.
[0035] Figure 5 This is a schematic diagram of the structural principle of the Actor network used in this embodiment.
[0036] Figure 6 This is a schematic diagram of the composition principle of the Actor network action parameters in this embodiment.
[0037] Figure 7 It is a schematic diagram of the implementation process of the joint control of the quadrupole lens in the radio frequency section of the high-energy ion implanter in the specific application embodiment of the present invention. DETAILED DESCRIPTION
[0038] The present invention will be further described below in conjunction with the accompanying drawings and specific preferred embodiments, but the scope of protection of the present invention is not limited thereby.
[0039] Electrostatic ion implanters such as beam ion implanters and large beam ion implanters use particle electrostatic acceleration. Particle electrostatic accelerators use high-voltage electrostatic fields to accelerate charged particles. The force exerted on particles in the electric field is proportional to their charge, so the particles will be accelerated along the direction of the electric field. Their dynamics are relatively simple, mainly involving the single force of the electric field on the particles. The dynamics of electrostatic accelerators are relatively simple, mainly focusing on the accelerating effect of the electric field on the particles, so only uniformity and stability need to be considered.
[0040] The high energy ion implanter uses radio frequency acceleration, such as Figure 1 As shown in Figure 1, particle radio frequency accelerators use alternating electric fields to accelerate particles, and therefore have a specific radio frequency range, which leads to great complexity in the beam debugging process. Figure 2 As shown in the figure, RF represents the RF barrel DTL structure, Q represents the quadrupole lens, and there are 15 quadrupole lenses (Q1 to Q15). In a radio frequency accelerator, while particles are being accelerated, the guiding and focusing effects of the magnetic field on the particles need to be considered. Therefore, its dynamics are more complex, and it is necessary to precisely control the synchronization and stability of the particles in the alternating electric field. The dynamic complexity of the radio frequency accelerator is even higher, and it is necessary to simultaneously consider the acceleration effect of the electric field and the guiding and focusing effect of the magnetic field. In the particle beam dynamics in the radio frequency accelerator, it is necessary to correctly reveal the evolution mechanism of the energy difference and position dispersion of the particles in the particle beam, and propose appropriate control measures to ensure that the beam is effectively and stably transmitted and focused along the central orbit. Therefore, it is necessary not only to consider uniformity and stability, but also to precisely control the intensity and distribution of the radio frequency and magnetic field to ensure the stability of the phase space of the particle beam. Therefore, particle control in high-energy ion implantation machines is more difficult.
[0041] The present invention takes into account the characteristic that the state of machines such as high-energy ion implanters may change over time. By combining the reinforcement learning method in machine learning, the strategy for jointly controlling each quadrupole lens is learned through pre-training, and a reinforcement learning model with the highest transmission efficiency is obtained. The model can be quickly migrated and applied to the control of the quadrupole lens in the RF section of the real-time high-energy ion implanter. Since it is a joint control strategy for each quadrupole lens, compared with the control strategy for a single quadrupole lens, it can fully consider the coupling effect and influence between different quadrupole lenses, which can not only speed up the calculation speed and shorten the actual debugging time, but also greatly improve the control effect, quickly obtain better transmission efficiency, and effectively solve the ion implanter's demand for fast, stable and automatic beam adjustment. At the same time, it can also have extremely strong migration capabilities, without relying on complex calculations, and without installing additional equipment such as beam diagnostic equipment, and can achieve effective improvement in control performance at a relatively low cost.
[0042] The present invention will be further described below with reference to specific embodiments.
[0043] like Figure 3 As shown, the steps of the method for jointly controlling the quadrupole lens in the RF section of the high energy ion implanter in this embodiment include:
[0044] Step S01. Obtain various operating parameters of the controlled high-energy ion implanter during its movement under different working conditions, wherein the operating parameters include the setting voltage, readback current and beam transmission efficiency of each quadrupole lens, and the readback current is generated by beam loss on the quadrupole lens electrode plate. The obtained operating parameters are classified into environmental parameters, action parameters, state parameters and constraints required for reinforcement learning and constructed into a data set, wherein the environmental parameters are parameters that characterize the state of the ion implanter equipment, the action parameters include the setting voltage of the quadrupole lens and the change in the setting voltage relative to the previous moment, the state parameters are the beam transmission efficiency and the readback current of the quadrupole lens, and the constraints include ion energy, beam size and beam uniformity, etc.
[0045] The transmission efficiency is the ratio of the beam current intensity at the RF outlet to the beam current intensity at the RF inlet. The operating parameters may include various parameters such as state parameters, setting parameters and process parameters. The state parameters are parameters that characterize the operating state of the ion implanter. The setting parameters are parameters for setting the ion implanter. The process parameters are process-related parameters such as ion energy, beam intensity, dose injection range, uniformity, and ion type. In this embodiment, when obtaining various operating parameters of the controlled high-energy ion implanter during movement under different working conditions, it also includes reconstructing the cavity pressure Uc and phase P based on the setting voltage V and the readback current Ir, and obtaining the RF segment inlet current intensity Iin and the RF segment outlet current intensity Iout by reading the RF segment inlet Faraday cup current intensity Iinj and the RF segment outlet Faraday cup current intensity Ifem. The transmission efficiency is calculated according to the formula TransportRatio(TR)=Iin / Iout.
[0046] In this embodiment, the operating parameters also include particle type, valence beam parameter intensity I, beam energy E, as well as ion source vacuum, beam line vacuum, target chamber vacuum, filament usage time, cathode cap usage time, etc., which can be configured according to actual needs.
[0047] In a specific application embodiment, Figure 2Taking the RF segment structure as an example, during operation of a high-energy ion implanter, the quadrupole lens set voltage V and readback current Ir are measured, calibrated, and standardized to obtain the reconstructed cavity pressure Uc and phase P, along with V1-V15, Ir1-Ir15, Uc1-Uc15, and P1-P15 values for a total of 15 quadrupole lenses. Simultaneously, the RF inlet Faraday cup current Iinj and the RF segment outlet Faraday cup current Ifem are read. Local deductions are performed based on the extracted data to obtain the inlet and outlet currents Iin and Iout, which are then used to calculate the transmission efficiency. Furthermore, relevant machine parameters, including particle type, valence beam parameters, current intensity I, and beam energy E, are obtained, as well as other relevant machine parameters such as ion source vacuum, beam line vacuum, target chamber vacuum, filament life, and cathode cap life, are obtained. By adjusting and controlling the RF quadrupole lens set voltage V, the beam motion is optimized for optimal transmission efficiency at both the inlet and outlet.
[0048] After obtaining various operating parameters during the movement of the controlled high-energy ion implanter, the operating parameters are classified into environmental parameters, action parameters, state parameters and constraints in reinforcement learning according to the characteristics of reinforcement learning, so as to be used for subsequent reinforcement learning model training.
[0049] In specific application embodiments, environmental parameters may include water temperature, water resistance, air temperature, air humidity, ion source vacuum, beamline vacuum, target chamber vacuum, filament life, and cathode cap life. Action parameters include the quadrupole lens's set voltage and its change relative to the previous moment. State parameters include the quadrupole lens's readback current and transmission efficiency. Constraints include ion energy, beam current, and beam uniformity. The specific parameter division method can be determined based on the actual control objectives and requirements.
[0050] In a specific application embodiment, data preprocessing is also included for the reinforcement learning training data, such as eliminating invalid data, standardizing data, and normalizing data, so as to ensure the validity of the data.
[0051] Step S02. Use the data set to train the reinforcement learning model using the Actor-Critic architecture. During the training process, the setting voltage of each quadrupole lens is adjusted according to the feedback change in transmission efficiency, and the setting voltage of the quadrupole lens and the change in the setting voltage relative to the previous moment are input into the Actor network during the adjustment process. After the Actor network, the transmission efficiency and the quadrupole lens readback current are output as state parameters (state). The Critic network further evaluates the Actor network results to find the optimal strategy, so that the intelligent agent learns the mapping relationship in the transmission efficiency adjustment process. After the training is completed, a reinforcement learning intelligent agent with the highest transmission efficiency is obtained.
[0052] This embodiment uses a reinforcement learning-based adaptive control strategy to enable the system to automatically adjust the settings of the quadrupole lens in real time based on the transmission efficiency feedback according to the debugging data target parameters. The learning-based strategy can help the system quickly adapt to and maintain optimal performance when facing unknown or changing environmental conditions. The reinforcement learning algorithm can learn how to automatically adjust the settings of the quadrupole lens according to the input beam characteristics (such as particle energy, beam intensity, transmission efficiency, etc.) to achieve the best focusing effect, realize the automation of quadrupole lens adjustment, and reduce manual intervention.
[0053] This embodiment uses the parameters obtained and divided in step S2 as a training data set to train a reinforcement learning model, so that compared with the relatively simple structure and low-dimensional data of the electrostatic ion implanter, a reinforcement learning model for complex RF segment quadrupole lens adjustment can be constructed. This model can handle complex physical processes and high-dimensional input data, and improves the system's robustness to abnormal situations through reinforcement learning, ensuring that the system can still maintain optimal performance when facing various challenges. At the same time, the model is made transferable and can be flexibly applied to the beam control of actual high-energy ion implanters.
[0054] like Figure 4 As shown, reinforcement learning consists of two parts: simulation training and on-device testing. The agent interacts with the simulation environment built by TraceWin software. The agent determines the settings for the electric quadrupole lens based on changes in transmission efficiency. During training, the agent's goal is to maximize the expected value of the reward under each operating condition. Once the model is trained in the simulation environment, it can be evaluated on a real ion implanter without the need for online data.
[0055] In this embodiment, the basic architecture of the Agent adopts the TD3 (Twin Delayed Deep Deterministic Policy Gradients) model, such as Figure 4As shown, it is based on the Actor-Critic framework and adopts a deep deterministic policy gradient algorithm, which can effectively solve continuous control problems. If the readout is directly used as the observation value in the TD3 network, the network trained in the simulation environment will not reflect the correct mapping relationship in the real ion implanter. In contrast, despite the differences between the real accelerator and the virtual accelerator environment, the main trends of the quadrupole lens, readback, and efficiency mapping are similar. Therefore, if the transmission efficiency trend is used as the observation of the TD3 network, the agent trained in the simulation environment can also describe the correct mapping relationship in the real accelerator to a certain extent, thereby reducing the difference between the virtual and real ion implanter environments and enhancing the robustness of the agent. This embodiment uses the TD3 model to implement RF segment electric quadrupole lens adjustment. At the same time, during the control process, the trend of quadrupole lens debugging (the trend of transmission efficiency change) is taken into account, so that the agent can learn the invariant relationship (fixed mapping relationship) during the transmission efficiency adjustment process. It enables the agent to understand the adjustment value and the adjustment gradient trend, rather than just the relationship between the values, so that the trained agent has sufficient robustness to overcome the differences between the simulation and real environments.
[0056] To verify the effectiveness of the method described above in this embodiment, experiments were conducted using two agents based on the TD3 model. These agents had different network structures: one agent used a network based on physical prior knowledge, while the other used the original TD3 network structure. The agents were first trained in a simulation environment constructed using TraceWin simulation software. This environment included the chamber pressures Uc1-Uc15, phases P1-P15, particle type, valence beam parameters I and E, and other parameters such as ion source vacuum, beamline vacuum, target chamber vacuum, filament life, and cathode cap life, obtained in step S1. The agents were then directly implanted into a real high-energy particle implanter. Experimental results demonstrated that the agents trained using the method described above in this embodiment can effectively improve the implanter's transmission efficiency. Furthermore, the agents in this embodiment remained effective when applied to different debugging recipes, demonstrating their applicability to diverse debugging targets. Furthermore, the method has been validated to bypass online training and data collection, improving the efficiency of ion implanter beam debugging.
[0057] There are two main types of RL methods: value-based methods and policy-based methods. Value-based methods learn policies by maximizing the Q-value function (s, a), which represents the long-term benefit of taking action a in state s by following the policy π, which can be expressed as:
[0058]
[0059] Where E is the expected value operator, T is the remaining steps from state s = st to the end of the event, γ is the discount factor, which indicates the importance we want to give to future rewards; R is the reward, which is the feedback from the environment when the agent performs action a in state s.
[0060] The policy-based approach learns an agent by directly searching for the optimal policy, without using a Q-value function. This embodiment uses the actor-critic framework for reinforcement learning model training. The actor-critic framework combines numerical and policy-based methods into a more efficient approach. The actor observes the state of the environment and outputs the optimal action, and the critic evaluates the action by calculating the Q-value function. As the policy network and value function network are trained, the critic gradually guides the actor to find the optimal policy through temporal difference error. In addition, to better simulate the actual machine environment, this embodiment adds noise to the actor step.
[0061] In the process of training the reinforcement learning model using the Actor-Critic architecture using the dataset, the total reward is composed of distance reward, trend reward and value reward. The distance reward is the Euclidean distance between the readback current and 0, the trend reward is the reward set for the change trend of the transmission efficiency, and the value reward is the reward set for the value of the readback current. The total reward for a single time step is r total Expressed as:
[0062] r total = r distance + r trend + r value (2)
[0063] Among them, r distance is the distance reward, r trend For trend rewards, r value Numerical rewards.
[0064] This embodiment follows the above method and sets a reward function by comprehensively considering the distance of the readback current, the trend of the transmission efficiency, and the numerical value. The distance reward represents the Euclidean distance of each position relative to the zero value when the plate has no beam loss. This distance reward only reflects the overall situation. By setting a trend reward, the direction of improving transmission efficiency can be clearly defined, encouraging the agent to make decisions towards higher transmission efficiency by observing the transmission efficiency trend. The numerical reward sets a reward mechanism for the value of the readback current. It can be used to determine whether a single position has a significant beam loss based on the readback current of a single position.
[0065] In this embodiment, the trend reward r trend The calculation steps are:
[0066] The beam transmission efficiency and the readback current value of the quadrupole lens are read to form a vector B = [Ir1, Ir2, ..., Iri ... IrN, TR], where Iri represents the readback current value of the i-th quadrupole lens, N is the number of quadrupole lenses, and TR is the transmission efficiency;
[0067] Traverse each element b in vector B to obtain the change Δb of each element b, that is, Δb=bt i -bt i-1 , bt i Indicates the value of element b at the current moment, bt i-1 Represents the value of element b at the previous moment. The final change Δb is obtained by combining the changes Δb of all elements b. The trend reward r is set according to the value of the change Δb. trend , where when the change Δb is less than 0, the trend reward r trend-single is a value greater than 0, otherwise the trend reward r trend-single is a value less than 0;
[0068] The trend reward r of each element in vector B trend-single The sum of the total trend reward r trend .
[0069] In this embodiment, the trend reward r is set trend In order to set up a reward mechanism for transmission efficiency, the transmission efficiency and the readback current value of the quadrupole lens form a vector B, and the trend reward r is set by traversing the change of each element in B. trend , when Δb is less than 0, that is, the trend is decreasing, the trend reward r trend If Δb is greater than 0, the change is rewarded, and if Δb is greater than or equal to 0, that is, the change trend increases, the change is punished accordingly, so that decisions can be made towards higher transmission efficiency based on the transmission efficiency trend, effectively improving the efficiency and accuracy of regulation.
[0070] As an optional implementation, the trend reward r of each element can be calculated according to the following formula: trend-single :
[0071]
[0072] Where N is the number of quadrupole lenses.
[0073] From formula (3), we can see that when Δb is less than 0, that is, the trend of change is decreasing, then 1 / N is used as the reward value to reward the change, and if Δb is greater than or equal to 0, that is, the trend of change is increasing, then -3 / 2N is used as the reward value to penalize the change. The number of quadrupole lenses can be combined to set a more accurate trend reward r. trend-single, which enables decisions to be made towards higher transmission efficiency based on the transmission efficiency trend, thereby effectively improving the efficiency and accuracy of the joint control of multiple quadrupole lenses.
[0074] In this embodiment, the numerical reward r of each element b in vector B is calculated according to the following formula: vallue-single :
[0075]
[0076] As shown in formula (5), the numerical reward r is calculated in this way value-single , we can judge whether there is a single location with a large beam loss by reading back the current at a single location, and then reward r by the numerical value of all elements b in vector B value-single The sum of the total numerical reward r value .
[0077] As Figure 2 As an example, the RF segment structure is adopted. Figure 5 The network structure shown in FIG1 is a graph showing a network structure in which the input is the transmission efficiency and readback current readings (B), including TR and Ir1-Ir15. The input layer is followed by two fully connected hidden layers, each containing 256 nodes. The output of the network is the quadrupole lens setting voltage (M) (including V1-V15). At the same time, in order to reduce the difference between the virtual and real ion implantation machine environments and enhance the robustness of the agent, this embodiment adds an optimization strategy: after changing (ΔM) the voltage (M) of the electric quadrupole lens, the readback current is combined with the transmission efficiency reading (B) and its change (ΔB), rather than just the reading, as shown in FIG1. Figure 6 As shown in the figure, the convolution layer between the input layer and the first hidden layer has four 3×1 convolution kernels for the x and y directions respectively. The output in each direction is a 16-dimensional vector, and then these two vectors are directly connected as the input of the next hidden layer.
[0078] Step S03. Obtain the readback current of each quadrupole lens and the beam transmission efficiency of the controlled high-energy ion implanter in real time, and input them into the trained reinforcement learning agent. The reinforcement learning agent controls and adjusts the setting voltage of each quadrupole lens until the beam transmission efficiency is maximized.
[0079] Once the reinforcement learning model is trained, it can be transferred and applied to the actual control of high-energy ion implanters. Furthermore, data can be accumulated during the debugging process to further optimize the agent's performance. Ultimately, an automated system for debugging the quadrupole lens beam in high-energy ion implanters will be built, enabling rapid control to maximize the ion implanter's transmission efficiency.
[0080] In a specific application embodiment, in order to implement the above method of the present invention, the following methods can be used: Figure 7The architecture shown, where:
[0081] The data and theory module records the state parameters, setup parameters, and process parameters of the controlled ion implanter during operation. The recorded parameters are categorized by environment, action, input, and output. Based on the basic theory of transverse dynamics and the relationships between data, physical formulas are introduced as physical constraints. Furthermore, this module preprocesses the recorded data, including but not limited to normalization and regularization. The preprocessed parameters are used as the training dataset for training the reinforcement learning model.
[0082] The debugging task agent research module is used to adjust the parameters according to the transmission efficiency of the ion implanter, such as the setting voltage of the ion implanter. It can also include the maximum magnetic field strength of the quadrupole lens, the adjustment range of the quadrupole lens power supply, the response time of the quadrupole lens power supply, etc. The reward mechanism is formed based on the influence of the beam state on the implant in beam dynamics. The total reward is calculated according to the reward function of formula (2). At the same time, based on the accelerator transverse dynamics and beam phase space theory, combined with reinforcement learning to form physical constraints, training is performed based on the training data set to obtain a reinforcement learning agent. The reinforcement learning agent is capable of completing the beam adjustment task based on the environment.
[0083] The validity verification module verifies the transfer validity of trained agents. Based on the characteristics of the target task, it selects appropriate transfer strategies and performs similarity assessments to ensure sufficient similarity between offline and online training results for effective knowledge transfer. It also considers negative transfer prevention to avoid negative impacts on online beam debugging under sudden changes in environmental parameters. It also incorporates continuous learning, continuously training and optimizing the model based on the large amount of available data generated by the actual machine to adapt to possible environmental changes.
[0084] The migration module is used to apply the intelligent agent that can be used for online beam adjustment to the machine, perform quadrupole lens debugging, and detect all parameters of the machine after debugging in real time. By briefly adjusting the action parameter value, the state parameter, that is, the beam target, can be quickly made to meet the machine requirements. Due to pre-training, the entire adjustment process can limit the debugging actions to a small number (such as 10) or less, which can significantly improve the transmission efficiency and adjustment speed of the ion implanter, greatly shorten the time for equipment to switch process parameters and save manpower. In addition, the stability and robustness of the program can be evaluated through long-term operation. At the same time, considering the possibility of unexpected problems, the unexpected problems can also be recorded (stable during actual operation, no abnormalities have occurred) to further improve stability and performance.
[0085] The present invention trains the model through reinforcement learning to obtain optimal parameters, can quickly find the optimal solution from complex calculations, and find the quadrupole lens parameters corresponding to the optimal transmission efficiency. Based on the machine learning method, time is mainly used in the pre-training process, and the trained model can be used directly, which can greatly improve the speed of beam debugging, so that higher transmission efficiency can be obtained in a short time. At the same time, by observing the transmission efficiency, the adjustment is directly guided by machine learning, without the need for other beam diagnostic equipment, and can also reduce the implementation cost and complexity.
[0086] This embodiment further provides a high-energy ion implanter RF section quadrupole lens joint control device, including a processor and a memory, the memory is used to store computer programs, and the processor is used to execute the computer program to perform the above method.
[0087] It is understandable that the above method of this embodiment can be executed by a single device, such as a computer or server, etc., and can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of a distributed scenario, one of the multiple devices can only execute one or more steps in the above method of this embodiment, and multiple devices interact to complete the above method. The processor can be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., for executing relevant programs to implement the above method of this embodiment. The memory can be implemented in the form of a read-only memory ROM, a random access memory RAM, a static storage device, and a dynamic storage device. The memory can store an operating system and other application programs. When the above method of this embodiment is implemented by software or firmware, the relevant program code is stored in the memory and called and executed by the processor.
[0088] This embodiment further provides a computer-readable storage medium storing a computer program, which implements the above method when executed by a processor.
[0089] Those skilled in the art will appreciate that the above-mentioned embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0090] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed above with reference to the preferred embodiment, it is not intended to limit the present invention. Therefore, any simple modifications, equivalent variations, and modifications to the above embodiment that do not depart from the technical solution of the present invention and are based on the technical essence of the present invention shall fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for joint control of quadrupole lens in the radio frequency section of a high energy ion implanter, characterized in that the steps include: Acquire various operating parameters of a controlled high-energy ion implanter during its motion under different operating conditions. The operating parameters include the set voltage, readback current, and beam transmission efficiency of each quadrupole lens. The readback current is generated by beam loss on the quadrupole lens electrode plate. The acquired operating parameters are classified into environmental parameters, action parameters, state parameters, and constraints required for reinforcement learning and constructed into a data set. The environmental parameters are parameters that characterize the state of the ion implanter equipment. The action parameters include the set voltage of the quadrupole lens and the change in the set voltage relative to the previous moment. The state parameters are the beam transmission efficiency and the readback current of the quadrupole lens. The data set is used to train a reinforcement learning model using an Actor-Critic architecture. During the training process, the setting voltage of each quadrupole lens is adjusted according to the change in the feedback transmission efficiency. During the adjustment process, the quadrupole lens setting voltage at the current moment and the change in the quadrupole lens setting voltage relative to the previous moment are input into the Actor network as action parameters. After the Actor network, the transmission efficiency and the quadrupole lens readback current are output as state parameters. The Critic network further evaluates the Actor network results based on the state parameters to find the optimal strategy, so that the intelligent agent learns the mapping relationship in the transmission efficiency adjustment process. After the training is completed, a reinforcement learning intelligent agent with the highest transmission efficiency is obtained; The readback current of each quadrupole lens and the beam transmission efficiency of the controlled high-energy ion implanter are obtained in real time and read into the trained reinforcement learning agent. The reinforcement learning agent controls and adjusts the setting voltage of each quadrupole lens until the beam transmission efficiency is maximized.
2. The method for joint control of quadrupole lens in the radio frequency section of a high energy ion implanter according to claim 1, characterized in that: When obtaining various operating parameters of the controlled high-energy ion implanter during movement under different working conditions, the RF segment inlet current intensity Iin and the RF segment outlet current intensity Iout are obtained by reading the RF segment inlet Faraday cup current intensity Iinj and the RF segment outlet Faraday cup current intensity Ifem, and the transmission efficiency is calculated according to TransportRatio(TR) = Iin / Iout.
3. The method for joint control of quadrupole lens in the radio frequency section of a high energy ion implanter according to claim 2, characterized in that: When obtaining various operating parameters of the controlled high-energy ion implanter during movement under different working conditions, it also includes reconstructing the cavity pressure Uc and phase P based on the set voltage V and the readback current Ir. The operating parameters also include particle type, valence beam parameter current intensity I, beam energy E, and any number of ion source vacuum, beam line vacuum, target chamber vacuum, filament usage time, and cathode cap usage time.
4. The method for joint control of quadrupole lens in the radio frequency section of a high energy ion implanter according to claim 1, characterized in that: The environmental parameters include any one or more of water temperature, water resistance, air temperature, air humidity, ion source vacuum, beam line vacuum, target chamber vacuum, filament usage time, and cathode cap usage time; the action parameters include the setting voltage of the quadrupole lens and the change in the setting voltage relative to the previous moment; and the constraints include any one or more of ion energy, beam size, and beam uniformity.
5. The method for joint control of quadrupole lenses in the radio frequency section of a high energy ion implanter according to any one of claims 1 to 4, characterized in that: In the process of using the data set to train the reinforcement learning model using the Actor-Critic architecture, the total reward is composed of the distance reward, the trend reward and the value reward. The distance reward is the Euclidean distance between the readback current and 0, the trend reward is the reward set for the change trend of the transmission efficiency, and the value reward is the reward set for the value of the readback current. The total reward for a single time step is r total Expressed as: r total = r distance + r trend + r value Among them, r distance is the distance reward, r trend For trend rewards, r value Numerical rewards.
6. The method for joint control of quadrupole lenses in the radio frequency section of a high energy ion implanter according to claim 5, characterized in that: Trend Rewards trend The calculation steps are: At any moment, the beam transmission efficiency and the readback current value of the quadrupole lens are read to form a vector B=[Ir1, Ir2,…, Iri…IrN, TR], where Iri represents the readback current value of the i-th quadrupole lens, N is the number of quadrupole lenses, and TR is the transmission efficiency; Traverse each element b in vector B to obtain the change Δb of each element b, that is, Δb=bt i -bt i-1 , bt i Indicates the value of element b at the current moment, bt i-1 Represents the value of element b at the previous moment. The final change Δb is obtained by combining the changes Δb of all elements b. The trend reward r is set according to the value of the change Δb. trend , where when the change Δb is less than 0, the trend reward r trend_single is a value greater than 0, otherwise the trend reward r trend_single is a value less than 0; The trend reward r of each element in vector B trend_single The sum of the trend reward r trend .
7. The method for joint control of quadrupole lenses in the radio frequency section of a high energy ion implanter according to claim 6, characterized in that: Calculate the trend reward r for each element as follows trend_single : in, N is the number of quadrupole lenses.
8. The method for joint control of quadrupole lenses in the radio frequency section of a high energy ion implanter according to claim 5, characterized in that: Calculate the numerical reward r for each element b in vector B according to the following formula value_single : The numerical reward r of all elements b in vector B value_single The sum of the total numerical reward r value .
9. A high-energy ion implanter radio frequency section quadrupole lens joint control device, comprising a processor and a memory, wherein the memory is used to store a computer program, characterized in that: The processor is configured to execute the computer program to perform the method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Enhanced learning method for calibrating beam deviation of accelerator
CN110278651A
Ion implanter beam adjusting system with self-learning function and adjusting method
CN116959944A