Temperature compensation method for constant temperature zone of horizontal liquid phase epitaxial growth equipment based on reinforcement learning
Through a temperature compensation method based on reinforcement learning, the temperature of the temperature zones of horizontal liquid phase epitaxial growth equipment is automatically adjusted, solving the complex manual operation problems in the existing technology, achieving fast and effective temperature uniformity control, improving equipment operating efficiency and reducing maintenance costs.
Patent Information
- Application Number
- CN202411486115.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-10-23
AI Technical Summary
The temperature compensation method of existing horizontal liquid phase epitaxial growth equipment requires manual operation, which is complicated and time-consuming. It is difficult to quickly achieve temperature uniformity in the constant temperature zone, affecting the equipment's production efficiency and maintenance costs.
A temperature compensation method based on reinforcement learning is adopted. By obtaining the temperature data of multiple measurement points in the constant temperature zone, the temperature compensation value is calculated using the reinforcement learning training strategy model to automatically adjust the set temperatures of the five temperature zones.
It simplifies the temperature compensation process, reduces equipment downtime, improves production efficiency, reduces maintenance costs, and is suitable for a variety of semiconductor equipment.
Smart Images

Figure CN119380871B_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the technical field of horizontal liquid phase epitaxial growth equipment, and in particular to a temperature compensation method for a constant temperature zone of a horizontal liquid phase epitaxial growth equipment based on reinforcement learning. Background Art
[0002] The horizontal liquid phase epitaxy process offers advantages such as easy composition change, strong thickness control, and excellent repeatability. It is an effective and mature method for developing and producing HgCdTe thin films. Horizontal liquid phase epitaxy equipment is currently the mainstream equipment for this process in China. Its temperature control accuracy, ramp rates, and temperature field uniformity directly impact the quality of the epitaxial material, the uniformity, repeatability, size, and production capacity of the epitaxial wafers, and are key performance indicators.
[0003] The working principle of the horizontal liquid phase epitaxial growth equipment is that the liquid phase epitaxial growth system adopts a Te-rich flux and a horizontal boat pulling method. The graphite boat carrying the substrate and solvent is transported through a load lock and an atmosphere protection glove box to a cantilever quartz rake fixed on the end cover of the reaction tube. It is pushed into the reaction chamber by the boat advance and retreat mechanism, and the end cover is closed. In the reaction chamber, the vacuum pump is first used to evacuate the chamber to create a background clean space, and then hydrogen is introduced through the air inlet pipe. A clean laminar reducing atmosphere is established in coordination with the exhaust pipe. A heat field is established through an electric heating furnace and reaches a stable state. After the graphite boat is heated thoroughly and the solvent is fully converted into a solution, the temperature begins to drop rapidly. When the temperature drops to a level slightly above the freezing point of the solution, the boat pulling mechanism drives the pull rod to move, thereby driving the pull rod. The slider connected to the rod moves, and the slider carries the solvent solution to smoothly cover the substrate surface. Then the temperature begins to slowly decrease, and the solution is in a critical state of solidification and evaporation, crystallizing and epitaxially growing. When the temperature drops to a certain temperature, the epitaxy is completed. The boat pulling mechanism drives the pull rod to move again, thereby driving the slider connected to the pull rod to move. The slider carries the solvent solution smoothly away from the substrate surface. The electric furnace is then moved away from the reaction tube. The fan is started, and the temperature is forced to drop to near room temperature by air cooling. The hydrogen is replaced, and the boat is withdrawn. Finally, the graphite boat is manually unloaded from the cantilever quartz rake into the load lock, and finally the graphite boat is taken out of the load lock. In the above process, from the graphite boat being pushed into the reaction tube during the above epitaxial process, the temperature control in the reaction tube runs through the entire process. The degree of temperature control accuracy directly determines the quality rate of the process product. At present, in order to ensure the temperature control accuracy and temperature uniformity of the temperature field, the furnace body adopts 5 independent heating constant temperature zones and independent temperature control.
[0004] In order to ensure the uniformity of temperature distribution in the constant temperature zone of the horizontal liquid phase epitaxial equipment, with the increase of the number of processes, the maintenance of quartz parts, internal and external thermocouple sensors and the replacement of graphite boat tooling, it is necessary to manually pull the constant temperature zone to ensure the temperature uniformity of the constant temperature zone, and then optimize the set temperature compensation value of the five temperature zones; the operating principle of pulling the constant temperature zone compensation is as follows: adopt the temperature measuring thermocouple moving method and single-end full-range measurement; use a temperature measuring thermocouple to extend to the position flush with the top of a single internal thermocouple sensor (five temperature measuring points) (the temperature measuring thermocouple and the internal thermocouple are aligned with each other) The thermocouple sensor is in the same thermocouple sleeve). After the temperature in the constant temperature zone stabilizes, the working end of the temperature measuring thermocouple is moved along the axis of the reaction tube or its parallel line from the closed end of the reaction tube 250 mm away from the center of the furnace tube to the open end of the reaction tube until the end of the temperature field in the constant temperature zone. The difference between the temperature measured by the moving thermocouple and the set temperature is calculated, that is, the temperature deviation. According to the temperature deviation trend mapped by the five temperature measuring points of the internal thermocouple, the set temperature compensation values of the five temperature zones are calculated. Finally, the temperature control accuracy of the constant temperature zone is within (450±0.5) degrees.
[0005] Currently, after running multiple processes in horizontal liquid-phase epitaxial growth equipment, a large amount of Hg will adhere to the reaction tube cavity, which will affect heat conduction and the temperature measurement sensitivity and accuracy of the thermocouple. Therefore, it directly leads to inaccurate temperature in the epitaxial process, which ultimately affects the process quality. At this time, on-site equipment personnel or process personnel are also required to re-draw the constant temperature zone and improve the temperature field uniformity of the constant temperature zone. The set temperatures of the five temperature zones are compensated according to the temperatures obtained by the temperature measuring thermocouples. A certain amount of experience and understanding of the corresponding positions of the five external thermocouple temperature measurement points and the five internal thermocouple temperature measurement points in the temperature field are required to quickly obtain the correct compensation value. This indirectly affects the online operation time and maintenance costs of the equipment.
[0006] The existing temperature compensation method has the following shortcomings:
[0007] First, the existing temperature compensation method for the constant temperature zone requires the operator to understand the position of the corresponding temperature zone of the thermocouple outside the furnace body, as well as the position and size diagram of the five temperature measurement points of the internal thermocouple sensor.
[0008] Second, the existing temperature compensation method for the constant temperature zone requires at least five times to control the temperature uniformity of the constant temperature zone within the technical indicators.
[0009] Third, the existing temperature compensation method for the constant temperature zone is complicated to operate and takes a long time to maintain and adjust, which is not conducive to improving the production efficiency of the equipment.
[0010] Fourth, the existing temperature compensation method for the constant temperature zone can only be applied to horizontal liquid phase epitaxial growth equipment.
[0011] Therefore, a constant temperature zone temperature compensation method is needed that can reduce the downtime of horizontal liquid phase epitaxial growth equipment. Summary of the Invention
[0012] In view of the technical problems existing in the prior art, the present invention provides a temperature compensation method for a constant temperature zone of a horizontal liquid phase epitaxial growth device based on reinforcement learning, which is simple to operate and has fast compensation.
[0013] In order to solve the above technical problems, the technical solution proposed by the present invention is:
[0014] A method for temperature compensation in a constant temperature zone of a horizontal liquid phase epitaxial growth device based on reinforcement learning comprises the following steps:
[0015] S1. Acquire temperature data of multiple measurement points in a constant temperature zone of a horizontal liquid phase epitaxial growth device;
[0016] S2. Obtain the maximum value, minimum value, average value, and standard deviation of each temperature data; and compare each temperature data with the initial set temperature to obtain a deviation value;
[0017] S3. Input the maximum value, minimum value, average value, standard deviation and deviation value of each temperature data as input into a pre-built and trained reinforcement learning training strategy model to obtain the temperature compensation value of each temperature zone; wherein the reinforcement learning training strategy model presets a mapping relationship between the input and the temperature compensation value.
[0018] Preferably, in step S3, the temperature compensation value of the middle temperature zone is first obtained, and then the temperature compensation values of each temperature zone are obtained one by one in sequence toward the outside.
[0019] Preferably, the process of obtaining the temperature compensation value of the intermediate temperature zone is:
[0020] According to the deviation between each temperature data and the reference temperature, the feedback reward value Reward is established for each position point in the intermediate temperature zone;
[0021] According to the feedback reward value Reward of each position point in the intermediate temperature zone, the weighted reward value of the intermediate temperature zone close to the adjacent temperature zone is obtained, and finally the temperature compensation value of the intermediate temperature zone is obtained, and then the temperature setting value of the intermediate temperature zone is obtained.
[0022] Preferably, in step S1, the total number of measuring points of the temperature measuring thermocouple is n, the single moving distance is l, the initial set temperature of each temperature zone is k degrees, the single point stabilization time is t, the initial temperature compensation value is 0, and the position relationship diagram of each temperature measuring point on the internal thermocouple sensor corresponding to each temperature zone of the furnace body is obtained.
[0023] Preferably, n=25; l=20 mm; k=450 degrees; and t=120 s.
[0024] The present invention also discloses a constant temperature zone temperature compensation system for a horizontal liquid phase epitaxial growth device based on reinforcement learning, comprising:
[0025] A temperature acquisition module is used to obtain temperature data of multiple measurement points in a constant temperature zone of a horizontal liquid phase epitaxial growth device;
[0026] The temperature analysis module is used to obtain the maximum value, minimum value, average value and standard deviation of each temperature data; at the same time, each temperature data is compared with the initial set temperature to obtain the deviation value;
[0027] The temperature compensation module is used to input the maximum, minimum, average, standard deviation and deviation values of each temperature data as input into a pre-built and trained reinforcement learning training strategy model to obtain the temperature compensation value for each temperature zone; wherein the reinforcement learning training strategy model presets the mapping relationship between the input and the temperature compensation value.
[0028] The present invention further discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method described above are executed.
[0029] The present invention also discloses a computer device, comprising a memory and a processor connected to each other, wherein a computer program is stored in the memory, and when the computer program is run by the processor, the steps of the above method are executed.
[0030] Compared with the prior art, the advantages of the present invention are:
[0031] The constant temperature zone temperature compensation method of the horizontal liquid phase epitaxial growth equipment based on reinforcement learning of the present invention can calculate the constant temperature zone temperature compensation value in a fewer number of times, thereby reducing the equipment downtime maintenance time; this constant temperature zone temperature compensation method has strong repeatability and simple implementation, and the obtained five temperature zone set temperature compensation values have significant effects, reducing the intermediate process of manual calculation, low maintenance cost, strong practicality, faster acquisition of the set temperature compensation value, and improved efficiency of pulling the constant temperature zone. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 FIG. 1 is a flow chart of a temperature compensation method according to an embodiment of the present invention.
[0033] Figure 2 This is a position correspondence diagram of the five external thermocouple sensors of the furnace body and the five temperature measuring points on the internal thermocouple sensor in the present invention.
[0034] Figure 3 It is a schematic diagram of the constant temperature zone of the horizontal liquid phase epitaxial growth equipment in the present invention.
[0035] Figure 4Schematic diagram of the reinforcement learning training process in the present invention.
[0036] Figure 5 This is a comparison diagram of the constant temperature zone before and after using the reinforcement learning training strategy in the present invention; (a) is before improvement; (b) is after improvement. DETAILED DESCRIPTION
[0037] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0038] like Figure 1 As shown, the constant temperature zone temperature compensation method of a horizontal liquid phase epitaxial growth device based on reinforcement learning in an embodiment of the present invention includes the following steps:
[0039] S1. Acquire temperature data of multiple measurement points in a constant temperature zone of a horizontal liquid phase epitaxial growth device;
[0040] S2. Obtain the maximum value, minimum value, average value, and standard deviation of each temperature data; and compare each temperature data with the initial set temperature to obtain a deviation value;
[0041] S3. Input the maximum value, minimum value, average value, standard deviation and deviation value of each temperature data as input into a pre-built and trained reinforcement learning training strategy model to obtain the temperature compensation value of each temperature zone; wherein the reinforcement learning training strategy model presets a mapping relationship between the input and the temperature compensation value.
[0042] The constant temperature zone temperature compensation method of the horizontal liquid phase epitaxial growth equipment based on reinforcement learning of the present invention can calculate the constant temperature zone temperature compensation value in a fewer number of times, thereby reducing the equipment downtime maintenance time; this constant temperature zone temperature compensation method has strong repeatability and simple implementation, and the obtained five temperature zone set temperature compensation values have significant effects, reducing the intermediate process of manual calculation, low maintenance cost, strong practicality, faster acquisition of the set temperature compensation value, and improved efficiency of pulling the constant temperature zone.
[0043] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0044] like Figure 1As shown, the constant temperature zone temperature compensation method of the horizontal liquid phase epitaxial growth equipment based on reinforcement learning provided by an embodiment of the present invention can be used for a semiconductor furnace tube equipment with five constant temperature zones. There are five temperature zones on the furnace body of the horizontal liquid phase epitaxial growth equipment and they are adjacent to each other. Each temperature zone has a DC power supply to heat the furnace wire in the corresponding temperature zone in the furnace body to form a temperature field. The external temperature is the temperature value measured by the external thermocouple obtained by heating the furnace wire through the DC power supply. The internal temperature is the temperature value measured by the internal thermocouple obtained by the heat radiation of the furnace wire into the process chamber. It is divided into five sections of independent temperature control. The length of the uniform temperature of the five temperature zones is the constant temperature zone length, and the constant temperature zone length is 500mm.
[0045] The specific steps of the above constant temperature zone temperature compensation method are as follows:
[0046] Step 1: Preset the total number of measuring points of the thermocouple to 25, the single movement distance to 20 mm, the initial set temperature of the five temperature zones to 450 degrees, the single point stabilization time to 120 seconds, and the initial temperature compensation value to 0;
[0047] Step 2: Obtain the position relationship diagram of the five temperature measurement points on the internal thermocouple sensor corresponding to the five temperature zones of the furnace body, such as Figure 2-Figure 3 As shown;
[0048] Step 3: After the constant temperature zone is automatically drawn, the maximum and minimum values of 25 temperature measurement data of the thermocouple within the constant temperature zone are obtained, and the average value and standard deviation of this set of temperature measurement data are calculated;
[0049] Step 4: Based on the technical specifications of this horizontal liquid phase epitaxial growth equipment, the upper and lower limits of the measurement temperature are divided into (450±0.5) degrees. Then, combined with the Q-Learning reinforcement learning training strategy model, when the temperature measurement point in a temperature zone exceeds the upper limit value (450+0.5) degrees, the set temperature of the corresponding temperature zone needs to be adjusted downward. When the temperature measurement point in a temperature zone is lower than the lower limit value (450-0.5) degrees, the set temperature of the corresponding temperature zone needs to be adjusted upward. The compensation values of the set temperatures of the five temperature zones are obtained in sequence.
[0050] Take the first constant temperature zone temperature measurement data (the set temperature of the five temperature zones is 450 degrees, the temperature compensation value is 0, and the index reference temperature is 450 degrees) as an example:
[0051] Table 1 Temperature measurement data of automatic constant temperature zone
[0052]
[0053] Since the core area of the process is in temperature zone 3, we first analyze how the measured temperature curve from temperature zone 3 to temperature zones 2 and 4 changes with the single-point temperature measurement distance. The temperature coupling boundary between temperature zones 3 and 4 is the temperature measurement position 80, and the temperature coupling boundary between temperature zones 3 and 2 is the temperature measurement position -80. Based on the deviation between the first temperature measurement data and the reference temperature of 450 degrees, positive or negative feedback reward values are established for each position point in temperature zones 2, 3, and 4.
[0054] Table 2 Feedback reward value weighted to obtain temperature compensation value for temperature zone 3
[0055]
[0056] Influencing Factor: The influence factor of the current temperature measurement point is inversely proportional to the distance from the current measurement point to the measurement point in a specific temperature zone, as the step size is related to the current temperature measurement point. However, in horizontal liquid phase epitaxial growth equipment, the constant temperature zones are distributed, with Zone 1 located at the furnace entrance and Zone 5 at the furnace tail. This provides excellent heat dissipation, and the temperature decreases from the furnace tail to the furnace entrance. Therefore, the discount factor decreases with distance for measurement points farther from Zones 2 and 4, but the discount factor is significantly smaller for measurement points farther from Zone 4 toward the furnace tail.
[0057] Feedback reward value: When the deviation value is higher than the reference temperature setting value, the value of the deviation value * the impact factor is negative; when the deviation value is lower than the reference temperature setting value, the value of the deviation value * the impact factor is positive.
[0058] According to the comparison between the weighted reward value of temperature zone 3 close to temperature zone 4 and the weighted reward value close to temperature zone 2, the temperature compensation value of temperature zone 3 is -2.5, and the temperature setting value of temperature zone 3 is 447.5 degrees.
[0059] Then analyze the temperature coupling changes of temperature zone 4 affected by temperature zone 5 and temperature zone 3 as the single point moves. The temperature coupling boundary position between temperature zone 4 and temperature zone 3 is the temperature measurement position 80, and the temperature coupling boundary position between temperature zone 4 and temperature zone 5 is the temperature measurement position 220. According to the deviation value between the first temperature measurement data and the reference temperature of 450 degrees, the positive feedback or negative feedback reward value Reward is established for each position point of temperature zone 5, temperature zone 3 and temperature zone 4 respectively.
[0060] Table 3 Feedback reward value weighted to obtain temperature compensation value for temperature zone 4
[0061]
[0062] According to the comparison between the weighted reward value of temperature zone 4 close to temperature zone 5 and the weighted reward value close to temperature zone 3, the temperature compensation value of temperature zone 4 is -3.0, and the temperature setting value of temperature zone 4 is 447.0 degrees.
[0063] Then analyze the temperature coupling changes of temperature zone 2 affected by temperature zone 3 and temperature zone 1 as the single point moves. The temperature coupling boundary position between temperature zone 2 and temperature zone 3 is the temperature measurement position -80, and the temperature coupling boundary position between temperature zone 2 and temperature zone 1 is the temperature measurement position -220. According to the deviation value between the first temperature measurement data and the reference temperature of 450 degrees, the positive feedback or negative feedback reward value Reward is established for each position point of temperature zone 3, temperature zone 1 and temperature zone 2 respectively.
[0064] Table 4 Feedback reward value weighted to obtain temperature compensation value for temperature zone 2
[0065]
[0066] According to the comparison between the weighted reward value of temperature zone 4 close to temperature zone 5 and the weighted reward value close to temperature zone 3, the temperature compensation value of temperature zone 4 is -1.6, and the temperature setting value of temperature zone 2 is 448.4 degrees.
[0067] The temperature compensation values for temperature zone 1 and temperature zone 5 can be obtained by analogy.
[0068] The constant temperature zone temperature compensation method for horizontal liquid phase epitaxial growth equipment based on reinforcement learning of the present invention does not depend on whether the operator understands the relative position and dimension diagram of the external galvanic couple of the furnace body and the position and dimension diagram of the internal galvanic sensor; the constant temperature zone temperature compensation value can be calculated in a relatively small number of times, reducing the time for equipment downtime maintenance; this constant temperature zone temperature compensation method has strong repeatability and simple implementation, and the obtained five temperature zone set temperature compensation values have significant effects, reducing the intermediate process of manual calculation, and is applicable to semiconductor equipment with five temperature zones.
[0069] Specifically, the processes in this embodiment include: HgCdTe epitaxial growth, high temperature annealing, high temperature oxidation, TEOS, POLY and other processes that require semiconductor process equipment with a constant temperature zone.
[0070] The process temperature of HgCdTe epitaxial growth equipment is 450 degrees Celsius and the pressure is 750 Torr, and a constant flow of hydrogen is required to maintain a uniform temperature field in the constant temperature zone.
[0071] Furthermore, since temperature zone 3 is located in the middle of the graphite boat, and temperature zones 2 and 4 are located on both sides of the graphite boat, in the constant temperature zone temperature compensation method, the temperature compensation value of temperature zone 3 is obtained first, and then the temperature compensation values of temperature zones 2 and 4 are obtained in sequence, and finally the temperature compensation values of temperature zones 1 and 5 are obtained. The initial temperature compensation value is iteratively updated to the temperature compensation value obtained through the trained strategy model.
[0072] As can be seen from the above, in this embodiment, the corresponding relationship between the temperature corresponding to the 25 total measurement points measured by the temperature measuring thermocouple and the position of the set temperature of 450 degrees in the five temperature zones, as well as the deviation value from the set temperature of the first group of five temperature zones, is combined with the Q-Learning reinforcement learning training strategy to calculate the compensation value of the set temperature of the five temperature zones. The calculation is simple and easy to implement. It is not limited to the total number of measurement points. There is no need for operators to participate in the constant temperature zone temperature compensation, nor is there a need to obtain a temperature compensation table for dynamic adjustment. It only needs to be implemented based on temperature data, which can save labor costs in the equipment maintenance process and improve equipment operation efficiency. Figure 5 Shown are before and after comparisons of the constant temperature zone temperature compensation method using a reinforcement learning trained policy model.
[0073] Specifically, the temperature control system includes a programmable logic controller (PLC), thermocouples, a DC power supply, a temperature controller, and a host computer. The thermocouples measure the temperature inside the reaction chamber or furnace. The PLC transmits the host computer's output signal to the electrical actuator to control heating and power on. The DC power supply provides the required current. The temperature controller performs closed-loop feedback control based on the temperatures of the furnace thermocouples and the thermocouples inside the reaction chamber. The host computer sends control signals to the slave computer, displays the set temperature and the actual temperature inside the reaction chamber, and provides feedback on the temperature measurement data from the constant temperature zone.
[0074] In this embodiment, the reinforcement learning training strategy refers to a computational method for understanding and automatically processing goal-oriented learning and decision-making problems. For example, at any time step t, the agent observes the state S of the current environment. t , get the corresponding reward value r t Based on these states and reward information, the agent decides how to act. The agent performs action a t Then get new feedback from the environment and obtain the state s of the next time step t+1 and reward r t+1 The reinforcement learning training strategy emphasizes that the agent learns through direct interaction with the environment, with the goal of maximizing the cumulative reward to obtain the optimal action strategy. The interaction process between the agent and the environment can be formalized as a Markov decision process (MDP). The Markov decision process is represented by a five-tuple<S,A,P,R,γ> To describe, where S is the state space set of the environment where the agent is located, A is the action space set, P: S×A×S→[0,1] is the state transition function, and R: is the reward function, and γ∈(0,1] is the discount factor. The initial state is determined by the initial state distribution ρ0. In general, the policy π is a probability distribution mapping from state to action: π: S→P(A=a|S). At each discrete time step t, the agent observes the current state s t ∈S, use the strategy to select action a t ∈A, the environment is represented by the transformation function P(s t+1|s t ,a t ) Transition to the next state s t+1 .
[0075] The reinforcement learning training strategy model of this embodiment is a value-based approach. The behavior of the agent when interacting with the environment is generated based on a certain strategy π. The value function corresponding to the strategy is used to estimate the expected reward of the agent in a given state or state action. It includes the following steps:
[0076] Value-based methods usually define the following action-value function to evaluate the quality of the strategy:
[0077]
[0078] Optimal value function Q * (s,a) gives the maximum expected return value among all strategies for each state-action pair, that is, it satisfies:
[0079] Q * (s,a)=maxπE[R t |s t =s,a t =a,π]
[0080] The policy with the best value function is the optimal policy. For a given MDP, although the optimal value function is unique, there may be multiple optimal policies. The optimal action-value function follows the Bellman equation:
[0081]
[0082] Q-Learning is a classic algorithm for learning action-value functions. It uses a Q-value table to record the action-value function, calculates the optimal value function, and then obtains the optimal policy. The core iterative update function of this algorithm is as follows:
[0083]
[0084] Where α∈(0,1) is the learning rate, which is used to weigh the influence of recently acquired knowledge and previously known knowledge. The optimal policy for each state selects the action that maximizes the Q value.
[0085] In the Q-Learning algorithm, a reinforcement learning training strategy, an agent starts from an arbitrary initial state and takes actions to move to new states until it reaches a target state. Each time the agent completes the process of moving from an initial state to a target state is called an episode. In each episode, the agent selects an action in its current state and applies it to the environment, thereby moving to a new state. This process continues until the agent reaches the target state or reaches a maximum number of steps. The Q-Learning algorithm can be divided into the following steps:
[0086] Step 1: Given the parameter γ and reward matrix R
[0087] Step 2: Set Q = 0
[0088] Step 3: For each episode:
[0089] 3.1 Randomly select an initial state s
[0090] 3.2 If the optimal target state is not reached, continue to perform the following steps
[0091] Select an action a from all possible behaviors in the current state s;
[0092] Using the selected behavior a, get the next state s′;
[0093] According to the formula
[0094] Let s = s′.
[0095] According to the above steps, the Q-Learning algorithm can be used to achieve automatic optimization and find the optimal solution to achieve the required objective function.
[0096] Since the optimization goal is to minimize TempBias, the smaller the TempBias, the larger the reward. This paper uses the following formula to calculate Reward. The final result can be calculated using a value table based on the number of states and actions. All actions that the agent can perform in Q-learning are listed as all states.
[0097] The reward formula is: Reward = 100000 / TempBias; Table 2 Q-Learning state / action value table
[0098]
[0099] Specifically, in the Q-Learning reinforcement learning training strategy model, such as Figure 4As shown, the agent's initial state is S0. The agent randomly selects an action (assuming it's A1). After completing the action, the temperature control system uses the changed variables to calculate the corresponding process parameters. After returning to the trained policy model, the parameters are substituted into the objective function to calculate TempBias. The reward mechanism compares the changes in TempBias and assigns a positive or negative reward to the agent, which then proceeds to the next iteration.
[0100] After completing the action, the agent enters state S1 and randomly selects another action (assuming it is A2). The training strategy model then gives the agent new feedback based on the reward mechanism's judgment of TempBias. The code can be used to obtain the agent's current position and the Q value of the current action.
[0101] The present invention aims at the fact that the uniformity of the temperature field in the constant temperature zone of the horizontal liquid phase epitaxial growth equipment directly affects the substrate coating effect in the epitaxial process stage. This constant temperature zone temperature compensation method has strong repeatability and is simple to implement. It only needs to set 5 input values (the upper and lower temperature limits of the temperature uniformity of the constant temperature zone, the initial temperature compensation value, the single movement distance, and the total number of measurement points) according to the existing constant temperature zone data in combination with the temperature compensation method to obtain the set temperature compensation values of the five temperature zones. The effect is significant and does not depend on whether the operator is experienced. The intermediate process of manual calculation is reduced, the maintenance cost is low, the practicability is strong, the compensation value of the set temperature is obtained more quickly, the efficiency of pulling the constant temperature zone is improved, and the maintenance time of equipment downtime is reduced.
[0102] The constant temperature zone temperature compensation method of the present invention is not limited to horizontal liquid phase epitaxial growth equipment, but can also be used in other semiconductor equipment with five temperature zones and different processes, such as diffusion, epitaxy, TEOS, and Poly process equipment.
[0103] This method measures the actual temperature of the constant temperature zones at equal intervals by using a single-point movement and single-point stabilization time method. Based on a first set of data, where the initial set temperature of the five zones is 450°C and the initial compensation value is 0, and a reinforcement learning training strategy model, the set temperature compensation values for the five zones are accurately derived. This constant temperature zone temperature compensation method is independent of the operator's familiarity with the equipment, saving both the labor costs of temperature compensation and the downtime costs associated with equipment maintenance. The core measurement area of this constant temperature zone is located within the graphite boat, closest to the epitaxial growth site, enabling sensitive monitoring of temperature changes during this process.
[0104] The present invention also discloses a constant temperature zone temperature compensation system for a horizontal liquid phase epitaxial growth device based on reinforcement learning, comprising:
[0105] A temperature acquisition module is used to obtain temperature data of multiple measurement points in a constant temperature zone of a horizontal liquid phase epitaxial growth device;
[0106] The temperature analysis module is used to obtain the maximum value, minimum value, average value and standard deviation of each temperature data; at the same time, each temperature data is compared with the initial set temperature to obtain the deviation value;
[0107] The temperature compensation module is used to input the maximum, minimum, average, standard deviation and deviation values of each temperature data as input into a pre-built and trained reinforcement learning training strategy model to obtain the temperature compensation value for each temperature zone; wherein the reinforcement learning training strategy model presets the mapping relationship between the input and the temperature compensation value.
[0108] The present invention further discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the above-described method. The present invention also discloses a computer device comprising a memory and a processor connected to each other, the memory having a computer program stored thereon, which, when executed by the processor, performs the steps of the above-described method. The constant temperature zone temperature compensation system, medium, and device of the present invention correspond to the above-described constant temperature zone temperature compensation method and similarly have the advantages described for the above-described constant temperature zone temperature compensation method.
[0109] The present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by hardware related to computer program instructions. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned method embodiment when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. Computer-readable storage media include: any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. The memory is used to store computer programs and / or modules, and the processor implements various functions by running or executing computer programs and / or modules stored in the memory, and calling data stored in the memory. The memory may include high-speed random access memory and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device, etc.
[0110] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A temperature compensation method for a constant temperature zone of a horizontal liquid phase epitaxial growth device based on reinforcement learning, characterized in that: Including steps: S1. Acquire temperature data of multiple measurement points in a constant temperature zone of a horizontal liquid phase epitaxial growth device; S2. Obtain the maximum value, minimum value, average value, and standard deviation of each temperature data; and compare each temperature data with the initial set temperature to obtain a deviation value; S3. Input the maximum value, minimum value, average value, standard deviation and deviation value of each temperature data as input into the pre-built and trained reinforcement learning training strategy model to obtain the temperature compensation value of each temperature zone; The reinforcement learning training strategy model presets a mapping relationship between the input quantity and the temperature compensation value.
2. The method for temperature compensation in a constant temperature zone of a horizontal liquid phase epitaxial growth device based on reinforcement learning according to claim 1, characterized in that: In step S3 , the temperature compensation value of the middle temperature zone is first obtained, and then the temperature compensation values of each temperature zone are obtained one by one in sequence toward the outside.
3. The method for temperature compensation in a constant temperature zone of a horizontal liquid phase epitaxial growth device based on reinforcement learning according to claim 2, characterized in that: The process of obtaining the temperature compensation value of the intermediate temperature zone is: According to the deviation between each temperature data and the reference temperature, the feedback reward value Reward is established for each position point in the intermediate temperature zone; According to the feedback reward value Reward of each position point in the intermediate temperature zone, the weighted reward value of the intermediate temperature zone close to the adjacent temperature zone is obtained, and finally the temperature compensation value of the intermediate temperature zone is obtained, and then the temperature setting value of the intermediate temperature zone is obtained.
4. The method for temperature compensation in a constant temperature zone of a horizontal liquid phase epitaxial growth device based on reinforcement learning according to claim 1, 2 or 3, characterized in that: In step S1, the total number of measuring points of the temperature measuring thermocouple is n, the single moving distance is l, the initial set temperature of each temperature zone is k degrees, the single point stabilization time is t, the initial temperature compensation value is 0, and the position relationship diagram of each temperature measuring point on the internal thermocouple sensor corresponding to each temperature zone of the furnace body is obtained.
5. The method for temperature compensation in a constant temperature zone of a horizontal liquid phase epitaxial growth device based on reinforcement learning according to claim 4, characterized in that: Wherein n=25; l=20 mm; k=450 degrees; t=120 s.
6. A constant temperature zone temperature compensation system for horizontal liquid phase epitaxial growth equipment based on reinforcement learning, characterized in that: include: A temperature acquisition module is used to obtain temperature data of multiple measurement points in a constant temperature zone of a horizontal liquid phase epitaxial growth device; The temperature analysis module is used to obtain the maximum value, minimum value, average value and standard deviation of each temperature data; at the same time, each temperature data is compared with the initial set temperature to obtain the deviation value; The temperature compensation module is used to input the maximum, minimum, average, standard deviation and deviation values of each temperature data into a pre-built and trained reinforcement learning training strategy model to obtain the temperature compensation value of each temperature zone; The reinforcement learning training strategy model presets a mapping relationship between the input quantity and the temperature compensation value.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program performs the steps of the method according to any one of claims 1 to 5.
8. A computer device comprising a memory and a processor connected to each other, wherein a computer program is stored in the memory, wherein: When the computer program is executed by a processor, the computer program performs the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Indoor space temperature and humidity regulation and control method and system
CN115717758A
Monolithic silicon epitaxial equipment reaction chamber temperature control method and device based on active disturbance rejection control
CN118034411A