Heat treatment method for manufacturing semiconductor device
By constructing and training a heat treatment model including the first processing module and the second control module, the problem of difficult to balance temperature deviation and time cost in traditional heat treatment methods is solved, and the efficient, energy-saving and high-quality heat treatment effects of the semiconductor device manufacturing process are achieved.
Patent Information
- Application Number
- CN202510607105.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional heat treatment methods cannot effectively balance the balance between heat treatment effect, temperature deviation and time cost, and it is difficult to achieve efficient and energy-saving production goals while ensuring device performance.
A heat treatment model including the first processing module and the second control module is constructed and trained, and multiple converters and networks are updated by obtaining historical heat treatment data, and the action generation network, effect evaluation network, temperature deviation evaluation network and time cost evaluation network are optimized to achieve accurate control of the heat treatment process.
Accurate control of the heat treatment process of semiconductor device manufacturing is achieved, the quality and efficiency of heat treatment is improved, energy consumption and time cost are reduced, and the performance and consistency of the device is improved.
Smart Images

Figure CN120492930A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semiconductor manufacturing, and more particularly, to a heat treatment method for manufacturing semiconductor devices. Background Art
[0002] In the manufacturing process of semiconductor devices, heat treatment is a key link. Its precise control of parameters such as the surface temperature distribution, heating rate, and cooling rate of the wafer directly affects the performance and quality of the device. Traditional heat treatment methods usually rely on fixed process parameters and empirical control methods. Although they can meet production needs to a certain extent, their limitations gradually become apparent when faced with complex process requirements and changing production environments. For example, it is difficult for traditional methods to accurately adjust the temperature distribution in real time to adapt to the process requirements of different stages, resulting in poor temperature uniformity on the wafer surface, which in turn affects the performance consistency of the device. In addition, traditional heat treatment processes often take a long time to achieve the desired effect and have high energy consumption, which not only increases production costs but also reduces production efficiency.
[0003] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the prior art: traditional heat treatment methods cannot effectively balance the heat treatment effect, temperature deviation and time cost, and it is difficult to achieve efficient and energy-saving production goals while ensuring device performance. Summary of the Invention
[0004] The present invention provides a heat treatment method for semiconductor device manufacturing, comprising: Constructing a thermal treatment model, the model comprising a first processing module and a second control module; Training the model, comprising: Step 1: Acquire historical heat treatment data, wherein the historical heat treatment data includes first temperature information, second temperature information, treatment action, heat treatment effect score, temperature deviation cost, and time cost, wherein the first temperature information and the second temperature information correspond to adjacent moments; Step 2: updating the first processing module, which includes a temperature converter, an action-temperature converter, a temperature deviation converter, and a time cost converter, including: using the temperature converter to process the first temperature information and the second temperature information respectively to obtain a first feature map and a second feature map, and normalizing the first feature map and the second feature map respectively to obtain a third feature map and a fourth feature map; using the action-temperature converter to process the action and the third feature map to obtain a fifth feature map; using the temperature deviation converter to convert the third feature map and the fourth feature map respectively to obtain a first temperature deviation prediction value and a second temperature deviation prediction value; using the time cost converter to convert the third feature map and the fourth feature map respectively to obtain a first time cost prediction value and a second time cost prediction value; constructing a first objective function based on the second feature map, the fifth feature map, the temperature deviation cost, the first temperature deviation prediction value, the second temperature deviation prediction value, the time cost cost, the first time cost prediction value, and the second time cost prediction value, and minimizing the first objective function to update each converter; Step 3: updating the second control module, wherein the second control module includes an action generation network, an effect evaluation network, a temperature deviation evaluation network, and a time cost evaluation network; Step 4: Repeat steps 1 to 3 until the preset number of training times is exceeded to obtain the trained model; The temperature information at the current moment is obtained and input into the trained action generation network to obtain the current processing action to control the heat treatment process.
[0005] Furthermore, updating the second control module includes: Inputting the first temperature information and the first feature map into the action generation network to obtain a first probability distribution; Inputting the first characteristic map and the second characteristic map into the effect evaluation network, the temperature deviation evaluation network, and the time cost evaluation network, respectively, to obtain corresponding heat treatment effect values, temperature deviation values, and time cost values, wherein the heat treatment effect value includes a first effect value corresponding to the first characteristic map, the temperature deviation value includes a first temperature deviation value corresponding to the first characteristic map, and the time cost value includes a first time cost value corresponding to the first characteristic map; Inputting the heat treatment effect score and the heat treatment effect value into the effect advantage evaluation function to obtain the effect advantage evaluation value; Inputting the temperature deviation cost and the temperature deviation value into the temperature deviation advantage evaluation function to obtain a temperature deviation evaluation value; Input the time cost price and the time cost value into the time cost advantage evaluation function to obtain the time cost evaluation value; The action generation network is optimized based on the effect advantage evaluation value, temperature deviation evaluation value, time cost evaluation value and the first probability distribution; the effect evaluation network is optimized based on the heat treatment effect score and the first effect value; the temperature deviation evaluation network is optimized based on the temperature deviation cost and the first temperature deviation value; and the time cost evaluation network is optimized based on the time cost cost and the first time cost value.
[0006] Furthermore, the temperature information includes wafer surface temperature distribution, heating rate, cooling rate, duration of the current processing stage, whether the current processing stage is an annealing stage, and whether the current processing stage reaches a preset temperature threshold; The heat treatment effect score is calculated by an effect score function, which is constructed based on wafer surface temperature uniformity, processing time, energy consumption and device performance. The temperature deviation cost is calculated by a temperature deviation cost function constructed based on temperature distribution and target temperature distribution. The time cost is calculated by a time cost function constructed based on processing time and a preset time threshold.
[0007] Furthermore, the expression of the first objective function is: in, represents the first objective function value, Represent the loss weight coefficients, represents the fifth eigenmap, represents the second feature map, represents the root mean square error, and Represent the predicted and actual temperature deviation values, and denote the predicted and actual time cost values, respectively. represents the KL divergence.
[0008] Furthermore, the action generation network is optimized based on the effect advantage evaluation value, the temperature deviation evaluation value, the time cost evaluation value, and the first probability distribution, including: Inputting the effect advantage evaluation value, the temperature deviation evaluation value, the time cost evaluation value, and the first probability distribution into a second objective function, and obtaining an optimized action generation network by minimizing the second objective function; Optimizing the effect evaluation network based on the heat treatment effect score and the first effect value, including: inputting the heat treatment effect score and the first effect value into a loss function of the effect evaluation network and optimizing to obtain an optimized effect evaluation network; Optimizing the temperature deviation evaluation network based on the temperature deviation cost and the first temperature deviation value, including: inputting the temperature deviation cost and the first temperature deviation value into a loss function of the temperature deviation evaluation network and optimizing to obtain an optimized temperature deviation evaluation network; Optimizing a time cost evaluation network based on the time cost price and the first time cost value includes: inputting the time cost price and the first time cost value into a loss function of the time cost evaluation network and optimizing to obtain an optimized time cost evaluation network.
[0009] Furthermore, it also includes: Execute the current processing action to obtain the temperature information at the next moment; The current heat treatment effect score is calculated according to the effect score function; The current temperature deviation cost is calculated according to the temperature deviation cost function; The current time cost is calculated according to the time cost function; Construct current state data based on the temperature information at the current moment, the temperature information at the next moment, the current processing action, the current heat treatment effect score, the current temperature deviation cost, and the current time cost; The first processing module and the second control module are updated based on the current state data, and the updated first processing module and the second control module are used to process the temperature information at the next moment.
[0010] Furthermore, the expression of the effect scoring function is: in, Indicates at time Heat treatment effect score, Represent the weight coefficients, Indicates time Temperature uniformity, Indicates processing efficiency, Indicates energy consumption, Indicates device performance; The expression of the temperature deviation cost function is: in, represents the temperature deviation cost, Indicates the The temperature of the measuring point, Indicates the target temperature, Indicates the number of measurement points; The expression of the time cost function is: in, Indicates the time cost, Indicates the actual processing time, Indicates the preset time threshold.
[0011] Furthermore, the expression of the effect advantage evaluation function is: in, Indicates at time The effect advantage evaluation value, represents the discount factor, represents the number of stages of advantage assessment, Indicates the The timing difference error at the moment, Indicates the The status value at the moment.
[0012] Furthermore, the expression of the second objective function is: in, represents the second objective function value, Indicates that the status Next select action The probability of represents the effect advantage evaluation value, and are the Lagrange multipliers for temperature deviation and time cost, respectively, and Represent the temperature deviation cost and time cost respectively.
[0013] Furthermore, it also includes: The root mean square error is used to construct the loss functions of the effect evaluation network, temperature deviation evaluation network and time cost evaluation network. The loss functions of each network are minimized through the gradient descent optimization algorithm, and the parameters of each network are updated to obtain the optimized effect evaluation network, temperature deviation evaluation network and time cost evaluation network.
[0014] According to the above-mentioned embodiment of the present invention, there are at least the following beneficial effects: the heat treatment method can realize precise control of the heat treatment process of semiconductor device manufacturing by constructing and training a model including a first processing module and a second control module. During the training process, the multiple converters in the first processing module are updated using historical heat treatment data, which can accurately map and predict the characteristics of temperature information, processing actions, etc., thereby providing a reliable reference basis for the heat treatment process. At the same time, by optimizing the action generation network, effect evaluation network, temperature deviation evaluation network and time cost evaluation network in the second control module, a better processing action can be generated, and the heat treatment effect, temperature deviation and time cost can be more accurately evaluated, thereby effectively improving the quality and efficiency of heat treatment, reducing energy consumption and time cost, and improving the performance and consistency of semiconductor devices.
[0015] Furthermore, in practical applications, this method can be used to input the trained action generation network based on the current temperature information, quickly obtaining the current processing action to control the heat treatment process, thereby achieving real-time dynamic regulation of the heat treatment process. Furthermore, by constructing current state data after executing the current processing action and updating the first processing module and second control module based on this data, the model performance can be continuously optimized to better adapt to various changes and requirements in the heat treatment process, further enhancing the stability and reliability of the heat treatment process, providing strong technical support for semiconductor device manufacturing, and helping to promote the development of the semiconductor industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily apparent by reading the following detailed description with reference to the accompanying drawings, in which several embodiments of the present invention are shown by way of example and not limitation, in which: Figure 1 A schematic flow chart of a heat treatment method for semiconductor device manufacturing provided in one embodiment of the present invention. DETAILED DESCRIPTION
[0017] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided solely to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. Rather, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0018] Those skilled in the art will appreciate that the embodiments of the present invention may be implemented as a system, apparatus, device, method, or computer program product. Therefore, the present invention may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software.
[0019] It should be noted that any number of elements in the drawings is for illustration only and not for limitation, and any naming is only for distinction and does not have any limiting meaning.
[0020] Reference below Figure 1 , Figure 1 Schematic diagram of a heat treatment method for semiconductor device manufacturing provided by one embodiment of the present invention. Figure 1 As shown, a heat treatment method for semiconductor device manufacturing includes: Constructing a thermal treatment model, the model comprising a first processing module and a second control module; Training the model, comprising: Step 1: Acquire historical heat treatment data, wherein the historical heat treatment data includes first temperature information, second temperature information, treatment action, heat treatment effect score, temperature deviation cost, and time cost, wherein the first temperature information and the second temperature information correspond to adjacent moments; Step 2: updating the first processing module, which includes a temperature converter, an action-temperature converter, a temperature deviation converter, and a time cost converter, including: using the temperature converter to process the first temperature information and the second temperature information respectively to obtain a first feature map and a second feature map, and normalizing the first feature map and the second feature map respectively to obtain a third feature map and a fourth feature map; using the action-temperature converter to process the action and the third feature map to obtain a fifth feature map; using the temperature deviation converter to convert the third feature map and the fourth feature map respectively to obtain a first temperature deviation prediction value and a second temperature deviation prediction value; using the time cost converter to convert the third feature map and the fourth feature map respectively to obtain a first time cost prediction value and a second time cost prediction value; constructing a first objective function based on the second feature map, the fifth feature map, the temperature deviation cost, the first temperature deviation prediction value, the second temperature deviation prediction value, the time cost cost, the first time cost prediction value, and the second time cost prediction value, and minimizing the first objective function to update each converter; Step 3: updating the second control module, wherein the second control module includes an action generation network, an effect evaluation network, a temperature deviation evaluation network, and a time cost evaluation network; Step 4: Repeat steps 1 to 3 until the preset number of training times is exceeded to obtain the trained model; The temperature information at the current moment is obtained and input into the trained action generation network to obtain the current processing action to control the heat treatment process.
[0021] It should be noted that the present invention proposes a heat treatment method for semiconductor device manufacturing. Its core lies in constructing and training a heat treatment model comprising a first processing module and a second control module. The first processing module primarily extracts and converts features from temperature information and other data, while the second control module generates processing actions and evaluates heat treatment effects, temperature deviation, and time costs. During the training process, the model is optimized and updated by acquiring historical heat treatment data, including first and second temperature information, processing actions, heat treatment effect scores, temperature deviation costs, and time costs. This data reflects key parameters and performance indicators during the heat treatment process. By analyzing and processing these data, the model's performance and the control accuracy of the heat treatment process can be effectively improved. Specifically, the first and second temperature information refer to temperature data at adjacent moments, describing temperature changes during the heat treatment process. Processing actions refer to operating instructions for equipment during the heat treatment process, such as adjusting heating power and controlling cooling rates. The heat treatment effect score is calculated using an effect scoring function and is used to comprehensively evaluate the quality of the heat treatment process. The temperature deviation costs and time costs are calculated using corresponding cost functions and are used to measure the impact of temperature deviation and time costs on the heat treatment process.
[0022] Specifically, the first processing module includes a temperature converter, an action-to-temperature converter, a temperature deviation converter, and a time-cost converter. The temperature converter converts input temperature information into a feature map for subsequent processing. The action-to-temperature converter associates processing actions with temperature features to generate action-related feature maps. The temperature deviation converter and the time-cost converter predict temperature deviation and time cost, respectively, providing a basis for subsequent optimization. The second control module includes an action generation network, an effect evaluation network, a temperature deviation evaluation network, and a time-cost evaluation network. The action generation network generates a probability distribution of processing actions based on the input temperature information; the effect evaluation network evaluates the heat treatment effect; the temperature deviation evaluation network and the time-cost evaluation network evaluate temperature deviation and time cost, respectively. During training, the converters in the first processing module are updated by minimizing a first objective function, which comprehensively considers the temperature deviation cost, time cost, and the error between the predicted value and the true value. When updating the second control module, by optimizing the action generation network, effect evaluation network, temperature deviation evaluation network, and time-cost evaluation network, the model can generate more optimal processing actions and more accurately evaluate various indicators of the heat treatment process. The first objective function is expressed as follows: The first objective function value is equal to the loss weight coefficient 1 multiplied by the root mean square error 1 plus the loss weight coefficient 2 multiplied by the root mean square error 2 plus the loss weight coefficient 3 multiplied by the KL divergence. The root mean square error measures the difference between the predicted value and the true value, while the KL divergence measures the similarity between probability distributions. These parameters and concepts together form the foundation of model training. By properly setting and optimizing them, model performance can be effectively improved.
[0023] Preferably, the thermal treatment model construction process can be further refined. First, based on the process requirements and historical data of semiconductor device manufacturing, the model's input parameters are determined, including wafer surface temperature distribution, heating rate, cooling rate, and the duration of the current processing phase. These parameters can comprehensively reflect key information during the thermal treatment process. Then, based on these input parameters, the network structures of the first processing module and the second control module are constructed. For example, the temperature converter can employ a multi-layer neural network structure, extracting features layer by layer to convert temperature information into a feature map. The action-temperature converter can fuse the processing action with temperature features to generate a feature map related to the action. The temperature deviation converter and the time cost converter can respectively employ regression models to predict temperature deviation and time cost. In the second control module, the action generation network can employ a reinforcement learning algorithm to generate an optimal probability distribution for the processing action based on the input temperature information. The effect evaluation network, the temperature deviation evaluation network, and the time cost evaluation network can respectively employ supervised learning algorithms to evaluate the thermal treatment effect, temperature deviation, and time cost using training data. During model training, a preset number of training runs, such as 1,000, can be set. Each training run randomly samples data from historical thermal processing data to improve the model's generalization capabilities. This approach allows for the construction of a high-performance, adaptable thermal processing model, effectively improving the quality and efficiency of thermal processing in semiconductor device manufacturing.
[0024] In some embodiments, updating the second control module includes: Inputting the first temperature information and the first feature map into the action generation network to obtain a first probability distribution; Inputting the first characteristic map and the second characteristic map into the effect evaluation network, the temperature deviation evaluation network, and the time cost evaluation network, respectively, to obtain corresponding heat treatment effect values, temperature deviation values, and time cost values, wherein the heat treatment effect value includes a first effect value corresponding to the first characteristic map, the temperature deviation value includes a first temperature deviation value corresponding to the first characteristic map, and the time cost value includes a first time cost value corresponding to the first characteristic map; Inputting the heat treatment effect score and the heat treatment effect value into the effect advantage evaluation function to obtain the effect advantage evaluation value; Inputting the temperature deviation cost and the temperature deviation value into the temperature deviation advantage evaluation function to obtain a temperature deviation evaluation value; Input the time cost price and the time cost value into the time cost advantage evaluation function to obtain the time cost evaluation value; The action generation network is optimized based on the effect advantage evaluation value, temperature deviation evaluation value, time cost evaluation value and the first probability distribution; the effect evaluation network is optimized based on the heat treatment effect score and the first effect value; the temperature deviation evaluation network is optimized based on the temperature deviation cost and the first temperature deviation value; and the time cost evaluation network is optimized based on the time cost cost and the first time cost value.
[0025] It should be noted that the process of updating the second control module in the present invention is a key link in the heat treatment method. The second control module includes an action generation network, an effect evaluation network, a temperature deviation evaluation network, and a time cost evaluation network. Its purpose is to achieve precise control of the heat treatment process by optimizing these networks. During the updating process, the first temperature information and the first feature map are first input into the action generation network to obtain the probability distribution of the treatment action. Then, the first feature map and the second feature map are respectively input into the effect evaluation network, the temperature deviation evaluation network, and the time cost evaluation network to obtain the corresponding heat treatment effect value, temperature deviation value, and time cost value. These values reflect the impact of the current treatment action on the heat treatment process. By inputting the heat treatment effect score and the heat treatment effect value into the effect advantage evaluation function, an effect advantage evaluation value can be obtained, which is used to measure the contribution of the treatment action to the heat treatment effect. Similarly, by using the temperature deviation advantage evaluation function and the time cost advantage evaluation function, the temperature deviation evaluation value and the time cost evaluation value can be obtained respectively. These evaluation values are combined with the probability distribution of the action generation network to optimize the action generation network so that it can generate more optimal treatment actions. At the same time, the effect evaluation network is optimized based on the heat treatment effect score and the first effect value, the temperature deviation evaluation network is optimized based on the temperature deviation cost and the first temperature deviation value, and the time cost evaluation network is optimized based on the time cost cost and the first time cost value, thereby improving the performance of the entire control module.
[0026] Specifically, the first temperature information refers to the temperature data at a specific moment during the heat treatment process, and the first feature map is the feature representation obtained by processing the first temperature information using a temperature converter. The second feature map is the feature representation obtained by processing the second temperature information. These feature maps are abstract representations of temperature information used within the model to facilitate subsequent network processing and analysis. The effect evaluation network is used to evaluate the heat treatment effect. Its inputs include the first and second feature maps, and its output is a heat treatment effect value, which reflects the quality of the current heat treatment process. The temperature deviation evaluation network and the time cost evaluation network are used to evaluate temperature deviation and time cost, respectively. Their inputs also include the first and second feature maps, and their outputs are temperature deviation and time cost values, respectively. The effect advantage evaluation function, temperature deviation advantage evaluation function, and time cost advantage evaluation function are functions used to measure the difference between the evaluation value and the target value. These functions can be used to generate evaluation values to guide network optimization. The loss function is crucial in the optimization process, measuring the difference between the network output and the target value. By minimizing the loss function, the network parameters can be adjusted to bring its output closer to the target value.
[0027] Preferably, the process of updating the second control module can be further refined. When constructing the action generation network, a deep reinforcement learning algorithm, such as a policy gradient method, can be used. The input of the network includes the first temperature information and the first feature map, and through the multi-layer structure of the neural network, the probability distribution of the processing action is output. During the training process, specific processing actions can be obtained by sampling, and the network can be optimized based on the effect advantage evaluation value, the temperature deviation evaluation value, and the time cost evaluation value. For the effect evaluation network, the temperature deviation evaluation network, and the time cost evaluation network, a supervised learning algorithm can be used for training. The input of these networks includes the first feature map and the second feature map, and the output is the heat treatment effect value, the temperature deviation value, and the time cost value, respectively. During the training process, the root mean square error can be used as the loss function, and the network parameters can be updated using the gradient descent algorithm. For example, the root mean square error is used to measure the difference between the predicted value and the true value. By minimizing this error, the network output can be made closer to the true value. In a specific implementation, multiple training cycles can be set, and the network can be evaluated and adjusted in each cycle to ensure that the performance of the network continues to improve.
[0028] In some embodiments, the temperature information includes wafer surface temperature distribution, heating rate, cooling rate, duration of the current processing stage, whether the current processing stage is an annealing stage, and whether the current processing stage reaches a preset temperature threshold; The heat treatment effect score is calculated by an effect score function, which is constructed based on wafer surface temperature uniformity, processing time, energy consumption and device performance. The temperature deviation cost is calculated by a temperature deviation cost function constructed based on temperature distribution and target temperature distribution. The time cost is calculated by a time cost function constructed based on processing time and a preset time threshold.
[0029] It should be noted that the temperature information mentioned in the present invention is an important parameter in the heat treatment process, which includes multiple aspects, such as wafer surface temperature distribution, heating rate, cooling rate, duration of the current processing stage, whether the current processing stage is the annealing stage, and whether the preset temperature threshold is reached. These parameters jointly determine the effect and quality of the heat treatment. The heat treatment effect score is calculated by the effect score function, which comprehensively considers multiple factors such as wafer surface temperature uniformity, processing time, energy consumption and device performance. The temperature deviation cost is calculated by the temperature deviation cost function, which reflects the difference between the actual temperature distribution and the target temperature distribution. The time cost is calculated by the time cost cost function, which measures the deviation between the actual processing time and the preset time threshold. The setting of these parameters and functions enables the heat treatment process to control temperature and time more accurately, thereby improving the manufacturing quality of semiconductor devices.
[0030] Specifically, wafer surface temperature distribution refers to the temperature distribution at different locations on the wafer surface during the thermal treatment process. The heating rate and cooling rate represent the speed of temperature rise and fall, respectively. These two parameters have a significant impact on the thermal stress and final performance of the wafer. The duration of the current process refers to the duration of a specific stage in the thermal treatment process, such as the duration of the annealing stage. The annealing stage is a critical step in the thermal treatment process, used to improve the physical and chemical properties of the wafer. The preset temperature threshold is a target temperature set based on process requirements, and the actual temperature should be as close to this threshold as possible. The thermal treatment effectiveness score is a comprehensive indicator calculated using an effectiveness score function. This function weights factors such as temperature uniformity, processing time, energy consumption, and device performance to produce a score. The temperature deviation cost function is calculated by calculating the difference between the actual temperature distribution and the target temperature distribution, reflecting the accuracy of temperature control. The time cost function is calculated by comparing the actual processing time with the preset time threshold, measuring the accuracy of time control. The settings of these parameters and functions provide a quantitative basis for optimizing the thermal treatment process.
[0031] Preferably, the collection of temperature information can be achieved through high-precision temperature sensors, which can be distributed at different positions on the surface of the wafer to monitor temperature changes in real time. The heating rate and cooling rate can be adjusted by controlling the power of the heating equipment and the flow rate of the cooling system. When constructing the effect scoring function, the weight coefficient can be set according to the actual process requirements. For example, if the temperature uniformity has a greater impact on the performance of the device, a higher weight can be given. The temperature deviation cost function can be implemented by calculating the square difference between the actual temperature and the target temperature at each measurement point, and then summing the square differences of all measurement points. The time cost cost function can be implemented by calculating the difference between the actual processing time and the preset time threshold. If the actual time exceeds the preset time, additional costs will be incurred. In practical applications, the heat treatment process can be optimized by adjusting the parameters of the heating equipment and the parameters of the cooling system to maximize the heat treatment effect score while minimizing the temperature deviation cost and time cost.
[0032] In some embodiments, the expression of the first objective function is: in, represents the first objective function value, Represent the loss weight coefficients, represents the fifth eigenmap, represents the second feature map, represents the root mean square error, and Represent the predicted and actual temperature deviation values, and denote the predicted and actual time cost values, respectively. represents the KL divergence.
[0033] It should be noted that the first objective function mentioned in the present invention is a key tool for optimizing the first processing module in the thermal treatment model. The first objective function minimizes the loss of the model by comprehensively considering multiple factors, including temperature deviation cost, time cost, and the error between the predicted value and the true value, thereby improving the performance of the model. In this process, the loss weight coefficient is used to balance the impact of different factors on the objective function, the root mean square error is used to measure the difference between the predicted value and the true value, and the KL divergence is used to measure the similarity between probability distributions. By minimizing the first objective function, the various converters in the first processing module can be effectively updated so that they can process temperature information and generate processing actions more accurately.
[0034] Specifically, the construction of the first objective function involves multiple key parameters and concepts. The first objective function value is obtained by calculating the weighted sum of the loss weight coefficient and the corresponding error or divergence. This parameter is used to adjust the weight of different error terms in the objective function. Its value can be set according to actual needs. For example, if temperature deviation has a significant impact on the heat treatment process, the corresponding weight coefficient can be increased. The fifth and second feature maps are feature representations generated in the first processing module, corresponding to the feature conversion results of the processing action and temperature information, respectively. The root mean square error (RMSE) is a commonly used error metric that calculates the square root of the mean of the squared differences between the predicted value and the true value. Here, it is used to measure the prediction accuracy of temperature deviation and time cost. The KL divergence (Kullback-Leibler Divergence) measures the difference between two probability distributions and is used here to assess the similarity between the probability distribution of the model output and the target distribution. By minimizing the first objective function, the parameters of the temperature converter, action-temperature converter, temperature deviation converter, and time cost converter in the first processing module can be optimized to better adapt to various changes in the heat treatment process.
[0035] Preferably, when constructing the first objective function, the parameter settings and processing procedures can be further refined. For example, the loss weight coefficient can be dynamically adjusted according to the actual needs during the heat treatment process. If the impact of temperature deviation on device performance at a certain stage is more significant, the weight coefficient corresponding to the temperature deviation cost can be increased. When calculating the root mean square error, the predicted value and the true value can be normalized to eliminate the influence of different dimensions and orders of magnitude, so as to measure the error more accurately. For the calculation of KL divergence, a numerical stability method can be used to avoid numerical problems caused by too small probability values. During the model training process, the first objective function can be minimized by the gradient descent algorithm. The specific steps include: first calculating the gradient of the objective function with respect to the model parameters, and then adjusting the parameter values according to the gradient direction. In each iteration, small batch data can be used for training to improve training efficiency and the generalization ability of the model. In this way, the first processing module can be effectively optimized so that it performs better during the heat treatment process.
[0036] In some embodiments, optimizing the action generation network based on the effect advantage evaluation value, the temperature deviation evaluation value, the time cost evaluation value, and the first probability distribution includes: Inputting the effect advantage evaluation value, the temperature deviation evaluation value, the time cost evaluation value, and the first probability distribution into a second objective function, and obtaining an optimized action generation network by minimizing the second objective function; Optimizing the effect evaluation network based on the heat treatment effect score and the first effect value, including: inputting the heat treatment effect score and the first effect value into a loss function of the effect evaluation network and optimizing to obtain an optimized effect evaluation network; Optimizing the temperature deviation evaluation network based on the temperature deviation cost and the first temperature deviation value, including: inputting the temperature deviation cost and the first temperature deviation value into a loss function of the temperature deviation evaluation network and optimizing to obtain an optimized temperature deviation evaluation network; Optimizing a time cost evaluation network based on the time cost price and the first time cost value includes: inputting the time cost price and the first time cost value into a loss function of the time cost evaluation network and optimizing to obtain an optimized time cost evaluation network.
[0037] It should be noted that the process of optimizing action generation network, effect evaluation network, temperature deviation evaluation network and time cost evaluation network mentioned in the present invention is an important link in the heat treatment method. Optimizing action generation network is completed based on effect advantage evaluation value, temperature deviation evaluation value, time cost evaluation value and the first probability distribution, and the purpose is to enable action generation network to generate more optimal processing action. Optimizing effect evaluation network is carried out based on heat treatment effect scoring and the first effect value, and the purpose is to improve the evaluation accuracy of effect evaluation network to heat treatment effect. Similarly, the optimization of temperature deviation evaluation network and time cost evaluation network is based on temperature deviation cost and the first temperature deviation value, time cost cost and the first time cost value respectively, and the purpose is to improve the evaluation accuracy of these two networks to temperature deviation and time cost. Through these optimization steps, the control accuracy and efficiency of the whole heat treatment process can be improved.
[0038] Specifically, the first probability distribution refers to the probability distribution of each treatment action generated by the action generation network under a given state. The effect advantage evaluation value is calculated using the effect advantage evaluation function and is used to measure the advantage of selecting a specific action under a given state. The temperature deviation evaluation value and time cost evaluation value are used to measure the impact of temperature deviation and time cost on the heat treatment process, respectively. The second objective function is a key tool for optimizing the action generation network. It adjusts the parameters of the action generation network by comprehensively considering the effect advantage evaluation value, temperature deviation evaluation value, time cost evaluation value, and the first probability distribution. During the optimization process, the loss function of the effect evaluation network is constructed by comparing the heat treatment effect score with the first effect value and is used to adjust the parameters of the effect evaluation network. The loss functions of the temperature deviation evaluation network and the time cost evaluation network are constructed by comparing the temperature deviation cost with the first temperature deviation value and the time cost cost with the first time cost value, respectively, and are used to adjust the parameters of these two networks. Through these optimization steps, each network can more accurately evaluate various indicators of the heat treatment process and generate more optimal treatment actions.
[0039] Preferably, when optimizing the action generation network, a policy gradient method can be used, which is a reinforcement learning algorithm that adjusts the parameters of the policy network by maximizing the expected return. The specific steps include: first calculating the effect advantage evaluation value, the temperature deviation evaluation value, and the time cost evaluation value, and then inputting these evaluation values into the second objective function together with the first probability distribution. By minimizing the second objective function, the parameters of the action generation network can be updated so that it can generate better processing actions. When optimizing the effect evaluation network, the temperature deviation evaluation network, and the time cost evaluation network, the root mean square error can be used as the loss function, and these loss functions can be minimized by the gradient descent algorithm. For example, when optimizing the effect evaluation network, the root mean square error between the heat treatment effect score and the first effect value can be calculated, and then the parameters of the effect evaluation network can be adjusted based on the error. Similarly, for the temperature deviation evaluation network and the time cost evaluation network, the network parameters can also be adjusted by calculating the corresponding root mean square error. In practical applications, the weight coefficients in the loss function can be adjusted according to the specific requirements of the heat treatment process to balance the importance of different evaluation indicators.
[0040] In some embodiments, further comprising: Execute the current processing action to obtain the temperature information at the next moment; The current heat treatment effect score is calculated according to the effect score function; The current temperature deviation cost is calculated according to the temperature deviation cost function; The current time cost is calculated according to the time cost function; Construct current state data based on the temperature information at the current moment, the temperature information at the next moment, the current processing action, the current heat treatment effect score, the current temperature deviation cost, and the current time cost; The first processing module and the second control module are updated based on the current state data, and the updated first processing module and the second control module are used to process the temperature information at the next moment.
[0041] It should be noted that the process of executing the current processing action and constructing the current state data mentioned in the present invention is a key step for dynamically optimizing the model in the heat treatment method. After executing the current processing action, the system can obtain the temperature information of the next moment according to actual measurement, and calculate the current heat treatment effect score, temperature deviation price and time cost price based on this. These data, together with the temperature information of the current moment, processing action, etc., constitute the current state data. Subsequently, the first processing module and the second control module are updated using the current state data so that the model can be optimized and adjusted according to the latest heat treatment state, and a more accurate control strategy is provided for the temperature information processing of the next moment. This process realizes real-time feedback and dynamic optimization of the heat treatment process, and improves the precision and efficiency of heat treatment.
[0042] Specifically, the current processing action is a processing instruction selected based on a probability distribution generated by the action generation network and is used to control the actual operation of the thermal treatment equipment. The temperature information at the next moment refers to data such as the wafer surface temperature distribution measured by devices such as temperature sensors after the current processing action is executed. The thermal treatment effect score is calculated using an effect score function, which comprehensively considers factors such as wafer surface temperature uniformity, processing time, energy consumption, and device performance. The temperature deviation cost is calculated using a temperature deviation cost function and reflects the difference between the actual temperature distribution and the target temperature distribution. The time cost is calculated using a time cost cost function and measures the deviation between the actual processing time and a preset time threshold. The current state data includes the current temperature information, the next temperature information, the current processing action, the thermal treatment effect score, the temperature deviation cost, and the time cost. This data provides comprehensive information for model updates. When updating the first processing module and the second control module, model parameters are adjusted based on the current state data to optimize model performance.
[0043] Preferably, when constructing the current state data, the temperature information can be preprocessed, for example, by removing noise through a filtering algorithm to improve the accuracy of the data. When calculating the heat treatment effect score, the weight coefficients of each factor in the effect score function can be adjusted according to actual needs. For example, if the temperature uniformity at a certain stage has a more significant impact on the device performance, the weight coefficient corresponding to the temperature uniformity can be increased. When updating the first processing module and the second control module, a small batch gradient descent algorithm can be used to randomly extract a portion of the data from the current state data for training each time to improve the generalization ability and training efficiency of the model. For example, for the temperature converter in the first processing module, its parameters can be adjusted according to the temperature information in the current state data so that it can extract temperature features more accurately; for the action generation network in the second control module, its parameters can be adjusted according to information such as the heat treatment effect score and the temperature deviation cost in the current state data so that it can generate better processing actions. In this way, dynamic optimization of the heat treatment process can be achieved and the quality and efficiency of the heat treatment can be improved.
[0044] In some embodiments, the expression of the effect scoring function is: in, Indicates at time Heat treatment effect score, Represent the weight coefficients, Indicates time Temperature uniformity, Indicates processing efficiency, Indicates energy consumption, Indicates device performance; The expression of the temperature deviation cost function is: in, represents the temperature deviation cost, Indicates the The temperature of the measuring point, Indicates the target temperature, Indicates the number of measurement points; The expression of the time cost function is: in, Indicates the time cost, Indicates the actual processing time, Indicates the preset time threshold.
[0045] It should be noted that the effect scoring function, temperature deviation cost function and time cost cost function mentioned in the present invention are key indicators for quantifying the heat treatment process. The effect scoring function is used to comprehensively evaluate the heat treatment effect. It is based on multiple factors such as wafer surface temperature uniformity, processing efficiency, energy consumption and device performance, and the heat treatment effect score is obtained by calculating the weighted sum of these factors. The temperature deviation cost function is used to measure the difference between the actual temperature distribution and the target temperature distribution, and the temperature deviation cost is obtained by calculating the square sum of the temperature deviations of each measuring point. The time cost cost function is used to measure the deviation between the actual processing time and the preset time threshold, and the time cost is obtained by calculating the difference between the two. The setting of these functions enables the heat treatment process to be optimized through quantitative indicators, thereby improving the quality and efficiency of the heat treatment.
[0046] Specifically, the weight coefficient in the effect scoring function (such as ) is a parameter used to adjust the weight of different factors in the thermal treatment effect score, and its value can be set according to actual needs. For example, if temperature uniformity has a greater impact on device performance, the weight coefficient corresponding to temperature uniformity can be increased. The number of measurement points (N) in the temperature deviation cost function refers to the number of points set on the wafer surface for measuring temperature, and the temperature data of these measurement points are used to calculate the temperature deviation cost. The target temperature refers to the ideal temperature value set according to the thermal treatment process requirements. The deviation between the actual temperature and the target temperature reflects the accuracy of the thermal treatment process. The actual processing time and the preset time threshold in the time cost function are parameters used to measure the time efficiency of the thermal treatment process. If the actual processing time exceeds the preset time threshold, additional time cost will be incurred. The settings of these functions provide a quantitative basis for the optimization of the thermal treatment process, so that the model can be dynamically adjusted according to these quantitative indicators.
[0047] Preferably, when constructing the effect scoring function, the weight coefficients of different factors can be dynamically adjusted according to the specific needs of the heat treatment process. For example, in certain stages, processing efficiency may be more important, and the weight coefficient corresponding to the processing efficiency can be increased. When calculating the temperature deviation cost, the temperature data of the measurement point can be normalized to eliminate the influence of different dimensions and orders of magnitude, so as to more accurately measure the temperature deviation. For the time cost cost function, different time thresholds can be set according to actual production needs. For example, in some urgent production tasks, the time threshold can be appropriately lowered to improve production efficiency. In practical applications, by adjusting the parameters in these functions, the heat treatment process can be made as efficient and energy-saving as possible while ensuring quality. For example, by optimizing the temperature deviation cost function, the temperature distribution in the heat treatment process can be made more uniform; by optimizing the time cost cost function, the heat treatment time can be shortened and production efficiency can be improved.
[0048] In some embodiments, the expression of the effect advantage evaluation function is: in, Indicates at time The effect advantage evaluation value, represents the discount factor, represents the number of stages of advantage assessment, Indicates the The timing difference error at the moment, Indicates the The status value at the moment.
[0049] It should be noted that the effect advantage evaluation function mentioned in this invention is a tool for measuring the long-term benefits of a specific action during the heat treatment process. It obtains an effect advantage evaluation value by performing a weighted sum of the temporal difference error and state value at different moments. This evaluation value reflects the contribution of the current action to the future heat treatment effect, thus providing a basis for optimizing the action generation network. In this way, it is ensured that the generated treatment action not only performs well at the current moment but also continuously optimizes the heat treatment effect at subsequent moments.
[0050] Specifically, the discount factor (γ) in the effect advantage evaluation function measures the discount rate for future rewards and typically ranges from 0 to 1. A discount factor closer to 1 places greater emphasis on future rewards; a discount factor closer to 0 places greater emphasis on current rewards. The number of stages (n) represents the number of future time steps considered when calculating the effect advantage evaluation value and is typically set based on the specific requirements and computational complexity of the heat treatment process. Temporal difference error (TD error) is a key concept in reinforcement learning, measuring the difference between actual and expected rewards. The state value (V) represents the expected future reward at a given state. When calculating the effect advantage evaluation value, the TD error and state value at different moments are weighted and summed, with the weights determined by the discount factor. This approach yields an evaluation value that comprehensively considers both current and future effects, which can be used to guide the optimization of the action generation network.
[0051] Preferably, when constructing the effect advantage evaluation function, the value of the discount factor can be adjusted based on the specific requirements of the heat treatment process. For example, if the heat treatment process requires high short-term results, the discount factor can be set smaller; if long-term results are more important, the discount factor can be set larger. When calculating the temporal difference error, a multi-step temporal difference method can be used to improve the accuracy of the evaluation by considering returns over multiple time steps. For example, the number of stages can be set to 5 or 10 and dynamically adjusted based on the actual heat treatment process. In practical applications, by optimizing the effect advantage evaluation function, the action generation network can generate more proactive treatment actions. For example, in the initial stage of heat treatment, rapid temperature increase to reach the preset temperature may be more important. In this case, the discount factor and number of stages can be adjusted to ensure that the network-generated actions respond quickly. During the stable stage of heat treatment, temperature uniformity and energy consumption control can be emphasized, and parameters can be adjusted to make the network-generated actions more precise. In this way, dynamic optimization of the heat treatment process can be achieved, improving the quality and efficiency of the heat treatment.
[0052] In some embodiments, the expression of the second objective function is: in, represents the second objective function value, Indicates that the status Next select action The probability of represents the effect advantage evaluation value, and are the Lagrange multipliers for temperature deviation and time cost, respectively, and Represent the temperature deviation cost and time cost respectively.
[0053] It should be noted that the second objective function is a key tool for optimizing the action generation network. This objective function optimizes the probability distribution of the action generation network by comprehensively considering factors such as the effect advantage evaluation value, temperature deviation cost, and time cost. Specifically, it adjusts the parameters of the action generation network by maximizing the effect advantage evaluation value while minimizing the impact of temperature deviation and time cost. This optimization process enables the generated processing actions to not only improve the heat treatment effect, but also achieve a better balance between temperature deviation and time cost, thereby achieving precise control of the heat treatment process.
[0054] Specifically, the probability distribution refers to the probability of selecting a certain action in a given state, and this probability is output by the action generation network. The effect advantage evaluation value is calculated using the effect advantage evaluation function, which reflects the long-term advantage of selecting a specific action in a certain state. The temperature deviation cost and time cost cost respectively measure the impact of temperature deviation and time cost on the heat treatment process. The Lagrange multiplier is used to balance the weights of the temperature deviation cost and time cost cost in the objective function. In the second objective function, the parameters of the action generation network are optimized by maximizing the negative value of the effect advantage evaluation value and the weighted sum of the temperature deviation cost and time cost cost. This optimization method enables the network to take into account the control of temperature deviation and time cost while pursuing the heat treatment effect. In terms of specific parameter settings, the Lagrange multiplier can be adjusted according to the importance of temperature deviation and time cost in the heat treatment process. For example, if the temperature deviation has a greater impact on the heat treatment quality, the value of the Lagrange multiplier corresponding to the temperature deviation can be increased.
[0055] Preferably, when constructing the second objective function, the value of the Lagrange multiplier can be adjusted according to the specific requirements of the heat treatment process. For example, if the temperature deviation has a greater impact on the heat treatment quality, the value of the Lagrange multiplier corresponding to the temperature deviation can be increased to more strictly control the temperature deviation; if the time cost is more critical, the value of the Lagrange multiplier corresponding to the time cost can be increased. When optimizing the action generation network, a policy gradient method can be used to calculate the gradient of the second objective function with respect to the network parameters and adjust the parameter values according to the gradient direction. The specific steps include: first calculating the effect advantage evaluation value and the temperature deviation and time cost; then calculating the gradient of the second objective function based on these values; and finally updating the network parameters using a gradient descent algorithm. In practical applications, the performance of the action generation network can be gradually improved through multiple iterative optimizations, enabling it to generate more optimal treatment actions, thereby improving the overall efficiency and quality of the heat treatment process. In addition, the value of the Lagrange multiplier can be dynamically adjusted according to the dynamic characteristics of the heat treatment process to adapt to different process requirements.
[0056] In some embodiments, further comprising: The root mean square error is used to construct the loss functions of the effect evaluation network, temperature deviation evaluation network and time cost evaluation network. The loss functions of each network are minimized through the gradient descent optimization algorithm, and the parameters of each network are updated to obtain the optimized effect evaluation network, temperature deviation evaluation network and time cost evaluation network.
[0057] It should be noted that the use of root mean square error in the present invention to construct the loss functions of the effect evaluation network, the temperature deviation evaluation network, and the time cost evaluation network, and the loss functions of each network are minimized by the gradient descent optimization algorithm, and the parameters of each network are updated in order to optimize the performance of these evaluation networks so that they can more accurately evaluate various indicators in the heat treatment process. Root mean square error is a commonly used error measurement indicator that can quantify the difference between the predicted value and the true value. By minimizing the root mean square error, the network parameters can be adjusted so that the output of the network is closer to the true value, thereby improving the accuracy of the evaluation. The gradient descent optimization algorithm is a commonly used optimization method that adjusts the network parameters by calculating the gradient of the loss function to minimize the value of the loss function, thereby optimizing the performance of the network.
[0058] Specifically, the effect evaluation network, temperature deviation evaluation network, and time cost evaluation network are three independent networks used to evaluate the heat treatment effect, temperature deviation, and time cost during the heat treatment process. The input of the effect evaluation network includes various parameters in the heat treatment process, such as temperature distribution, processing time, etc., and the output is the heat treatment effect score; the input of the temperature deviation evaluation network is the actual temperature distribution and the target temperature distribution, and the output is the temperature deviation value; the input of the time cost evaluation network is the actual processing time and the preset time threshold, and the output is the time cost value. The root mean square error is obtained by calculating the square root of the mean of the squared difference between the predicted value and the true value, and is used to measure the difference between the predicted value and the true value. The gradient descent optimization algorithm calculates the gradient of the loss function with respect to the network parameters, and adjusts the parameter value according to the gradient direction to minimize the value of the loss function. During the optimization process, the step size of the parameter update can be controlled by setting the learning rate. The value of the learning rate is usually adjusted according to actual needs.
[0059] Preferably, when constructing the effect evaluation network, temperature deviation evaluation network, and time cost evaluation network, a multi-layer neural network structure can be employed, with each layer consisting of multiple neurons. Nonlinear properties are introduced through nonlinear activation functions, enabling the network to learn complex mapping relationships. For example, the effect evaluation network may comprise an input layer, a hidden layer, and an output layer. The input layer receives various parameters from the heat treatment process, the hidden layer extracts and transforms features from the learning data, and the output layer outputs the heat treatment effect score. During training, a mini-batch gradient descent algorithm can be used, with a random portion of the training data randomly sampled for training each time, to improve training efficiency and the network's generalization ability. Furthermore, different learning rates and optimizers (such as Adam and RMSprop) can be used to accelerate the training process and enhance optimization results. In practical applications, the network structure and parameters can be adjusted according to the specific requirements of the heat treatment process, such as increasing the number of hidden layers or neurons, to improve network performance and evaluation accuracy.
[0060] The above-mentioned various embodiments of the present invention have the following beneficial effects: the present invention can improve the accuracy and efficiency of the semiconductor heat treatment process. By constructing a collaborative optimization model including a first processing module and a second control module, the temperature change of the wafer can be accurately predicted and controlled, while optimizing time cost and energy consumption. Components such as temperature converters and action-temperature converters trained based on historical data can accurately extract temperature characteristics and generate control actions, while modules such as effect evaluation networks and temperature deviation evaluation networks can dynamically adjust parameters to achieve optimal heat treatment effects. Through multiple rounds of iterative training and real-time feedback mechanisms, the model performance can be continuously optimized, ultimately achieving more stable temperature uniformity, lower energy consumption and a more efficient production rhythm. In addition, this method can flexibly adapt to the needs of different heat treatment stages. Through the comprehensive calculation of the effect scoring function, the temperature deviation cost function, and the time cost cost function, the relationship between temperature control accuracy and production efficiency can be balanced. The action generation network can dynamically adjust the heating or cooling strategy based on real-time temperature information, while the advantage evaluation function can further optimize the long-term benefits of the control strategy. This intelligent control method based on multi-objective optimization can significantly improve the performance consistency of semiconductor devices while reducing resource waste in the production process, providing reliable technical support for the heat treatment process of advanced processes.
[0061] Furthermore, the storage medium of the embodiment of the present application stores program instructions that can implement all the above methods, wherein the program instructions can be stored in the above storage medium in the form of a software product, including a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or a terminal device such as a computer, server, mobile phone, or tablet.
[0062] The above descriptions merely illustrate some preferred embodiments of the present invention and the underlying technical principles. Those skilled in the art should understand that the scope of the invention encompassed by the embodiments of the present invention is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned inventive concept. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present invention.
Claims
1. A heat treatment method for semiconductor device manufacturing, characterized in that: include: Constructing a thermal treatment model, the model comprising a first processing module and a second control module; Training the model, comprising: Step 1: Acquire historical heat treatment data, wherein the historical heat treatment data includes first temperature information, second temperature information, treatment action, heat treatment effect score, temperature deviation cost, and time cost, wherein the first temperature information and the second temperature information correspond to adjacent moments; Step 2: updating the first processing module, which includes a temperature converter, an action-temperature converter, a temperature deviation converter, and a time cost converter, including: using the temperature converter to process the first temperature information and the second temperature information respectively to obtain a first feature map and a second feature map, and normalizing the first feature map and the second feature map respectively to obtain a third feature map and a fourth feature map; using the action-temperature converter to process the action and the third feature map to obtain a fifth feature map; using the temperature deviation converter to convert the third feature map and the fourth feature map respectively to obtain a first temperature deviation prediction value and a second temperature deviation prediction value; using the time cost converter to convert the third feature map and the fourth feature map respectively to obtain a first time cost prediction value and a second time cost prediction value; constructing a first objective function based on the second feature map, the fifth feature map, the temperature deviation cost, the first temperature deviation prediction value, the second temperature deviation prediction value, the time cost cost, the first time cost prediction value, and the second time cost prediction value, and minimizing the first objective function to update each converter; Step 3: updating the second control module, wherein the second control module includes an action generation network, an effect evaluation network, a temperature deviation evaluation network, and a time cost evaluation network; Step 4: Repeat steps 1 to 3 until the preset number of training times is exceeded to obtain the trained model; The temperature information at the current moment is obtained and input into the trained action generation network to obtain the current processing action to control the heat treatment process.
2. A heat treatment method for semiconductor device manufacturing according to claim 1, characterized in that: Updating the second control module includes: Inputting the first temperature information and the first feature map into the action generation network to obtain a first probability distribution; Inputting the first characteristic map and the second characteristic map into the effect evaluation network, the temperature deviation evaluation network, and the time cost evaluation network, respectively, to obtain corresponding heat treatment effect values, temperature deviation values, and time cost values, wherein the heat treatment effect value includes a first effect value corresponding to the first characteristic map, the temperature deviation value includes a first temperature deviation value corresponding to the first characteristic map, and the time cost value includes a first time cost value corresponding to the first characteristic map; Inputting the heat treatment effect score and the heat treatment effect value into the effect advantage evaluation function to obtain the effect advantage evaluation value; Inputting the temperature deviation cost and the temperature deviation value into the temperature deviation advantage evaluation function to obtain a temperature deviation evaluation value; Input the time cost price and the time cost value into the time cost advantage evaluation function to obtain the time cost evaluation value; The action generation network is optimized based on the effect advantage evaluation value, temperature deviation evaluation value, time cost evaluation value and the first probability distribution; the effect evaluation network is optimized based on the heat treatment effect score and the first effect value; the temperature deviation evaluation network is optimized based on the temperature deviation cost and the first temperature deviation value; and the time cost evaluation network is optimized based on the time cost cost and the first time cost value.
3. A heat treatment method for semiconductor device manufacturing according to claim 1, characterized in that: Temperature information includes wafer surface temperature distribution, heating rate, cooling rate, duration of the current processing stage, whether the current processing stage is an annealing stage, and whether the current processing stage reaches a preset temperature threshold; The heat treatment effect score is calculated by an effect score function, which is constructed based on wafer surface temperature uniformity, processing time, energy consumption and device performance. The temperature deviation cost is calculated by a temperature deviation cost function constructed based on temperature distribution and target temperature distribution. The time cost is calculated by a time cost function constructed based on processing time and a preset time threshold.
4. A heat treatment method for semiconductor device manufacturing according to claim 1, characterized in that: The expression of the first objective function is: ; in, represents the first objective function value, Represent the loss weight coefficients, represents the fifth eigenmap, represents the second feature map, represents the root mean square error, and Represent the predicted and actual temperature deviation values, and denote the predicted and actual time cost values, respectively. represents the KL divergence.
5. A heat treatment method for semiconductor device manufacturing according to claim 1, characterized in that: The action generation network is optimized based on the effect advantage evaluation value, the temperature deviation evaluation value, the time cost evaluation value and the first probability distribution, including: Inputting the effect advantage evaluation value, the temperature deviation evaluation value, the time cost evaluation value, and the first probability distribution into a second objective function, and obtaining an optimized action generation network by minimizing the second objective function; Optimizing the effect evaluation network based on the heat treatment effect score and the first effect value, including: inputting the heat treatment effect score and the first effect value into a loss function of the effect evaluation network and optimizing to obtain an optimized effect evaluation network; Optimizing the temperature deviation evaluation network based on the temperature deviation cost and the first temperature deviation value, including: inputting the temperature deviation cost and the first temperature deviation value into a loss function of the temperature deviation evaluation network and optimizing to obtain an optimized temperature deviation evaluation network; Optimizing a time cost evaluation network based on the time cost price and the first time cost value includes: inputting the time cost price and the first time cost value into a loss function of the time cost evaluation network and optimizing to obtain an optimized time cost evaluation network.
6. A heat treatment method for semiconductor device manufacturing according to claim 3, characterized in that: Also includes: Execute the current processing action to obtain the temperature information at the next moment; The current heat treatment effect score is calculated according to the effect score function; The current temperature deviation cost is calculated according to the temperature deviation cost function; The current time cost is calculated according to the time cost function; Construct current state data based on the temperature information at the current moment, the temperature information at the next moment, the current processing action, the current heat treatment effect score, the current temperature deviation cost, and the current time cost; The first processing module and the second control module are updated based on the current state data, and the updated first processing module and the second control module are used to process the temperature information at the next moment.
7. A heat treatment method for semiconductor device manufacturing according to claim 6, characterized in that: The expression of the effect scoring function is: ; in, Indicates at time Heat treatment effect score, Represent the weight coefficients, Indicates time Temperature uniformity, Indicates processing efficiency, Indicates energy consumption, Indicates device performance; The expression of the temperature deviation cost function is: ; in, represents the temperature deviation cost, Indicates the The temperature of the measuring point, Indicates the target temperature, Indicates the number of measurement points; The expression of the time cost function is: ; in, Indicates the time cost, Indicates the actual processing time, Indicates the preset time threshold.
8. A heat treatment method for semiconductor device manufacturing according to claim 1, characterized in that: The expression of the effect advantage evaluation function is: ; in, Indicates at time The effect advantage evaluation value, represents the discount factor, represents the number of stages of advantage assessment, Indicates the The timing difference error at the moment, Indicates the The status value at the moment.
9. A heat treatment method for semiconductor device manufacturing according to claim 4, characterized in that: The expression of the second objective function is: ; in, represents the second objective function value, Indicates that the status Next select action The probability of represents the effect advantage evaluation value, and are the Lagrange multipliers for temperature deviation and time cost, respectively, and Represent the temperature deviation cost and time cost respectively.
10. A heat treatment method for semiconductor device manufacturing according to claim 1, characterized in that: Also includes: The root mean square error is used to construct the loss functions of the effect evaluation network, temperature deviation evaluation network and time cost evaluation network. The loss functions of each network are minimized through the gradient descent optimization algorithm, and the parameters of each network are updated to obtain the optimized effect evaluation network, temperature deviation evaluation network and time cost evaluation network.