Semiconductor production equipment dynamic bottleneck prediction method and device and electronic equipment
Through machine learning methods combined with multi-source data sets and improved stacking integration model, the problem of low prediction accuracy of semiconductor production equipment in the prior art is solved, and high-precision identification of bottlenecks of semiconductor production equipment is achieved.
Patent Information
- Application Number
- CN202510530850.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-01
AI Technical Summary
The existing bottleneck prediction methods for semiconductor production equipment lack comprehensive consideration of multiple types of dynamic production factors, resulting in low prediction accuracy and difficulty in accurately identifying bottleneck equipment.
Using machine learning methods, multiple iterative training is carried out by obtaining multi-source operation data sets on the semiconductor production line, including in-product water level data, equipment energy efficiency data and equipment capacity data, and using improved stacked integration models to predict equipment bottlenecks, combining time series prediction, long and short-term memory networks and random forest models, multiple iterative training is carried out to improve prediction accuracy.
Effectively identifying bottleneck equipment on the production line has improved the accuracy of dynamic bottleneck prediction of semiconductor production equipment and can more accurately identify normal, potential and serious bottleneck equipment.
Smart Images

Figure CN120409811A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of semiconductor manufacturing technology. More specifically, it relates to a method, device, and electronic device for predicting dynamic bottlenecks of semiconductor production equipment. Background Art
[0002] During the semiconductor production process, the emergence of bottlenecks in semiconductor production equipment will seriously affect production efficiency and product quality. Timely and accurately predicting production equipment bottlenecks is crucial for reasonably arranging production plans, optimizing resource allocation, and improving overall production efficiency.
[0003] Currently, there are many deficiencies in the existing bottleneck prediction methods in semiconductor production. Traditional methods based on experience and simple statistics are highly subjective and lack comprehensive consideration of various dynamic production factors in the production process, making it difficult to accurately predict the production equipment with bottlenecks.
[0004] In addition, in the prior art, the patent document with the publication number CN113283223A discloses a data visualization manufacturing production monitoring report system, which is used to monitor records such as inputs, outputs, work-in-progress, throughput, and actual throughput of bottleneck machines on the production line. This report can be used to analyze the actual production status on the line, but it cannot predict the production equipment bottlenecks of the production line; the patent document with the publication number CN105988439A discloses a production capacity optimization method based on dynamically predicting the load of machines, which compares based on the relationship between rated production capacity, Master Production Schedule (MPS), and preset production capacity, and sets the machines with preset production capacity greater than the rated production capacity as bottleneck machines. This method only determines the equipment bottlenecks on the production line from a single factor of machine production capacity and also lacks comprehensive consideration of other dynamic production factors in the production process, resulting in low accuracy in predicting dynamic bottlenecks of semiconductor production equipment.
[0005] Therefore, how to better achieve the prediction of dynamic bottlenecks of semiconductor production equipment has become an urgent technical problem in the industry. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the purpose of this application is to better achieve the prediction of dynamic bottlenecks of semiconductor production equipment, aiming to solve the problems in the prior art that lack comprehensive consideration of various dynamic production factors in the production process, resulting in low accuracy in predicting dynamic bottlenecks of semiconductor production equipment and difficulty in accurately predicting the production equipment with bottlenecks.
[0007] To achieve the above objective, in the first aspect, this application provides a method for predicting dynamic bottlenecks of semiconductor production equipment, including: Obtaining multi-source operation data sets of each semiconductor production equipment on the semiconductor production line; Standardize various types of operation data in each of the multi-source operation data sets to obtain each standardized multi-source operation data set; Input each of the standardized multi-source operation data sets into the equipment bottleneck prediction model to obtain the bottleneck prediction results of each of the semiconductor production equipment output by the equipment bottleneck prediction model; The equipment bottleneck prediction model is trained according to multi-source operation data set samples and corresponding equipment bottleneck labels; the multi-source operation data set includes in-process water level data, equipment energy efficiency data, equipment normal operation duration data, and equipment production capacity data.
[0008] Optionally, the equipment bottleneck prediction model is an improved stacked ensemble model. The base models of the improved stacked ensemble model include a time series prediction model, a long short-term memory network model, and a random forest model, and the meta-model is a linear regression model; the step of inputting each of the standardized multi-source operation data sets into the equipment bottleneck prediction model to obtain the bottleneck prediction results of each of the semiconductor production equipment output by the equipment bottleneck prediction model includes: Input the in-process water level data, equipment energy efficiency data, and equipment normal operation duration data in each of the standardized multi-source operation data sets into the time series prediction model to obtain the first bottleneck prediction data of each of the semiconductor production equipment output by the time series prediction model; Input the in-process water level data and equipment energy efficiency data in each of the standardized multi-source operation data sets into the long short-term memory network model to obtain the second bottleneck prediction data of each of the semiconductor production equipment output by the long short-term memory network model; Input the equipment production capacity data in each of the standardized multi-source operation data sets into the random forest model to obtain the third bottleneck prediction data of each of the semiconductor production equipment output by the random forest model; Perform feature vector splicing according to the first bottleneck prediction data, second bottleneck prediction data, and third bottleneck prediction data of each semiconductor production equipment to obtain the meta-feature vectors corresponding to each semiconductor production equipment; Input the meta-feature vectors corresponding to each semiconductor production equipment into the linear regression model to obtain the bottleneck prediction results of each of the semiconductor production equipment output by the linear regression model.
[0009] Optionally, the equipment energy efficiency data includes equipment operation time, equipment overall utilization rate, and unit production capacity data; the equipment production capacity data includes a production capacity compliance rate, which is determined based on the daily actual production capacity and the daily expected production capacity.
[0010] Optionally, the device bottleneck label is determined based on the in-process water level data, the overall equipment effectiveness, the unit production capacity data, and the production capacity compliance rate in the multi-source operation dataset sample.
[0011] Optionally, before inputting each of the standardized multi-source operation datasets into the device bottleneck prediction model to obtain the bottleneck prediction results of each semiconductor production device output by the device bottleneck prediction model, the method further includes: Obtain multiple multi-source operation dataset samples; Perform standardization processing on various types of operation data in each of the multi-source operation dataset samples to obtain each standardized multi-source operation dataset sample; Take each standardized multi-source operation dataset sample and the corresponding device bottleneck label as a set of training samples to obtain multiple sets of the training samples; Use multiple sets of the training samples to train the device bottleneck prediction model.
[0012] Optionally, the device bottleneck prediction model is an improved stacked ensemble model. The base models of the improved stacked ensemble model include a time series prediction model, a long short-term memory network model, and a random forest model, and the meta-model is a linear regression model. The step of using multiple sets of the training samples to train the device bottleneck prediction model includes: Divide multiple sets of the training samples into a training set and a test set according to a preset ratio; Use the training set to perform multi-round iterative training on the time series prediction model, the long short-term memory network model, and the random forest model respectively until a trained time series prediction model, a trained long short-term memory network model, and a trained random forest model are obtained; Use each training sample in the test set, the trained time series prediction model, the trained long short-term memory network model, and the trained random forest model to determine the meta-feature vector corresponding to each training sample in the test set; Use the device bottleneck label and the corresponding meta-feature vector of each training sample in the test set to perform multi-round iterative training on the linear regression model until a trained linear regression model is obtained; According to the trained time series prediction model, the trained long short-term memory network model, the trained random forest model, and the trained linear regression model, obtain a trained stacked ensemble model to complete the training of the device bottleneck prediction model.
[0013] In a second aspect, the present application provides a semiconductor production equipment dynamic bottleneck prediction device, including: A data acquisition module, configured to acquire multi-source operation data sets of various semiconductor production equipment on a semiconductor production line; A normalization module, configured to perform normalization processing on various types of operation data in each of the multi-source operation data sets to obtain each normalized multi-source operation data set; A bottleneck prediction module, configured to input each of the normalized multi-source operation data sets into an equipment bottleneck prediction model to obtain bottleneck prediction results of each of the semiconductor production equipment output by the equipment bottleneck prediction model; The equipment bottleneck prediction model is trained according to multi-source operation data set samples and corresponding equipment bottleneck labels; the multi-source operation data set includes in-process water level data, equipment energy efficiency data, equipment normal operation duration data, and equipment production capacity data.
[0014] In a third aspect, the present application provides an electronic device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is configured to execute the method described in the first aspect or any one of the possible implementation manners of the first aspect.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program. When the computer program runs on a processor, the processor is caused to execute the method described in the first aspect or any one of the possible implementation manners of the first aspect.
[0016] In a fifth aspect, the present application provides a computer program product. When the computer program product runs on a processor, the processor is caused to execute the method described in the first aspect or any one of the possible implementation manners of the first aspect.
[0017] It can be understood that the beneficial effects of the above second aspect to fifth aspect can refer to the relevant descriptions in the first aspect above, and will not be elaborated here.
[0018] Generally speaking, compared with the prior art through the above technical solutions conceived by the present application, the following beneficial effects are achieved: A method, device and electronic device for dynamically predicting bottlenecks of semiconductor production equipment provided by the present application, by comprehensively considering various dynamic production factors on a semiconductor production line, adopting a machine learning method, using multi-source operation data set samples including in-process water level data, equipment energy efficiency data, equipment normal operation duration data and equipment production capacity data of semiconductor production equipment and equipment bottleneck labels to train an equipment bottleneck prediction model, learning the internal relationship between the production equipment bottleneck and the multi-source dynamic production data of the production equipment, can effectively identify bottleneck equipment on the production line and improve the accuracy of dynamically predicting bottlenecks of semiconductor production equipment. Description of the Drawings
[0019] Figure 1 It is one of the schematic flowcharts of the dynamic bottleneck prediction method for semiconductor manufacturing equipment provided by an embodiment of the present application; Figure 2 It is the second of the schematic flowcharts of the dynamic bottleneck prediction method for semiconductor manufacturing equipment provided by an embodiment of the present application; Figure 3 It is the schematic structural diagram of the dynamic bottleneck prediction device for semiconductor manufacturing equipment provided by an embodiment of the present application; Figure 4 It is the schematic structural diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners
[0020] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0021] The terms "first" and "second" in the description and claims of the present application are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first bottleneck prediction data and the second bottleneck prediction data are used to distinguish different bottleneck prediction data, rather than to describe the specific order of the bottleneck prediction data.
[0022] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Exactly speaking, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0023] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units, etc.; a plurality of elements refers to two or more elements, etc.
[0024] With the continuous development of semiconductor manufacturing processes and the continuous expansion of production scales, the amount of data in the production process has increased sharply, and the complexity and dynamics of the data have also increased day by day. The existing bottleneck prediction methods can no longer meet the needs of actual production, and there is an urgent need for a method that can adaptively process multi-stage dynamic data and accurately predict equipment bottlenecks. For this reason, the present application proposes a dynamic bottleneck prediction method for semiconductor manufacturing equipment.
[0025] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.
[0026] Figure 1 is one of the schematic flowcharts of the method for predicting dynamic bottlenecks of semiconductor manufacturing equipment provided by the embodiments of the present application. As Figure 1 shown, the method includes: Step S1, obtaining multi-source operation data sets of various semiconductor manufacturing equipment on a semiconductor production line; Step S2, performing standardization processing on various types of operation data in each multi-source operation data set to obtain each standardized multi-source operation data set; Step S3, inputting each standardized multi-source operation data set into an equipment bottleneck prediction model to obtain bottleneck prediction results of each semiconductor manufacturing equipment output by the equipment bottleneck prediction model; The equipment bottleneck prediction model is trained according to multi-source operation data set samples and corresponding equipment bottleneck labels; the multi-source operation data set includes in-process water level data, equipment energy efficiency data, equipment normal operation duration data, and equipment production capacity data.
[0027] Specifically, the multi-source operation data set described in the embodiments of the present application refers to operation data during the production process of semiconductor manufacturing equipment collected from a multi-dimensional production perspective. Specifically, it may include in-process (Work In Process, WIP) water level data, equipment energy efficiency data, equipment normal operation duration (Uptime) data, and equipment production capacity data.
[0028] The equipment bottleneck prediction model described in the embodiments of the present application is obtained by training a model using machine learning methods according to multi-source operation data set samples and corresponding equipment bottleneck labels, and is used to learn the internal relationship between the production equipment bottleneck and the multi-source dynamic production data of the production equipment. By identifying the multi-source operation data set of the semiconductor manufacturing equipment, bottleneck prediction results of each semiconductor manufacturing equipment are output.
[0029] Among them, the equipment bottleneck prediction model can be specifically constructed based on a neural network model or a stacking integration model. Among them, the neural network model can specifically adopt a convolutional neural network CNN model, etc.; the stacking integration model can specifically adopt a model with a tree model, a support vector machine (SVM), and a K-nearest neighbor (KNN) algorithm model as the base model and a logistic regression model as the meta-model, or can also adopt an improved base model and meta-model structure for the characteristics of semiconductor production, or can also be other deep neural networks that can realize equipment bottleneck prediction, which are not specifically limited in the present application.
[0030] It can be understood that the multi-source operation data set samples can specifically include WIP water level data samples, equipment energy efficiency data samples, equipment normal operation duration data samples, and equipment production capacity data samples.
[0031] In the embodiments of the present application, the model training samples are composed of multi-source operation dataset samples carrying equipment bottleneck labels. Among them, the equipment bottleneck labels are determined in advance according to the multi-source operation dataset samples and are in one-to-one correspondence with the multi-source operation dataset samples. That is to say, for each multi-source operation dataset sample in the training samples, a corresponding equipment bottleneck label is preset to be carried.
[0032] More specifically, in the embodiments of the present application, in step S1, the multi-dimensional data streams of each semiconductor production equipment on the semiconductor production line can be collected through the Manufacturing Execution System (MES) data interface, and the multi-source operation datasets of each semiconductor production equipment can be obtained therefrom, including the in-process work-in-progress (WIP) water level data, equipment energy efficiency data, equipment normal operation duration data, and equipment production capacity data.
[0033] Among them, the WIP water level data is used to show the number of wafers being processed (work-in-progress) on each production equipment machine platform, which can help track the load situation and production capacity utilization rate of each machine platform; the equipment energy efficiency data is used to reflect the situation of equipment operation and operation efficiency, such as equipment operation time (operation time), overall utilization rate (Util), etc.; the equipment normal operation duration data is used to show the normal operation time of each production equipment machine platform, that is, the time when the equipment is in a producible state, which can reflect the reliability and stability of the equipment.
[0034] Based on the content of the above embodiments, as an optional embodiment, the equipment energy efficiency data includes equipment operation time, equipment overall utilization rate, and unit production capacity data; the equipment production capacity data includes production capacity compliance rate, and the production capacity compliance rate is determined based on the daily actual production capacity and the daily expected production capacity.
[0035] Specifically, in the embodiments of the present application, the equipment operation time is used to show the operation time of each production equipment machine platform invested in production processing operations, which reflects the actual production status and efficiency level of the equipment.
[0036] In the embodiments of the present application, the equipment overall utilization rate is used to show the utilization rate of each production equipment machine platform, that is, the production utilization rate of the equipment within a given time. It represents the ratio between the actual used time of the machine platform and the total available time.
[0037] In the embodiments of the present application, the equipment production capacity data includes production capacity compliance rate, and the production capacity compliance rate is determined based on the daily actual production capacity (OUTPUT) and the daily expected production capacity (EXPECTED_OUTPUT), and can be specifically obtained by calculating the ratio of the daily actual production capacity to the daily expected production capacity.
[0038] The method according to the embodiment of the present application considers multi-dimensional production factors such as equipment operation time, comprehensive utilization rate, and production capacity, mines multi-source operation data affecting the bottleneck of semiconductor production equipment for data analysis, which is beneficial to improving the accuracy and effect of semiconductor production equipment bottleneck prediction.
[0039] In the embodiment of the present application, in step S2, the Z-Score model can be specifically used to standardize various operation data in each multi-source operation dataset. Specifically, for each operation data in the multi-source operation dataset of semiconductor production equipment , after being standardized according to the following calculation formula, the operation data is obtained, so as to obtain each standardized multi-source operation dataset. That is: ; where represents the mean value of the operation data within the historical preset period, and represents the standard deviation corresponding to the operation data .
[0040] In the embodiment of the present application, in step S3, a device bottleneck prediction model is obtained by pre-training with multi-source operation dataset samples and corresponding device bottleneck labels. Furthermore, each of the aforementioned standardized multi-source operation datasets can be input into the pre-trained device bottleneck prediction model. Through the identification and prediction calculation of the device bottleneck prediction model, the bottleneck prediction results of each semiconductor production device can be obtained.
[0041] The semiconductor production equipment dynamic bottleneck prediction method according to the embodiment of the present application comprehensively considers various dynamic production factors on the semiconductor production line, adopts a machine learning method, uses multi-source operation dataset samples including in-process water level data, equipment energy efficiency data, equipment normal operation duration data, and equipment production capacity data of semiconductor production equipment and device bottleneck labels to train a device bottleneck prediction model, and learns the internal relationship between the production equipment bottleneck and the multi-source dynamic production data of the production equipment, which can effectively identify the bottleneck equipment on the production line and improve the accuracy of semiconductor production equipment dynamic bottleneck prediction.
[0042] In the embodiment of the present application, the device bottleneck labels can include three types of labels: normal equipment, potential bottleneck equipment, and severe bottleneck equipment. That is to say, the device bottleneck prediction model can predict the bottleneck prediction results of each semiconductor production device on the production line as normal equipment, or potential bottleneck equipment, or severe bottleneck equipment.
[0043] It should be noted that in the embodiments of the present application, the equipment bottleneck label can be used to label the multi-source operation data set samples of the equipment by manual marking according to the actual working conditions of the semiconductor equipment; it can also be determined by the computer through calculation according to the bottleneck index calculation model.
[0044] Based on the content of the above embodiments, as an optional embodiment, the equipment bottleneck label is determined based on the in-process water level data, equipment overall utilization rate, unit production capacity data, and production capacity compliance rate in the multi-source operation data set samples.
[0045] Specifically, in the embodiments of the present application, the WIP water level data (unit: piece), equipment overall utilization rate (Util) (value range is [0,1]), unit production capacity data ( )(theoretical hourly production capacity, unit: piece / hour) and production capacity compliance rate in the multi-source operation data set samples are used to construct the bottleneck index LPI. Based on the comparison between output and demand, bottleneck classification judgment is carried out, and the bottleneck index LPI is calculated. The specific calculation process can be expressed as follows: ; where UTIL represents the normalized value of the equipment overall utilization rate which can be specifically scaled to the [0,1] interval through min-max standardization; represents the load factor of in-process products (dimensionless); represents the WIP water level data; represents the production capacity compliance rate, and the range is in the [0,1] interval.
[0046] When 1 (default 1 = 0.85), it can be determined as a severely bottlenecked device; When 1 (default = 0.65), it can be determined as a potentially bottlenecked device; When it can be determined as a normal device.
[0047] In the embodiments of the present application, by considering various types of operation data affecting equipment bottlenecks for comprehensive analysis and calculation of the equipment bottleneck index, the accuracy and reliability of equipment bottleneck judgment can be improved.
[0048] Figure 2 is the second schematic flow chart of the semiconductor production equipment dynamic bottleneck prediction method provided by the embodiments of the present application, as shown in Figure 2As shown in the figure, in the embodiment of the present application, the device bottleneck prediction model is an improved Stacking ensemble model. The base models of the Stacking ensemble model include the time series prediction model Prophet model, the LSTM network model, and the random forest model, and the meta-model is a linear regression model. Each standardized multi-source operation dataset is input into the device bottleneck prediction model to obtain the bottleneck prediction results of each semiconductor production device output by the device bottleneck prediction model, including: The WIP water level data, device energy efficiency data, and device normal operation duration data in each standardized multi-source operation dataset are input into the Prophet model to obtain the first bottleneck prediction data of each semiconductor production device output by the Prophet model. The WIP water level data and device energy efficiency data in each standardized multi-source operation dataset are input into the LSTM network model to obtain the second bottleneck prediction data of each semiconductor production device output by the LSTM network model. The device production capacity data in each standardized multi-source operation dataset is input into the random forest model to obtain the third bottleneck prediction data of each semiconductor production device output by the random forest model. According to the first bottleneck prediction data, the second bottleneck prediction data, and the third bottleneck prediction data of each semiconductor production device, feature vector splicing is performed to obtain the meta-feature vector corresponding to each semiconductor production device. The meta-feature vectors corresponding to each semiconductor production device are input into the linear regression model to obtain the bottleneck prediction results of each semiconductor production device output by the linear regression model.
[0049] Specifically, the first bottleneck prediction data described in the embodiment of the present application refers to the results obtained by the Prophet model in the base model for device bottleneck prediction for the input WIP water level data, device energy efficiency data, and device normal operation duration data.
[0050] The second bottleneck prediction data described in the embodiment of the present application refers to the results obtained by the LSTM network model in the base model for device bottleneck prediction for the input WIP water level data and device energy efficiency data.
[0051] The third bottleneck prediction data described in the embodiment of the present application refers to the results obtained by the random forest model in the base model for device bottleneck prediction for the input device production capacity data.
[0052] In an embodiment of the present application, the base model includes a Prophet model, which is used to capture the trends, seasonality, and impact of external variables in semiconductor equipment production data. Semiconductor production data often exhibits periodic changes over time, and the Prophet model is well suited to handle these situations. The input data for the Prophet model includes standardized WIP level data, equipment energy efficiency data, and equipment uptime data. Equipment energy efficiency data includes equipment operating time, equipment comprehensive utilization rate, and unit production capacity data.
[0053] In this way, by inputting the WIP water level data, equipment operation time, equipment comprehensive utilization rate, unit production capacity data and equipment normal operation time data in each standardized multi-source operation data set into the Prophet model, the first bottleneck prediction data of each semiconductor production equipment output by the Prophet model can be obtained.
[0054] In the embodiments of this application, the base model also includes an LSTM network model. Specifically, the LSTM network model uses a 7-day time window to learn the long-term dependencies of the time series of relevant equipment operating data. The input data includes standardized work-in-process (WIP) level data, equipment operating time, equipment utilization rate, and unit production capacity data. The semiconductor production process has distinct time series characteristics, and the LSTM network model can effectively capture these characteristics, improving prediction accuracy.
[0055] In this way, the WIP water level data, equipment operation time, equipment comprehensive utilization rate and unit production capacity data in each standardized multi-source operation data set are input into the LSTM network model, and the second bottleneck prediction data of each semiconductor production equipment output by the LSTM network model can be obtained.
[0056] In an embodiment of the present application, the base model also includes a random forest model. The random forest model is used to extract nonlinear patterns in the high-dimensional features of the input data. These features can reflect the quality and target achievement of the semiconductor production process. The random forest model can mine the nonlinear relationships therein. The input data of the random forest model includes standardized equipment capacity data, namely, the capacity achievement rate.
[0057] In this way, by inputting the capacity achievement rates in each standardized multi-source operation data set into the random forest model, the third bottleneck prediction data of each semiconductor production equipment output by the random forest model can be obtained.
[0058] Furthermore, the feature vectors corresponding to the first bottleneck prediction data, the second bottleneck prediction data, and the third bottleneck prediction data of each semiconductor production equipment can be used to perform feature vector splicing to obtain the meta-feature vector corresponding to each semiconductor production equipment.
[0059] Further, in the embodiments of the present application, the meta-feature vectors corresponding to each semiconductor production device are input into a linear regression model. Through the regression analysis calculation of the linear regression model, the bottleneck prediction results of each semiconductor production device are obtained.
[0060] Among them, the linear regression model, as the meta-model in the Stacking ensemble model, can specifically adopt the Lasso regression model to replace the traditional logistic regression model, eliminate redundant features through L1 regularization (such as excluding unit production capacity data highly correlated with the daily actual production capacity), and use 10-fold cross-validation to select the optimal hyperparameter value, which can effectively avoid model overfitting and improve the generalization ability of the model.
[0061] The method of the embodiments of the present application, for the specific scenario of a semiconductor production line, by considering the trend and seasonality of semiconductor equipment production data, the long-term dependence relationship of equipment operation data time series, and the quality and target achievement in the semiconductor production process, etc., uses the Prophet model, the LSTM network model, and the random forest model to improve the base models of the Stacking ensemble model, and uses the Lasso regression model to replace the traditional meta-model, which can further improve the accuracy of dynamic bottleneck prediction of semiconductor production equipment.
[0062] Based on the content of the above embodiments, as an optional embodiment, before inputting each standardized multi-source operation dataset into the equipment bottleneck prediction model to obtain the bottleneck prediction results of each semiconductor production device output by the equipment bottleneck prediction model, the method further includes: Obtain multiple multi-source operation dataset samples; Perform standardization processing on various types of operation data in each multi-source operation dataset sample to obtain each standardized multi-source operation dataset sample; Take each standardized multi-source operation dataset sample and the corresponding equipment bottleneck label as a group of training samples to obtain multiple groups of training samples; Use multiple groups of training samples to train the equipment bottleneck prediction model.
[0063] Specifically, in the embodiments of the present application, before inputting each standardized multi-source operation dataset into the equipment bottleneck prediction model, it is also necessary to train the equipment bottleneck prediction model to obtain a trained equipment bottleneck prediction model.
[0064] In the embodiments of the present application, historical production data of each semiconductor device on the production line is obtained through a pre-arranged MES system, and then multiple multi-source operation dataset samples can be obtained. After that, the above-mentioned LPI bottleneck index calculation and classification judgment can be performed on the multi-source operation dataset samples, so as to obtain each multi-source operation dataset sample and its corresponding device bottleneck label.
[0065] Further, in the embodiments of the present application, a Z-Score model can be used to standardize various operation data in each multi-source operation dataset sample, and thus each standardized multi-source operation dataset sample can be obtained. Furthermore, each standardized multi-source operation dataset sample and the corresponding calculated device bottleneck label can be used as a set of training samples, and finally multiple sets of training samples can be obtained.
[0066] In the embodiments of the present invention, each standardized multi-source operation dataset sample and the device bottleneck label it carries are in one-to-one correspondence.
[0067] Exemplarily, the device bottleneck prediction model is constructed using a neural network model, and the above-mentioned multiple sets of training samples are used to train the device bottleneck prediction model. The specific training process is as follows: After obtaining multiple sets of training samples, the multiple sets of training samples are sequentially input into the device bottleneck prediction model, and the device bottleneck prediction model is trained using the multiple sets of training samples, that is: The standardized multi-source operation dataset sample in each set of training samples and the device bottleneck label it carries are simultaneously input into the device bottleneck prediction model. According to each output result in the device bottleneck prediction model, by calculating the loss function value, the model parameters in the device bottleneck prediction model are adjusted. When the preset training termination condition is met, the entire training process of the device bottleneck prediction model is finally completed, and the trained device bottleneck prediction model is obtained.
[0068] The method of the embodiments of the present invention is beneficial to improving the model accuracy of the trained device bottleneck prediction model by using the standardized multi-source operation dataset sample and the corresponding device bottleneck label as a set of training samples and training the device bottleneck prediction model using multiple sets of training samples.
[0069] Based on the content of the above embodiments, as an alternative embodiment, the device bottleneck prediction model is an improved Stacking ensemble model. The base models of the improved Stacking ensemble model include the Prophet model, the LSTM network model, and the random forest model, and the meta-model is a linear regression model; training the device bottleneck prediction model using multiple sets of training samples includes: Dividing the multiple sets of training samples into a training set and a test set according to a preset ratio; Use the training set to perform multiple rounds of iterative training on the Prophet model, LSTM network model, and random forest model respectively until the trained Prophet model, trained LSTM network model, and trained random forest model are obtained; Use each training sample in the test set, the trained Prophet model, the trained LSTM network model, and the trained random forest model to determine the meta-feature vector corresponding to each training sample in the test set; Use the device bottleneck labels and corresponding meta-feature vectors in each training sample of the test set to perform multiple rounds of iterative training on the linear regression model until the trained linear regression model is obtained; According to the trained Prophet model, trained LSTM network model, trained random forest model, and trained linear regression model, obtain the trained Stacking ensemble model to complete the training of the device bottleneck prediction model.
[0070] Specifically, in the embodiments of the present application, the device bottleneck prediction model is an improved Stacking ensemble model. The base models of the improved Stacking ensemble model include the Prophet model, LSTM network model, and random forest model, and the meta-model is a linear regression model. Among them, the linear regression model can specifically adopt the Lasso regression model. <00>
[0071] Further, in the embodiments of the present application, multiple groups of training samples are divided into a training set and a test set according to a preset ratio, such as 7:3.
[0072] More specifically, in a specific embodiment of the present application, for the training of the base models, the relevant parameters of the three base models are specifically set as follows: For the Prophet model, configure the periodic term parameters as follows: Annual periodic term ; Weekly periodic term 。
[0073] For the LSTM network model, configure the parameters as follows: The dimension of the input data of the input layer : ∈ , where Specifically, it can include four types of data: WIP water level data, equipment comprehensive utilization rate, unit production capacity data, and equipment operation time within the historical 7-day window; The network structure adopts a double-layer LSTM network structure, specifically including 128 neurons, and the hyperparameter Dropout takes a value of 0.3; The output layer is used to output the bottleneck probability prediction value ∈[0, 1].
[0074] For the random forest model, the configuration parameters are as follows: The number of decision trees is 100, the maximum depth is 8, and the minimum number of samples in a leaf is 5.
[0075] Furthermore, in the embodiments of the present application, each training sample in the training set is used to perform multiple rounds of iterative training on the Prophet model, the LSTM network model, and the random forest model until the trained Prophet model, the trained LSTM network model, and the trained random forest model are obtained.
[0076] For example, by inputting the WIP water level data samples, equipment operation time samples, equipment comprehensive utilization rate samples, unit production capacity data samples, and equipment normal operation duration data samples in each training sample in the training set into the Prophet model, the bottleneck prediction data output by the Prophet model can be obtained. Furthermore, using a preset loss function, the loss value is calculated based on the equipment bottleneck label corresponding to the training sample and the bottleneck prediction data. Further, after the loss value is calculated, the current training process ends. Then, based on the loss value, the model parameters of the Prophet model are adjusted, and then the next training is performed, and the model training is repeated iteratively in this way.
[0077] During the training process, if the training result for a certain group of training samples meets the preset training termination condition, such as the calculated loss value is less than the preset threshold, or when the current number of iterations reaches the preset number, and the loss value of the model can be controlled within the convergence range, the model training ends. At this time, the obtained model parameters can be used as the model parameters of the trained Prophet model, and the Prophet model training is completed, and thus the trained Prophet model is obtained.
[0078] Similarly, in the same way as described above, based on the foregoing embodiments, the input data for the LSTM network model and the random forest model is determined, and each training sample in the training set is used to perform multiple rounds of iterative training on the LSTM network model and the random forest model, and the trained LSTM network model and the trained random forest model can be obtained.
[0079] Furthermore, the same training sample in the test set is sequentially input into the trained Prophet model, the trained LSTM network model, and the trained random forest model to obtain the bottleneck prediction data output by the trained Prophet model and the bottleneck prediction data output by the trained LSTM network model , and the bottleneck prediction data output by the trained random forest model , and then The eigenvectors corresponding to the three are concatenated, and thus the meta-eigenvector corresponding to the training sample in the test set can be determined.
[0080] Furthermore, a Lasso regression model can be used for meta-learning. By performing linear regression analysis on the meta-eigenvector corresponding to the training sample, the bottleneck prediction data output by the Lasso regression model can be obtained. Then, using a preset loss function, the loss value is calculated based on the device bottleneck label corresponding to the training sample and the bottleneck prediction data, and this training process ends. Then, based on this loss value, the model parameters of the Lasso regression model are adjusted. After that, the next training sample is used for the next training, and so on, iteratively training the model repeatedly until the preset training termination condition is met and the loss value of the model can be controlled within the convergence range, then the model training ends. At this time, the obtained model parameters can be used as the model parameters of the trained Lasso regression model, and thus the trained Lasso regression model is obtained.
[0081] It can be understood that according to the trained base models, including the trained Prophet model, the trained LSTM network model, the trained random forest model, and the trained meta-model, that is, the trained Lasso regression model, a trained Stacking ensemble model can be combined to complete the training of the device bottleneck prediction model.
[0082] The method of the embodiments of the present application, by considering the trend and seasonality of semiconductor device production data, the long-term dependence relationship of device operation data time series, and the quality and target achievement in the semiconductor production process, etc., uses the Prophet model, the LSTM network model, and the random forest model to improve the base models of the Stacking ensemble model, uses the Lasso regression model to replace the traditional meta-model, and uses multiple groups of training samples to iteratively train the improved Stacking ensemble model repeatedly, controlling the loss value of the Stacking ensemble model within the convergence range, thereby obtaining a trained device bottleneck prediction model, which is beneficial to improving the accuracy of the prediction data output by the device bottleneck prediction model and enhancing the accuracy of dynamic bottleneck prediction of semiconductor production equipment.
[0083] Next, the semiconductor production equipment dynamic bottleneck prediction device provided by the present application will be described. The semiconductor production equipment dynamic bottleneck prediction device described below can be correspondingly referred to the semiconductor production equipment dynamic bottleneck prediction method described above.
[0084] Figure 3 is a schematic structural diagram of the semiconductor production equipment dynamic bottleneck prediction device provided by the embodiments of the present application, as Figure 3 shown, including: A data acquisition module 10, configured to acquire a multi-source operation dataset of each semiconductor production device on a semiconductor production line; A normalization module 20, configured to perform normalization processing on various operation data in each multi-source operation dataset to obtain each normalized multi-source operation dataset; A bottleneck prediction module 30, configured to input each normalized multi-source operation dataset into a device bottleneck prediction model to obtain a bottleneck prediction result of each semiconductor production device output by the device bottleneck prediction model; The device bottleneck prediction model is trained according to a multi-source operation dataset sample and a corresponding device bottleneck label; the multi-source operation dataset includes in-process water level data, device energy efficiency data, device normal operation duration data, and device production capacity data.
[0085] It can be understood that for the detailed function implementation of each of the above units / modules, reference can be made to the introduction in the foregoing method embodiments, and details are not described herein again.
[0086] It should be understood that the above device is used to execute the method in the above embodiment. For the corresponding program modules in the device, their implementation principles and technical effects are similar to those described in the above method. The working process of this device can refer to the corresponding process in the above method, and details are not described herein again.
[0087] The semiconductor production device dynamic bottleneck prediction device according to the embodiment of the present application, by comprehensively considering various dynamic production factors on the semiconductor production line, adopting a machine learning method, using a multi-source operation dataset sample including in-process water level data, device energy efficiency data, device normal operation duration data, and device production capacity data of semiconductor production devices and device bottleneck labels to train a device bottleneck prediction model, and learning the internal relationship between the production device bottleneck and the multi-source dynamic production data of the production device, can effectively identify bottleneck devices on the production line and improve the accuracy of semiconductor production device dynamic bottleneck prediction.
[0088] Based on the method in the above embodiment, an embodiment of the present application provides an electronic device, as Figure 4 shown. The electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute the method in the above embodiment.
[0089] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.
[0090] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a processor, the processor is caused to execute the method in the above embodiment.
[0091] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor is caused to execute the method in the above embodiment.
[0092] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0093] The method steps in the embodiments of the present application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0094] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server, data center, etc. that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0095] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.
[0096] It should be understood that expressions such as "including" and "may include" that can be used in this application indicate the existence of the disclosed functions, operations, or components, and do not limit one or more additional functions, operations, and components. In this application, terms such as "including" and / or "having" may be interpreted as indicating a specific characteristic, number, operation, component, assembly, or a combination thereof, but do not exclude the existence or possibility of addition of one or more other characteristics, numbers, operations, components, assemblies, or a combination thereof.
[0097] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for predicting dynamic bottlenecks in semiconductor manufacturing equipment, characterized in that, Including: Obtain multi-source operation data sets of each semiconductor production equipment on the semiconductor production line; Perform standardization processing on various types of operation data in each of the multi-source operation data sets to obtain each standardized multi-source operation data set; Input each of the standardized multi-source operation data sets into the equipment bottleneck prediction model to obtain the bottleneck prediction results of each of the semiconductor production equipment output by the equipment bottleneck prediction model; The equipment bottleneck prediction model is trained according to multi-source operation data set samples and corresponding equipment bottleneck labels; the multi-source operation data set includes in-process water level data, equipment energy efficiency data, equipment normal operation duration data, and equipment production capacity data.
2. The method for predicting dynamic bottlenecks of semiconductor manufacturing equipment according to claim 1, wherein The equipment bottleneck prediction model is an improved stacked ensemble model. The base models of the improved stacked ensemble model include a time series prediction model, a long short-term memory network model, and a random forest model, and the meta-model is a linear regression model; the step of inputting each of the standardized multi-source operation data sets into the equipment bottleneck prediction model to obtain the bottleneck prediction results of each of the semiconductor production equipment output by the equipment bottleneck prediction model includes: Input the in-process water level data, equipment energy efficiency data, and equipment normal operation duration data in each of the standardized multi-source operation data sets into the time series prediction model to obtain the first bottleneck prediction data of each of the semiconductor production equipment output by the time series prediction model; Input the in-process water level data and equipment energy efficiency data in each of the standardized multi-source operation data sets into the long short-term memory network model to obtain the second bottleneck prediction data of each of the semiconductor production equipment output by the long short-term memory network model; Input the equipment production capacity data in each of the standardized multi-source operation data sets into the random forest model to obtain the third bottleneck prediction data of each of the semiconductor production equipment output by the random forest model; Perform feature vector splicing according to the first bottleneck prediction data, second bottleneck prediction data, and third bottleneck prediction data of each semiconductor production equipment to obtain the meta-feature vectors corresponding to each semiconductor production equipment; Input the meta-feature vectors corresponding to each semiconductor production equipment into the linear regression model to obtain the bottleneck prediction results of each semiconductor production equipment output by the linear regression model.
3. The method for predicting dynamic bottlenecks of semiconductor manufacturing equipment according to claim 1, wherein The equipment energy efficiency data includes equipment operation time, equipment overall utilization rate, and unit production capacity data; the equipment production capacity data includes production capacity compliance rate, and the production capacity compliance rate is determined based on the daily actual production capacity and the daily expected production capacity.
4. The dynamic bottleneck prediction method for semiconductor production equipment according to claim 3, characterized in that The equipment bottleneck label is determined based on the in-process water level data, equipment overall utilization rate, unit production capacity data, and production capacity compliance rate in the multi-source operation data set sample.
5. The method for predicting dynamic bottlenecks of semiconductor production equipment according to any one of claims 1-4, characterized in that, Before the step of inputting each of the standardized multi-source operation data sets into the equipment bottleneck prediction model to obtain the bottleneck prediction results of each of the semiconductor production equipment output by the equipment bottleneck prediction model, the method further includes: Obtain multiple multi-source operation data set samples; Standardize various types of operation data in each of the multi-source operation dataset samples to obtain standardized multi-source operation dataset samples; Take each of the standardized multi-source operation dataset samples and the corresponding equipment bottleneck labels as a set of training samples to obtain multiple sets of the training samples; Use multiple sets of the training samples to train the equipment bottleneck prediction model.
6. The dynamic bottleneck prediction method for semiconductor manufacturing equipment according to claim 5, wherein The equipment bottleneck prediction model is an improved stacked ensemble model. The base models of the improved stacked ensemble model include a time series prediction model, a long short-term memory network model, and a random forest model, and the meta-model is a linear regression model; The training of the equipment bottleneck prediction model using multiple sets of the training samples includes: Divide multiple sets of the training samples into a training set and a test set according to a preset ratio; Use the training set to perform multiple rounds of iterative training on the time series prediction model, the long short-term memory network model, and the random forest model respectively until a trained time series prediction model, a trained long short-term memory network model, and a trained random forest model are obtained; Use each training sample in the test set, the trained time series prediction model, the trained long short-term memory network model, and the trained random forest model to determine the meta-feature vector corresponding to each training sample in the test set; Use the equipment bottleneck labels and the corresponding meta-feature vectors of each training sample in the test set to perform multiple rounds of iterative training on the linear regression model until a trained linear regression model is obtained; According to the trained time series prediction model, the trained long short-term memory network model, the trained random forest model, and the trained linear regression model, obtain a trained stacked ensemble model to complete the training of the equipment bottleneck prediction model.
7. A dynamic bottleneck prediction device for semiconductor production equipment, characterized in that Include: A data acquisition module for acquiring multi-source operation datasets of various semiconductor production equipment on a semiconductor production line; A standardization module for standardizing various types of operation data in each of the multi-source operation datasets to obtain standardized multi-source operation datasets; A bottleneck prediction module for inputting each of the standardized multi-source operation datasets into the equipment bottleneck prediction model to obtain the bottleneck prediction results of each of the semiconductor production equipment output by the equipment bottleneck prediction model; The equipment bottleneck prediction model is trained based on multi-source operation dataset samples and the corresponding equipment bottleneck labels; the multi-source operation datasets include in-process water level data, equipment energy efficiency data, equipment normal operation duration data, and equipment production capacity data.
8. An electronic device, characterized in that, Include: At least one memory for storing a computer program; At least one processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1-6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program runs on the processor, the processor is caused to execute the method according to any one of claims 1-6.
10. A computer program product, characterized in that, When the computer program product runs on the processor, the processor is caused to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Capacity optimization method based on dynamic stock load prediction
CN105988439A
Data visualization manufacturing production monitoring report system
CN113283223A
Cited By
Real-time judgment method, device and equipment for bottleneck equipment and storage medium
CN121146726A