Method, apparatus, and semiconductor processing equipment for determining process parameters
By using a target network model based on capsule network, the process parameters are quickly determined, and the problem of long-term process parameter adjustment in the prior art is solved, and efficient and accurate process parameter determination is achieved.
Patent Information
- Application Number
- CN202411219024.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-08-30
AI Technical Summary
The prior art takes a long time to determine process parameters, and it is necessary to continuously adjust the process parameters to meet process needs.
The target network model built based on the capsule network is used to predict the target process information input by the user to quickly determine the process parameters that can meet the process needs.
It significantly shortens the time for determining process parameters, quickly obtain parameters that meet process requirements, and improves the accuracy of process parameters.
Smart Images

Figure CN119180199B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the technical field of production processes. Specifically, they relate to process formula processing technology in the technical field of production processes, and more specifically, to a method and device for determining process parameters and semiconductor process equipment. Background Art
[0002] A process formula can be simply understood as a combination of various process parameters under the working state of process equipment, and its accuracy directly affects the process result. During the process generation, usually the first to appear is the process requirement, that is, the desired process result. Then, based on this process requirement, the process parameters are continuously adjusted so that the adjusted process parameters can meet the process requirement.
[0003] However, the above process is not achieved overnight. It is necessary to continuously adjust the process parameters so that the process results of the adjusted process parameters continuously approach the desired process result, and finally obtain the process parameters that meet the process requirement. The whole process takes a long time. Summary of the Invention
[0004] Multiple embodiments in this specification provide a method and device for determining process parameters and semiconductor process equipment to achieve the purpose of quickly determining process parameters that can meet process requirements.
[0005] In a first aspect, an embodiment of this specification provides a method for determining process parameters. The method for determining process parameters includes:
[0006] In response to the target process information input by the user, input the target process information into the target network model; wherein, the target process information includes the process result that the process formula to be determined can achieve;
[0007] Predict the process parameters required to achieve the process result through the target network model; wherein, the target network model includes: a capsule network layer constructed based on a capsule network;
[0008] Determine the prediction result output by the target network model as the process parameters of the process formula to be determined.
[0009] In some embodiments, the capsule network layer includes: an input layer, a convolutional layer, a primary capsule layer, a digital capsule layer, and an output layer;
[0010] Predicting the process parameters required to achieve the process result through the target network model includes:
[0011] Preprocess the target process information through the input layer to generate input features;
[0012] Feature extraction is performed on the input features through the convolutional layer to generate an initial feature vector;
[0013] Feature extraction is performed on the initial feature vector through the primary capsule layer to generate an initial capsule vector;
[0014] Feature processing is performed on the initial capsule vector through the digital capsule layer to generate a target number of result capsule vectors; wherein, each result capsule vector corresponds to a process parameter, and the result capsule vector includes multiple predicted values of its corresponding process parameter and the prediction probability of each predicted value;
[0015] Based on the result capsule vectors, the output layer outputs the predicted value with the highest prediction probability for each process parameter.
[0016] In some embodiments, performing feature extraction on the initial feature vector through the primary capsule layer to generate an initial capsule vector includes:
[0017] Feature extraction is performed on the initial feature vector through the primary capsule layer in a dynamic routing manner to generate an initial capsule vector.
[0018] In some embodiments, before performing feature processing on the initial capsule vector through the digital capsule layer to generate a target number of result capsule vectors, the method further includes:
[0019] The execution chamber is determined based on the target process information through the target network model, wherein the execution chamber is the chamber in the process equipment that executes the to-be-determined process recipe;
[0020] Based on the correspondence between each chamber in the process equipment and the number of process parameters, the number of process parameters corresponding to the execution chamber is determined as the target number through the target network model.
[0021] In some embodiments, the target network model further includes a long short-term memory network layer constructed based on a long short-term memory network, wherein the input of the long short-term memory network layer is the output of the capsule network layer, and the output of the long short-term memory network layer is the output of the target network model.
[0022] In some embodiments, predicting the process parameters required to achieve the process result through the target network model includes:
[0023] The target process information is processed through the capsule network layer to generate a target number of result capsule vectors; wherein, each result capsule vector corresponds to a process parameter, and the result capsule vector includes multiple predicted values of its corresponding process parameter and the prediction probability of each predicted value;
[0024] Process the result capsule vector through the forget gate, input gate, output gate, and output layer of the long short-term memory network layer to obtain the process parameters required to achieve the process result.
[0025] In some embodiments, the target network model further includes: a long short-term memory network layer constructed based on the long short-term memory network, and an attention module constructed based on the attention mechanism; wherein, the input of the attention module is the target process information, the input of the long short-term memory network layer is the fusion result of the output of the attention module and the output of the capsule network layer, and the output of the long short-term memory network layer is the output of the target network model.
[0026] In some embodiments, predicting the process parameters required to achieve the process result through the target network model includes:
[0027] Process the target process information through the capsule network layer to generate a target number of result capsule vectors; wherein, each result capsule vector corresponds to a process parameter, and the result capsule vector includes multiple predicted values of its corresponding process parameter and the prediction probability of each predicted value.
[0028] Process the target process information through the attention module to generate a context vector corresponding to each result capsule vector.
[0029] Process the fusion result of the result capsule vector and the context vector through the forget gate, input gate, output gate, and output layer of the long short-term memory network layer to obtain the process parameters required to achieve the process result.
[0030] In some embodiments, in response to the target process information input by the user, before inputting the target process information into the target network model, the method further includes:
[0031] Obtain training data; wherein, the training data includes: historical process recipes and the process results that can be achieved by the historical process recipes.
[0032] Use the training data to train the network model constructed by the capsule network, and determine the trained network model as the target network model.
[0033] In some embodiments, during the process of using the training data to train the network model constructed by the capsule network, the sum of the interval loss and the reconstruction loss is used as the model loss.
[0034] In a second aspect, an embodiment of the present specification provides a device for determining process parameters, and the device for determining process parameters includes:
[0035] A response module, configured to input the target process information into a target network model in response to the target process information input by a user; wherein, the target process information includes a process result that can be achieved by a process recipe to be determined.
[0036] A model processing module, configured to predict process parameters required to achieve the process result through the target network model; wherein, the target network model includes: a capsule network layer constructed based on a capsule network.
[0037] A determination module, configured to determine the prediction result output by the target network model as the process parameters of the process recipe to be determined.
[0038] In some embodiments, the capsule network layer includes: an input layer, a convolutional layer, a primary capsule layer, a digital capsule layer, and an output layer.
[0039] The model processing module includes:
[0040] A preprocessing unit, configured to preprocess the target process information through the input layer to generate input features.
[0041] A first feature extraction unit, configured to extract features from the input features through the convolutional layer to generate an initial feature vector.
[0042] A second feature extraction unit, configured to extract features from the initial feature vector through the primary capsule layer to generate an initial capsule vector.
[0043] A model processing unit, configured to perform feature processing on the initial capsule vector through the digital capsule layer to generate a target number of result capsule vectors; wherein, each result capsule vector corresponds to a process parameter, and the result capsule vector includes multiple predicted values of its corresponding process parameter and the prediction probability of each predicted value.
[0044] A model output unit, configured to output, through the output layer, the predicted value with the maximum prediction probability of each process parameter based on the result capsule vectors.
[0045] In some embodiments, the second feature extraction unit is specifically configured to extract features from the initial feature vector through the primary capsule layer in a dynamic routing manner to generate an initial capsule vector.
[0046] In some embodiments, the apparatus further includes:
[0047] A chamber determination module, configured to determine an execution chamber through the target network model based on the target process information, where the execution chamber is a chamber in a process device that executes the process recipe to be determined.
[0048] A quantity determination module, configured to determine the quantity of process parameters corresponding to the execution chamber as the target quantity based on the corresponding relationship between each chamber in the process equipment and the quantity of process parameters through the target network model.
[0049] In some embodiments, the target network model further includes: a long short-term memory network layer constructed based on a long short-term memory network, wherein the input of the long short-term memory network layer is the output of the capsule network layer, and the output of the long short-term memory network layer is the output of the target network model.
[0050] In some embodiments, the model processing module is specifically configured to:
[0051] Process the target process information through the capsule network layer to generate result capsule vectors of the target quantity; wherein each result capsule vector corresponds to a process parameter, and the result capsule vector includes multiple predicted values of its corresponding process parameter and the prediction probability of each predicted value.
[0052] Process the result capsule vectors through the forget gate, input gate, output gate, and output layer of the long short-term memory network layer to obtain the process parameters required to achieve the process result.
[0053] In some embodiments, the target network model further includes: a long short-term memory network layer constructed based on a long short-term memory network, and an attention module constructed based on an attention mechanism; wherein the input of the attention module is the target process information, the input of the long short-term memory network layer is the fusion result of the output of the attention module and the output of the capsule network layer, and the output of the long short-term memory network layer is the output of the target network model.
[0054] In some embodiments, the model processing module is specifically configured to:
[0055] Process the target process information through the capsule network layer to generate result capsule vectors of the target quantity; wherein each result capsule vector corresponds to a process parameter, and the result capsule vector includes multiple predicted values of its corresponding process parameter and the prediction probability of each predicted value.
[0056] Process the target process information through the attention module to generate context vectors corresponding to each result capsule vector.
[0057] Process the fusion result of the result capsule vectors and the context vectors through the forget gate, input gate, output gate, and output layer of the long short-term memory network layer to obtain the process parameters required to achieve the process result.
[0058] In some embodiments, the device further includes:
[0059] a model training module, configured to:
[0060] obtain training data; wherein the training data includes: historical process recipes and process results achievable by the historical process recipes;
[0061] use the training data to perform model training on a network model constructed by a capsule network, and determine the trained network model as the target network model.
[0062] In some embodiments, during the process of using the training data to perform model training on a network model constructed by a capsule network, the sum of the margin loss and the reconstruction loss is used as the model loss.
[0063] In a third aspect, an embodiment of the present specification provides a semiconductor processing device, including: a host computer and a slave computer connected to the host computer;
[0064] The host computer is configured to execute the method for determining process parameters according to any one of the first aspect to determine the process parameters of a process recipe to be determined, and send the process parameters of the process recipe to be determined to the slave computer;
[0065] The slave computer is configured to receive the process parameters of the process recipe to be determined and perform process production according to the process parameters of the process recipe to be determined.
[0066] In a fourth aspect, an embodiment of the present specification provides an electronic device, including a processor and a memory;
[0067] wherein the memory is connected to the processor, and the memory is configured to store a computer program;
[0068] The processor is configured to implement the method for determining process parameters according to any one of the first aspect by running the computer program stored in the memory.
[0069] In a fifth aspect, an embodiment of the present specification provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, the method for determining process parameters according to any one of the first aspect is implemented.
[0070] Sixth aspect, an embodiment of this specification provides a computer program product or a computer program. The computer program product includes a computer program which is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and when the processor executes the computer program, the steps of the method for determining process parameters as described in any item of the first aspect are implemented.
[0071] For the process parameters of the process recipe to be determined in multiple embodiments provided in this specification, first, determine the process result that can be achieved or is expected to be achieved, that is, the process requirement. Then, predict the process parameters that can achieve this process result with the help of the target network model. Finally, use the predicted process parameters as the process parameters of the process recipe to be determined. Compared with the method of inversely inferring and finding the optimal process parameters by self-consistent iteration, this application can quickly determine the process parameters that can meet the process requirements and save a lot of time. At the same time, the target network model includes a capsule network layer constructed based on the capsule network, enabling the target network model to integrate the characteristics of the capsule network, better considering the correlation between different features, and thus improving the accuracy of the process parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 It is a flowchart of a traditional method used to determine process parameters provided by an embodiment of this specification;
[0073] Figure 2 It is a schematic flowchart of the method for determining process parameters provided by an embodiment of this specification;
[0074] Figure 3 It is a schematic diagram of the process of dynamic routing in the capsule network in an embodiment of this specification;
[0075] Figure 4 It is a network structure diagram of a long short-term memory network in an embodiment of this specification;
[0076] Figure 5 It is a schematic flowchart of the method for determining process parameters provided by another embodiment of this specification;
[0077] Figure 6 It is an actual application flowchart of the method for determining process parameters provided by an embodiment of this specification;
[0078] Figure 7 It is one of the application architecture diagrams of the method for determining process parameters provided by an embodiment of this specification;
[0079] Figure 8 It is another application architecture diagram of the method for determining process parameters provided by an embodiment of this specification;
[0080] Figure 9 Schematic structural diagram of a device for determining process parameters provided for an embodiment of this specification;
[0081] Figure 10 Schematic structural diagram of an electronic device provided for an embodiment of this specification. Detailed implementation manners
[0082] Unless otherwise defined, the technical terms or scientific terms used in the embodiments of this specification shall have the ordinary meanings understood by those of ordinary skill in the art to which this specification pertains. The "first", "second" and similar terms used in the embodiments of this specification do not denote any order, quantity or importance, but are merely used to avoid confusion of components.
[0083] Unless otherwise required by the context, throughout this specification, "a plurality of" means "at least two", and "including" is interpreted as open and inclusive, that is, "including, but not limited to". In the description of this specification, the terms "one embodiment", "some embodiments", "exemplary embodiments", "examples", "specific examples" or "some examples", etc. are intended to indicate that specific features, structures, materials or characteristics related to the embodiment or example are included in at least one embodiment or example of this specification. The schematic representations of the above terms do not necessarily refer to the same embodiment or example.
[0084] Next, the technical solutions in the embodiments of this specification will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0085] Overview
[0086] As described in the background art, the process recipe is very important in the process production process. How to determine appropriate process parameters to meet the process requirements has become an inevitable problem in the process generation process. As Figure 1 shown, after inputting the process requirements, a set of process parameters is usually randomly given, and then it is judged whether the current process parameters meet the requirements. If not, the process parameters need to be updated according to the degree of deviation from the expected result until the requirements are met, and then the process production will be executed according to the latest process parameters.
[0087] Although a neural network model is used to determine process parameters, the principle of this model is still to slowly find the optimal process parameters through self-consistent iteration. In order to shorten the time for determining process parameters, the inventors found through research that the process parameters required can be directly predicted based on the desired process results. At the same time, combined with the characteristics of the capsule network, the target network model can finally quickly obtain the process parameters that meet the process requirements.
[0088] Based on the above concept, the method for determining process parameters provided in the embodiments of this specification will be described exemplarily below.
[0089] Exemplary Method
[0090] As Figure 2 shown, the embodiments of this specification provide a method for determining process parameters, and this method for determining process parameters includes:
[0091] S201: In response to the target process information input by the user, input the target process information into the target network model.
[0092] In this step, the target process information includes the process results that the process recipe to be determined can achieve. Among them, the process recipe to be determined is the process recipe expected to be obtained or the process recipe that can meet the process requirements and is expected to be obtained. The process results can be understood as process requirements or process requirement results. Specifically, in the process generation process, if it is necessary to determine a process recipe that can achieve a certain process result or meet a certain process requirement, then this process recipe can be regarded as the process recipe to be determined, and this process result or process requirement can be regarded as the target process information. For example, for semiconductor process equipment, assume that it is necessary to determine the specific parameter values of process temperature, process pressure, and gas flow rate when the deposition rate reaches 10 nanometers per minute. The deposition rate reaching 10 nanometers per minute is the target process information, and the specific parameter values of process temperature, process pressure, and gas flow rate are the process recipe to be determined or the process parameters in the process recipe. It can be understood that the target process information and the process recipe are related to the process production process involved in the process equipment, and this process production process includes but is not limited to the process production process of semiconductor process equipment.
[0093] It should be noted that the method provided in this specification can be directly applied in the process equipment to determine the process parameters. Thus, after determining the process parameters, the process production can be directly carried out according to these process parameters. In this case, the user only needs to input the target process information in the process equipment. Of course, it is also possible to use other electronic devices with relatively powerful computing power to determine the process parameters through the method provided in this embodiment. Then input the process parameters into the process equipment so that the process equipment can carry out the process production according to the input process parameters.
[0094] S202: Predict the process parameters required to achieve the process result through the target network model.
[0095] S203: Determine the prediction result output by the target network model as the process parameters of the process recipe to be determined.
[0096] It should be noted that the target network model is a pre-trained network model, whose input is the process requirement or process result, and the output is the predicted process parameters that can meet the process requirement or achieve the process result, that is, the prediction result. Therefore, with the help of the target network model, the process parameters corresponding to the target process information can be obtained, that is, the process parameters that can achieve the process result in the target process information can be obtained. After determining the process parameters, the process recipe to be determined can be determined. Furthermore, the process equipment can use the determined process recipe to be determined for process production.
[0097] The target network model includes: a capsule network layer constructed based on the capsule network. Among them, the capsule network is Capsule Net, and its advantages are more obvious in some aspects compared with the convolutional neural network. For example, through hierarchical feature extraction, the capsule network can obtain richer object features. In contrast, the convolutional neural network mainly extracts features through tiling and pooling operations, and may lose some detailed information. The capsule network represents objects in the form of vectors, and this distributed representation can better capture the internal structure and relationships of objects. The convolutional neural network generally uses scalars to represent features and is difficult to represent the complex structure of objects.
[0098] In this embodiment, for the process parameters of the process recipe to be determined, first determine the process result that can be achieved or expected to be achieved, that is, the process requirement. Furthermore, with the help of the target network model, predict the process parameters that can achieve this process result, and finally use the predicted process parameters as the process parameters of the process recipe to be determined. Compared with the method of using self-consistent iteration to inversely deduce and find the optimal process parameters, the present application can quickly determine the process parameters that can meet the process requirements and save a lot of time. At the same time, the target network model includes a capsule network layer constructed based on the capsule network, so that the target network model integrates the characteristics of the capsule network, can better consider the correlation between different features, and thus improves the accuracy of the process parameters.
[0099] In some embodiments, the capsule network layer includes: an input layer, a convolutional layer, a primary capsule layer, a digital capsule layer, and an output layer;
[0100] Predicting the process parameters required to achieve the process result through the target network model includes:
[0101] Preprocess the target process information through the input layer to generate input features;
[0102] Feature extraction is performed on the input features through a convolutional layer to generate an initial feature vector;
[0103] Feature extraction is performed on the initial feature vector through a primary capsule layer to generate an initial capsule vector;
[0104] Feature processing is performed on the initial capsule vector through a digit capsule layer to generate a result capsule vector with the target number; wherein, each result capsule vector corresponds to a process parameter, and the result capsule vector includes multiple predicted values of its corresponding process parameter and the prediction probability of each predicted value;
[0105] Based on the result capsule vector, the output layer outputs the predicted value with the maximum prediction probability for each process parameter.
[0106] It should be noted that the input of the input layer is the input of the target network model, which is used to preprocess the input data of the input target network model, so as to process it into features that can be directly recognized and processed. The input of the convolutional layer is the output of the input layer, which is used to perform convolutional processing on the output of the input layer to generate an initial feature vector. The convolutional processing will not be elaborated here. Optionally, the convolutional layer in this embodiment can adopt a three-layer convolutional structure to avoid the network structure from being too complex and can be applied to the parameter adjustment of large equipment. The input of the primary capsule layer is the output of the convolutional layer, which is used to process the output of the convolutional layer to generate an initial capsule vector. Among them, the primary capsule layer is the PrimaryCaps layer in the capsule network, and its data processing can be realized, which will not be elaborated here. The input of the digit capsule layer is the output of the primary capsule layer, which is used to process the output of the primary capsule layer to generate a result capsule vector. Among them, the digit capsule layer is the DigitCaps layer of the capsule network, and its data processing can be realized, which will not be elaborated here. The input of the output layer is the output of the digit capsule layer, which is used to process the output of the digit capsule layer to obtain the output of the final target network model. The output of the output layer is the output of the target network model.
[0107] Specifically, the input layer preprocesses the target process information, including: unifying the length of the target process information and performing IDX conversion. Among them, unifying the length includes: uniformly adjusting the input data to the same length, which can be achieved by filling null values or truncating, etc. IDX is a binary file format, usually used to store multi-dimensional array data. Through IDX conversion, it is converted into a tensor or matrix format that can adapt to the network model for input.
[0108] The processing process of the convolutional layer includes: processing the input features using Equation 1.
[0109] Equation 1: Conv i =σ(Wconv · x i + b conv );
[0110] Among them, Conv i represents the initial feature vector, W conv represents the convolution kernel, b conv represents the bias term, σ is a non-linear activation function, for example, it can be the RELU function, but not limited to this, x i represents the input feature.
[0111] It can be understood that the larger the convolution kernel, the greater the amount of information contained in the initial feature vector.
[0112] The processing process of the primary capsule layer includes: processing the initial feature vector using Equation 2.
[0113] Equation 2: Cap conv = ρ(W cap · Conv i + b cap );
[0114] Among them, Cap conv represents the initial capsule vector, W cap represents the convolution kernel of the primary capsule layer, Conv i represents the initial feature vector, b cap represents the bias term, and ρ is the squash function in the capsule network.
[0115] The primary capsule layer can convert the multi-dimensional data output by the convolutional layer into several capsules or capsule vectors. Specifically, it performs deep feature extraction on the target process information.
[0116] The processing process of the digital capsule layer includes: processing the initial capsule vector using Equation 3.
[0117] Equation 3: Cap full = φ(W full · Cap conv + b full );
[0118] Among them, Cap full represents the resulting capsule vector, Cap conv represents the initial capsule vector, W full and b full are the weight and bias respectively. φ represents the transformation matrix.
[0119] It can be understood that each capsule is associated with the previous capsule network through a transformation matrix. The input is predicted using the transformation matrix. Through the calculation of the digital capsule layer, multiple capsule vectors containing feature information, that is, the result capsule vectors, can be obtained, and the number of capsule vectors is equal to the number of process parameters in the process recipe to be determined.
[0120] In this embodiment, based on the input layer, convolutional layer, primary capsule layer, digital capsule layer, and output layer to form the capsule network layer, the characteristics of the capsule network can be well integrated in the target network model, while avoiding the network structure of the capsule network layer from being too complex.
[0121] In some embodiments, the primary capsule layer extracts features from the initial feature vectors to generate initial capsule vectors, including:
[0122] The primary capsule layer extracts features from the initial feature vectors in a dynamic routing manner to generate initial capsule vectors.
[0123] It should be noted that in the capsule network, the primary capsule layer can be regarded as consisting of multiple layers of capsules, and each capsule in each layer of capsules is connected to each capsule in the next layer. Moreover, each capsule performs the same dynamic routing process on the output of the previous layer of capsules. Finally, the output of the last layer of capsules is used as the initial capsule vector. Among them, when the lower-layer capsules transfer the input vector to the higher-layer capsules, the corresponding parameters are continuously adjusted so that the higher-layer capsules can output more appropriate data. This process can be called dynamic routing. Since the dynamic routing process of each capsule is the same, only one capsule is taken as an example for illustration here. As Figure 3 shown, assume that the previous layer of capsules includes two capsules, and the outputs of the two capsules are V1 and V2 respectively, and the weights are W1 and W2 respectively. Then, the coupling coefficients used by this capsule in the first processing are and The result obtained is S1. After being processed by the squashing function, the coupling coefficients used in the update process are obtained as and Then, a second processing is performed in this capsule, and the updated coupling coefficients and are used to obtain a new result S2. After being processed by the squashing function again, the coupling coefficients used in the update process are obtained as and Then, a third processing is performed in this capsule, and the updated coupling coefficients and Get a new result S3. When determining that S3 is appropriate data, the true output V of the capsule is obtained through the squash function. It can be understood that when representing the probability of process parameters based on the length of the capsule, the norm length of the capsule can represent the significance of the feature, and the longer the norm length, the more significant the feature. At the same time, when the norm length is close to 0, there is also an amplification effect, and there is no need to use the squash function for global compression. Among them, the squash function is as follows:
[0124]
[0125] Among them, the squash function squash(v j ) is mainly divided into two parts. The first part changes the length of the input vector s j through the squash function without changing the direction of the input vector s j . The second part controls the size of the first part through a vector with a length of 1. In this way, both the direction of the input vector s j is retained and the size can be controlled. For the input vector s j , it can be determined by the following formula:
[0126]
[0127] Among them, u i represents the output of the lower-level capsule, W ij represents the weight parameter, c ij represents the coupling coefficient. Among them, there is a coupling coefficient between the higher-level capsule and the lower-level capsule, and the sum of the coupling coefficients is 1.
[0128] During the process of updating the coupling coefficient, the intermediate parameter b can be used. Specifically as follows:
[0129] Among them,
[0130]
[0131] Among them, b ij represents the intermediate parameter, u i represents the output of the lower-level capsule, v j represents the output of the higher-level capsule. When and v j are in the same direction, it means that the features expressed by the lower-level capsule and the higher-level capsule are highly similar, and the weight of b ij will increase; on the contrary, it means that the features of the lower-level capsule most likely do not belong to the higher-level capsule, and the weight of b ij will decrease. The dynamic routing algorithm establishes the connection between the lower-level capsule and the higher-level capsule through the weight update method.
[0132] Of course, in addition to dynamic routing, clustering algorithms can also be used to generate initial capsule vectors. It can be understood that when using clustering algorithms, the process from low-level capsules to high-level capsules is similar to the process of generating cluster centroids in clustering. Assume that the features of low-level capsules are u1, u2, …, u n , and they are divided into k categories. The process of determining the initial capsule vectors can be understood as finding the central features v1, v2, …, v k of k high-level capsules, and minimizing the within-class separation. The within-class separation can be calculated based on Equation Four.
[0133] Equation Four:
[0134] where L represents the minimum within-class separation, u i represents the output of the low-level capsules, and v j represents the output of the high-level capsules. d represents the measure of the distance between the low-level capsules and the high-level capsules. The result of clustering depends on the specific form of d.
[0135] Adding up the distances between all categories of high-level capsules, with the goal of minimizing the distance, then where v1, v2, …, v k represent the central features of k high-level capsules, and L represents the minimum within-class separation. Then, with the help of the approximation formula, L can be solved. Specifically,
[0136] Approximation Formula One:
[0137] Approximation Formula Two: min(λ1, λ2,... λ n ) = -max(-λ1, -λ2,... -λ n );
[0138] where λ1, λ2 …… λ n represent n numerical values.
[0139] Based on the above approximation formulas, the calculation formula for the minimum within-class separation can be obtained:
[0140]
[0141] Through mathematical calculations, it can be obtained that:
[0142]
[0143] Let represent the result of the r-th iteration of v j , <,> represents the scalar product operation, c ij represents the coupling coefficient, u iRepresents the output of the low-level capsule, v i Represents the output of the high-level capsule, from which the relationship between the low-level capsule and the high-level capsule can be obtained. That is:
[0144] Among them, Represents the coupling coefficient in the r-th iteration, u i Represents the output of the low-level capsule, Represents the output of the high-level capsule in the (r + 1)-th iteration.
[0145] Among them, the transformation matrix W is introduced ij Can arbitrarily change the dimensions of the centers of the low-level capsule and the high-level capsule to ensure the feature expression ability of the high-level capsule. Correspondingly, the relationship between the low-level capsule and the high-level capsule will change to:
[0146] Among them, W ij Represents the transformation matrix, Represents the coupling coefficient in the r-th iteration, u i Represents the output of the low-level capsule, Represents the output of the high-level capsule in the (r + 1)-th iteration.
[0147] In this embodiment, by virtue of the characteristics of dynamic routing, the connection between various parameters can be strengthened, so that the internal structure and relationship of the processing object can be better captured. Compared with other clustering algorithms, the network structure in this embodiment is more lightweight.
[0148] To adapt to the process parameters in different execution chambers, in some implementation methods, before generating the result capsule vectors with the target quantity by performing feature processing on the initial capsule vectors through the digital capsule layer, the method further includes:
[0149] Determining the execution chamber through the target network model based on the target process information, where the execution chamber is the chamber in the process equipment that executes the process recipe to be determined;
[0150] Determining the quantity of process parameters corresponding to the execution chamber as the target quantity through the target network model based on the corresponding relationship between each chamber in the process equipment and the quantity of process parameters.
[0151] It should be noted that the process equipment includes multiple execution chambers, and different execution chambers can perform different process productions according to different process parameters or process recipes. For example, semiconductor process equipment includes: MOW chamber, CVDW chamber. Among them, the process parameters executed by the MOW chamber include: process temperature, process pressure, gas flow rate, etc. The process parameters executed by the CVDW chamber include: Soak time, process temperature, etc.
[0152] The target network model can be applicable to multiple execution chambers. Since the process results of different execution chambers are different, the different execution chambers can be distinguished by the process results, so as to determine which process parameters need to be output. Of course, the target network model can also be only applicable to one execution chamber. No matter which execution chamber's target process information is the input of the target network model, its output is the process parameters of a certain fixed execution chamber.
[0153] In this embodiment, the target network model can determine the execution chamber based on the target process information, so as to predict the process parameters for the determined execution chamber, enabling the target network model to adapt to the process parameters in different execution chambers.
[0154] To further improve the accuracy of the model output and avoid the problems of gradient vanishing or gradient explosion, in some implementation methods, the target network model further includes: a long short-term memory network layer constructed based on the long short-term memory network, where the input of the long short-term memory network layer is the output of the capsule network layer, and the output of the long short-term memory network layer is the output of the target network model.
[0155] It should be noted that the long short-term memory network is the LSTM (Long Short-Term Memory) network. The long short-term memory network layer in the target network model integrates the advantages of the LSTM network in terms of gradients, and can avoid the problems of gradient vanishing and gradient explosion. In this embodiment, the long short-term memory network layer is located behind the capsule network layer, and the input of the long short-term memory network layer is the output of the capsule network layer. The output of the long short-term memory network layer is the output of the target network model.
[0156] In this embodiment, by integrating the advantages of the LSTM network in terms of gradients and setting the long short-term memory network layer in the target network model, not only can the problems of gradient vanishing and gradient explosion be avoided; at the same time, the temporal characteristics of the target process information can be considered to improve the accuracy of the model output.
[0157] In some embodiments, predicting the process parameters required to achieve the process result through the target network model includes:
[0158] Processing the target process information through the capsule network layer to generate a target number of result capsule vectors; where each result capsule vector corresponds to a process parameter, and the result capsule vector includes multiple predicted values of its corresponding process parameter and the prediction probability of each predicted value;
[0159] Processing the result capsule vectors through the forget gate, input gate, output gate, and output layer of the long short-term memory network layer to obtain the process parameters required to achieve the process result.
[0160] It should be noted that the network architecture of the long short-term memory network layer is as shown in Figure 4 . Among them, the forget gate is f t , the input gate is i t , and the output gate is o t . The input of the input gate i t is determined by the historical state h t-1 and the current node state x t . σ is the sigmoid function, which controls the output between 0 and 1. 1 means retaining historical information, and 0 means forgetting historical information. The input gate can be calculated by the following formula:
[0161] i t = σ(W ix · x t + W ih · h t-1 + b i );
[0162] Among them, i t represents the output result of the input gate, W ix represents the weight matrix of the current state in the input gate, W ih represents the weight matrix of the historical state in the input gate, and b i represents the weight bias for the input gate.
[0163] Similarly, the forget gate and the output gate can be calculated by the following formula:
[0164] f t = σ(W fx · x t + W fh · h t-1 + b f ); o t = σ(W ox · x t + W oh · h t-1 + b o );
[0165] Among them, f t represents the output result of the forget gate, W fx represents the weight matrix of the current state in the forget gate, W fh represents the weight matrix of the historical state in the forget gate, and b f represents the weight bias for the forget gate; o t represents the output result of the output gate, W ox represents the weight matrix of the current state in the output gate, W oh represents the weight matrix of the historical state in the output gate, and b o represents the weight bias for the output gate.
[0166] Then, according to the previous state g t and the current state s t select new information to be passed to the next neuron. The specific process is as follows:
[0167] g t = σ(W ox ·x t + W oh ·h t-1 + b o ); s t = g t ·i t + s t-1 ·f t ; h t = tanh(s t )·o t ; The function tanh is used to process the current state and multiply it with the output gate to achieve the hidden vector h t , and then connect it to the next unit to extract the associated time feature information of the front and back data packets during the transmission process. Finally, the softmax function is used to obtain the final predicted output result.
[0168] In this embodiment, a long short-term memory network layer is composed of a forget gate, an input gate, an output gate, and an output layer to process the result capsule vector, taking into account the temporal characteristics of the data, and more accurate process parameters can be obtained.
[0169] To further improve the accuracy of the model output and avoid the problems of gradient disappearance or gradient explosion, in some implementation methods, the target network model further includes: a long short-term memory network layer constructed based on the long short-term memory network, and an attention module constructed based on the attention mechanism; wherein, the input of the attention module is the target process information, the input of the long short-term memory network layer is the fusion result of the output of the attention module and the output of the capsule network layer, and the output of the long short-term memory network layer is the output of the target network model.
[0170] It should be noted that the situation of introducing the LSTM network in the target network model is similar to that of introducing the LSTM network in the above embodiments, and will not be elaborated here. Among them, the difference between the two is that the present application also introduces an attention module constructed based on the attention mechanism. Among them, the attention mechanism can be understood as a method that mimics the human visual and cognitive systems, allowing the neural network model to focus on relevant parts when processing input data. By introducing the attention mechanism, the neural network model can automatically learn and selectively focus on important information in the input, thereby improving the performance and generalization ability of the model. After the attention module is trained, its output can indicate the importance degree of each feature output by the capsule network layer. Furthermore, by fusing the output of the attention module and the output of the capsule network layer, the proportion of features with higher importance degree can be increased, so as to focus on important information. Then, the fusion result is input into the LSTM network, and by continuously adjusting the weights between features, a more accurate result can be obtained.
[0171] In this embodiment, the advantages of the LSTM network in terms of gradients are integrated. By setting a long short-term memory network layer in the target network model, not only can the problems of gradient disappearance and gradient explosion be avoided; at the same time, the temporal characteristics of the target process information can be considered, and the attention mechanism is introduced, which can further improve the accuracy of the model output.
[0172] In some embodiments, the process parameters required to achieve the process result are predicted through the target network model, including:
[0173] The target process information is processed by the capsule network layer to generate a target number of result capsule vectors; wherein, each result capsule vector corresponds to a process parameter, and the result capsule vector includes multiple predicted values of its corresponding process parameter and the predicted probability of each predicted value;
[0174] The target process information is processed by the attention module to generate a context vector corresponding to each result capsule vector;
[0175] The fusion result of the result capsule vector and the context vector is processed by the forget gate, input gate, output gate and output layer of the long short-term memory network layer to obtain the process parameters required to achieve the process result.
[0176] It should be noted that the network architecture of the long short-term memory network layer is as Figure 4As shown. The situation of the long short-term memory network layer in the target network model processing the fusion result is similar to the situation of the long short-term memory network layer processing the result capsule vector in the above embodiment, which will not be elaborated here. Among them, the object processed by this application is different, that is, the fusion result of the result capsule vector and the context vector. Regarding the attention mechanism, it will not be elaborated here. It can be understood that the context vector is the output of the attention module. The context vector corresponding to each result capsule vector, and the numerical size of the context vector is positively correlated with the importance degree of the result capsule vector. That is to say, the higher the importance degree of a certain result capsule vector in the process of predicting process parameters, the larger its corresponding context vector. This is the ability learned by the attention module based on the attention mechanism during the model training process. Specifically, multiple partially overlapping consecutive subsequences can be intercepted from the original input sequence, that is, the target process information, in the input of the input layer as the input of the attention module. The context vector output by the attention module can be regarded as a significant feature, which can identify the importance degree of each output feature of the digital capsule layer. The higher the importance degree of the output feature of the digital capsule layer, the closer the output of its corresponding attention module is to 1; on the contrary, the lower the importance degree of the output feature of the digital capsule layer, the closer the output value of its corresponding attention module is to 0. Then, the output feature of the digital capsule layer is fused with the significant feature output by its corresponding attention module. Among them, the fusion method is to multiply the output feature of the digital capsule layer and the significant feature output by its corresponding attention module element by element. Since the level of the value output by the attention module reflects the importance degree of the output feature of the digital capsule layer, the discrimination of important features is completed, which is more conducive to the subsequent processing of the LSTM network layer. Therefore, the accuracy of the model can be further improved. During the model training process, introducing the attention mechanism can focus on important information with high weight, continuously adjust the weights between features, and improve the accuracy of the prediction result compared with the existing algorithms.
[0177] In this embodiment, the long short-term memory network layer is composed of a forget gate, an input gate, an output gate, and an output layer to process the fusion result, taking into account the temporal characteristics of the data, and more accurate process parameters can be obtained. At the same time, introducing the attention mechanism can reduce the interference of useless parameters, solve the problem of information overload, weight the data with greater influence on the result among many input information, and improve the accuracy of the final prediction.
[0178] In some implementation methods, in response to the target process information input by the user, before inputting the target process information into the target network model, the method further includes:
[0179] Obtain training data; wherein, the training data includes: historical process recipes and the process results that can be achieved by the historical process recipes;
[0180] The network model constructed by the capsule network is trained using training data, and the trained network model is determined as the target network model.
[0181] It should be noted that the historical process recipe is the process recipe used during the generation of the historical process. The corresponding process result is the achievable process result. For example, during the generation of the historical process, when using historical process recipe A, process result a can be achieved. When using historical process recipe B, process result b can be achieved. Then both historical process recipe A and historical process recipe B can be used as training data.
[0182] For the training of the network model constructed by the capsule network, its training process is the same as that of the neural network model, which is to use the pre-prepared training data for training, so that the trained network model can accurately predict process parameters based on the process result. In some embodiments, a dataset required for training can be prepared in advance, and then the dataset is divided into 80% offline training data and 20% real-time test data to achieve model training and testing. During the testing process, the parameters are adjusted according to the accuracy of the prediction results. If overfitting occurs in the loss function and gradient descent function, the network architecture and parameters are continuously adjusted to improve the generalization ability, and the predicted process result is obtained through the finally adjusted model.
[0183] In this embodiment, the training of the network model constructed by the capsule network can be achieved by using the training method of the neural network model.
[0184] In some embodiments, during the process of training the network model constructed by the capsule network using training data, the sum of the margin loss and the reconstruction loss is used as the model loss.
[0185] It should be noted that the margin loss can be determined as the sum of the K-class margin losses. Specifically,
[0186] where, L margin represents the margin loss. T k represents the loss value of the k-th class prediction result. If the prediction result is consistent with the label, T k is 1, otherwise it is 0. Usually, λ is taken as 0.5, m + = 0.9, m - = 0.9, v k represents the output of the k-th class high-level capsule.
[0187] For the reconstruction loss, the reconstruction loss can be obtained by calculating the Euclidean distance between the prediction and the input. Therefore, in some embodiments, when the intermediate margin loss is dominant, the model loss is as follows:
[0188] L = Lmargin +0.005*L reconstruct , where L represents the model loss, and L margin represents the margin loss, and L reconstruct represents the reconstruction loss.
[0189] In this embodiment, taking the sum of the margin loss and the reconstruction loss as the model loss can combine the advantages of the two loss functions and prevent overfitting.
[0190] In another embodiment of this specification, the model is continuously adjusted and optimized during training using offline data, and the modified model is synchronized to the online in real time for real-time data processing. Then, the data generated during the real-time data processing is used as offline training data to expand the training library, thereby improving the prediction accuracy of the model. Specifically, as Figure 5 shown, it includes an offline data processing stage and a real-time data processing stage. The offline data processing stage includes: offline data training, adjusting the gradient optimization method and the objective function, tuning the model parameters, outputting and analyzing the correct rate, and modifying the model according to the analysis. In this way, an optimized model can be obtained and updated to the real-time data processing stage. The real-time data processing stage: inputting real-time data, making predictions on the real-time data, outputting and analyzing the correct rate, and optimizing the model according to the results. At the same time, the generated data is synchronized to the offline training data to expand the training library.
[0191] The following uses a specific example to illustrate the method for determining process parameters provided by this application. As Figure 6 shown, this method is applied to a process device and specifically includes:
[0192] Step S601: Input the original process requirement data into the process device, where the original process requirement data is equivalent to the target process information in the foregoing embodiment and will not be elaborated here.
[0193] Step S602: Data processing. This is equivalent to the input layer of the foregoing capsule network layer, and preprocesses the original process requirement data.
[0194] Step S603: Use the capsule long short-term memory network model to predict the processed data, and process parameters that meet the original process requirement data can be obtained.
[0195] Step S604: Determine the process parameters output by the capsule long short-term memory network model as the process parameters that can meet the original process requirement data.
[0196] It can be understood that the application architecture of this embodiment is as Figure 7As shown in the figure, it includes: an input layer, a convolutional layer, a primary capsule layer, a digit capsule layer, a long short-term memory network layer, and an output layer. Of course, the application architecture of this embodiment can also be as Figure 8 As shown in the figure, it includes: an input layer, a convolutional layer, a primary capsule layer, a digit capsule layer, an attention module, a long short-term memory network layer, and an output layer. Among them, the input layer is used to execute step S602. The convolutional layer, the primary capsule layer, the digit capsule layer, the attention module, and the long short-term memory network layer can refer to the convolutional layer, the primary capsule layer, the digit capsule layer, the attention module, and the long short-term memory network layer in the foregoing embodiments. The output layer is composed of a softmax function. For the introduction of each part in this application architecture, reference can be made to the description of the corresponding part in the foregoing embodiments, and details will not be described here.
[0197] In this embodiment, the efficiency of obtaining process parameters is improved. Process personnel do not need to manually try one by one or guess the approximate range value of process parameters. At the same time, it can adapt to different inputs of different chambers to obtain different outputs, and is widely used. In addition, the capsule long short-term memory network can obtain the characteristics of process results in all directions through a unique dynamic routing algorithm, making the prediction results more accurate.
[0198] Exemplary Apparatus
[0199] Some embodiments of this specification also provide a device for determining process parameters, such as Figure 9 As shown in the figure, the device for determining process parameters includes:
[0200] A response module 91, configured to input the target process information into the target network model in response to the target process information input by the user; where the target process information includes the process result that the process recipe to be determined can achieve;
[0201] A model processing module 92, configured to predict the process parameters required to achieve the process result through the target network model; where the target network model includes: a capsule network layer constructed based on a capsule network;
[0202] A determination module 93, configured to determine the prediction result output by the target network model as the process parameters of the process recipe to be determined.
[0203] In some embodiments, the capsule network layer includes: an input layer, a convolutional layer, a primary capsule layer, a digit capsule layer, and an output layer;
[0204] The model processing module 92 includes:
[0205] A preprocessing unit, configured to preprocess the target process information through the input layer to generate input features;
[0206] The first feature extraction unit is used to extract features from the input features through a convolutional layer to generate an initial feature vector;
[0207] The second feature extraction unit is used to extract features from the initial feature vector through a primary capsule layer to generate an initial capsule vector;
[0208] The model processing unit is used to perform feature processing on the initial capsule vector through a digital capsule layer to generate a target number of result capsule vectors; wherein, each result capsule vector corresponds to a process parameter, and the result capsule vector includes multiple predicted values of its corresponding process parameter and the prediction probability of each predicted value;
[0209] The model output unit is used to output, through an output layer based on the result capsule vectors, the predicted value with the highest prediction probability for each process parameter.
[0210] In some embodiments, the second feature extraction unit is specifically used to extract features from the initial feature vector in a dynamic routing manner through a primary capsule layer to generate an initial capsule vector.
[0211] In some embodiments, the device further includes:
[0212] The chamber determination module is used to determine an execution chamber based on target process information through a target network model, where the execution chamber is the chamber in the process equipment that executes the process recipe to be determined;
[0213] The quantity determination module is used to determine the quantity of process parameters corresponding to the execution chamber as the target quantity based on the correspondence between each chamber in the process equipment and the quantity of process parameters through the target network model.
[0214] In some embodiments, the target network model further includes a long short-term memory network layer constructed based on a long short-term memory network, where the input of the long short-term memory network layer is the output of the capsule network layer, and the output of the long short-term memory network layer is the output of the target network model.
[0215] In some embodiments, the model processing module 92 is specifically used to:
[0216] Process the target process information through a capsule network layer to generate a target number of result capsule vectors; wherein, each result capsule vector corresponds to a process parameter, and the result capsule vector includes multiple predicted values of its corresponding process parameter and the prediction probability of each predicted value;
[0217] Process the result capsule vectors through the forget gate, input gate, output gate and output layer of the long short-term memory network layer to obtain the process parameters required to achieve the process result.
[0218] In some embodiments, the target network model further includes: a long short-term memory network layer constructed based on a long short-term memory network, and an attention module constructed based on an attention mechanism; wherein, the input of the attention module is the target process information, the input of the long short-term memory network layer is the fusion result of the output of the attention module and the output of the capsule network layer, and the output of the long short-term memory network layer is the output of the target network model.
[0219] In some embodiments, the model processing module 92 is specifically configured to:
[0220] Process the target process information through the capsule network layer to generate a target number of result capsule vectors; wherein, each result capsule vector corresponds to a process parameter, and the result capsule vector includes multiple predicted values of its corresponding process parameter and the prediction probability of each predicted value;
[0221] Process the target process information through the attention module to generate a context vector corresponding to each result capsule vector;
[0222] Process the fusion result of the result capsule vector and the context vector through the forget gate, input gate, output gate, and output layer of the long short-term memory network layer to obtain the process parameters required to achieve the process result.
[0223] In some embodiments, the device further includes:
[0224] A model training module, configured to:
[0225] Obtain training data; wherein, the training data includes: historical process recipes and the process results that can be achieved by the historical process recipes;
[0226] Use the training data to perform model training on the network model constructed by the capsule network, and determine the trained network model as the target network model.
[0227] In some embodiments, during the process of performing model training on the network model constructed by the capsule network using the training data, the sum of the margin loss and the reconstruction loss is used as the model loss.
[0228] The device for determining process parameters provided in the embodiments of this specification belongs to the same inventive concept as the method for determining process parameters provided in the above embodiments of this specification. For technical details not described in detail in this embodiment, reference may be made to the specific processing content of the method for determining process parameters provided in the above embodiments of this specification, which will not be elaborated here.
[0229] Exemplary Semiconductor Processing Equipment and Electronic Equipment
[0230] Another embodiment of the present application further provides a semiconductor process equipment, including: a host computer and a slave computer connected to the host computer;
[0231] The host computer is used to execute the method for determining process parameters of various embodiments of this specification to determine the process parameters of the process recipe to be determined, and send the process parameters of the process recipe to be determined to the slave computer;
[0232] The slave computer is used to receive the process parameters of the process recipe to be determined and perform process production according to the process parameters of the process recipe to be determined.
[0233] Another embodiment of this application also proposes an electronic device. Refer to Figure 10 As shown, an exemplary embodiment of this specification also provides an electronic device, including: a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it executes the steps in the method for determining process parameters according to various embodiments of this specification described in the above embodiments of this specification.
[0234] The internal structure of the electronic device can be as Figure 10 As shown, the electronic device includes a processor, a memory, a network interface, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it performs the steps in the method for determining process parameters according to various embodiments of this specification described in the above embodiments of this specification.
[0235] The processor may include a main processor, and may also include a baseband chip, a modem, etc.
[0236] The memory stores a computer program for implementing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the computer program may include program code, and the program code includes computer operation instructions. More specifically, the memory may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash, etc.
[0237] The processor can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0238] The input device can include devices for receiving data and information input by the user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer or a gravity sensor, etc.
[0239] The output device can include devices for allowing information to be output to the user, such as a display screen, a printer, a speaker, etc.
[0240] The communication interface can include devices of any transceiver type for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), etc.
[0241] The processor executes the computer program stored in the memory and calls other devices, which can be used to implement each step of the method for determining any process parameter provided in the above embodiments of the present application.
[0242] The electronic device can also include a display component and a voice component. The display component can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display component, or a button, a trackball or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad or mouse, etc.
[0243] Those skilled in the art can understand that Figure 10 the structure shown is only a block diagram of some structures related to the solution of this specification, and does not constitute a limitation on the electronic device to which the solution of this specification is applied. The specific electronic device can include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0244] Exemplary Computer Program Product and Computer Readable Storage Medium
[0245] In addition to the above methods and devices, the method for determining process parameters provided in the embodiments of this specification can also be a computer program product, which includes a computer program. When the computer program is run by a processor, the processor is caused to execute the steps in the method for determining process parameters according to various embodiments of this specification described in the "Exemplary Method" section of this specification.
[0246] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of this specification. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0247] In addition, the embodiments of this specification also provide a computer-readable storage medium, on which a computer program is stored. The computer program is executed by a processor to perform the steps in the method for determining process parameters according to various embodiments of this specification described in the "Exemplary Method" section above of this specification.
[0248] It can be understood that the specific examples herein are only for helping those skilled in the art to better understand the embodiments of this specification, rather than limiting the scope of this specification.
[0249] It can be understood that in various embodiments of this specification, the magnitudes of the sequence numbers of the various processes do not mean the order of execution. The order of execution of the various processes should be determined by their functions and internal logics, and should not constitute any limitation to the implementation process of the embodiments of this specification.
[0250] It can be understood that the various embodiments described in this specification can be implemented alone or in combination, and the embodiments of this specification do not limit this.
[0251] Unless otherwise specified, all technical and scientific terms used in the embodiments of this specification have the same meaning as commonly understood by those skilled in the technical field of this specification. The terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the scope of this specification. The term "and / or" used in the embodiments of this specification and the appended claims includes any and all combinations of one or more of the related listed items. The singular forms "a", "above", and "the" used in the embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0252] It can be understood that the processor in the embodiments of this specification can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of this specification. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of this specification can be directly embodied as being executed and completed by a hardware decoding processor, or can be executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0253] It can be understood that the memory in the embodiments of this specification can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.
[0254] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this specification.
[0255] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0256] In the several embodiments provided in this specification, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0257] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0258] In addition, in each embodiment of this specification, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0259] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this specification, in essence, or the part that contributes to the prior art or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this specification. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0260] The above is only the specific embodiment of this specification, but the protection scope of this specification is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this specification and should be covered by the protection scope of this specification. Therefore, the protection scope of this specification should be subject to the protection scope of the claims.
Claims
1. A method for determining process parameters, characterized in that: The method comprises: In response to target process information input by a user, the target process information is input into a target network model; wherein the target process information includes a process result that can be achieved by a process equipment performing process production according to a process recipe to be determined; the process equipment includes a semiconductor process equipment; Predicting the process parameters required to achieve the process result through the target network model; wherein the target network model includes: a capsule network layer constructed based on a capsule network; Determine the prediction result output by the target network model as the process parameter of the process recipe to be determined, so that the process equipment uses the determined process recipe to be determined to perform process production; The target network model further includes: a long short-term memory network layer constructed based on a long short-term memory network, and an attention module constructed based on an attention mechanism; wherein the input of the attention module is the target process information, the input of the long short-term memory network layer is the fusion result of the output of the attention module and the output of the capsule network layer, and the output of the long short-term memory network layer is the output of the target network model; Predicting the process parameters required to achieve the process result through the target network model includes: Processing the target process information through the capsule network layer to generate a target number of result capsule vectors; wherein each of the result capsule vectors corresponds to a process parameter, and the result capsule vector includes a plurality of predicted values of the corresponding process parameter and a predicted probability of each of the predicted values; Processing the target process information through the attention module to generate a context vector corresponding to each of the result capsule vectors; The fusion result of the result capsule vector and the context vector is processed through the forget gate, input gate, output gate and output layer of the long short-term memory network layer to obtain the process parameters required to achieve the process result.
2. The method according to claim 1, characterized in that The capsule network layer includes: an input layer, a convolutional layer, a main capsule layer, a digital capsule layer and an output layer; Predicting the process parameters required to achieve the process result through the target network model includes: Preprocessing the target process information through the input layer to generate input features; Extracting features from the input features through the convolution layer to generate an initial feature vector; Performing feature extraction on the initial feature vector through the main capsule layer to generate an initial capsule vector; Performing feature processing on the initial capsule vector through the digital capsule layer to generate a target number of result capsule vectors; wherein each of the result capsule vectors corresponds to a process parameter, and the result capsule vector includes a plurality of predicted values of the corresponding process parameter and a predicted probability of each of the predicted values; The output layer outputs a predicted value with the maximum predicted probability for each process parameter based on the result capsule vector.
3. The method according to claim 2, characterized in that Extracting features from the initial feature vector through the main capsule layer to generate an initial capsule vector includes: The initial feature vector is extracted by the main capsule layer in a dynamic routing manner to generate an initial capsule vector.
4. The method according to claim 2, characterized in that: Before performing feature processing on the initial capsule vector through the digital capsule layer to generate a target number of result capsule vectors, the method further includes: Determining an execution chamber based on the target process information through the target network model, wherein the execution chamber is a chamber in a process equipment that executes the process recipe to be determined; The target network model is used to determine the number of process parameters corresponding to the execution chamber as the target number based on the correspondence between each chamber in the process equipment and the number of process parameters.
5. The method according to claim 1, characterized in that The target network model also includes: a long short-term memory network layer constructed based on a long short-term memory network, wherein the input of the long short-term memory network layer is the output of the capsule network layer, and the output of the long short-term memory network layer is the output of the target network model.
6. The method according to claim 5, characterized in that Predicting the process parameters required to achieve the process result through the target network model includes: Processing the target process information through the capsule network layer to generate a target number of result capsule vectors; wherein each of the result capsule vectors corresponds to a process parameter, and the result capsule vector includes a plurality of predicted values of the corresponding process parameter and a predicted probability of each of the predicted values; The result capsule vector is processed through the forget gate, input gate, output gate and output layer of the long short-term memory network layer to obtain the process parameters required to achieve the process result.
7. The method according to claim 1, characterized in that In response to the target process information input by the user, before inputting the target process information into the target network model, the method further includes: Acquire training data; wherein the training data includes: historical process recipes and process results that can be achieved by the historical process recipes; The training data is used to perform model training on the network model constructed by the capsule network, and the trained network model is determined as the target network model.
8. The method according to claim 7, characterized in that In the process of training the network model constructed by the capsule network using the training data, the sum of the interval loss and the reconstruction loss is used as the model loss.
9. A device for determining process parameters, characterized in that: The device comprises: A response module, for responding to target process information input by a user, inputting the target process information into a target network model; wherein the target process information includes a process result that can be achieved by a process equipment performing process production according to a process recipe to be determined; the process equipment includes a semiconductor process equipment; A model processing module, used to predict the process parameters required to achieve the process result through the target network model; wherein the target network model includes: a capsule network layer constructed based on a capsule network; A determination module, used for determining the prediction result output by the target network model as the process parameter of the process recipe to be determined, so that the process equipment uses the determined process recipe to be determined to perform process production; The target network model further includes: a long short-term memory network layer constructed based on a long short-term memory network, and an attention module constructed based on an attention mechanism; wherein the input of the attention module is the target process information, the input of the long short-term memory network layer is the fusion result of the output of the attention module and the output of the capsule network layer, and the output of the long short-term memory network layer is the output of the target network model; Model processing module, specifically used for: The target process information is processed through the capsule network layer to generate a target number of result capsule vectors; wherein each result capsule vector corresponds to a process parameter, and the result capsule vector includes multiple predicted values of the corresponding process parameter and the predicted probability of each predicted value; The target process information is processed through the attention module to generate a context vector corresponding to each result capsule vector; The fusion result of the result capsule vector and the context vector is processed through the forget gate, input gate, output gate and output layer of the long short-term memory network layer to obtain the process parameters required to achieve the process results.
10. A semiconductor process equipment, characterized in that: It includes: a host computer and a slave computer connected to the host computer; The upper computer is used to execute the method for determining the process parameters according to any one of claims 1 to 8 to determine the process parameters of the process recipe to be determined, and send the process parameters of the process recipe to be determined to the lower computer; The lower computer is used to receive the process parameters of the process recipe to be determined, and perform process production according to the process parameters of the process recipe to be determined.
11. An electronic device, characterized in that: include: Processor and memory; Wherein, the memory is connected to the processor, and the memory is used to store a computer program; The processor is used to implement the method for determining the process parameters according to any one of claims 1 to 8 by running the computer program stored in the memory.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for determining the process parameters according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Capsule neural network integrated with multi-scale feature attention and text classification method
CN111897957A
Financial text recognition method and device, computer equipment and storage medium
CN115019331A