A virtual measurement method, product, device and medium
By constructing a virtual measurement model, using mapped features and interpretable dot product attention mechanism combined with regression sampling convolutional interaction model, the problem of mismatch between feature selection and regression prediction is solved, and efficient and interpretable material removal rate prediction is achieved.
Patent Information
- Application Number
- CN202411721863.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-11-28
AI Technical Summary
In the virtual measurement method based on feature extraction and feature selection, the two steps of feature selection and regression prediction use different algorithms, resulting in the selected features not exactly match the regression model, introducing uncertainty.
By constructing a virtual measurement model, the model maps polishing process variables into mapped features, and uses an interpretable dot product attention mechanism and a regression sampling convolutional interaction model, combining feature mapping and numerical prediction, to solve the mismatch problem of feature selection and regression prediction.
The combination of feature screening and numerical prediction is realized, which improves the credibility and interpretability of the model, quantifies the relative importance of the input data, and solves the uncertainty problem.
Smart Images

Figure CN119203799B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of measurement technologies, and particularly to a virtual measurement method, product, device, and medium. Background Art
[0002] In recent years, the development and competition of Chemical Mechanical Polishing (CMP) technology have become the focus of attention in the scientific research community and the industrial community. As a key step in many precision manufacturing processes, CMP technology is widely used in production and life. In the CMP process, the Material Removal Rate (MRR) is an important indicator for measuring the process quality. MRR represents the amount of material removed from the material surface per unit time and is a key parameter for evaluating the polishing effect. However, it is challenging to measure MRR accurately in real time because direct measurement may damage the material surface. Therefore, it is feasible to use Virtual Measurement (VM) technology to indirectly estimate MRR.
[0003] In the application field of CMP technology, there are VM methods that use machine learning and statistical analysis techniques to establish a prediction model for MRR. By analyzing a large amount of historical data to extract features related to MRR and using a regression algorithm to predict MRR. However, in the VM method based on feature extraction and feature selection, different algorithms are used in the two steps of feature selection and regression prediction, which may lead to incomplete matching between the selected features and the regression model, introducing uncertainty.
[0004] In view of the above problems, how to solve the problem that in the current VM method based on feature extraction and feature selection, different algorithms are used in the two steps of feature selection and regression prediction, resulting in incomplete matching between the selected features and the regression model and introducing uncertainty is an urgent problem for those skilled in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a virtual measurement method, product, device, and medium to solve the problem that in the current VM method based on feature extraction and feature selection, different algorithms are used in the two steps of feature selection and regression prediction, resulting in incomplete matching between the selected features and the regression model and introducing uncertainty.
[0006] To solve the above technical problems, the present invention provides a virtual measurement method, including:
[0007] Obtain each target polishing process variable that affects the polishing quality of the target material;
[0008] Input each of the target polishing process variables into a virtual measurement model to output the target material removal rate corresponding to the target material;
[0009] Among them, the construction process of the virtual measurement model includes: mapping each polishing process variable into a mapping feature; the mapping feature retains the mathematical meaning corresponding to the polishing process variable; determining the output result of the dot product attention structure according to the mapping feature and the interpretable dot product attention mechanism; constructing a regression sampling convolution interaction model according to the output result of the dot product attention structure.
[0010] On the one hand, mapping each of the polishing process variables into the mapping feature includes:
[0011] Mapping each of the polishing process variables in the training dataset into the feature space;
[0012] Determining the positions corresponding to each of the polishing process variables in the feature space to generate the mapping feature.
[0013] On the other hand, determining the output result of the dot product attention structure according to the mapping feature and the interpretable dot product attention mechanism includes:
[0014] Obtaining the query, key, and value of the interpretable dot product attention mechanism according to the mapping feature; wherein, the query, key, and value of the interpretable dot product attention mechanism retain the mathematical meaning of the polishing process variable;
[0015] Obtaining the interpretable attention weight distribution matrix and the output result of the dot product attention structure according to the query, key, and value of the interpretable dot product attention mechanism.
[0016] On the other hand, constructing the regression sampling convolution interaction model according to the output result of the dot product attention structure includes:
[0017] Constructing the regression sampling convolution interaction model based on the binary tree structure according to the output result of the dot product attention structure; wherein, each node of the binary tree structure contains complete input features;
[0018] Determining a single value output by the regression sampling convolution interaction model to complete the construction of the virtual measurement model.
[0019] On the other hand, determining the positions corresponding to each of the polishing process variables in the feature space to generate the mapping feature includes:
[0020] Extracting primary features and their corresponding feature dimensions based on each of the polishing process variables;
[0021] Obtaining the positions of each of the primary features in the corresponding feature dimensions according to each of the primary features and their corresponding feature dimensions;
[0022] Sum each of the primary features with its position in the corresponding feature dimension to obtain the mapped feature of each of the polishing process variables.
[0023] On the other hand, obtaining the query, key, and value of the interpretable dot-product attention mechanism according to the mapped feature includes:
[0024] Set three one-dimensional convolutional filters; the kernel width and stride of the one-dimensional convolutional filter are both 1;
[0025] Input the mapped feature into the three one-dimensional convolutional filters respectively to obtain the query, key, and value of the interpretable dot-product attention mechanism respectively.
[0026] On the other hand, obtaining the interpretable attention weight distribution matrix and the output result of the dot-product attention structure according to the query, key, and value of the interpretable dot-product attention mechanism includes:
[0027] Determine the product of the query and key of the interpretable dot-product attention mechanism to obtain the interpretable attention weight distribution matrix;
[0028] Determine the product of the interpretable attention weight distribution matrix and the value of the interpretable dot-product attention mechanism to obtain the output result of the dot-product attention structure.
[0029] On the other hand, constructing the regression sampling convolutional interaction model based on the binary tree structure according to the output result of the dot-product attention structure includes:
[0030] Copy the output result of the dot-product attention structure to obtain a first copy and a second copy as the input of the first-layer binary tree node;
[0031] Determine the weighted feature output result of the current-layer binary tree node according to the first copy and the second copy;
[0032] Copy the weighted feature output result to obtain a first copy and a second copy as the input of the next-layer binary tree node;
[0033] Return to the step of determining the weighted feature output result of the current-layer binary tree node according to the first copy and the second copy until the binary tree structure reaches the preset number of layers to obtain the regression sampling convolutional interaction model.
[0034] On the other hand, determining the weighted feature output result of the current-layer binary tree node according to the first copy and the second copy includes:
[0035] Perform convolutional operations on the first copy and the second copy respectively to generate a convolutional result corresponding to the first copy and a convolutional result corresponding to the second copy;
[0036] Perform exponential operations on the convolution results corresponding to the first copy and the convolution results corresponding to the second copy respectively to generate the exponential result corresponding to the first copy and the exponential result corresponding to the second copy;
[0037] Generate the weighted feature output result of the binary tree node of the current layer based on the first copy and its corresponding exponential result, and the second copy and its corresponding exponential result.
[0038] On the other hand, perform convolution operations on the first copy and the second copy respectively to generate the convolution result corresponding to the first copy and the convolution result corresponding to the second copy, including:
[0039] Perform a one-dimensional convolution operation on the first copy through the first one-dimensional convolution module to obtain the first convolution result;
[0040] Perform a one-dimensional convolution operation on the second copy through the second one-dimensional convolution module to obtain the second convolution result;
[0041] Correspondingly, perform exponential operations on the convolution results corresponding to the first copy and the convolution results corresponding to the second copy respectively to generate the exponential result corresponding to the first copy and the exponential result corresponding to the second copy, including:
[0042] Apply the natural exponential function to the first convolution result to obtain the first exponential result;
[0043] Apply the natural exponential function to the second convolution result to obtain the second exponential result.
[0044] On the other hand, generate the weighted feature output result of the binary tree node of the current layer based on the first copy and its corresponding exponential result, and the second copy and its corresponding exponential result, including:
[0045] Perform a Hadamard product operation on the first copy and the second exponential result to obtain the first process quantity;
[0046] Perform a Hadamard product operation on the second copy and the first exponential result to obtain the second process quantity;
[0047] Perform a one-dimensional convolution operation on the second process quantity through the fourth one-dimensional convolution module to obtain the fourth convolution result, and add the first process quantity and the fourth convolution result to obtain the first weighted feature output result;
[0048] Perform a one-dimensional convolution operation on the first process quantity through the third one-dimensional convolution module to obtain the third convolution result, and subtract the third convolution result from the second process quantity to obtain the second weighted feature output result.
[0049] On the other hand, the structures of the first one-dimensional convolution module, the second one-dimensional convolution module, the third one-dimensional convolution module, and the fourth one-dimensional convolution module are the same, and they are successively composed of a hyperbolic tangent activation function, a one-dimensional convolution layer, a stochastic loss layer, a leaky rectified linear unit activation function, and a one-dimensional convolution layer.
[0050] On the other hand, determining a single value output by the regression sampling convolution interaction model to complete the construction of the virtual measurement model includes:
[0051] Concatenating multiple feature matrices corresponding to the regression sampling convolution interaction model to obtain a concatenated matrix; wherein, the sizes of the feature matrices corresponding to the regression sampling convolution interaction model are equal to the size of the mapped feature;
[0052] Stacking the concatenated matrix through fully connected layers and then calculating the matrix average value to obtain the single value, thus completing the construction of the virtual measurement model.
[0053] On the other hand, the training process of the virtual measurement model includes:
[0054] Obtaining a training data set composed of multiple polishing process variables and their corresponding material removal rates;
[0055] Training the virtual measurement model according to the data in the training data set until the virtual measurement model converges to obtain the virtual measurement model.
[0056] On the other hand, after training the virtual measurement model according to the data in the training data set until the virtual measurement model converges, it further includes:
[0057] Obtaining a test data set composed of multiple polishing process variables and their corresponding material removal rates;
[0058] Obtaining an interpretable attention weight distribution matrix corresponding to the current virtual measurement model;
[0059] Testing the current virtual measurement model according to the data in the test data set to obtain the virtual measurement accuracy;
[0060] Judging whether the virtual measurement accuracy reaches a preset accuracy value;
[0061] If not, screening data in the training data set according to the interpretable attention weight distribution matrix to obtain a new training data set, and returning to the step of training the virtual measurement model according to the data in the training data set;
[0062] If so, taking the current virtual measurement model as the final virtual measurement model.
[0063] On the other hand, after inputting each of the target polishing process variables into the virtual measurement model to output the target material removal rate corresponding to the target material, the following steps are further included:
[0064] Storing the target polishing process variables and the corresponding target material removal rates in a training data set to update the training data set;
[0065] Retraining the virtual measurement model based on the updated training data set.
[0066] To solve the above technical problems, the present invention also provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above virtual measurement method are implemented.
[0067] To solve the above technical problems, the present invention also provides a virtual measurement device, including:
[0068] A memory for storing a computer program;
[0069] A processor for implementing the steps of the above virtual measurement method when executing the computer program.
[0070] To solve the above technical problems, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above virtual measurement method are implemented.
[0071] The virtual measurement method provided by the present invention obtains each target polishing process variable that affects the polishing quality of the target material; inputs each target polishing process variable into a virtual measurement model to output the target material removal rate corresponding to the target material; wherein, the construction process of the virtual measurement model includes: mapping each polishing process variable into a mapping feature; the mapping feature retains the mathematical meaning of the corresponding polishing process variable; determining the output result of the dot product attention structure according to the mapping feature and the interpretable dot product attention mechanism; constructing a regression sampling convolution interaction model according to the output result of the dot product attention structure.
[0072] The beneficial effects of the present invention are as follows: a virtual measurement model for material removal rate prediction is provided, which is specifically constructed and trained by successively performing feature mapping, feature output based on an interpretable dot product attention mechanism, and numerical prediction based on a regression sampling convolution interaction mechanism according to multiple polishing process variables and their corresponding material removal rates; during the process of constructing the virtual measurement model, the mathematical meaning of the polishing process variables is retained in the mapped features, and at the same time, the features output by the interpretable dot product attention mechanism are used as the input of the regression sampling convolution interaction mechanism, so that both the mathematical meanings of all input features are retained, and feature selection and regression prediction are combined. Only one model, the virtual measurement model, can complete the two steps of feature screening and numerical prediction, solving the potential mismatch problem between the current feature selection algorithm and the regression numerical prediction algorithm. It can play the advantage of model interpretability while regressing and predicting the virtual measurement target value, quantitatively measure the relative importance of the input data, greatly improve the credibility of the model, and facilitate real-time and effective supervision of the model in industrial production.
[0073] On the other hand, the present invention specifically maps each polishing process variable in the training dataset to a mapped feature; obtains the query, key, and value of the interpretable dot product attention mechanism according to the mapped feature; obtains the interpretable attention weight distribution matrix and the dot product attention structure output result according to the query, key, and value of the interpretable dot product attention mechanism; constructs a regression sampling convolution interaction model based on a binary tree structure according to the dot product attention structure output result; determines a single value output by the regression sampling convolution interaction model, thereby realizing the construction of the virtual measurement model. Obtain a training dataset composed of multiple polishing process variables and their corresponding material removal rates; train the virtual measurement model according to the data in the training dataset until the virtual measurement model converges, realizing the training of the virtual measurement model. If the virtual measurement accuracy does not reach the preset accuracy value, then screen the data in the training dataset according to the interpretable attention weight distribution matrix to obtain a new training dataset, and return to the step of training the virtual measurement model according to the data in the training dataset, so as to find more relevant new training data through data screening and train the model again based on the new training data to improve the model accuracy.
[0074] In addition, the present invention also provides a computer program product, a virtual measurement device, and a medium, and the effects are the same as above. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0076] Figure 1 Flow chart of a virtual measurement method provided by an embodiment of the present invention;
[0077] Figure 2 Schematic diagram of the framework of a virtual measurement model provided by an embodiment of the present invention;
[0078] Figure 3 Schematic diagram of the framework of an interpretable dot product attention mechanism provided by an embodiment of the present invention;
[0079] Figure 4 Schematic diagram of the structure of a sampling convolutional interaction block provided by an embodiment of the present invention;
[0080] Figure 5 Schematic diagram of a virtual measurement device provided by an embodiment of the present invention;
[0081] Figure 6 Schematic diagram of a virtual measurement device provided by an embodiment of the present invention. Detailed implementation manners
[0082] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0083] The core of the present invention is to provide a virtual measurement method, product, device and medium to solve the problem that in the current VM method based on feature extraction and feature selection, different algorithms are used in the two steps of feature selection and regression prediction, resulting in incomplete matching between the selected features and the regression model and introducing uncertainty.
[0084] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0085] Currently, in the application field of CMP technology, there is a VM method that uses machine learning and statistical analysis techniques to establish a prediction model for MRR. By analyzing a large amount of historical data to extract features related to MRR and using a regression algorithm to predict MRR. However, in the VM method based on feature extraction and feature selection, different algorithms are used in the two steps of feature selection and regression prediction, which may lead to incomplete matching between the selected features and the regression model and introduce uncertainty. Therefore, to solve the above problems, the present invention provides a virtual measurement method.
[0086] Figure 1 Flow chart of a virtual measurement method provided by an embodiment of the present invention. AsFigure 1 As shown, the method includes:
[0087] S10: Obtain each target polishing process variable that affects the polishing quality of the target material.
[0088] Specifically, in order to predict the MRR of the target material, it is first necessary to obtain each target polishing process variable that affects the polishing quality of the target material.
[0089] It should be noted that the target material in this embodiment is the material for chemical mechanical polishing, and the specific type of the target material in this embodiment is not limited. In addition, the polishing process variable refers to various parameters and conditions that affect the polishing effect and quality during the polishing process, and the target polishing process variable changes with the different target materials; the polishing process variables generally include pressure, speed, flow rate, swing parameters, and other parameters. The target polishing process variables involved in the target material in this embodiment are not limited and are determined according to the specific implementation situation.
[0090] S11: Input each target polishing process variable into the virtual measurement model to output the target material removal rate corresponding to the target material.
[0091] Among them, the construction process of the virtual measurement model includes: mapping each polishing process variable into a mapping feature; the mapping feature retains the mathematical meaning of the corresponding polishing process variable; determining the output result of the dot product attention structure according to the mapping feature and the interpretable dot product attention mechanism; constructing a regression sampling convolution interaction model according to the output result of the dot product attention structure.
[0092] Furthermore, inputting each target polishing process variable into the virtual measurement model, thereby predicting and outputting the target material removal rate corresponding to the target material through the virtual measurement model, realizes the prediction of the MRR of the target material.
[0093] It should be noted that the virtual measurement model in this embodiment is an interpretable sample convolution and interaction former (SCI-Former) model. It can numerically measure the importance of all input features while outputting with high precision. Figure 2 This is a schematic diagram of the framework of the virtual measurement model provided by the embodiment of the present invention. As Figure 2 shown, the virtual measurement model includes two parts. One part is the interpretable dot product attention module as the encoder, and the other part is the regression sampling convolution and interaction network module as the decoder. The interpretable dot product attention measures the importance of the embedded input features and outputs weighted features according to the attention weights. Compared with the traditional dot product attention structure, the information flow of the interpretable dot product attention provided by the present invention is more directly interpretable. And the regression sampling convolution and interaction network module is improved based on the sample convolution and interaction network model and is an efficient regressor for MRR prediction based on weighted features.
[0094] In addition, in this embodiment, each polishing process variable is specifically mapped to a mapping feature; the mapping feature retains the mathematical meaning of the corresponding polishing process variable; the dot product attention structure output result is determined according to the mapping feature and the interpretable dot product attention mechanism; a regression sampling convolution interaction model is constructed according to the dot product attention structure output result, realizing the construction of the virtual measurement model. Therefore, the virtual measurement model can complete two steps of feature screening and numerical prediction, solving the potential mismatch problem between the feature screening algorithm and the numerical prediction algorithm. At the same time, the virtual measurement model is also interpretable, and can numerically measure the importance of all input features while outputting with high precision, greatly improving the credibility of the model. In this embodiment, the specific construction method and training process of the virtual measurement model are not limited and are determined according to specific implementation situations.
[0095] In this embodiment, it is constructed and trained by successively performing feature mapping, feature output based on the interpretable dot product attention mechanism, and numerical prediction based on the regression sampling convolution interaction mechanism according to multiple polishing process variables and their corresponding material removal rates; in the process of constructing the virtual measurement model, the mathematical meaning of the polishing process variables is retained in the mapping feature, and at the same time, the features output by the interpretable dot product attention mechanism are used as the input of the regression sampling convolution interaction mechanism, so that both the mathematical meanings of all input features are retained, and feature selection and regression prediction are combined. Only using this one virtual measurement model can complete two steps of feature screening and numerical prediction, solving the potential mismatch problem between the current feature selection algorithm and the regression numerical prediction algorithm, and can exert the advantage of model interpretability while regressing and predicting the virtual measurement target value, quantitatively measuring the relative importance of the input data, greatly improving the credibility of the model, and facilitating real-time and effective supervision of the model in industrial production.
[0096] The construction of the virtual measurement model in the present invention is divided into three stages:
[0097] First, the first stage is the mapping layer, which maps the data to the feature space while retaining the independence of the input data. Therefore, on the basis of the above embodiment, in some embodiments, each polishing process variable is mapped to a mapping feature, specifically including:
[0098] S12: Map each polishing process variable in the training data set to the feature space;
[0099] S13: Determine the positions corresponding to each polishing process variable in the feature space to generate mapping features.
[0100] In this embodiment, specifically mapping each polishing process variable in the training dataset to the feature space can preserve the mathematical meaning of each polishing process variable. It can be understood that the training set data consists of multiple polishing process variables and their corresponding material removal rates. Further determine the positions corresponding to each polishing process variable in the feature space, so as to mark each polishing process variable, and finally generate mapped features. It should be noted that in this embodiment, there is no limitation on the specific process of determining the positions corresponding to each polishing process variable in the feature space.
[0101] Furthermore, the second stage is the interpretable dot product attention mechanism, which extracts necessary feature information, shields invalid feature information, outputs high-order features, and simultaneously accurately measures the importance of each feature. Therefore, based on the above embodiments, in some embodiments, determining the output result of the dot product attention structure according to the mapped features and the interpretable dot product attention mechanism includes:
[0102] S14: Obtain the query, key, and value of the interpretable dot product attention mechanism according to the mapped features;
[0103] Among them, the query, key, and value of the interpretable dot product attention mechanism retain the mathematical meaning of the polishing process variables;
[0104] S15: Obtain the interpretable attention weight distribution matrix and the output result of the dot product attention structure according to the query, key, and value of the interpretable dot product attention mechanism.
[0105] In this embodiment, specifically obtain the query, key, and value of the interpretable dot product attention mechanism according to the mapped features, and obtain the interpretable attention weight distribution matrix and the output result of the dot product attention structure according to the query, key, and value of the interpretable dot product attention mechanism.
[0106] It should be noted that the query, key, and value of the interpretable dot product attention mechanism in this embodiment retain the mathematical meaning of the polishing process variables. At the same time, the interpretable dot product attention mechanism no longer adds the output of the attention mechanism to the original input features at the element level. Discarding the residual connection can prevent the original features without attention weight adjustment from entering the deep network, ensuring the importance of the variables measured by the attention weights. In this embodiment, there is no limitation on the specific process of obtaining the query, key, and value, the interpretable attention weight distribution matrix, and the output result of the dot product attention structure, which depends on the specific implementation situation.
[0107] Finally, the third stage is the regression sampling convolution interaction, which consists of convolutional layers and makes predictions based on the high-order features of the previous stage. Therefore, based on the above embodiments, in some embodiments, constructing a regression sampling convolution interaction model according to the output result of the dot product attention structure includes:
[0108] S16: Construct a regression sampling convolutional interaction model based on a binary tree structure according to the output result of the dot product attention structure;
[0109] Among them, each node of the binary tree structure contains complete input features;
[0110] S17: Determine the single value output by the regression sampling convolutional interaction model to complete the construction of the virtual measurement model.
[0111] In this embodiment, specifically construct a regression sampling convolutional interaction model based on the binary tree structure according to the output result of the dot product attention structure; determine the single value output by the regression sampling convolutional interaction model to complete the construction of the virtual measurement model.
[0112] It should be noted that in this embodiment, there is no limitation on the specific construction process of constructing a regression sampling convolutional interaction model based on the output result of the dot product attention structure, nor on the determination method of the single value output by the regression sampling convolutional interaction model, which depends on the specific implementation situation.
[0113] The following combines specific embodiments to elaborate on the construction process of the virtual measurement model in detail:
[0114] (1) Feature mapping layer;
[0115] Based on the above embodiments, in some embodiments, determine the positions corresponding to each polishing process variable in the feature space to generate mapped features, including:
[0116] S131: Extract primary features and their corresponding feature dimensions based on each polishing process variable;
[0117] S132: According to each primary feature and its corresponding feature dimension, obtain the positions of each primary feature in the corresponding feature dimension;
[0118] S133: Add each primary feature to its position in the corresponding feature dimension to obtain the mapped features of each polishing process variable.
[0119] Specifically, the feature mapping layer consists of multiple one-dimensional convolutional layers, the kernel width of these convolutional layers is 1, and the stride is also 1. Based on each polishing process variable, the feature mapping layer extracts primary features and their corresponding feature dimensions . A one-dimensional convolutional layer that only calculates a single input variable can effectively retain the mathematical meaning of the input variable.
[0120] Furthermore, according to each primary feature and its corresponding feature dimension, obtain the positions of each primary feature in the corresponding feature dimension. Specifically, assume primary features ; The positions of each primary feature on the corresponding feature dimension are specifically as follows:
[0121] ;
[0122] ;
[0123] Among them, and respectively represent the even position and the odd position of the primary feature on the feature dimension. .
[0124] Finally, add each primary feature to its position on the corresponding feature dimension to obtain the mapped feature of each polishing process variable, specifically as follows:
[0125] ;
[0126] Among them, is the mapped feature, is the primary feature, is the sum of the primary feature and its position on the corresponding feature dimension.
[0127] (2) Interpretable dot product attention mechanism;
[0128] Based on the above embodiments, in some embodiments, obtaining the query, key, and value of the interpretable dot product attention mechanism according to the mapped feature includes:
[0129] S141: Set three one-dimensional convolutional filters; the kernel width and stride of the one-dimensional convolutional filters are both 1;
[0130] S142: Input the mapped feature into the three one-dimensional convolutional filters respectively to obtain the query, key, and value of the interpretable dot product attention mechanism respectively.
[0131] It is understandable that the attention values of the Transformer structure are positively correlated with the importance of variables. The attention mechanism of the neural network multiplies the non-zero weights with the input, and the sum of the results is 1. The attention value of a variable numerically represents the relative importance of that variable. However, for the sake of efficiency and accuracy, the Transformer structure breaks the mathematical meaning of the input features, making it impossible to measure the importance of the original input features. The classical dot-product attention mechanism also cannot be used as a feature filter. There are two specific reasons: First, the queries, keys, and values in the traditional Transformer structure are obtained through fully connected layers, which destroys the mathematical meaning of the features. The queries, keys, and values obtained through fully connected layers are a mixture of all input variables, and it is very cumbersome to trace back to the original input variables they represent. Second, the traditional Transformer structure adds its original input to the output of the attention mechanism at the element level. This addition preserves the original data features, resulting in the original features continuing to participate in the subsequent calculation process without being adjusted by the attention weights.
[0132] Figure 3 This is a schematic diagram of the framework of the interpretable dot-product attention mechanism provided by the embodiments of the present invention. Aiming at the shortcomings of the traditional Transformer structure, such as Figure 3 As shown, this embodiment provides an interpretable multi-head dot-product attention mechanism. Specifically, first, three one-dimensional convolutional filters are set; the kernel width and stride of the one-dimensional convolutional filter are both 1. The mapped features are respectively input into the three one-dimensional convolutional filters to obtain the query, key, and value of the interpretable dot-product attention mechanism respectively.
[0133] Furthermore, after obtaining the query, key, and value of the interpretable dot-product attention mechanism, according to the query, key, and value of the interpretable dot-product attention mechanism, an interpretable attention weight distribution matrix and the output result of the dot-product attention structure are obtained, which specifically includes:
[0134] S151: Determine the product of the query and the key of the interpretable dot-product attention mechanism to obtain the interpretable attention weight distribution matrix. The formula is as follows:
[0135] ;
[0136] where, is the interpretable attention weight distribution matrix, is the query, is the key.
[0137] S152: Determine the product of the interpretable attention weight distribution matrix and the value of the interpretable dot-product attention mechanism to obtain the output result of the dot-product attention structure. The formula is as follows:
[0138] ;
[0139] Among them, is the output result of the dot - product attention structure, is the interpretable attention weight distribution matrix, is the column vector of is the value, is the row vector of
[0140] Therefore, compared with the classical attention mechanism, the main improvement of the interpretable multi - head dot - product attention mechanism provided by this solution is: queries, keys, and values are obtained through a one - dimensional convolutional filter. Such a filter ensures that the mathematical meanings of the obtained queries, keys, and values are clear and directly inherited from the input variables. At the same time, the residual connection structure is discarded. The interpretable dot - product attention mechanism no longer adds the output of the attention mechanism to the original input features at the element level. Discarding the residual connection can prevent the original features without attention - weight adjustment from entering the deep network, ensuring the importance of the variables measured by the attention weights.
[0141] (3) Regression sampling convolutional interaction model;
[0142] Based on the above - mentioned embodiments, in some embodiments, a regression sampling convolutional interaction model based on a binary - tree structure is constructed according to the output result of the dot - product attention structure, including:
[0143] S161: Duplicate the output result of the dot - product attention structure to obtain a first copy and a second copy as the inputs of the first - layer binary - tree nodes;
[0144] S162: Determine the weighted feature output result of the current - layer binary - tree nodes according to the first copy and the second copy;
[0145] S163: Duplicate the weighted feature output result to obtain a first copy and a second copy as the inputs of the next - layer binary - tree nodes; return to step S162 until the binary - tree structure reaches the preset number of layers to obtain the regression sampling convolutional interaction model.
[0146] Figure 4 is the structural schematic diagram of the sampling convolutional interaction block provided by the embodiments of the present invention. Since the original sampling convolutional interaction block (SCINet Block) downsamples the input sequence according to the features of the time series to retain basic information, this operation is not recommended when processing non - time - series inputs. Therefore, in this embodiment, as Figure 4As shown, specifically, the output result of the dot product attention structure is copied to obtain the first copy and the second copy as the input of the nodes of the first layer binary tree. Further, the weighted feature output result of the current layer binary tree node is determined according to the first copy and the second copy, and the weighted feature output result is copied to obtain the first copy and the second copy as the input of the next layer binary tree node; return to the step of determining the weighted feature output result of the current layer binary tree node according to the first copy and the second copy, until the binary tree structure reaches the preset number of layers, and a regression sampling convolution interaction model is obtained.
[0147] It should be noted that each binary tree node is a sampling convolution interaction block. The regression sampling convolution interaction model constructed by the sampling convolution interaction block as nodes has a total of L layers of binary tree structures, where L is a positive integer, and the specific value of L is not limited in this embodiment. Compared with the general convolution-based neural network model, the most significant advantage of the regression sampling convolution interaction model is that it has a rich number of independent convolution layers, which improves the ability of the model to extract features. The binary tree with sampling convolution interaction blocks as nodes is the core part of the regression sampling convolution interaction model. Thanks to the replication operation of the sampling convolution interaction block, each sampling convolution interaction block node contains a complete set of input features.
[0148] In addition, to enhance the nonlinear fitting ability of the model and enhance the interaction between data, in some embodiments, the sampling convolution interaction block includes four groups of one-dimensional convolution modules and two exponential nonlinear modules. Among them, is the first one-dimensional convolution module, is the second one-dimensional convolution module, is the third one-dimensional convolution module, is the fourth one-dimensional convolution module. It should be noted that the structures of the four groups of one-dimensional convolution modules are the same, and are successively composed of a hyperbolic tangent (Tanh) activation function, a one-dimensional convolution layer, a random loss layer, a leaky rectified linear unit (LeakyRelu) activation function, and a one-dimensional convolution layer. Based on this, the process of determining the weighted feature output result of the current layer binary tree node according to the first copy and the second copy specifically includes:
[0149] S164: Perform convolution operations on the first copy and the second copy respectively to generate the convolution result corresponding to the first copy and the convolution result corresponding to the second copy;
[0150] S165: Perform exponential operations on the convolution result corresponding to the first copy and the convolution result corresponding to the second copy respectively to generate the exponential result corresponding to the first copy and the exponential result corresponding to the second copy;
[0151] S166: Generate the weighted feature output result of the current-layer binary tree node based on the first copy and its corresponding exponential result, and the second copy and its corresponding exponential result.
[0152] To determine the weighted feature output result, specifically perform convolution operations on the first copy and the second copy respectively to generate the convolution result corresponding to the first copy and the convolution result corresponding to the second copy; further perform exponential operations on the convolution result corresponding to the first copy and the convolution result corresponding to the second copy respectively to generate the exponential result corresponding to the first copy and the exponential result corresponding to the second copy. Finally, generate the weighted feature output result of the current-layer binary tree node based on the first copy and its corresponding exponential result, and the second copy and its corresponding exponential result. The following specifically explains the calculation process with specific formulas:
[0153] Perform convolution operations on the first copy and the second copy respectively to generate the convolution result corresponding to the first copy and the convolution result corresponding to the second copy. Specifically, perform a one-dimensional convolution operation on the first copy through the first one-dimensional convolution module to obtain the first convolution result ; perform a one-dimensional convolution operation on the second copy through the second one-dimensional convolution module to obtain the second convolution result ;
[0154] Correspondingly, perform exponential operations on the convolution result corresponding to the first copy and the convolution result corresponding to the second copy respectively to generate the exponential result corresponding to the first copy and the exponential result corresponding to the second copy. Specifically, apply the natural exponential function to the first convolution result to obtain the first exponential result ; apply the natural exponential function to the second convolution result to obtain the second exponential result .
[0155] Furthermore, generate the weighted feature output result of the current-layer binary tree node based on the first copy and its corresponding exponential result, and the second copy and its corresponding exponential result, including:
[0156] S167: Perform a Hadamard product operation on the first copy and the second exponential result to obtain the first process quantity; the formula is as follows:
[0157] ;
[0158] where is the first process quantity, is the first copy, is the second copy, is the second convolution result, is the second exponential result, is the Hadamard product.
[0159] S168: Perform a Hadamard product operation on the second copy and the first exponential result to obtain a second process quantity;
[0160] ;
[0161] wherein, is the second process quantity, is the first copy, is the second copy, is the first convolution result, is the first exponential result, is the Hadamard product.
[0162] S169: Perform a one-dimensional convolution operation on the second process quantity through a fourth one-dimensional convolution module to obtain a fourth convolution result, and add the first process quantity and the fourth convolution result to obtain a first weighted feature output result; The formula is as follows:
[0163] ;
[0164] wherein, is the first weighted feature output result, is the fourth convolution result.
[0165] S170: Perform a one-dimensional convolution operation on the first process quantity through a third one-dimensional convolution module to obtain a third convolution result, and subtract the third convolution result from the second process quantity to obtain a second weighted feature output result. The formula is as follows:
[0166] ;
[0167] wherein, is the second weighted feature output result, is the third convolution result.
[0168] In summary, the calculation of the weighted feature output result is realized.
[0169] Based on the above embodiments, in some embodiments, determining a single value output by the regression sampling convolution interaction model to complete the construction of the virtual measurement model includes:
[0170] S171: Concatenate multiple feature matrices corresponding to the regression sampling convolution interaction model to obtain a concatenated matrix; wherein, the sizes of the feature matrices corresponding to the regression sampling convolution interaction model are equal to the size of the mapped features;
[0171] S172: Stack the concatenated matrix through fully connected layers and then calculate the matrix average value to obtain a single value, thus completing the construction of the virtual measurement model.
[0172] Finally, a regression sampling convolutional interaction model based on an L-layer binary tree was generated, and 2 feature matrices of the same size as the mapping features were obtained. Further, multiple feature matrices corresponding to the regression sampling convolutional interaction model were concatenated to obtain a concatenated matrix. The concatenated matrix passed through the stacking of two fully connected layers to obtain a matrix, and the average value of this matrix was calculated. The mean value was taken as the single value of the final output of the virtual measurement model. The average calculation enhanced the robustness and generalization ability of the model. In this way, the construction of the virtual measurement model was completed. L feature matrices of the same size as the mapping features were obtained. Further, multiple feature matrices corresponding to the regression sampling convolutional interaction model were concatenated to obtain a concatenated matrix. The concatenated matrix passed through the stacking of two fully connected layers to obtain a matrix, and the average value of this matrix was calculated. The mean value was taken as the single value of the final output of the virtual measurement model. The average calculation enhanced the robustness and generalization ability of the model. In this way, the construction of the virtual measurement model was completed.
[0173] Based on the above embodiments, in some embodiments, the training process of the virtual measurement model includes:
[0174] S18: Obtain a training data set composed of multiple polishing process variables and their corresponding material removal rates;
[0175] S19: Train the virtual measurement model according to the data in the training data set until the virtual measurement model converges to obtain the virtual measurement model.
[0176] To implement the training of the virtual measurement model, specifically obtain a training data set composed of multiple polishing process variables and their corresponding material removal rates, and train the virtual measurement model according to the data in the training data set until the virtual measurement model converges to obtain the virtual measurement model.
[0177] To determine whether the virtual measurement model converges, in specific implementation, the change of the loss function can be observed. Specifically, during the training process, by observing the change of the loss function on the training set and the validation set, it can be judged whether the model converges. For example, if both the training loss and the validation loss decrease, it indicates that the network is still learning; if the training loss decreases but the validation loss is stable, it may indicate that the network starts to overfit; if the training loss is stable but the validation loss decreases, it may indicate that there is a problem with the dataset; if both the training loss and the validation loss are stable, it may indicate that the model has converged, or the learning encounters a bottleneck, and the learning rate can be tried to be adjusted smaller; if both the training loss and the validation loss increase, it may indicate that there is a problem with the network structure design or the training parameter setting is improper, and the learning needs to be stopped in time and the code needs to be adjusted. It can also be judged according to the training curve. By plotting the training curve, the convergence situation of the model can be observed more intuitively. For example, underfitting usually shows that both the training loss and the validation loss are relatively high and there is no obvious downward trend, indicating that the model does not fully learn the data features; overfitting usually shows that the training loss is relatively low, but the validation loss starts to increase after a certain point, indicating that the model fits the training data too well and the generalization ability becomes poor; model convergence usually shows that both the training loss and the validation loss decrease to a stable point and the difference between the two is not large. In addition, the change of the gradient can also be observed. Specifically, by calculating the norm of the gradient, it can be judged whether the gradient is close to zero, so as to judge whether the model converges; if the norm of the gradient is very small and the change is not large, it indicates that the model has converged. Finally, the above methods can be used to comprehensively judge whether the virtual measurement model has converged, and corresponding adjustments and optimizations can be made according to needs, so as to ensure the prediction accuracy of the model.
[0178] On the basis of the above embodiments, in order to improve the prediction accuracy of the virtual measurement model, in some embodiments, after training the virtual measurement model according to the data in the training dataset until the virtual measurement model converges, it further includes:
[0179] S20: Obtain a test dataset composed of multiple polishing process variables and their corresponding material removal rates;
[0180] S21: Obtain the interpretable attention weight distribution matrix corresponding to the current virtual measurement model;
[0181] S22: Test the current virtual measurement model according to the data in the test dataset to obtain the virtual measurement accuracy;
[0182] S23: Judge whether the virtual measurement accuracy reaches a preset accuracy value; if not, go to step S24; if so, go to step S25;
[0183] S24: Screen data in the training dataset according to the interpretable attention weight distribution matrix to obtain a new training dataset, and return to step S19;
[0184] S25: Take the current virtual measurement model as the final virtual measurement model.
[0185] Specifically, after the model converges, obtain a test dataset composed of multiple polishing process variables and their corresponding material removal rates, and obtain the interpretable attention weight distribution matrix corresponding to the current virtual measurement model. It can be understood that the interpretable attention weight distribution matrix here is the interpretable attention weight distribution matrix obtained in the above embodiment.
[0186] Further, test the current virtual measurement model according to the data in the test dataset to obtain the virtual measurement accuracy. Determine whether the virtual measurement accuracy reaches the preset accuracy value. It should be noted that in this embodiment, there is no limit to the preset accuracy value, which depends on the specific implementation situation.
[0187] If the virtual measurement accuracy does not reach the preset accuracy value, directly take the current virtual measurement model as the final virtual measurement model. If the virtual measurement accuracy does not reach the preset accuracy value, screen data in the training dataset according to the interpretable attention weight distribution matrix to obtain a new training dataset, and return to the step of training the virtual measurement model according to the data in the training dataset, so as to find more relevant new training data through data screening, and based on the new training data, train the model again, improving the model accuracy.
[0188] Based on the above embodiment, in some embodiments, after inputting each target polishing process variable into the virtual measurement model to output the target material removal rate corresponding to the target material, it further includes:
[0189] S26: Store the target polishing process variable and the corresponding target material removal rate in the training dataset to update the training dataset;
[0190] S27: Retrain the virtual measurement model based on the updated training dataset.
[0191] In specific implementation, in order to maintain the prediction accuracy of the virtual measurement model, after the virtual measurement model finishes predicting the target material removal rate each time, the target polishing process variable and the corresponding target material removal rate can be stored in the training dataset to update the training dataset; further, retrain the virtual measurement model based on the updated training dataset. In this way, the virtual measurement model after iterative training can more accurately predict the target material removal rate corresponding to the target material.
[0192] In the above embodiments, the virtual measurement method has been described in detail. The present invention also provides embodiments corresponding to the virtual measurement device.
[0193] Figure 5 It is a schematic diagram of a virtual measurement device provided by an embodiment of the present invention. As Figure 5 shown, the device includes:
[0194] An acquisition module 10, configured to acquire each target polishing process variable that affects the polishing quality of the target material.
[0195] A prediction module 11, configured to input each target polishing process variable into the virtual measurement model to output the target material removal rate corresponding to the target material.
[0196] Among them, the construction process of the virtual measurement model includes: mapping each polishing process variable into a mapping feature; the mapping feature retains the mathematical meaning of the corresponding polishing process variable; determining the output result of the dot product attention structure according to the mapping feature and the interpretable dot product attention mechanism; constructing a regression sampling convolution interaction model according to the output result of the dot product attention structure.
[0197] In some embodiments, in order to construct the virtual measurement model, each polishing process variable is mapped into a mapping feature. Specifically, each polishing process variable in the training dataset is mapped to the feature space; the positions corresponding to each polishing process variable in the feature space are determined to generate the mapping feature. Determining the output result of the dot product attention structure according to the mapping feature and the interpretable dot product attention mechanism, specifically obtaining the query, key, and value of the interpretable dot product attention mechanism according to the mapping feature; among them, the query, key, and value of the interpretable dot product attention mechanism retain the mathematical meaning of the polishing process variable; according to the query, key, and value of the interpretable dot product attention mechanism, an interpretable attention weight distribution matrix and the output result of the dot product attention structure are obtained. Constructing a regression sampling convolution interaction model according to the output result of the dot product attention structure, specifically constructing a regression sampling convolution interaction model based on a binary tree structure according to the output result of the dot product attention structure; each node of the binary tree structure includes complete input features; determining a single value output by the regression sampling convolution interaction model to complete the construction of the virtual measurement model.
[0198] In some embodiments, determining the positions corresponding to each polishing process variable in the feature space to generate the mapping feature specifically includes extracting primary features and their corresponding feature dimensions based on each polishing process variable; obtaining the positions of each primary feature in the corresponding feature dimension according to each primary feature and its corresponding feature dimension; adding each primary feature and its position in the corresponding feature dimension to obtain the mapping feature of each polishing process variable.
[0199] In some embodiments, queries, keys, and values of the interpretable dot-product attention mechanism are obtained according to mapping features. Specifically, three one-dimensional convolutional filters are set; the kernel width and stride of the one-dimensional convolutional filters are both 1; the mapping features are respectively input into the three one-dimensional convolutional filters to obtain the queries, keys, and values of the interpretable dot-product attention mechanism respectively.
[0200] In some embodiments, an interpretable attention weight distribution matrix and a dot-product attention structure output result are obtained according to the queries, keys, and values of the interpretable dot-product attention mechanism. Specifically, the product of the queries and keys of the interpretable dot-product attention mechanism is determined to obtain the interpretable attention weight distribution matrix; the product of the interpretable attention weight distribution matrix and the values of the interpretable dot-product attention mechanism is determined to obtain the dot-product attention structure output result.
[0201] In some embodiments, a regression sampling convolutional interaction model based on a binary tree structure is constructed according to the dot-product attention structure output result. Specifically, the dot-product attention structure output result is copied to obtain a first copy and a second copy that are used as inputs to the nodes of the first layer of the binary tree; the weighted feature output result of the nodes of the current layer of the binary tree is determined according to the first copy and the second copy; the weighted feature output result is copied to obtain a first copy and a second copy that are used as inputs to the nodes of the next layer of the binary tree; return to the step of determining the weighted feature output result of the nodes of the current layer of the binary tree according to the first copy and the second copy until the binary tree structure reaches a preset number of layers to obtain the regression sampling convolutional interaction model.
[0202] In some embodiments, the weighted feature output result of the nodes of the current layer of the binary tree is determined according to the first copy and the second copy. Specifically, convolution operations are respectively performed on the first copy and the second copy to generate a convolution result corresponding to the first copy and a convolution result corresponding to the second copy; exponential operations are respectively performed on the convolution result corresponding to the first copy and the convolution result corresponding to the second copy to generate an exponential result corresponding to the first copy and an exponential result corresponding to the second copy; the weighted feature output result of the nodes of the current layer of the binary tree is generated according to the first copy and its corresponding exponential result, and the second copy and its corresponding exponential result.
[0203] In some embodiments, convolution operations are respectively performed on the first copy and the second copy to generate a convolution result corresponding to the first copy and a convolution result corresponding to the second copy. Specifically, a one-dimensional convolution operation is performed on the first copy through a first one-dimensional convolution module to obtain a first convolution result; a one-dimensional convolution operation is performed on the second copy through a second one-dimensional convolution module to obtain a second convolution result. Correspondingly, exponential operations are respectively performed on the convolution result corresponding to the first copy and the convolution result corresponding to the second copy to generate an exponential result corresponding to the first copy and an exponential result corresponding to the second copy, including: applying the natural exponential function to the first convolution result to obtain a first exponential result; applying the natural exponential function to the second convolution result to obtain a second exponential result.
[0204] In some embodiments, based on the first copy and its corresponding exponential result, and the second copy and its corresponding exponential result, a weighted feature output result of the current layer binary tree node is generated. Specifically, a Hadamard product operation is performed on the first copy and the second exponential result to obtain a first process quantity; a Hadamard product operation is performed on the second copy and the first exponential result to obtain a second process quantity; a one-dimensional convolution operation is performed on the second process quantity through a fourth one-dimensional convolution module to obtain a fourth convolution result, and the first process quantity is added to the fourth convolution result to obtain a first weighted feature output result; a one-dimensional convolution operation is performed on the first process quantity through a third one-dimensional convolution module to obtain a third convolution result, and the second process quantity is subtracted from the third convolution result to obtain a second weighted feature output result.
[0205] In some embodiments, the structures of the first one-dimensional convolution module, the second one-dimensional convolution module, the third one-dimensional convolution module, and the fourth one-dimensional convolution module are the same, and are successively composed of a hyperbolic tangent activation function, a one-dimensional convolution layer, a random loss layer, a leaky rectified linear unit activation function, and a one-dimensional convolution layer.
[0206] In some embodiments, a single value output by the regression sampling convolution interaction model is determined to complete the construction of the virtual measurement model. Specifically, a plurality of feature matrices corresponding to the regression sampling convolution interaction model are concatenated to obtain a concatenated matrix; wherein, the sizes of the feature matrices corresponding to the regression sampling convolution interaction model are equal to the size of the mapped features; the concatenated matrix is stacked with fully connected layers and then the matrix average value is obtained to obtain a single value, thereby completing the construction of the virtual measurement model.
[0207] In some embodiments, the training process of the virtual measurement model specifically obtains a training data set composed of a plurality of polishing process variables and their corresponding material removal rates; the virtual measurement model is trained according to the data in the training data set until the virtual measurement model converges to obtain the virtual measurement model.
[0208] In some embodiments, after training the virtual measurement model according to the data in the training dataset until the virtual measurement model converges, a test dataset composed of multiple polishing process variables and their corresponding material removal rates is further obtained; an interpretable attention weight distribution matrix corresponding to the current virtual measurement model is obtained; the current virtual measurement model is tested according to the data in the test dataset to obtain the virtual measurement accuracy; it is determined whether the virtual measurement accuracy reaches a preset accuracy value; if not, data is screened from the training dataset according to the interpretable attention weight distribution matrix to obtain a new training dataset, and the process returns to the step of training the virtual measurement model according to the data in the training dataset; if so, the current virtual measurement model is used as the final virtual measurement model.
[0209] In some embodiments, it further includes:
[0210] An update module, configured to store the target polishing process variable and the corresponding target material removal rate into the training dataset to update the training dataset;
[0211] An iterative training module, configured to retrain the virtual measurement model based on the updated training dataset.
[0212] Since the embodiments of the device part correspond to the embodiments of the method part, for the embodiments of the device part, please refer to the description of the embodiments of the method part, which will not be elaborated here for the time being.
[0213] In addition, the present invention further provides a computer program product, including a computer program or instruction, which when executed by a processor implements the steps of the above virtual measurement method.
[0214] Figure 6 As shown in the schematic diagram of a virtual measurement device provided by an embodiment of the present invention, Figure 6 the virtual measurement device includes:
[0215] A memory 20, configured to store a computer program;
[0216] A processor 21, configured to implement the steps of the virtual measurement method as mentioned in the above embodiments when executing the computer program.
[0217] The virtual measurement device provided in this embodiment may include, but is not limited to, a smart phone, a tablet computer, a laptop computer, or a desktop computer, etc.
[0218] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one hardware form of a Digital Signal Processor (DSP), a Field-Programmable Gate Array (FPGA), or a Programmable Logic Array (PLA). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the Central Processing Unit (CPU); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may be integrated with a Graphics Processing Unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may further include an Artificial Intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.
[0219] The memory 20 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 20 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 20 is at least used to store the following computer program 201. After the computer program is loaded and executed by the processor 21, it can implement the relevant steps of the virtual measurement method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 20 may further include an operating system 202 and data 203, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 202 may include Windows, Unix, Linux, etc. The data 203 may include, but is not limited to, the data involved in the virtual measurement method.
[0220] In some embodiments, the virtual measurement device may further include a display screen 22, an input / output interface 23, a communication interface 24, a power supply 25, and a communication bus 26.
[0221] Those skilled in the art can understand that Figure 6 the structure shown in
[0222] Finally, the present invention also provides an embodiment corresponding to a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps recorded in the above method embodiment are implemented.
[0223] It can be understood that if the method in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0224] The above has introduced in detail a virtual measurement method, product, device, and medium provided by the present invention. The various embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
[0225] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the said element.
Claims
1. A virtual measurement method, characterized in that: include: Obtaining various target polishing process variables that affect the polishing quality of the target material; Inputting each of the target polishing process variables into a virtual measurement model to output a target material removal rate corresponding to the target material; Among them, the construction process of the virtual measurement model includes: mapping each polishing process variable into a mapping feature; the mapping feature retains the mathematical meaning of the corresponding polishing process variable; determining the output result of the dot product attention structure according to the mapping feature and the interpretable dot product attention mechanism; constructing a regression sampling convolution interaction model according to the output result of the dot product attention structure; the structure of the interpretable dot product attention mechanism is: setting three one-dimensional convolution filters; the kernel width and step size of the one-dimensional convolution filter are both 1; the mapping features are respectively input into the three one-dimensional convolution filters to obtain the query, key and value of the interpretable dot product attention mechanism respectively; the interpretable dot product attention measures the importance of the embedded input features and outputs the weighted features according to the attention weights.
2. The virtual measurement method according to claim 1, characterized in that: Mapping each of the polishing process variables into the mapping feature comprises: Mapping each of the polishing process variables in the training data set to a feature space; The position corresponding to each of the polishing process variables in the feature space is determined to generate the mapping feature.
3. The virtual measurement method according to claim 2, characterized in that: Determining the output result of the dot product attention structure according to the mapping features and the interpretable dot product attention mechanism includes: Obtaining a query, key, and value of an interpretable dot product attention mechanism according to the mapping feature; wherein the query, key, and value of the interpretable dot product attention mechanism retain the mathematical meaning of the polishing process variable; According to the query, key and value of the interpretable dot product attention mechanism, an interpretable attention weight distribution matrix and the output result of the dot product attention structure are obtained.
4. The virtual measurement method according to claim 3, characterized in that: The regression sampling convolution interaction model is constructed according to the output result of the dot product attention structure, including: Constructing the regression sampling convolution interaction model based on the binary tree structure according to the output result of the dot product attention structure; wherein each node of the binary tree structure contains complete input features; A single value output by the regression sampling convolution interaction model is determined to complete the construction of the virtual measurement model.
5. The virtual measurement method according to any one of claims 2 to 4, characterized in that: Determining the position corresponding to each of the polishing process variables in the feature space to generate the mapping feature includes: Extracting primary features and their corresponding feature dimensions based on each of the polishing process variables; According to each of the primary features and the corresponding feature dimensions, obtaining a position of each of the primary features on the corresponding feature dimensions; Each of the primary features and its position on the corresponding feature dimension are added together to obtain the mapping feature of each of the polishing process variables.
6. The virtual measurement method according to claim 3 or 4, characterized in that: According to the query, key and value of the interpretable dot product attention mechanism, an interpretable attention weight distribution matrix and the output result of the dot product attention structure are obtained, including: Determine the product of the query and the key of the interpretable dot product attention mechanism to obtain the interpretable attention weight distribution matrix; Determine the product of the interpretable attention weight distribution matrix and the value of the interpretable dot product attention mechanism to obtain the dot product attention structure output result.
7. The virtual measurement method according to claim 4, characterized in that: Constructing the regression sampling convolution interaction model based on the binary tree structure according to the output result of the dot product attention structure, including: The output result of the dot product attention structure is copied to obtain a first copy and a second copy as inputs of the first-layer binary tree nodes; Determine the weighted feature output result of the binary tree node at the current layer according to the first copy and the second copy; The weighted feature output result is copied to obtain a first copy and a second copy as inputs to a binary tree node of a next layer; Return to the step of determining the weighted feature output results of the binary tree nodes of the current layer according to the first copy and the second copy, until the binary tree structure reaches a preset number of layers to obtain the regression sampling convolution interaction model.
8. The virtual measurement method according to claim 7, characterized in that: Determine the weighted feature output result of the binary tree node at the current layer according to the first copy and the second copy, including: Performing convolution operations on the first copy and the second copy respectively to generate a convolution result corresponding to the first copy and a convolution result corresponding to the second copy; Performing exponential operations on the convolution result corresponding to the first copy and the convolution result corresponding to the second copy respectively to generate an exponential result corresponding to the first copy and an exponential result corresponding to the second copy; The weighted feature output result of the binary tree node at the current layer is generated according to the first copy and its corresponding index result, and the second copy and its corresponding index result.
9. The virtual measurement method according to claim 8, characterized in that: Performing convolution operations on the first copy and the second copy respectively to generate a convolution result corresponding to the first copy and a convolution result corresponding to the second copy, including: Performing a one-dimensional convolution operation on the first copy through a first one-dimensional convolution module to obtain a first convolution result; Performing a one-dimensional convolution operation on the second copy through a second one-dimensional convolution module to obtain a second convolution result; Correspondingly, performing exponential operations on the convolution result corresponding to the first copy and the convolution result corresponding to the second copy respectively to generate an exponential result corresponding to the first copy and an exponential result corresponding to the second copy, including: Applying a natural exponential function to the first convolution result to obtain a first exponential result; Applying a natural exponential function to the second convolution result obtains a second exponential result.
10. The virtual measurement method according to claim 9, characterized in that: According to the first copy and its corresponding index result, and the second copy and its corresponding index result, the weighted feature output result of the binary tree node of the current layer is generated, including: Performing a Hadamard product operation on the first copy and the second exponential result to obtain a first process quantity; Performing a Hadamard product operation on the second copy and the first exponential result to obtain a second process quantity; Performing a one-dimensional convolution operation on the second process quantity through a fourth one-dimensional convolution module to obtain a fourth convolution result, and adding the first process quantity and the fourth convolution result to obtain a first weighted feature output result; A third one-dimensional convolution module is used to perform a one-dimensional convolution operation on the first process quantity to obtain a third convolution result, and the second process quantity is subtracted from the third convolution result to obtain a second weighted feature output result.
11. The virtual measurement method according to claim 10, characterized in that: The structures of the first one-dimensional convolution module, the second one-dimensional convolution module, the third one-dimensional convolution module and the fourth one-dimensional convolution module are the same, and are composed of a hyperbolic tangent activation function, a one-dimensional convolution layer, a random loss layer, a leaky linear rectification activation function and a one-dimensional convolution layer in sequence.
12. The virtual measurement method according to claim 4, characterized in that: Determining a single value output by the regression sampling convolution interaction model to complete the construction of the virtual measurement model includes: Splicing multiple feature matrices corresponding to the regression sampling convolution interaction model to obtain a spliced matrix; wherein the size of each feature matrix corresponding to the regression sampling convolution interaction model is equal to the size of the mapping feature; The concatenated matrix is stacked with fully connected layers and the matrix average value is calculated to obtain the single value, thereby completing the construction of the virtual measurement model.
13. The virtual measurement method according to claim 1, characterized in that: The training process of the virtual measurement model includes: Obtaining a training data set consisting of a plurality of polishing process variables and their corresponding material removal rates; The virtual measurement model is trained according to the data in the training data set until the virtual measurement model converges to obtain the virtual measurement model.
14. The virtual measurement method according to claim 13, characterized in that: After the virtual measurement model is trained according to the data in the training data set until the virtual measurement model converges, the method further includes: Acquire a test data set consisting of a plurality of polishing process variables and their corresponding material removal rates; Obtaining an interpretable attention weight distribution matrix corresponding to the current virtual measurement model; Testing the current virtual measurement model according to the data in the test data set to obtain the virtual measurement accuracy; Determining whether the virtual measurement accuracy reaches a preset accuracy value; If not, filtering data in the training data set according to the interpretable attention weight distribution matrix to obtain a new training data set, and returning to the step of training the virtual measurement model according to the data in the training data set; If so, the current virtual measurement model is used as the final virtual measurement model.
15. The virtual measurement method according to claim 1, characterized in that: After inputting each of the target polishing process variables into the virtual measurement model to output the target material removal rate corresponding to the target material, the method further includes: Storing the target polishing process variables and the corresponding target material removal rates in a training data set to update the training data set; The virtual measurement model is retrained based on the updated training data set.
16. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the virtual measurement method according to any one of claims 1 to 15 are implemented.
17. A virtual measuring device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the virtual measurement method according to any one of claims 1 to 15 when executing the computer program.
18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the virtual measurement method according to any one of claims 1 to 15 are implemented.
Citation Information
Patent Citations
Wafer CMP material removal rate prediction method of GMDH neural network
CN112257337A
Expression recognition method based on causal separation convolution of time sequence attention mechanism
CN116645713A