A method and system for generating data of soft sensor technology based on diffusion model
By adopting a regression-enhanced diffusion model based on diffusion model (REDM) in soft sensor technology, the problem of insufficient data quality and diversity in the existing technology is solved, and high-quality virtual samples are generated, which significantly improves the prediction performance of the soft sensor model.
Patent Information
- Application Number
- CN202510132533.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2045-02-06
AI Technical Summary
The existing soft sensor technology data generation methods have problems such as insufficient data quality and diversity, insufficient learning of complex mapping relationships and data distribution, and lack of effective optimization and selection mechanisms, which leads to the inability to fully simulate actual production conditions, limiting the performance improvement of the soft sensor model.
A regression-enhanced diffusion model (REDM) based on diffusion model is adopted, which includes a diffusion noise addition module, a back-pass noise removal module and a regression-enhanced module based on multi-head attention mechanism. Representative high-quality virtual samples are generated through iterative training, and the diversity and distribution of virtual samples are optimized through non-dominant sorting and crowding screening.
It significantly improves the prediction performance of the soft sensor model, enhances the model's ability to capture complex mapping relationships, and ensures that the generated virtual samples achieve the best balance in distribution and diversity.
Smart Images

Figure CN119558209B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of soft sensor technology data generation, and particularly relates to a method and system for generating soft sensor technology data based on a diffusion model. Background Art
[0002] In the field of semiconductor manufacturing, precise control of the production process is crucial for ensuring high product quality and operational efficiency. A common industrial practice is to perform metrology steps, including sampling and measuring physical characteristics (such as film thickness) after key process steps. These measurements are used to guide subsequent process adjustments aimed at optimizing production results. However, this method typically relies on limited sample data, usually only sampling 2 to 3 wafers per batch. This data limitation results in less precise production process control and may lead to quality fluctuations. In addition, frequent physical measurement steps are not only time-consuming but also require expensive metrology equipment, increasing production costs and cycle times and reducing overall production efficiency. Soft sensor technology or virtual metrology provides an innovative solution by estimating critical dimensions and material properties without direct physical measurement. This reduces the dependence on expensive metrology tools and decreases the frequency of expensive measurement steps. Nevertheless, reducing the number of physical metrologies limits the availability of labeled training samples required to develop effective soft sensor models.
[0003] Recent research has explored data generation methods to address these challenges, but existing solutions still have the following problems:
[0004] 1. Insufficient quality and diversity of generated data: Existing data generation methods usually lack diversity, and the generated samples have low quality and cannot fully simulate actual production conditions. This limits the effectiveness of the generated data in training soft sensor models, resulting in limited improvement in model performance. At the same time, due to the insufficient quality and diversity of the generated data, the generalization ability of the trained soft sensor model when facing new data is poor, and it is difficult to adapt to changes and new conditions in the production process.
[0005] 2. Insufficient learning of complex mapping relationships and data distributions: Traditional generation models perform poorly in capturing the complex mapping relationships between process variables and target variables, resulting in the generated virtual samples being unable to accurately reflect the changes in actual production. In addition, existing models are difficult to simultaneously learn the data distributions of features and target variables, which leads to the generated data lacking realism and representativeness and affecting the prediction accuracy of the model.
[0006] 3. Lack of effective optimization and selection mechanisms: Existing technologies lack effective mechanisms for optimizing the selection of generated data and cannot ensure that the selected data achieves the best balance in terms of distribution and diversity. This deficiency affects the training effect of the model and cannot fully utilize the generated virtual samples to improve the performance of the soft sensor model. Summary of the Invention
[0007] To address the deficiencies in the existing data generation methods of soft sensor technology, the present invention proposes a data generation method and system for soft sensor technology based on a diffusion model. This diffusion model is a regression-enhanced diffusion model, named REDM (Regression-Enhanced Diffusion Model), which is applicable to the generation of virtual data in soft sensor technology and can widely improve the prediction performance of existing soft sensor models.
[0008] The technical solution adopted by the present invention is as follows:
[0009] In a first aspect, the present invention proposes a data generation method for soft sensor technology based on a diffusion model, including the following steps:
[0010] (1) Collect data during the wafer manufacturing process, including process data and measurement data of key semiconductor manufacturing steps, and use the collected process data and measurement data as the training set;
[0011] (2) Iteratively train the regression-enhanced diffusion model using the training set;
[0012] The regression-enhanced diffusion model includes a diffusion noise-adding module, a backward denoising module, and a regression enhancement module based on a multi-head attention mechanism. The diffusion noise-adding module gradually adds random noise to the process data of the training set to obtain pure Gaussian noise data. The backward denoising module gradually denoises and restores the pure Gaussian noise data to generate virtual process data. The regression enhancement module generates virtual measurement data based on the real process data and the virtual process data respectively, and updates the parameters of the regression-enhanced diffusion model according to the loss between the virtual measurement data and the real measurement data, the backward denoising loss, and the Lipschitz constraint loss;
[0013] During the process of iteratively training the regression-enhanced diffusion model, after each round of training, the training set is updated. The update method is as follows: divide the virtual process data generated by the backward denoising module into virtual sample subsets, and select some virtual sample subsets and add them to the training set according to the non-dominated sorting result and the crowding degree result of the virtual sample subsets;
[0014] (3) Using random Gaussian noise as the input, generate virtual process data using the backward denoising module of the trained regression-enhanced diffusion model, and use the generated result as the soft sensor technology data for key semiconductor manufacturing steps.
[0015] Further, the key semiconductor manufacturing steps include chemical vapor deposition steps, etching steps, and chemical mechanical polishing steps.
[0016] Further, the process data of the key semiconductor manufacturing steps are a series of sensor values.
[0017] Furthermore, the regression enhancement module includes an encoder, a pre-trained multi-head attention module, and a linear prediction layer. After the real process data and the virtual process data are encoded by the encoder respectively, the pre-trained multi-head attention module extracts features, and finally the linear prediction layer generates virtual measurement data.
[0018] Furthermore, during the iterative training of the regression-enhanced diffusion model, the loss functions of the backpropagation denoising module and the regression enhancement module are respectively:
[0019]
[0020]
[0021] Among them, and are the total losses of the backpropagation denoising module and the regression enhancement module respectively, is the loss between the predicted noise and the actually added noise during the backpropagation denoising process, is the loss between the virtual measurement data predicted from the real process data and the real measurement data, is the loss between the virtual measurement data predicted from the virtual process data and the real measurement data, is the Lipschitz constraint loss, is the loss coefficient.
[0022] Furthermore, the screening of some virtual sample subsets according to the non-dominated sorting results and crowding degree results of the virtual sample subsets includes:
[0023] Calculate the Sliced-Wasserstein distance between each virtual sample subset and the training set;
[0024] Calculate the sample similarity score within each virtual sample subset;
[0025] Based on the Sliced-Wasserstein distance and sample similarity score of each virtual sample subset, perform non-dominated sorting on the virtual sample subsets to obtain a Pareto front sequence, where each virtual sample subset corresponds to a Pareto front level;
[0026] Calculate the crowding distance of each virtual sample subset in the Pareto front sequence;
[0027] Start selecting virtual sample subsets from the first level of the Pareto front. If the number of virtual sample subsets at the current level does not meet the screening quantity requirement, continue to select from the next level. For virtual sample subsets with the same Pareto front level, preferentially select subsets with a larger crowding distance until the screening quantity requirement is met.
[0028] In a second aspect, the present invention proposes a data generation system for soft sensor technology based on a diffusion model, which is used to implement the data generation method for soft sensor technology based on the diffusion model described above.
[0029] In a third aspect, the present invention proposes an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the data generation method for soft sensor technology based on the diffusion model described above.
[0030] In a fourth aspect, the present invention proposes a machine-readable storage medium, which stores machine-executable instructions that, when called and executed by a processor, are used to implement the data generation method for soft sensor technology based on the diffusion model described above.
[0031] The beneficial effects of the present invention are as follows:
[0032] (1) The present invention proposes a regression-enhanced diffusion model, which can be used to generate representative and high-quality synthetic training samples. Using the generated soft sensor technology data for key semiconductor manufacturing steps to train the soft sensor model can significantly improve the performance of the existing soft sensor model.
[0033] (2) The present invention introduces a multi-head attention mechanism into the diffusion model to enhance the model's ability to capture complex mapping relationships.
[0034] (3) The present invention divides the virtual process data generated by the backpropagation denoising module into virtual sample subsets, and screens virtual samples that have both distribution similarity and diversity according to the non-dominated sorting results and crowding degree results of the virtual sample subsets. The optimized synthetic samples are incorporated into the training set to participate in the training, ensuring that the selected data achieves the best balance in terms of distribution and diversity. Description of the Drawings
[0035] Figure 1 It is a schematic flowchart of an unsupervised distortion correction matching method for scanning electron microscope images shown in an embodiment of the present invention;
[0036] Figure 2 It is a schematic diagram of the training process of the REDM model shown in an embodiment of the present invention;
[0037] Figure 3 It is a schematic diagram of the sample screening process shown in an embodiment of the present invention;
[0038] Figure 4 It is a comparison chart of the target distributions of the generated virtual samples and real samples. Detailed Embodiments
[0039] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention.
[0040] The accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0041] The flowcharts shown in the accompanying drawings are only exemplary illustrations and do not necessarily include all steps. For example, some steps can be further decomposed, while some steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.
[0042] As Figure 1 shown, the distortion correction matching method for unsupervised scanning electron microscope images mainly includes the following steps:
[0043] S1, collecting real-time data of the fault detection and classification system (abbreviated as FDC system) during the wafer manufacturing process, including process data and measurement data of key semiconductor manufacturing steps, regarding the process data as process variables and the measurement data as target variables.
[0044] The key semiconductor manufacturing steps include chemical vapor deposition steps (abbreviated as CVD), etching steps (abbreviated as ETCH), and chemical mechanical polishing steps (abbreviated as CMP).
[0045] The process data (process variables) of the key semiconductor manufacturing steps are a series of sensor data.
[0046] In a specific implementation of the present invention, for example, during the CVD process, parameters such as lateral / top silane ( ) gas flow rate, top / bias / lateral radio frequency reflection power, deposition time, and frontline pressure are collected as process data, and the final deposition thickness is used as measurement data; for another example, during the ETCH process, taking the polysilicon etching process as an example, high-energy plasma impacts the surface of the semiconductor material to initiate a chemical reaction to remove the material, mainly through the chemical reaction between the fluorine-containing gas and polysilicon, including , and In the process of etching, for example, the participation of sensors is involved, and all sensor data collected during the process are used as process data, including parameters such as gas type, gas flow rate, pressure and power during etching, etc. Then, the critical dimensions measured after etching are collected as measurement data. Another example is the CMP process. Similarly, all sensor data collected during the process are used as process data, including parameters such as chamber ID, wafer ID, chamber pressure, wafer rotation rate, slurry flow rate, processing stage, etc. Then, the average material removal rate of this process is collected as measurement data.
[0047] S2. Use the process data and measurement data of the key semiconductor manufacturing steps as the initial training sample set, and iteratively train the regression enhanced diffusion model (abbreviated as REDM model). Update the training sample set after each round of iterative training.
[0048] As Figure 2 shown, the REDM model mentioned above includes four parts, namely the diffusion noise addition module, the backpropagation denoising module, the regression enhancement module, and the sample screening module.
[0049] (1) Diffusion noise addition module:
[0050] During the diffusion noise addition process, the diffusion noise addition module gradually adds noise to the process data matrix in the training sample . Starting from the data distribution , sample noise from the predefined Gaussian distribution . The process data matrix will gradually become blurred. When t is large enough, the process data matrix will become a pure Gaussian noise data matrix. This process is defined as:
[0051]
[0052] Among them, represents the sequence of all process data matrices from time step 1 to T, represents the probability distribution of the matrix sequence given the initial state , represents the time step, represents the total number of time steps, represents the process data matrix at time step t, represents the process data matrix at time step t - 1.
[0053] Gradually generate the noise-added process data matrix through the diffusion process.
[0054] (2) Backpropagation denoising module:
[0055] In the backpropagation denoising process, the random noise data matrix is used as the latent variable to perform step-by-step denoising, predict the noise components added at each step of the diffusion process, and gradually generate a noise feature matrix as virtual process data . Noise Distribution is usually unknown, which is passed with parameters The output of this neural network is , used to predict the virtual process data matrix The "real" noise component of . The loss function of the backhaul denoising process can be described as:
[0056]
[0057] Among them, represents the loss between the predicted noise and the actual added noise in the backhaul denoising process, Indicates the initial data 、Noise 、Time step Expectations, represents the real noise component, indicates the virtual process data according to step t The predicted noise component.
[0058] In the present invention, the feedback denoising module is regarded as a soft sensor technology data generation model. The trained feedback denoising module can be used to generate process data of virtual semiconductor key manufacturing steps, that is, virtual samples are used as soft sensor technology data for subsequent training of soft sensor models.
[0059] (3) Regression enhancement module:
[0060] Traditional generative models perform poorly in capturing the complex mapping relationship between process variables and target variables in key semiconductor manufacturing steps, resulting in the inability of the generated virtual samples to accurately reflect changes in actual production. The present invention modifies the loss function of the generative model, hoping that the generated samples will be more conducive to the soft sensor model.
[0061] In the present invention, the regression enhancement module is regarded as a soft sensor technology data discrimination model, which receives the virtual process data matrix generated by the feedback denoising module and real process data matrix , generate virtual measurement data.
[0062] In a specific implementation of the present invention, the regression enhancement module includes a feature encoder, a pre-trained multi-head attention module, and a linear prediction layer. The encoder can adopt a Transformer encoder, and the linear prediction layer adopts a multi-layer perceptron. The real process data and the virtual process data are respectively encoded by the encoder to obtain their feature embedding representations in the feature space, and then the feature embedding representations are sent into the pre-trained multi-head attention module to extract features. By introducing the multi-layer head attention mechanism, the ability of the model to capture complex mapping relationships is enhanced. Finally, the virtual measurement data is generated by the linear prediction layer and . By comparing the virtual measurement data with the real measurement data to calculate the loss value:
[0063]
[0064]
[0065] where is the loss between the virtual measurement data predicted and generated from the real process data and the real measurement data , is the loss between the virtual measurement data predicted and generated from the virtual process data and the real measurement data , is the expectation of the real process data and the real measurement data y, is the expectation of the virtual process data and the real measurement data y
[0066] In order to prevent performance limitation and reduce stability, the present invention adds gradient penalty to enforce the Lipschitz constraint, and the loss function is calculated as follows:
[0067]
[0068] where the gradient of the discriminator network is denoted as , represents the probability distribution between the virtual process data and the real process data, represents the expectation sampled from this distribution .
[0069] The loss functions of the backpropagation denoising module and the regression enhancement module are finally obtained as:
[0070]
[0071]
[0072] Among them, and are the total losses of the feedback denoising module and the regression enhancement module respectively, is the loss between the predicted noise and the actually added noise during the feedback denoising process, is the loss between the virtual measurement data predicted and generated based on the real process data and the real measurement data, is the loss between the virtual measurement data predicted and generated based on the virtual process data and the real measurement data, is the Lipschitz constraint loss, is the loss coefficient.
[0073] (4) Sample screening module:
[0074] In an actual scenario, even when using a well-trained generation model, it is challenging to ensure the quality of all virtual samples. Virtual samples that deviate from the original sample distribution may degrade the performance of subsequent soft sensor models, and samples that are too similar are also not beneficial for enhancing the prediction performance of the soft sensor model. Maintaining the diversity of the samples themselves is also a key aspect. After each round of generation training is completed, the generated virtual samples with the same number as the training set samples are divided into subsets to form a virtual sample set and stored in the historical sample library, where represents the th virtual sample subset, represents the training set, is the number of virtual samples contained in the virtual sample subset . Here, the number of samples in each virtual sample subset can be the same or different. In this embodiment, the same is taken as an example.
[0075] Calculate the Sliced-Wasserstein distance between each virtual sample subset and the training set:
[0076]
[0077] Among them, represents the Sliced-Wasserstein distance, represents the unit sphere in the dimensional space, and respectively represent the original multi-dimensional probability distributions of the virtual sample subset and the training set , represents the linear projection, represents the conjectured measure, means that the distribution μ is mapped by Push to a new space, similarly, means it will be distributed Through mapping Push to a new space, represents the one-dimensional Wasserstein distance, indicates From the unit sphere Expected sampling in .
[0078] Since it is not feasible to project in all directions on the sphere, this embodiment uses 50 random projections to approximately calculate the SWD.
[0079] Calculate the cosine similarity within each virtual sample subset:
[0080]
[0081] Among them, respectively represent the The jth and kth virtual samples in the virtual sample subset.
[0082] Based on the SWD and SIM scores of each virtual sample subset, a non-dominated sort (NDS) is performed in the historical sample library to identify the Pareto frontier of each level and obtain a Pareto frontier sequence, where each virtual sample subset corresponds to a Pareto frontier level. For each subset, its crowding distance in the target space is calculated to measure the sparsity of the subset in the target space. The larger the crowding distance, the more sparse the subset is.
[0083] When screening virtual sample subsets, start selecting virtual sample subsets from the first level of the Pareto frontier. If the number of virtual sample subsets at the current level does not meet the screening number requirement, continue selecting from the next level. For virtual sample subsets at the same Pareto frontier level, give priority to subsets with larger crowding distances, thereby enhancing the diversity of virtual sample sets until the screening number requirement is met. In each training iteration, the selected virtual sample subset is added to the training set to form a synthetic sample.
[0084] (3) Using random Gaussian noise as input, the return denoising module of the trained regression enhanced diffusion model is used to generate virtual process data, and the generated results are used as soft sensor technology data for key semiconductor manufacturing steps.
[0085] In order to quantitatively evaluate the performance of the REDM model, this paper compares it with existing models. Due to the large number of generative models, only the leading methods in each generative model paradigm are selected for evaluation, including TVAE, CTGAN, SMOTE and TabDDPM as baseline models.
[0086] REDM model parameter settings: Set the number of iterations to 30 generations, and each generated virtual sample is divided into 10 parts for sample selection. By Figure 3 It can be observed the evolution of the selected subset during the iteration process. The red color represents the currently selected subset, the yellow color represents the historically selected subset, and the blue color represents the unselected subset. It can be seen that as the iteration progresses, the SWD and SIM scores of the virtual sample subset gradually evolve towards the lower left corner. At the same time, the red samples are scattered on the Pareto boundary of the final virtual sample subset, indicating that as the iteration progresses, the generated virtual samples are continuously learning the data distribution of the real samples, and the distribution distance between the two is gradually shrinking. At the same time, the finally selected virtual sample subset is more diverse and is not just a replication result of the real samples.
[0087] Machine learning utility evaluation: Use machine learning utility as the main evaluation index to quantify the performance of the regression model trained on synthetic data on the real test set. Three evaluation protocols, namely Random Forest, Xgboost, and Catboost, are used for calculation. In this embodiment, datasets of CVD, ETCH, and CMP processes are collected and named CVD-1, CVD-2, CVD-3, CVD-4, ETCH-1, CMP-1, CMP-2, and CMP-3 respectively.
[0088] Quality evaluation of generated virtual samples: Use the kernel density estimation algorithm to obtain the probability density curve of the virtual samples to effectively compare the probability density curves of the virtual samples and the real samples. By Figure 4 It can be observed the distribution of the target variables of the data generated by REDM, CTGAN, TVAE, and SMOTE and the real data. The first 8 columns represent the distribution differences of the target variables of the generated samples and the real samples under the training of 4 generation models for a single dataset, and the last column represents the average of the distribution distance WD (Wasserstein Distance) between the generated samples and the real samples of the target variables in 8 datasets for this generation model. It can be seen that the WD of SMOTE is the smallest and the distribution similarity is the highest. This is because SMOTE is based on the principle of interpolation, which is equivalent to a replication result of the real data. The WD of REDM is smaller than that of CTGAN and TVAE, indicating that the samples generated by REDM have higher similarity compared to CTGAN and TVAE and higher diversity compared to SMOTE.
[0089] Comparison of generation models: In this experiment, REDM is divided into two variants: REDM-I and REDM-II. REDM-I does not include a sample screening module, while REDM-II includes a complete regression enhanced diffusion model. By sampling synthetic datasets for each generation model, the R² score of the regression model is evaluated using the real test set.
[0090] Tables 1, 2, and 3 show the machine learning utility values under different evaluation protocols. The results are averaged over five random seeds.
[0091] Table 1
[0092]
[0093] Table 2
[0094]
[0095] Table 3
[0096]
[0097] In each evaluation protocol, REDM-I outperforms TVAE, CTGAN, SMOTE, and TabDDPM on most datasets, indicating that incorporating the mapping relationship between features and target variables into the generative model can improve the machine learning utility of synthetic samples. REDM-II further enhances this improvement through sample screening. As a lightweight generative method, SMOTE is competitive in data similarity, but the generated samples often "copy" real data and it is difficult to further improve the machine learning utility. In contrast, the data generated by REDM has certain noise and transformations, which helps the regression model to be more robust to noise and transformations, resulting in better performance in real tests. In the baseline regression model test, CatBoost performs the best. In the actual semiconductor manufacturing scenario, we usually only consider this best regression model. REDM-II increases the R² score of this model by 4.53% on average.
[0098] The present invention also provides a soft sensor technology data generation system based on a diffusion model, which is used to implement the above embodiments. The following terms "module", "unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible.
[0099] A soft sensor technology data generation system based on a diffusion model provided in this embodiment includes:
[0100] A wafer manufacturing process data collection module, which is used to collect data in the wafer manufacturing process, including process data and measurement data of key semiconductor manufacturing steps, and use the collected process data and measurement data as a training set;
[0101] Regression enhanced diffusion model module, which includes a diffusion noise addition module, a backpropagation denoising module, and a regression enhancement module based on a multi-head attention mechanism. The diffusion noise addition module gradually adds random noise to the process data of the training set to obtain pure Gaussian noise data. The backpropagation denoising module gradually denoises and restores the pure Gaussian noise data to generate virtual process data. The regression enhancement module generates virtual measurement data based on the real process data and the virtual process data respectively.
[0102] Regression enhanced diffusion model training module, which is used to iteratively train the regression enhanced diffusion model using the training set, and update the parameters of the regression enhanced diffusion model according to the losses between the virtual measurement data and the real measurement data, the backpropagation denoising loss, and the Lipschitz constraint loss. Moreover, during the iterative training of the regression enhanced diffusion model, after each round of training, the training set is updated. The update method is as follows: divide the virtual process data generated by the backpropagation denoising module into virtual sample subsets, and select some virtual sample subsets according to the non-dominated sorting results and crowding degree results of the virtual sample subsets and add them to the training set.
[0103] Soft sensor technology data generation module, which takes random Gaussian noise as input, uses the backpropagation denoising module of the trained regression enhanced diffusion model to generate virtual process data, and takes the generation result as the soft sensor technology data for key semiconductor manufacturing steps.
[0104] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. For example, the regression enhancement module includes:
[0105] An encoder module, which is used to encode the real process data and the virtual process data.
[0106] A pre-trained multi-head attention module, which is used to extract the features of the encoded real process data and virtual process data.
[0107] A linear prediction layer module, which is used to predict virtual measurement data based on the features of the encoded real process data and virtual process data.
[0108] The implementation methods of the remaining modules will not be elaborated here. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. One can select some or all of the modules according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0109] Embodiments of the system of the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware, or by a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation.
[0110] Embodiments of the present invention further provide an electronic device, including a memory and a processor;
[0111] The memory is used to store a computer program;
[0112] The processor is used to implement the above-mentioned data generation method of soft sensor technology based on the diffusion model when executing the computer program.
[0113] Embodiments of the present invention further provide a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the above-mentioned data generation method of soft sensor technology based on the diffusion model is implemented.
[0114] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or will be output.
[0115] Obviously, the above-described embodiments and drawings are only some examples of the present application. For those of ordinary skill in the art, the present application can also be applied to other similar situations based on these drawings without creative labor. Additionally, it can be understood that although the work done during this development process may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be regarded as insufficient disclosure of the present application. Without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for generating data of soft sensor technology based on diffusion model, characterized in that: The following steps are involved: (1) Collect data from the wafer manufacturing process, including process data and measurement data of key semiconductor manufacturing steps, and use the collected process data and measurement data as training sets; The key semiconductor manufacturing steps include chemical vapor deposition step, etching step, and chemical mechanical polishing step; (2) Iteratively train the regression enhanced diffusion model using the training set; The regression enhancement diffusion model includes a diffusion denoising module, a return denoising module and a regression enhancement module based on a multi-head attention mechanism. The diffusion denoising module gradually adds random noise to the process data of the training set to obtain pure Gaussian noise data, and the return denoising module gradually denoises and restores the pure Gaussian noise data to generate virtual process data. The regression enhancement module generates virtual measurement data according to the real process data and virtual process data, and updates the parameters of the regression enhancement diffusion model according to the loss of the virtual measurement data and the real measurement data, the return denoising loss and the Lipschitz constraint loss; In the process of iterative training of the regression enhanced diffusion model, the training set is updated after each round of training. The updating method is as follows: the virtual process data generated by the feedback denoising module is divided into virtual sample subsets, and some virtual sample subsets are selected according to the non-dominated sorting results and crowding results of the virtual sample subsets and added to the training set; (3) Using random Gaussian noise as input, the feedback denoising module of the trained regression enhanced diffusion model is used to generate virtual process data, and the generated results are used as soft sensor technology data for key semiconductor manufacturing steps.
2. The method for generating data of soft sensor technology based on diffusion model according to claim 1, characterized in that: The process data of the key semiconductor manufacturing steps are a series of sensor values.
3. The method for generating data of soft sensor technology based on diffusion model according to claim 1, characterized in that: The regression enhancement module includes an encoder, a pre-trained multi-head attention module, and a linear prediction layer; the real process data and the virtual process data are encoded by the encoder respectively, and then the features are extracted by the pre-trained multi-head attention module, and finally the virtual measurement data is generated by the linear prediction layer.
4. The method for generating data of soft sensor technology based on diffusion model according to claim 1, characterized in that: During the iterative training of the regression enhanced diffusion model, the loss functions of the backhaul denoising module and the regression enhancement module are: ; ;in, and are the total losses of the backhaul denoising module and the regression enhancement module, respectively. is the loss between the predicted noise and the actual added noise in the backhaul denoising process, It is the loss between the virtual measurement data generated based on the prediction of the real process data and the real measurement data. It is the loss between the virtual measurement data generated by the virtual process data prediction and the real measurement data. is the Lipschitz constraint loss, is the loss coefficient.
5. The method for generating data of soft sensor technology based on diffusion model according to claim 1, characterized in that: The method of selecting a portion of the virtual sample subset according to the non-dominated sorting result and the crowding result of the virtual sample subset includes: Calculate the Sliced-Wasserstein distance between each virtual sample subset and the training set; Calculate the sample similarity score within each virtual sample subset; Based on the Sliced-Wasserstein distance and sample similarity score of each virtual sample subset, the virtual sample subsets are non-dominated sorted to obtain a Pareto front sequence, where each virtual sample subset corresponds to a Pareto front level; Calculate the crowding distance of each virtual sample subset in the Pareto front sequence; Starting from the first level of the Pareto frontier, virtual sample subsets are selected. If the number of virtual sample subsets at the current level does not meet the screening quantity requirement, the selection continues from the next level. For virtual sample subsets at the same Pareto frontier level, subsets with larger crowding distances are selected first until the screening quantity requirement is met.
6. A soft sensor technology data generation system based on a diffusion model, characterized in that: include: A wafer manufacturing process data collection module is used to collect data in the wafer manufacturing process, including process data and measurement data of key semiconductor manufacturing steps, and use the collected process data and measurement data as a training set; The key semiconductor manufacturing steps include chemical vapor deposition, etching, and chemical mechanical polishing; A regression enhanced diffusion model module, which includes a diffusion denoising module, a backhaul denoising module and a regression enhancement module based on a multi-head attention mechanism. The diffusion denoising module gradually adds random noise to the process data of the training set to obtain pure Gaussian noise data, and the backhaul denoising module gradually denoises and restores the pure Gaussian noise data to generate virtual process data. The regression enhancement module generates virtual measurement data based on real process data and virtual process data respectively; A regression enhanced diffusion model training module is used to iteratively train the regression enhanced diffusion model using a training set, and update the parameters of the regression enhanced diffusion model according to the loss of virtual measurement data and real measurement data, the return denoising loss, and the Lipschitz constraint loss; and in the process of iteratively training the regression enhanced diffusion model, the training set is updated after each round of training, and the updating method is: the virtual process data generated by the return denoising module is divided into virtual sample subsets, and some virtual sample subsets are selected according to the non-dominated sorting results and congestion results of the virtual sample subsets and added to the training set; The soft sensor technology data generation module is used to generate virtual process data using random Gaussian noise as input and the return denoising module of the trained regression enhanced diffusion model, and the generated results are used as soft sensor technology data for key semiconductor manufacturing steps.
7. The soft sensor technology data generation system based on diffusion model according to claim 6 is characterized in that: The regression enhancement module includes: An encoder module for encoding real process data and virtual process data; A pre-trained multi-head attention module, which is used to extract features of the encoded real process data and virtual process data; The linear prediction layer module is used to predict the virtual measurement data according to the characteristics of the encoded real process data and the virtual process data.
8. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the method for generating data of soft sensor technology based on a diffusion model as described in any one of claims 1 to 5.
9. A machine-readable storage medium, characterized in that: The machine-readable storage medium stores machine-executable instructions, which, when called and executed by a processor, are used to implement the method for generating data using soft sensor technology based on a diffusion model as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Automatic elimination of noise for big data analytics
CN112149375A
Intelligent sampling semiconductor virtual measurement system based on SLGBM model
CN117272781A