Modeling for indexing and semiconductor defect image retrieval

Extracting semiconductor defect image features through deep learning and vision converter models, solving the problems of low efficiency and relying on manual experience in the prior art, achieving fast and accurate defect detection and flexible image retrieval, which is suitable for substrate processing defect analysis in semiconductor manufacturing.

CN120266155APending Publication Date: 2025-07-04APPLIED MATERIALS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380081682.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-29
Filing Date
2023-10-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has inefficient, dependant on manual experience, difficulty in processing new data and multiple defects in the defect detection and classification process in semiconductor manufacturing, and lacks flexibility to effectively index and retrieve image data of substrate processing defects.

Method used

Deep learning technology is adopted to extract feature representations of defective images using computer learning modeling and vision transformer models. By cropping and occluding image parts, flexible image retrieval and indexing methods are provided to quickly process new data and update databases.

Benefits of technology

Improves the efficiency and accuracy of defect detection, reduces dependence on manual experience, enables rapid retrieval and index similar images, supports the processing of new data and multi-defect analysis, and provides flexibility to target image areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120266155A_ABST
    Figure CN120266155A_ABST
Patent Text Reader

Abstract

The object of this specification may be implemented, inter alia, in methods, systems, computer-readable storage media. A method may include a processing device storing a plurality of feature vectors representing previously processed image frames corresponding to various substrate processing defects. The method further includes receiving first image data including one or more image frames indicative of a first substrate processing defect. The method further includes determining a first feature vector corresponding to the first image data. The method further includes determining a selection of the plurality of feature vectors based on a proximity between the first feature vector and each of the selection of the plurality of feature vectors. The method further includes determining second image data, the second image data including the selected one or more image frames corresponding to the plurality of embedding vectors, and performing an action based on determining the second image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present specification generally relate to modeling for semiconductor defect image indexing and retrieval. More specifically, embodiments of the present specification relate to performing image search on semiconductor defect images using a novel combination of deep learning and vector search techniques. Background Art

[0002] In manufacturing, such as in semiconductor device manufacturing, the wafer fab throughput depends on the device quality that can be directly measured using metrology tools and indirectly measured by monitoring process equipment sensors. This information is collected at different times during the product manufacturing lifecycle. When a manufacturing engineer needs to identify problems with a process tool or a final product, he or she must go through a laborious and expensive process to analyze a large number of data points (e.g., metrology data of many samples with various measurement parameters). For example, when an engineer is informed that there is a potential problem with a product, the engineer must check the corresponding metrology data to find the alarm characteristics of the product. A common method for identifying metrology violations is to use image analysis.

[0003] The semiconductor industry generates images for failure analysis between processing steps from various tools to identify defect locations, characteristics, and classifications stored in various databases. The obstacle to defect detection / classification is searching for index data in the database that helps to characterize substrate processing defects corresponding to image processing. Searching for defect information in conventional systems is often limited by text queries, which may be subjective, manual, tedious, indirect, inefficient, and limited by the knowledge, background, relevant experience, and specific interests of the evaluator. Summary of the Invention

[0004] A method, system, and computer readable media (CRM) facilitate modeling for semiconductor defect image indexing and retrieval. In some embodiments, a method executed by a processing device may include storing a plurality of feature vectors that represent image frames of previous processing corresponding to various substrate processing defects. The method further includes receiving first image data including one or more image frames indicating a first substrate processing defect. The method further includes determining a first feature vector corresponding to the first image data. The method further includes determining a selection of the plurality of feature vectors based on proximity between the first feature vector and each of the selections of the plurality of feature vectors. The method further includes determining second image data that includes one or more image frames corresponding to the selection of the plurality of embedded vectors, and performing an action based on the determination of the second image data.

[0005] In some embodiments, a method may include a processing device receiving a first video frame indicative of a substrate processing defect. The method may further include generating a second video frame by cropping a first region of the first video frame. The method may further include generating a third video frame by cropping a second region of the first video frame, wherein the first region encompasses the second region. The method may further include using the second video frame as an input to a first machine learning (ML) model. The method may further include obtaining one or more outputs of the first ML model, the one or more outputs indicative of a first feature vector corresponding to the second video frame. The method may further include using the third video frame as an input to a second ML model and obtaining one or more outputs of the second ML model indicative of a second feature vector corresponding to the third video frame. The method may further include updating one or more parameters of at least one of the first ML model or the second ML model based on a comparison between the one or more outputs of the first ML model and the one or more outputs of the second ML model.

[0006] In some embodiments, a non-transitory machine-readable storage medium includes instructions that, when executed by a processing device, cause the processing device to perform operations. The operations may include storing, in a data storage device, a plurality of feature vectors that represent previously processed video frames corresponding to various substrate processing defects. The operations may further include receiving first image data that includes one or more video frames indicative of a first substrate processing defect. The operations may further include determining a first feature vector corresponding to the first image data. The operations may further include determining a selection of the plurality of feature vectors based on a proximity between the first feature vector and each of the selections of the plurality of feature vectors. The operations may further include determining second image data that includes one or more video frames corresponding to the selection of the plurality of embedded vectors. The operations may further include performing an action based on determining the second image data. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Aspects and embodiments of the present disclosure will be more fully understood from the detailed description and the accompanying drawings given below, which are intended to illustrate, by way of example and without limitation, aspects and embodiments.

[0008] Figure 1 is a block diagram illustrating an example system architecture in which embodiments of the present disclosure may operate.

[0009] Figure 2 is a block diagram illustrating a substrate defect image indexing and retrieval system in which embodiments of the present disclosure may operate.

[0010] Figure 3 is a block diagram of a defect size determination system according to aspects of the present disclosure.

[0011] Figure 4 is a block diagram showing a process for training a machine learning model to generate an output according to an aspect of the present disclosure.

[0012] Figure 5 Shows a model training workflow and a model application workflow for substrate defect image indexing and retrieval according to an aspect of the present disclosure.

[0013] Figures 6A to 6B Shows a model architecture for substrate defect image indexing and retrieval according to an aspect of the present disclosure.

[0014] Figure 7 Depicts a flowchart of an example method for substrate defect image indexing and retrieval according to some embodiments of the present disclosure.

[0015] Figure 8 Depicts a block diagram of an example computing device operating in accordance with one or more aspects of the present disclosure. Detailed Description

[0016] The embodiments described herein relate to semiconductor defect image indexing and retrieval. In manufacturing, such as in semiconductor device manufacturing, product quality can be directly measured using metrology tools and indirectly measured by monitoring process equipment sensors. This information is collected at different times during the product manufacturing lifecycle. When a manufacturing engineer needs to identify problems with a processing tool or a final product, he or she must go through a laborious and expensive process of analyzing a large number of data points (e.g., metrology data of many samples with various measurement parameters). For example, when an engineer is informed that a product has a potential problem, the engineer must verify the corresponding metrology data to find the alert characteristics of the product. One form of analyzing metrology data is through the use of substrate imaging, such as, for example, by using an electron microscope (e.g., a scanning electron microscope).

[0017] The semiconductor industry generates images from various tools for failure analysis and troubleshooting between processing steps to identify defect locations, characteristics, and classifications stored in various databases. The obstacle to defect detection / classification is searching for index data in the database that helps characterize substrate processing defects corresponding to image processing. Searching for defect information in conventional systems is often limited by text queries, which may be subjective, manual, tedious, indirect, inefficient, and limited by the knowledge, background, relevant experience, and specific interests of the evaluator. Conventional systems further cannot handle images with new or out-of-distribution (OOD) data and combinations of multiple defects. Conventional substrate defect analysis also lacks flexibility in providing a specific region of the targeted substrate and in selectively removing the effects of other defects that may interfere with a certain defect classification.

[0018] Well-known defect classification algorithms often need to learn manufacturing processes, such as substrate device processing parameters. Images of defects (such as microscopic images) often require years of experience from engineers to understand the symptoms within the images (e.g., the size, orientation, shape, texture, topography, type, etc. of the defects) and determine one or more abnormalities of the substrate processing procedure or substrate processing equipment based on the symptoms, such as, for example, where the defect comes from, how the defect is caused (e.g., how the defect is generated or transported, etc.). Additionally, when multiple defects overlap each other, the challenge of identifying the defects is further exacerbated.

[0019] In addition to employing domain-specific filters, aspects and embodiments of the present disclosure address these other drawbacks of the prior art by providing a framework for indexing and retrieving digital images of semiconductor-based particles on a substrate that can index and retrieve based on the content of the image and / or the defect size and the option to focus on or ignore specific regions of the defect. The present disclosure utilizes computer learning modeling to determine a representation of the image and establish a repository of the image and / or a feature representation of the image. In some embodiments, the present disclosure provides a search mechanism that uses image features extracted using computational modeling (e.g., a deep learning-based Visual Transformer (ViT) model). In some embodiments, the present disclosure provides the option to crop certain portions of the image and / or ignore (e.g., mask) certain portions of the image. In some embodiments, the present disclosure provides a module for extracting size information (e.g., magnification, image scaling, defect size, etc.) and making it available for identifying similar images and / or defects.

[0020] The proposed solution utilizes a learning model (e.g., ViT) to extract a good representation of the defect image. The proposed solution employs an index optimized for vector similarity search to retrieve similar images in a fast manner. The proposed solution effectively alleviates the difficulty of processing new data because the index and database can be updated with new data. Compared with well-known text-based solutions, the proposed solution provides faster and more focused results. Additionally, the image-based search solution is not limited by defect keywords, class labels, and / or other mechanism knowledge required in well-known systems. The proposed solution is further capable of performing dynamic searches, where new information can be retrieved, processed, and indexed, and the ability to handle multiple defects in the same image is provided. In addition to providing the flexibility to target specific regions of the image (e.g., using the crop and / or mask features), the proposed solution further provides an improved image representation of the substrate defect.

[0021] A method, system, and computer-readable medium (CRM) facilitate modeling for semiconductor defect image indexing and retrieval. In an example implementation, a method executed by a processing device may include storing a plurality of feature vectors that represent image frames of previous processing corresponding to various substrate processing defects. The method may further include receiving first image data that includes one or more image frames indicating a first substrate processing defect. The method may further include determining a first feature vector corresponding to the first image data. The method may further include determining a selection of the plurality of feature vectors based on a proximity between the first feature vector and each of the selections of the plurality of feature vectors. The method may further include determining second image data that includes one or more image frames corresponding to the selection of the plurality of embedded vectors, and performing an action based on determining the second image data.

[0022] In an example implementation, a method may include a processing device receiving a first image frame indicating a substrate processing defect. The method may further include generating a second image frame by cropping a first region of the first image frame. The method may further include generating a third image frame by cropping a second region of the first image frame, where the first region includes the second region. The method may further include using the second image frame as an input to a first machine (ML) model. The method may further include obtaining one or more outputs of the first ML model, the one or more outputs indicating a first feature vector corresponding to the second image frame. The method may further include using the third image frame as an input to a second ML model and obtaining one or more outputs of the second ML model indicating a second feature vector corresponding to the third image frame. The method may further include updating one or more parameters of at least one of the first ML model or the second ML model based on a comparison between the one or more outputs of the first ML model and the one or more outputs of the second ML model.

[0023] In an example implementation, a non-transitory machine-readable storage medium includes instructions that, when executed by a processing device, cause the processing device to perform operations. The operations may include storing a plurality of feature vectors in a data storage device, the plurality of feature vectors representing image frames of previous processing corresponding to various substrate processing defects. The operations may further include receiving first image data that includes one or more image frames indicating a first substrate processing defect. The operations may further include determining a first feature vector corresponding to the first image data. The operations may further include determining a selection of the plurality of feature vectors based on a proximity between the first feature vector and each of the selections of the plurality of feature vectors. The operations may further include determining second image data that includes one or more image frames corresponding to the selection of the plurality of embedded vectors. The operations may further include performing an action based on determining the second image data.

[0024] Figure 1 is a block diagram showing an example system architecture 100 in which embodiments of the present disclosure may operate. As Figure 1 shown, the system architecture 100 includes a manufacturing system 102, a metrology system 110, a client device 150, a data store 140, a server 120, and a machine learning system 170. The machine learning system 170 may be part of the server 120. In some embodiments, one or more components of the machine learning system 170 may be fully or partially integrated into the client device 150. The manufacturing system 102, the metrology system 110, the client device 150, the data store 140, the server 120, and the machine learning system 170 may each be hosted by one or more computing devices, including server computers, desktop computers, portable computers, tablet computers, notebook computers, personal digital assistants (PDAs), mobile communication devices, cellular phones, handheld computers, cloud servers, cloud-based systems (e.g., cloud service devices, cloud network devices, or similar computing devices).

[0025] The manufacturing system 102, the metrology system 110, the client device 150, the data store 140, the server 120, and the machine learning system 170 may be coupled to each other via a network 160 (e.g., for performing the methods described herein). In some embodiments, the network 160 is a private network that provides access for each element of the system architecture 100 to each other and other privately available computing devices. The network 160 may include one or more wide area networks (WANs), local area networks (LANs), wired networks (e.g., Ethernet networks), wireless networks (e.g., 802.11 networks or Wi-Fi networks), cellular networks (e.g., Long Term Evolution (LTE) networks), routers, hubs, switches, server computers, and / or any combination thereof. In some embodiments, the network 160 is a cloud-based network capable of performing cloud-based functionality (e.g., providing cloud service functionality to one or more devices in the system). Alternatively or additionally, any of the elements of the system architecture 100 may be integrated together or otherwise coupled without using the network 160.

[0026] The client device 150 may be or include any personal computer (PC), laptop computer, mobile phone, tablet computer, notebook computer, network-connected television (“smart TV”), network-connected media player (e.g., Blu-ray player), set-top box, over-the-top (OTT) streaming device, operator box, etc. The client device is capable of performing cloud-based operations (e.g., leveraging the server 120, data storage 140, manufacturing system 102, machine learning system 170, metrology system 110, etc.). The client device 150 may include a browser 152, an application 154, and / or other tools described and executed by other systems of the system architecture 100. In some embodiments, the client device 150 is capable of accessing the manufacturing system 102, metrology system 110, data storage 140, server 120, and / or machine learning system 170 and communicating (e.g., sending and / or receiving) instructions regarding: metrology data, processed data (e.g., enhanced image data, embedding vectors, and the like), process result data (e.g., critical dimension data, thickness data), and / or inputs and outputs of various processing tools (e.g., imaging tool 114, data preparation tool 116, image enhancement tool 124, embedding tool 126, search tool 128, defect tool 130, and / or retrieval component 194) at various stages of the processing system architecture 100, as described herein.

[0027] As Figure 1 shown, the manufacturing system 102 includes processing tools 104, processing procedures 106, and a processing controller 108. The processing controller 108 may coordinate the operation of the processing tools 104 to execute one or more processing procedures 106. For example, each processing tool may include specialized chambers such as etch chambers, deposition chambers (including chambers for atomic layer deposition, chemical vapor deposition, sputtering chambers, physical vapor deposition, or plasma-enhanced versions thereof), annealing chambers, implantation chambers, electroplating chambers, processing chambers, and / or the like. In another example, the machine may incorporate a sample transport system (e.g., a selective compliance assembly robot arm (SCARA) robot, transfer chamber, front opening pod (FOUP), side storage pod (SSP), and / or the like) to transport samples between the machine and the processing steps.

[0028] The process recipe 106, or sometimes referred to as a process recipe or process step, may include various specifications for performing operations by the processing tool 104. For example, the process recipe 106 may include process specifications such as the start duration of a process operation of a machine (e.g., a chamber), the processing tool for the operation, temperature, flow rate, pressure, etc., deposition sequence, and the like. In another example, the process recipe may include transfer instructions for transporting a sample to a further processing step or for measurement by the metrology system 110.

[0029] The process controller 108 may include means designed to manage and coordinate the actions of the processing tool 104. In some embodiments, the process controller 108 is associated with a process recipe or a series of process recipe 106 instructions that, when applied in a designed manner, result in the desired processing outcome for substrate processing. For example, a process recipe may be associated with processing a substrate to produce a target processing result (e.g., critical dimension, thickness, uniformity criteria, etc.).

[0030] As Figure 1 shown, the metrology system 110 includes an imaging tool 114 and a data preparation tool 116. The imaging tool 114 may include various sensors to measure processing results (e.g., critical dimension, thickness, uniformity, etc.) within the manufacturing system 102. For example, the imaging tool 114 may include a scanning tunneling microscope (STM) or a scanning electron microscope (SEM). In another example, wafers processed in one or more processing chambers may be used to measure critical dimensions. The imaging tool 114 may also include means for measuring the processing results of substrates processed using the manufacturing system. For example, the processing results of substrates processed according to a process recipe and / or actions executed by the process controller 108, such as critical dimensions, thickness measurements (e.g., of film layers from etching, deposition, etc.), may be estimated. These measurements may also be used to measure chamber conditions throughout the substrate processing procedure.

[0031] The data preparation tool 116 may include processing methods for extracting features and / or generating synthetic / engineering data associated with data measured by the imaging tool 114. In some embodiments, the data preparation tool 116 may identify the relevance, patterns, and / or anomalies of metrology or process performance data. For example, the data preparation tool 116 may perform feature extraction, where the data preparation tool 116 uses a combination of measurement data to determine whether criteria are met. For example, the data preparation tool 116 may analyze multiple data points of associated parameters to determine whether rapid changes occur across multiple processing chambers during a substrate processing procedure. In some embodiments, the data preparation tool 116 performs normalization on various sensor data associated with various processing chamber conditions. Normalization may include processing the input sensor data to make it look similar across the individual chambers and sensors used to acquire the data.

[0032] In some embodiments, the data preparation tool 116 may perform one or more of process control analysis, univariate limit violation analysis, or multivariate limit violation analysis on metrology data (e.g., metrology data obtained by the imaging tool 114). For example, the data preparation tool 116 may perform statistical process control (SPC) by monitoring and controlling the process controller 108 using a statistics-based method. For example, SPC may improve the efficiency and accuracy of the substrate processing procedure (e.g., by identifying data points that fall within and / or outside the control limits).

[0033] In some embodiments, the extracted features, generated synthetic / engineering data, and statistical analysis may be used in association with the machine learning system 170 (e.g., for training, validating, and / or testing the machine learning model 190). Additionally and / or alternatively, the data preparation tool 116 may output data to the server 120 for use by any of the image enhancement tool 124, the embedding tool 126, the search tool 128, and the defect tool 130.

[0034] The data storage 140 may be a memory (e.g., random access memory), a drive (e.g., hard drive, flash drive), a database system, a cloud-based system, or another type of component or device capable of storing data. The data storage 140 may store one or more historical data 142, including an image repository 144 that houses images of previous processes (e.g., substrate defects) and corresponding vectorized image features and metadata 146. The vectorized image data may include feature vectors or embedding data representing imaging data (e.g., image frames, images acquired using the imaging tool 114).

[0035] Server 120 may include one or more computing devices, such as rack servers, router computers, server computers, personal computers, mainframe computers, portable computers, tablet computers, desktop computers, etc. Server 120 may include an image enhancement tool 124, an embedding tool 126, a search tool 128, and / or a defect tool 130. Server 120 includes a cloud server or a server capable of performing one or more cloud-based functions. For example, one or more operations of the image enhancement tool 124, the embedding tool 126, the search tool 128, and / or the defect tool 130 may be provided to a remote device (e.g., client device 150) using a cloud environment.

[0036] The image enhancement tool 124 receives image data (e.g., image frames) indicating substrate processing defects and performs image enhancement on the received image data. In some embodiments, the image enhancement tool 124 executes a filtering program, where processing logic processes the selection of the received image frames and removes image frames that do not meet specific criteria by identifying the characteristics of the images. For example, images of substrate defects may be included in an image set that includes graphs, data tables, and / or metrology-related images. The filtering program identifies which images depict substrate defects and removes images that do not directly indicate substrate processing defects.

[0037] In some embodiments, the image enhancement tool 124 performs a cropping program on one or more received image frames. Cropping is the removal of unwanted external areas from a photo or a shown image. This process typically consists of removing some peripheral areas of the image to remove irrelevant junk from the picture, improve its framing, change the aspect ratio, or highlight or isolate the object from its background. This can be performed by using image editing software and / or an algorithm that emulates an image editing program. For example, the image enhancement 124 may receive a selection of one or more image frames (e.g., from the metrology system 110 and / or the client device 150) and crop one or more image frames according to the selection (e.g., generating a second image frame by cropping a first image frame according to the selection).

[0038] In some embodiments, the image enhancement tool 124 performs a masking program on one or more received image frames. Image masking is a technique for isolating different parts of an image. For example, image masking may include compositing multiple images in a photo, hiding all or part of an application with selective adjustments, performing a cutout (e.g., removing the background), adjusting transparency, and / or the like. For example, the image enhancement 124 may receive a selection of one or more image frames (e.g., from the metrology system 110 and / or the client device 150) and mask one or more image frames according to the selection (e.g., generating a second image frame by cropping a first image frame according to the selection).

[0039] The embedding tool 126 receives image data including one or more image frames (e.g., from the metrology system 110) and determines embedding data representative of the image data. The embedding tool 126 includes a process for extracting features in the form of feature data (e.g., feature vectors) and / or generating synthetic / engineering data associated with the data measured by the imaging tool 114. In some embodiments, the embedding tool 116 can identify correlations, patterns, and / or anomalies in metrology or process performance data. Embedding is a relatively low-dimensional space into which high-dimensional representations (e.g., images) can be transformed. The embedding data (e.g., feature vectors) captures the semantics of the received image frames. The embedding tool 126 outputs the embedding data (e.g., for use by the search tool 128). The output of the embedding layer can be further passed to other machine learning techniques such as clustering, k-nearest neighbor analysis, etc.

[0040] The search tool 128 receives the embedding data (e.g., feature vectors, feature embeddings) and determines other feature embeddings (e.g., vectorized image features and metadata) corresponding to previously processed images (e.g., an image repository). The search tool 128 performs a proximity search between the received feature embeddings and one or more feature embeddings of the vectorized images and the features and metadata 146 of the historical data 142.

[0041] In some embodiments, the search tool 128 uses a vector search and / or nearest neighbor solution method to determine a set of similar images in the image repository 144. The search tool can identify the vector that is closest (e.g., most similar) to the received and / or provided feature vector.

[0042] In some embodiments, the search tool 128 scans the historical data 142 and retrieves the image repository 144 as well as the vectorized image features and metadata 146. The search tool 128 can adopt and index the vectorized image features to quickly resolve the feature vectors. In some embodiments, the search tool utilizes Euclidean distance and / or cosine similarity to determine the distance between the feature vectors.

[0043] The defect tool 130 receives one or more similar image frames indicating substrate processing defects. The defect tool 130 identifies substrate processing defects based on the selection of the similar image frames. The defect tool 130 identifies an anomalous instance of the manufacturing process based on a comparison between the current image and each of the selected similar image frames. In some embodiments, the defect tool 138 receives the similarity from the pattern mapping tool 137 and identifies anomalous instances based on the sample pattern.

[0044] The defect tool 130 can retrieve failure mode and effect analysis (FMEA) data. The FMEA data can include a list of known problems and root causes for a given device, and such known problems and root causes have known symptoms associated with each. The identified defects and / or similar images received by the defect tool 138 are applied to the list of known problems, and a report identifying the common causes of the identified defects is generated. For example, the defect tool 130 can determine a defect and identify the operation of a tool, machine, or manufacturing process corresponding to the identified defect.

[0045] In some embodiments, the defect tool 138 can be used with process dependency data to identify a tool, machine, or process in an upstream operation (e.g., an operation step that occurred prior to the current manufacturing step of the same manufacturing process) from the current machine operation being performed on the current sample. For example, the current sample may have recently undergone a first operation through a first machine. In some embodiments, the defect tool 130 can utilize a combination of process dependency data and failure mode and effect analysis data to find past operations of the sample, such as a second operation through a second machine or tool.

[0046] Once an abnormal instance is identified, the defect tool 130 can proceed by performing at least one of the following: changing at least one of the operation or the implementation of the process of the machine associated with the abnormal instance and / or providing a graphical user interface (GUI) that presents a visual indicator of the machine or process associated with the abnormal instance. The GUI can be sent over the network 160 and presented on the client device 150. In some embodiments, changing the operation or the implementation of the process of the machine can include sending instructions to the manufacturing execution system 102 to change the processing tool 104, the processing program 106, and / or the processing controller 108.

[0047] As previously described, some embodiments of the image enhancement tool 124, the embedding tool 126, the search tool 128, and / or the defect tool 130 can use machine learning models to perform the methods described. The associated machine learning models can be generated (e.g., trained, validated, and / or tested) using the machine learning system 170. The following example description of the machine learning system 170 will be described in the context of generating a machine learning model 190 associated with the embedding tool 126 using the machine learning system 170. However, it should be noted that this description is merely an example. Similar processing levels and methods can be used for the generation and execution of machine learning models associated with the image enhancement tool 124, the embedding tool 126, the search tool 128, and / or the defect tool 130, and the methods described can be performed using machine learning models.

[0048] The machine learning system 170 may include one or more computing devices, such as rack servers, router computers, server computers, personal computers, mainframe computers, portable computers, tablet computers, desktop computers, cloud computers, cloud servers, systems stored on one or more clouds, etc. The machine learning system 170 may include an embedding component 194 and a retrieval component 196. In some embodiments, the embedding component 194 may receive one or more image frames indicating substrate processing defects as input and use the historical data 142 and the trained machine learning model 190 to determine a feature embedding corresponding to the image frame. In some embodiments, the retrieval component 196 may use the trained machine learning model 190 to search the image repository 144 using the vectorized image features and the metadata 146.

[0049] In some embodiments, the machine learning system 170 further includes a server machine 172 and a server machine 180. The server machines 172 and 180 may be one or more computing devices (such as rack servers, router computers, server computers, personal computers, mainframe computers, portable computers, tablet computers, desktop computers, cloud computers, cloud servers, systems stored on one or more clouds, etc.), data storage (e.g., hard disk, in-memory database), network, software components, or hardware components.

[0050] The server machine 172 may include a dataset generator 174 that is capable of generating a dataset (e.g., a set of data inputs and a set of target outputs) to train, validate, or test a machine learning model. The dataset generator 174 may divide the historical data 142 into a training set (e.g., 60% of the historical data, or any other portion of the historical data), a validation set (e.g., 20% of the historical data, or some other portion of the historical data), and a test set (e.g., 20% of the historical data). In some embodiments, the dataset generator 174 generates multiple sets of training data. For example, one or more sets of training data may include each of the datasets (e.g., the training set, the validation set, and the test set).

[0051] Server machine 180 includes a training engine 182, a validation engine 184, and a testing engine 186. The training engine 182 is capable of training a machine learning model 190 using one or more images of the image repository 144, vectorized image features and metadata 146, and / or historical process result data 148 of the historical data 142 (of the data store 140). In some embodiments, the machine learning model 190 may be trained using one or more outputs of the data preparation tool 116, the image enhancement tool 124, the embedding tool 126, the search tool 128, and / or the defect tool 130. For example, the machine learning model 190 may be a hybrid machine learning model that uses image data and / or embedding features (such as feature extraction, mechanical modeling, and / or statistical modeling). The training engine 182 may generate multiple trained machine learning models 190, where each trained machine learning model 190 corresponds to a different set of features of each training set.

[0052] The validation engine 184 may determine the accuracy of each of the trained machine learning models 190 based on the corresponding set of features of each training set. The validation engine 184 may discard the trained machine learning models 190 that have an accuracy that does not meet the threshold accuracy. The testing engine 186 may determine the trained machine learning model 190 with the highest accuracy among all the trained machine learning models based on the test (and optionally, validation) set.

[0053] In some embodiments, training data is provided to train the machine learning model 190 such that the trained machine learning model can receive a new input with new image data indicating a new substrate processing defect. The new output may indicate a new feature embedding (e.g., a feature vector). In some embodiments, the training data may be further used such that the new output further includes a selection of similar feature vectors corresponding to images of similar substrate processing defects.

[0054] The machine learning model 190 may refer to a model created by the training engine 182 using a training set that includes data inputs and corresponding target outputs (image frames and corresponding vectorized image features and metadata). Patterns that map the data inputs to the target outputs in the dataset (e.g., identifying the connection between parts of the sensor data and the resulting chamber state) can be found and mappings that capture such patterns are provided to the machine learning model 190. The machine learning model 190 may use one or more of logistic regression, parsing, decision trees, or support vector machine (SVM). The machine learning may consist of single-stage linear or non-linear operations (e.g., SVM) and / or may be a neural network.

[0055] The embedding component 194 can provide current data (e.g., an image frame indicating a substrate processing defect) as an input to the trained machine learning model 190, and can run the trained machine learning model 190 on the input to obtain one or more outputs including a set of vectorized image features and metadata. The embedding component 194 is capable of identifying confidence data from the outputs, which confidence data indicates the confidence level of the predicted vectorized image features and metadata. In one non-limiting example, the confidence level is a real number between 0 and 1 (including 0 and 1), where 0 indicates no confidence in one or more chamber states, and 1 represents absolute confidence in the chamber states.

[0056] For purposes of illustration and not limitation, aspects of the present disclosure describe the training of machine learning models and the use of information regarding historical data 142 in the use of trained learning models. In other embodiments, heuristic models or rule-based models are used to determine chamber states.

[0057] In some embodiments, the functionality of the client device 150, the server 120, the data store 140, and the machine learning system 170 can be provided by a smaller number of machines compared to Figure 1 that shown. For example, in some embodiments, the server machines 172 and 180 can be integrated into a single machine, and in some other embodiments, the server machines 172, 180, and 192 can be integrated into a single machine. In some embodiments, the machine learning system 170 can be provided entirely or in part by the server 120.

[0058] Generally, functions described in one embodiment as being performed by the client device 150, the data store 140, the metrology system 110, the manufacturing system 102, and the machine learning system 170 can also be performed on the server 120 in other embodiments, if appropriate. Additionally, functionality attributed to a particular component can be performed by different or multiple components operating together.

[0059] In embodiments, a "user" can represent a single individual. However, other embodiments of the present disclosure encompass a "user" being an entity controlled by multiple users and / or automated sources. For example, a collection of independent users united as a group of administrators can be considered a "user".

[0060] Figure 2 is a block diagram showing a substrate defect image indexing and retrieval system 200 in which embodiments of the present disclosure can operate. The substrate defect image indexing and retrieval system 200 can include aspects and / or features of the system architecture 100.

[0061] As Figure 2As shown, the substrate defect indexing and retrieval system 200 receives an input image 202. The input image may include one or more image frames indicating substrate processing defects. The input image 202 may include an image of a substrate processing result having a substrate processing defect. For example, the input image 202 may include a scanning tunneling microscope (STM) or a scanning electron microscope (SEM) image.

[0062] As Figure 2 shown, the substrate defect indexing and retrieval system 200 includes image enhancement logic, including cropping logic 204, masking logic 206, and size detection logic 208. The cropping logic 204 performs a cropping procedure on one or more received image frames. Cropping is the removal of unwanted outer regions from a photograph or the presented image. The process typically consists of removing some of the peripheral regions of the image, for removing irrelevant junk from the picture, improving its framing, changing the aspect ratio, or highlighting or isolating the object from its background. This can be performed by using image editing software and / or algorithms that emulate image editing procedures. For example, the cropping logic 204 may receive a selection of one or more image frames and crop the one or more image frames according to the selection (e.g., crop the first image frame based on the selection to produce a second image frame).

[0063] The masking logic 206 performs a masking procedure on one or more received image frames. Image masking is a technique for isolating different parts of an image. For example, image masking may include photo-compositing multiple images, hiding all or part of an image, applying selective adjustments, performing shearing (e.g., removing the background), adjusting transparency, and / or the like. For example, the masking logic 206 may receive a selection of one or more image frames and mask the one or more image frames according to the selection (e.g., crop the first image frame based on the selection to produce a second image frame).

[0064] The size detection logic 208 processes the input image 202 and extracts content from the input image 202 that indicates an image scaling factor associated with the input image 202. The image scaling factor indicates the relative size of the depiction within the image relative to the size of the image. For example, the image scaling factor may include image scaling, magnification factor, the indicated size of the depicted object, and the like. The size detection logic 208 extracts size identification information (e.g., text) and determines the size of the depicted defect. The defect size is further passed to one or more downstream processes (e.g., embedding logic 210, search logic 212, defect detection logic 214, and the like). Further details regarding the size detection logic 208 are discussed Figure 3 in connection therewith.

[0065] As Figure 2As shown, the substrate defect indexing and retrieval system 200 includes embedded logic 210. The embedded logic 210 receives image data including an input image 202. The embedded logic 210 includes process methods for extracting features in the form of feature data (e.g., feature vectors) and / or generating synthetic / engineering data associated with data measured by an imaging tool. In some embodiments, the embedded logic 210 may identify correlations, patterns, and / or anomalies in metrology or process performance data. Embedding is a relatively low-dimensional space into which high-dimensional representations (e.g., images) can be transformed. The embedded data (e.g., feature vectors) captures the semantics of the received image frames. The embedded logic 210 outputs the embedded data (e.g., for use by the search tool 128). The output of the embedding layer can be further fed into other machine learning techniques such as clustering, k-nearest neighbor analysis, etc.

[0066] As Figure 2 shown, the substrate defect indexing and retrieval system 200 includes search logic 212. The search logic 212 receives the embedded data (e.g., feature vectors, feature embeddings) and determines other feature embeddings (e.g., vectorized image features and metadata) corresponding to previously processed images (e.g., an image repository). The search logic 212 performs a proximity search between the received feature embeddings and one or more feature embeddings of the vectorized images and the features and metadata of previously processed images of other substrate processing defects.

[0067] In some embodiments, the search logic 212 uses vector search and / or nearest neighbor solution methods to determine a set of similar images in the image repository. The search tool can identify the vector that is closest (e.g., most similar) to the received and / or provided feature vector.

[0068] In some embodiments, the search logic 212 scans the data structure and retrieves the image repository as well as the corresponding vectorized image features and metadata. The search logic 212 can employ an indexing method for the vectorized image features to quickly resolve the feature vectors. In some embodiments, the search tool uses Euclidean distance and / or cosine similarity to determine the distance between the feature vectors.

[0069] As Figure 2 shown, the substrate defect indexing and retrieval system 200 includes defect detection logic 214. The defect detection logic 214 receives one or more similar image frames indicating substrate processing defects. The defect detection logic 214 identifies substrate processing defects based on the selection of the similar image frames. The defect detection logic 214 identifies abnormal instances of the manufacturing process based on a comparison between the current image and each of the selected similar image frames. In some embodiments, the defect detection logic 214 receives the similar images from the search logic 212 and identifies the abnormal instances based on the similar images.

[0070] The defect detection logic 214 can retrieve Failure Mode and Effects Analysis (FMEA) data. The FMEA data can include a list of known problems and root causes for a given device, and such known problems and root causes have known symptoms associated with each. The identified defects and / or similar images received by the defect detection logic 214 are applied to the list of known problems, and a report identifying the common causes of the identified defects is generated. For example, the defect detection tool 214 can determine a defect and identify the operation of a tool, machine, or manufacturing process corresponding to the identified defect.

[0071] In some embodiments, the defect detection logic 214 can be used with process dependency data to identify a tool, machine, or process in an upstream operation (e.g., an operation step that occurred before the current manufacturing step of the same manufacturing process) from the current machine operation being performed on the current sample. For example, the current sample may have recently undergone a first operation through a first machine. In some embodiments, the defect detection tool 214 can utilize a combination of process dependency data and failure mode and effects analysis data to find past operations of the sample, such as a second operation through a second machine or tool.

[0072] Once an abnormal instance is identified, the defect tool 130 can proceed by performing at least one of the following: changing at least one of the operation or the implementation of the process of the machine associated with the abnormal instance and / or providing a graphical user interface (GUI) 216 that presents a visual indicator of the machine or process associated with the abnormal instance. The GUI can be sent over the network 160 and presented on the client device 150. In some embodiments, changing the operation or the implementation of the process of the machine can include sending instructions to the manufacturing execution system 102 to change the processing entities (e.g., Figure 1 processing tool 104, processing program 106, and / or processing controller 108) of the manufacturing system 102.

[0073] In some embodiments, the search logic 212 outputs a selection of similar vectorized image features or an image having vectorized image features similar to the input image 202. The image can be sent directly to the GUI 216 without determining a specific defect. For example, the GUI 216 can include a diagram showing the location of an image that has been determined to be similar to the input image 202.

[0074] Figure 3 A block diagram of a defect size determination system 300 in accordance with aspects of the present disclosure is shown. One or more features discussed in connection with Figure 3 can be performed by the size detection logic 208 of Figure 2 As Figure 3As shown, the defect size determination system 300 may include an input image receiving unit 302, a line masking logic unit 304, a line removal logic unit 306, a text recognition logic unit 308, and / or a post-processing logic unit 310. The input image 302 may include one or more image frames indicating substrate processing defects. The input image 302 may include an image of the substrate processing result with substrate processing defects. For example, the input image 302 may include a scanning tunneling microscope (STM) or a scanning electron microscope (SEM) image.

[0075] The line masking logic 304 determines the placement of edges within the image frame. The line masking logic may determine the boundaries of the image frame or global edges. For example, some images may include vertical or horizontal lines on the edges of the image frame. In some embodiments, the line masking logic 304 performs gamma correction threshold processing to identify one or more edges within the image frame. For example, vertical and / or horizontal lines may be used within the image frame to identify data stored in the image, such as, for example, text overlaid on the image. In some embodiments, the line masking logic 304 uses segmentation to perform localization and generate a mask for the repair of the image frame.

[0076] The line removal logic 306 receives the line masking data from the line masking logic 304. The line removal tool applies an image enhancement method (e.g., cropping, masking) to a portion of the image frame based on the detected lines. For example, the image may include a brand logo or other artificial markings identified using the line masking logic 304. The line removal logic 306 may apply a mask to remove the lines detected by the line masking logic 304.

[0077] The text recognition logic 308 identifies the text within the image frame. For example, the image may include a relative image scaling factor, such as a scale bar indicating the distance represented on the image. The text recognition logic 308 may identify the numbers and units for a given image scaling factor (e.g., magnification, relative depiction distance, etc.). The text recognition logic 308 may isolate the text from the remainder of the input image 302.

[0078] The post-processing logic 310 takes actions based on the identified text. For example, the post-processing log may perform data cleaning and normalization. This may include providing a cleaned depiction of the area of the input image 302 identified by the text recognition logic 308. The post-processing logic 310 may further include making the size data available for other processes of the system. In some embodiments, the post-processing logic 310 determines the size of the defect based on the scaling and provides the size of the defect (e.g., as metadata) for further embedding and / or classification procedures related to the input image 302.

[0079] Figure 4FIG. 400 is a block diagram illustrating a process 400 for training a machine learning model to produce an output, according to some embodiments. Process 400 includes receiving training data in the form of input image 402. Input image 402 may include one or more image frames indicative of substrate processing defects. Input image 402 may include an image of a substrate processing result having a substrate processing defect. For example, input image 402 may include a scanning tunneling microscope (STM) or scanning electron microscope (SEM) image.

[0080] As Figure 4 shown, process 400 includes cropping augmentation of input image 402. A copy of input image 402 may undergo local cropping 404 and a copy of input image 402 may undergo global cropping 406. Local cropping may be associated with a selection of input image 402 set within a selection of input image 402. The locally cropped image is used as an input to student model 408. The globally cropped image is used as an input to teacher model 410. The student model outputs student model prediction 412. The student model prediction 412 and / or the teacher model prediction 414 includes a profile (e.g., a distribution) indicative of different embedding vectors corresponding to the input image 402 and a corresponding confidence level (e.g., probability) associated with each of the outputs.

[0081] In operation, an example training procedure may include processing logic receiving a first image frame (e.g., input image 402) indicative of a substrate processing defect. The processing logic further produces a second image frame by cropping (e.g., global cropping 406) a first region of the first image frame. The processing logic further produces a third image frame by cropping (e.g., local cropping 404) a second region of the first image frame. The first region (associated with global cropping 406) encompasses the second region (associated with local cropping 404). The processing logic uses the second image frame as an input to a first machine learning (ML) model (e.g., teacher model 410). The processing logic obtains one or more outputs of the first ML model. The one or more outputs (teacher model prediction 414) indicate a first feature vector corresponding to the second image frame. The processing logic uses the third image frame as an input to a second ML model (e.g., student model 408). The processing logic obtains one or more outputs (e.g., student model prediction 412) of the second ML model. The one or more outputs indicate a second feature vector corresponding to the third image frame. The processing logic updates (e.g., model tuning 416) one or more parameters of at least one of the first ML model or the second ML model based on a comparison between the one or more outputs of the first ML model and the one or more outputs of the second ML model.

[0082] In some embodiments, the student model 408 can be a machine learning model that is similar to the trained teacher model 410 but contains fewer layers and / or nodes compared to each of the trained teacher models 410, resulting in a more compressed machine learning model. In some embodiments, multiple student models 408 can be trained, where each student model 408 can be trained to predict the embedding vectors of different subsets of the input image 402 in an input image cluster, and the teacher model 410 is trained to output an error prediction for the input image cluster. In some embodiments, different learning models 408 can be trained for each cluster of input images 402 and / or substrate processing defects. Each of the teacher models can then be used to train multiple student models.

[0083] In some embodiments, model tuning 416 includes determining the error or outcome of a loss function, such as, for example, the hierarchical loss backpropagated to the student model 408. The hierarchical loss represents the difference between the probability of an error (from the student model) and the probability of a perceived truth (from the teaching model). For example, if the teacher model 410 determines that a first prediction has a first probability and the student model 408 determines that the first prediction has a second probability, the error is related to the difference between the two probabilities. In some embodiments, the hierarchical loss function can be a categorical cross-entropy function, a Kullback-Leibler divergence function, or any suitable loss function.

[0084] In some embodiments, training can be performed by inputting the input image 402 into the machine learning model one at a time. In some embodiments, after one or more rounds of training, the processing logic can determine whether a stopping criterion has been met. The stopping criterion can be a target accuracy level, a target number of processed images from the training dataset, a target change amount of parameters at one or more previous data points, a combination thereof, and / or other criteria. In one embodiment, the stopping criterion is met when at least a minimum number of data points have been processed and the loss value has stabilized and / or stopped decreasing. The loss value can represent the total error in the machine learning model. For example, the loss value can represent the sum of the increments between the modeled value and the actual value. In one embodiment, the stopping criterion is met if the accuracy of the machine learning model has stopped improving. If the stopping criterion has not been met, further training is performed. If the stopping criterion has been met, training can be completed. Once the machine learning model has been trained, the retained portion of the training dataset can be used to test the model.

[0085] Figure 5Illustrated is a model training workflow 505 and a model application workflow 517 for substrate processing result prediction according to aspects of the present disclosure. In some embodiments, the model training workflow 505 may be executed at a server, which may or may not include a processing defect image indexing and retrieval application, and provide the trained model to a processing defect image indexing and retrieval application that can execute the model application workflow 517. The model training workflow 505 and the model application workflow 517 may be executed by processing logic executed by a processor of a computing device (e.g., Figure 1 server 120). One or more of such workflows 505, 517 may be implemented, for example, by processing devices implemented by one or more machine learning modules and / or other software and / or firmware executed on the processing device.

[0086] The model training workflow 505 is used to train one or more machine learning models (e.g., regression models, boosted regression models, principal component analysis models, deep learning models, vision transformers) to perform one or more tasks (e.g., feature extraction, image retrieval) associated with a processing result predictor such as determination, prediction, modification, etc. The model application workflow 517 is used to apply one or more trained machine learning models to perform tasks such as determination and / or tuning of image data (e.g., image frames indicating substrate processing defects). One or more machine learning models may receive image data (e.g., image frames indicating substrate processing defects).

[0087] Various machine learning outputs are described herein. A specific number and arrangement of machine learning models are described and illustrated. However, it should be understood that the number and type of machine learning models used and the arrangement of such machine learning models may be modified to achieve the same or similar end results. Thus, the arrangement of machine learning models described and illustrated is merely an example and should not be construed as limiting.

[0088] In some embodiments, one or more machine learning models are trained to perform one or more of the following tasks. Each task may be performed by a separate machine learning model. Alternatively, a single machine learning model may perform each or a subset of the tasks. Additionally or alternatively, different machine learning models may be trained to perform different combinations of tasks. In one example, one or several machine learning models may be trained, where the trained machine learning (ML) model is a single shared neural network having multiple shared layers and multiple higher-order different output layers, where each of the output layers outputs a different prediction, classification, identification, etc. The tasks that one or more trained machine learning models may be trained to perform are as follows:

[0089] a. Feature Extractor 567 - As previously discussed, the feature extractor receives image data (e.g., raw and / or enhanced image frames indicative of substrate processing defects). The feature extractor includes process methods for extracting features in the form of feature data (e.g., feature vectors) and / or generating synthetic / engineering data associated with data measured by the imaging tool. In some embodiments, the feature extractor may identify correlations, patterns, and / or anomalies in metrology or process performance data. Embedding is a relatively low-dimensional space into which high-dimensional representations (e.g., images) can be transformed. Embedded data (e.g., feature vectors) capture the semantics of the received image frames.

[0090] b. Image Retriever 564 - The image retriever receives embedded data (e.g., feature vectors, feature embeddings) and determines other feature embeddings (e.g., vectorized image features and metadata) corresponding to previously processed images (e.g., an image repository). The image retriever performs a proximity search between the received feature embeddings and one or more feature embeddings of the vectorized images and features and metadata. In some embodiments, the image retriever uses vector search and / or nearest neighbor solution methods to determine a set of similar images (e.g., within a threshold proximity). In some embodiments, the image retriever utilizes Euclidean distance and / or cosine similarity to determine the distance between feature vectors.

[0091] To complete training, the processing logic inputs the training dataset 536 into one or more untrained machine learning models. (See Figure 4 , for further details on training.) The machine learning model can be initialized before the first input is input into the machine learning model. The processing logic trains the untrained machine learning model based on the training dataset to generate one or more trained machine learning models that perform the various operations as described above.

[0092] Once one or more trained machine learning models 538 are generated, these machine learning models can be stored in the model memory 545 and added to the processing defect image indexing and retrieval application. The processing defect image indexing and retrieval application can then use one or more trained ML models 538 and additional processing logic to implement an automated model, where user manual input of information is minimized or even eliminated in some instances.

[0093] For model application workflow 517, according to one embodiment, input data 562 (e.g., an image frame indicating a substrate processing defect) can be used as the input to a feature extractor 567, which may include a trained machine learning model. Based on the input data 562, the feature extractor 567 outputs feature data 569 and metadata representing the input data 562. The feature data 569 is input into an image retriever 564, which may include a trained machine learning model. Based on the feature data 569, the image retriever 564 identifies one or more other images (e.g., similar image data 566) identified as similar to the input data 562 through the feature data 569.

[0094] Figures 6A to 6B A model architecture 600 for substrate defect image indexing and retrieval according to aspects of the present disclosure is shown. One or more ML models can be performed using the model architecture 600. Generally, the model architecture 600 consists of an embedding layer 608, an encoder 614, and a final header classifier 616. Initially, the image is subdivided into non-overlapping patches. Each patch is treated by the architecture as an independent token. For an image of size c×h×w (where h is the height, w is the width, and c represents the number of channels). Patches of size p×p are extracted for each dimension. This forms a sequence of patches (x1, x2, …, x n ), where n = . In some embodiments, the patch size p is selected to be 16×16 or 32×32.

[0095] As Figure 6A shown, at 604, the input image 602 is divided (e.g., flattened) into independent image patches. For example, the input image 602 is divided into a fixed number of patches or embedding tokens of equal size. The input image can be transformed (e.g., flattened) into a sequence of token embeddings 606 indicating the content of the image patches. In some embodiments, the model architecture uses a constant latent vector size n through all its layers, and the patches are flattened and mapped to n dimensions using a linear embedding layer 608 with a trainable linear projection.

[0096] Before feeding the patch sequence into the encoder 614, it is linearly projected into a vector of model dimension d using a learned embedding matrix. The embedding representation is then concatenated with a learnable classification token for performing the classification task. The embedded image patches are treated by the transformer as a set of patches without any concept of order. To maintain the spatial arrangement of the patches as in the original image, positional information 610 is encoded and appended to the patch representation 612 (e.g., the linear embedding representing the content of the corresponding image patch). The resulting embedding sequence for the patch with token 0 is given by:

[0097]

[0098] The resulting sequence embedding patch z0 is passed to the Transformer encoder 614. As Figure 6B shown, the encoder 614 consists of L equal layers. Each has two main sub-components: (1) a multihead self-attention block (MSA) 656, and (2) a feed-forward dense block (MLP) 660. Each of the two sub-components of the encoder employs a residual skip connection and is before a normalization layer (e.g., normalization layer 654 and normalization layer 658). At the last layer of the encoder 614, we take the first element in the sequence and pass it to an external head classifier for predicting the class token represented by

[0099] The MSA block 656 in the encoder 614 is the central component of the Transformer. The MSA block 656 determines the relative importance of a single patch embedding relative to the other embeddings in the sequence. This block has four layers: a linear layer, a self-attention layer, a concatenation layer (concatenating the outputs of multiple attention heads), and a final linear layer. At a high level, attention can be represented by attention weights, which are calculated by taking a weighted sum of all values of the sequence z. The MSA block employs an attention function that maps query and key-value pairs to a set of outputs, where the query, key, value, and output are all vectors. The output is calculated as a weighted sum of the values, where the weight assigned to each value is calculated by a compatibility function of the query with the corresponding key. The results of all attention heads are concatenated together and then projected to the desired dimension by a feed-forward layer with learnable weights. The MLP 616 makes a class prediction 618 based on the data received from the encoder 614.

[0100] Figure 7 FIG. depicts a flowchart of an example method 700 for substrate defect image indexing and retrieval according to some embodiments of the present disclosure. The method 700 is executed by processing logic that may include hardware (e.g., circuitry, dedicated logic, etc.), software (such as running on a general-purpose computer system or a dedicated machine), or any combination thereof. In one embodiment, the method uses Figure 1 server 120 and a trained machine learning model 190 to execute, while in some other embodiments, Figure 7 one or more of the

[0101] ​Method 700 may include receiving image data (e.g., associated with substrate processing defects), and processing the received image data using a trained machine learning model 190. The trained model may be configured to generate a feature embedding of the image based on the image data. Method 700 further includes identifying similar images that may show similar defects.

[0102] At block 702, the processing logic stores, in a data storage device, a plurality of feature vectors that represent previously processed image frames corresponding to various substrate processing defects. At block 704, the processing logic receives first image data that includes one or more image frames indicating a first substrate processing defect. The first image data may include one or more image frames indicating a substrate processing defect. The input image may include an image of a substrate processing result having a substrate processing defect. For example, the input image may include a scanning tunneling microscope (STM) or a scanning electron microscope (SEM) image.

[0103] In some embodiments, the processing logic receives a first selection of a first image frame of the first image data and generates a second image frame by cropping a region of the first image frame based on the first selection. The first feature vector is determined using the second image frame. In some embodiments, the processing logic receives a first selection of a first image frame of the first image data. The processing logic generates a second image frame by masking a region of the first image frame based on the first selection. The first feature vector is determined using the second image frame.

[0104] In some embodiments, the processing logic extracts a selection of text from a first image frame of the first image data. The processing logic further determines an image scaling factor associated with the first image frame. The processing logic further determines a size associated with the first substrate processing defect based on the image scaling factor. A selection of the plurality of feature vectors is determined using the size.

[0105] At block 706, the processing logic determines a first feature vector corresponding to the first image data. In some embodiments, the processing logic further divides a first image frame of the first image data into a set of image patches corresponding to the first image frame. The processing logic further determines a set of linear embeddings of the content of each image patch corresponding to the set of image patches. The processing logic further determines a set of position embeddings of the relative position of each of the image patches corresponding to the set of image patches. The first feature vector is determined based on the set of linear embeddings and the set of position embeddings.

[0106] At block 708, the processing logic determines a selection of the plurality of feature vectors based on proximity between the first feature vector and each of the selected ones of the plurality of feature vectors. At block 710, the processing logic determines second image data that includes one or more image frames corresponding to the selection of the plurality of embedding vectors. The processing logic receives embedding data (e.g., feature vectors, feature embeddings) and determines additional feature embeddings (e.g., vectorized image features and metadata) corresponding to previously processed images (e.g., an image repository). The processing logic performs a proximity search between the received feature embeddings and one or more feature embeddings of the vectorized image and the features and metadata. In some embodiments, the processing logic uses vector search and / or nearest neighbor solution methods to determine a set of images close (e.g., within a threshold proximity) to the image. In some embodiments, the processing logic utilizes Euclidean distance and / or cosine similarity to determine distances between feature vectors.

[0107] At block 712, the processing logic optionally performs an action based on determining the second image data. In some embodiments, the processing logic optionally prepares the second image data for presentation on a graphical user interface (GUI). For example, the second image data is a set of image frames similar to the first image frame and indicates similar substrate processing defects. In another example, the second image data can be displayed on the GUI by showing an indication of the first substrate processing defect.

[0108] In some embodiments, the processing logic optionally changes the operation of the processing chamber and / or processing tool based on the second image data. For example, the processing logic can determine a substrate processing defect based on the second image data. The processing logic further determines an anomalous instance of the manufacturing process associated with the first substrate processing defect. The processing logic can further send instructions (e.g., to perform a corrective action associated with the manufacturing process equipment) to one or more process controllers to change one or more operations of the processing apparatus associated with the anomalous instance (e.g., change a processing recipe and / or processing parameters, final substrate processing of one or more processing tools and / or processing chambers, initiate preventive maintenance associated with one or more processing chambers and / or processing tools, etc.).

[0109] Figure 8 A block diagram of an example computing device 800 operating in accordance with one or more aspects of the present disclosure is depicted. In various illustrative examples, the various components of the computing device 800 can represent Figure 1 the various components of the client device 150, metrology system 110, server 120, data store 140, and machine learning system 170 shown in

[0110] The exemplary computing device 800 may be connected to other processing devices in a LAN, an intranet, an extranet, and / or the Internet. The computing device 800 may operate in the capacity of a server in a client-server network environment. The computing device 800 may be a personal computer (PC), a set-top box (STB), a server, a network router, a switch or a bridge, or any device capable of executing a set of instructions (sequentially or otherwise) that specify actions to be taken by that device. Additionally, although only a single exemplary computing device is shown, the term "computer" shall also be considered to include any collection of computers that individually or jointly execute a set of instructions (or multiple sets of instructions) to perform any one or more of the methods discussed herein.

[0111] The exemplary computing device 800 may include a processing device 802 (also referred to as a processor or CPU), a main memory 804 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM), etc.), a static memory 806 (e.g., flash memory, static random access memory (SRAM), etc.), and an auxiliary memory (e.g., a data storage device 818), which may communicate with each other via a bus 830.

[0112] The processing device 802 represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, or the like. More specifically, the processing device 802 may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. The processing device 802 may also be one or more special-purpose processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. According to one or more aspects of the present disclosure, the processing device 802 may be configured to execute instructions implementing Figure 7 the method 700 shown in

[0113] The example computing device 800 may further include a network interface device 808 that may be communicatively coupled to a network 820. The example computing device 800 may further include a video display 810 (e.g., a liquid crystal display (LCD), a touch screen, or a cathode ray tube (CRT)), an alphanumeric input device 812 (e.g., a keyboard), a cursor control device 814 (e.g., a mouse), and an audio signal generating device 816 (e.g., a speaker).

[0114] The data storage device 818 may include a machine-readable storage medium (or more specifically, a non-transitory machine-readable storage medium) 828 having one or more sets of executable instructions 822 stored thereon. In accordance with one or more aspects of the present disclosure, the executable instructions 822 may include executable instructions associated with performing Figure 7 the method 700 shown therein.

[0115] The executable instructions 822 may also be fully or at least partially resident in the main memory 804 and / or the processing device 802 during their execution by the example computing device 800, the main memory 804, and the processing device 802, which also constitutes a computer-readable storage medium. The executable instructions 822 may further be sent or received over a network via the network interface device 808.

[0116] Although the computer-readable storage medium 828 is shown as a single medium in Figure 8 , the term "computer-readable storage medium" should be considered to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) having one or more sets of stored operational instructions. The term "computer-readable storage medium" should also be considered to include any medium that is capable of storing or encoding a set of instructions for execution by a machine and that causes the machine to perform any one or more of the methods described herein. The term "computer-readable storage medium" should thus be considered to include, but not be limited to, solid-state memory, as well as optical and magnetic media.

[0117] Some portions of the detailed descriptions above have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. For the sake of generality, it has proven convenient at times to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0118] However, it should be borne in mind that all such and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically stated otherwise, as will be apparent from the following discussion, it will be understood that throughout the description, discussions using terms such as "identifying," "determining," "storing," "adjusting," "causing," "returning," "comparing," "creating," "stopping," "loading," "copying," "throwing," "replacing," "executing," or the like, refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities within the memory or registers of the computer system or other such information storage, transmission, or display devices.

[0119] Examples of the present disclosure also relate to an apparatus for performing the methods described herein. This apparatus may be specially constructed for the required purposes, or it may be a general-purpose computer system selectively programmed by a computer program stored in the computer system. Such a computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk, including optical disks, compact disk read-only memory (CD-ROM), and magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic disk storage media, optical storage media, flash memory devices, other types of machine-accessible storage media, or any type of medium suitable for storing electronic instructions, each coupled to the computer system bus.

[0120] The methods and displays provided herein are not inherently related to any particular computer or other device. A variety of general-purpose systems may be used with the programs in accordance with the teachings herein, or it may prove convenient to construct more specialized devices to perform the desired method steps. The structure required for various such systems will be apparent as described hereinafter. Additionally, the scope of the present disclosure is not limited to any particular programming language. It will be understood that various programming languages may be used to implement the teachings of the present disclosure.

[0121] It will be understood that the foregoing description is intended to be illustrative and not restrictive. After reading and understanding the foregoing description, many other implementation examples will be apparent to those skilled in the art. Although the present disclosure describes specific examples, it will be recognized that the systems and methods of the present disclosure are not limited to the examples described herein, but may be practiced with modifications within the scope of the appended claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. Accordingly, the scope of the present disclosure should be determined with reference to the appended claims along with the full scope of equivalents to such claims.

Claims

1. A method, the method comprising: Storing, in a data storage device, a plurality of feature vectors that represent image frames of previous processing corresponding to various substrate processing defects; Receiving, by a processing device, first image data that includes one or more image frames indicating a first substrate processing defect; Determining, by the processing device, a first feature vector corresponding to the first image data; Determining, by the processing device, the selection of the plurality of feature vectors based on a proximity between the first feature vector and each of the selected plurality of feature vectors; Determining, by the processing device, second image data that includes one or more image frames corresponding to the selection of the plurality of embedded vectors; Performing, by the processing device, an action based on determining the second image data.

2. The method of claim 1, the method further comprising: Identifying, by the processing device, the first substrate processing defect based on the second image data, wherein the action is further based on an identification of the first substrate processing defect.

3. The method of claim 2, the method further comprising: Identifying, by the processing device, an abnormal instance of a manufacturing process associated with the first substrate processing defect; And Causing, by the processing device, a correction action associated with manufacturing processing equipment to be performed based on the abnormal instance.

4. The method of claim 1, the method further comprising: Preparing, by the processing device, one or more image frames of the second image data for presentation on a graphical user interface (GUI).

5. The method of claim 1, the method further comprising: Receiving, by the processing device, a first selection of a first image frame of the first image data; And Generating, by the processing device, a second image frame by cropping a region of the first image frame based on the first selection, wherein the first feature vector is determined using the second image frame.

6. The method of claim 1, the method further comprising: Receiving, by the processing device, a first selection of a first image frame of the first image data; And Generating, by the processing device, a second image frame by masking a region of the first image frame based on the first selection, wherein the first feature vector is determined using the second image frame.

7. The method of claim 1, the method further comprising: Extracting, by the processing device, a selection of text from a first image frame of the first image data. Determining, by the processing device, an image scaling factor associated with the first image frame; And Determining, by the processing device, a size associated with the first substrate processing defect based on the image scaling factor; wherein the selection of the plurality of feature vectors is further determined using the size.

8. The method of claim 1, the method further comprising: Dividing, by the processing device, a first image frame of the first image data into a set of image patches corresponding to the first image frame; Determine, by the processing device, a set of linear embeddings corresponding to the content of each image patch of the set of image patches; Determine, by the processing device, a set of position embeddings corresponding to the relative position of each of the image patches of the set of image patches, wherein the first feature vector is determined based on the set of linear embeddings and the set of position embeddings.

9. A method, the method comprising: Receive, by a processing device, a first image frame indicative of a substrate processing defect; Generate, by the processing device, a second image frame by cropping a first region of the first image frame; Generate, by the processing device, a third image frame by cropping a second region of the first image frame, wherein the first region includes the second region; Use the second image frame as an input to a first machine learning (ML) model; Obtain one or more outputs of the first ML model, the one or more outputs indicative of a first feature vector corresponding to the second image frame; Use the third image frame as an input to a second ML model; Obtain one or more outputs of the second ML model, the one or more outputs indicative of a second feature vector corresponding to the third image frame; Update one or more parameters of at least one of the first ML model or the second ML model based on a comparison between the one or more outputs of the first ML model and the one or more outputs of the second ML model.

10. The method of claim 9, wherein: The one or more outputs of the first ML model further indicate a confidence level associated with the first feature vector; and The one or more outputs of the second ML model further indicate a confidence level associated with the second feature vector.

11. The method of claim 9, the method further comprising: Use the first ML model to divide the second image frame into a set of image patches corresponding to the second image frame; Use the first ML model to determine a set of linear embeddings corresponding to the content of each image patch of the set of image patches; Use the first ML model to determine a set of position embeddings corresponding to the positions of the corresponding image patches of the set of image patches, wherein the first feature vector is determined based on the set of linear embeddings and the set of position embeddings.

12. The method of claim 9, wherein: The one or more outputs of the first ML model include (i) a first set of predictions and (ii) a first set of probabilities, each prediction corresponding to the first set of predictions; and The one or more outputs of the second ML model include (i) a second set of predictions and (ii) a second set of probabilities, each prediction corresponding to the second set of predictions.

13. The method of claim 9, wherein at least one of the first ML model or the second ML model includes a vision transformer (ViT).

14. A non-transitory machine-readable storage medium, the non-transitory machine-readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to perform operations including the following steps: Store a plurality of feature vectors in a data storage device, the feature vectors representing image frames of previous processing corresponding to various substrate processing defects; Receive first image data, the first image data including one or more image frames indicating a first substrate processing defect; Determine a first feature vector corresponding to the first image data; Determine the selection of the plurality of feature vectors based on a proximity between the first feature vector and each of the selected ones of the plurality of feature vectors; Determine second image data, the second image data including one or more image frames corresponding to the selection of the plurality of embedded vectors; Perform an action based on determining the second image data.

15. The non-transitory machine-readable storage medium of claim 14, the operation further comprising: Identifying the first substrate processing defect based on the second image data, wherein the action is further based on an identification of the first substrate processing defect.

16. The non-transitory machine-readable storage medium of claim 15, the operation further comprising: Identifying an abnormal instance of a manufacturing process associated with the first substrate processing defect; and Causing a correction action associated with a manufacturing processing device to be performed based on the abnormal instance.

17. The non-transitory machine-readable storage medium of claim 14, the operation further comprising: Preparing one or more image frames of the second image data for presentation on a graphical user interface (GUI).

18. The non-transitory machine-readable storage medium of claim 14, the operation further comprising: Receiving a first selection of a first image frame of the first image data; and Generating a second image frame by cropping or masking at least one of regions of the first image frame based on the first selection, wherein the first feature vector is determined using the second image frame.

19. The non-transitory machine-readable storage medium of claim 14, the operation further comprising: Extracting a selection of text from a first image frame of the first image data; Determining an image scaling factor associated with the first image frame; and Determining a size associated with the first substrate processing defect based on the image scaling factor, wherein the selection of the plurality of feature vectors is further determined using the size.

20. The non-transitory machine-readable storage medium of claim 14, the operation further comprising: Dividing a first image frame of the first image data into a set of image patches corresponding to the first image frame; Determining a set of linear embeddings of the content of each image patch corresponding to the set of image patches; Determining a set of position embeddings of the position of each of the image patches corresponding to the set of image patches, wherein the first feature vector is determined based on the set of linear embeddings and the set of position embeddings.