Automatic measurement of similar structures
By training machine learning models and using non-rigid registration technology, the position to be measured in structurally similar images is automatically determined, which solves the time-consuming, labor-intensive and inaccurate problems of manual measurement in the prior art, and improves the efficiency and accuracy of measurement.
Patent Information
- Application Number
- CN202380033780.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-13
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art requires manual marking of the position to be measured in each image when performing measurements between images of similar structures, resulting in time-consuming and labor-intensive and error-prone, and risk of inconsistent and inaccurate measurements.
By using a training dataset including historical image data of the structure to train machine learning models, non-rigid registration techniques are used to automatically determine the position to be measured in the image, reducing manual intervention.
Automatic measurement of position in images with similar structures is achieved, improving the efficiency and accuracy of measurement, and reducing the workload and error rate of technicians.
Smart Images

Figure CN120019408A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to image registration. More particularly, the present disclosure relates to deformable image registration. Background Art
[0002] Manufacturing equipment is used to produce products. For example, substrate processing equipment is used to produce substrates (e.g., wafers, semiconductors). The manufactured substrates have properties that can be measured. Products will be produced with specific structures suitable for the target application. Measurements can be made on products with similar structures. Measurements can be made on different images that share similar structures (e.g., images of manufactured substrates). Summary of the invention
[0003] The following is a simplified overview of the present disclosure in order to provide a basic understanding of some aspects of the present disclosure. The overview is not an extensive review of the present disclosure. It is neither intended to identify the key or important elements of the present disclosure, nor to describe any scope of a specific implementation of the present disclosure or any scope of the claims. Its sole purpose is to present some concepts of the present disclosure in a simplified form as a preface to a more detailed description presented later.
[0004] Aspects of the present disclosure include a method comprising training a machine learning model using a training data set comprising historical image data of a structure, the historical image data comprising multiple images of the structure, wherein the training comprises marking regions of a source image of the multiple images to create a source mask for the source image. The training further comprises providing a source image and a target image of the multiple images as training inputs to the machine learning model. The training further comprises receiving a deformation field between the source image and the target image as an output from the machine learning model. The training further comprises applying the deformation field to the source mask to create a distorted source mask. The training further comprises calculating a loss associated with the distorted source mask. The training further comprises updating a weight of the machine learning model based on the loss.
[0005] Further aspects of the present disclosure include a non-transitory computer-readable storage medium comprising instructions that, when executed by a processing device operably coupled to a memory, perform operations. The operations include receiving an input comprising a template image and a test image, wherein the template image comprises markings (e.g., manually placed) at first and second locations associated with a first measurement of the template image. The operations further include providing the template image and the test image as input to a trained machine learning model, wherein the trained machine learning model outputs a deformation field based on a non-rigid registration of the template image to the test image. The operations further include determining, based on applying the deformation field to the template image, a third location and a fourth location on the test image corresponding to the first location and the second location on the template image, respectively, wherein the third location and the fourth location can be used to determine the first measurement of the test image.
[0006] A further aspect of the present disclosure includes a system comprising a memory and a processing device coupled to the memory. The processing device is used to train a machine learning model using a training data set including historical image data of a structure, the historical image data including multiple images of the structure, wherein the training includes marking regions of a source image of the multiple images to create a source mask for the source image. The training further includes providing a source image and a target image from the multiple images as training inputs to the machine learning model. The training further includes receiving a deformation field between the source image and the target image as an output from the machine learning model. The training further includes applying the deformation field to the source mask to create a distorted source mask. The training further includes calculating a loss associated with the distorted source mask. The training further includes updating a weight of the machine learning model based on the loss. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The present disclosure is illustrated by way of example and not by way of limitation in the figures of the accompanying drawings.
[0008] Figure 1 is a block diagram illustrating an example system architecture according to some embodiments.
[0009] Figure 2 A dataset generator associated with measurement repetition between structurally similar images is shown in accordance with some embodiments.
[0010] Figure 3 is a block diagram illustrating determining prediction data associated with measured repetitions between structurally similar images in accordance with some embodiments.
[0011] Figure 4A is a block diagram illustrating training a machine learning model to perform image registration and / or measurement position determination according to some embodiments.
[0012] Figure 4Bis a block diagram illustrating the use of a machine learning model to perform image registration between input source and target images in accordance with some embodiments.
[0013] FIG. 5A to FIG. 5C is an example of precise point placement according to some embodiments.
[0014] FIG. 6A to FIG. 6B is an example of precise point placement using reference points according to some embodiments.
[0015] 7A to 7C is a flow chart of a method of associating measurement repetitions between structurally similar images according to some embodiments.
[0016] Figure 8 is a block diagram illustrating a computer system, according to some embodiments. DETAILED DESCRIPTION
[0017] Techniques for measurement duplication between images with similar structures (e.g., using deformable image registration and / or precise point placement) are described herein. Techniques for training and using one or more machine learning models to automate point selection to measure images of similar structures are also described herein. In an embodiment, a measurement point may be selected in a template image, and a template image with the selected image may be provided to the trained machine learning model together with a test image to be measured. The trained machine learning model may output a point on the test image to be measured. The template image and the test image may have the same type of structure (e.g., the same structure of a semiconductor device), but in an embodiment may have different instances of the same type of structure. Thus, the specific locations to be measured in different images may not be the same. Historically, this meant that a technician manually marked the locations to be measured in each image. An embodiment enables marking the locations to be measured in a single image, and automatically measuring the locations to be measured in one or more additional images of the same type of structure without manually marking the locations in the additional images. The specific locations automatically selected in the additional images may not correspond to the same coordinates (e.g., the same X, Y, and / or Z coordinates) as the coordinates of the points to be measured in the template image. In an embodiment, the test image is automatically measured using the automatically marked locations. Embodiments increase the speed at which critical dimensions of devices (eg, semiconductor devices) and / or device layers are measured and reduce the amount of time a technician spends making such measurements.
[0018] Embodiments also encompass techniques for training a machine learning model to receive labeled template images and test images and output locations on the test image to be measured.
[0019] In many applications in the semiconductor industry, many different images of the same structure may be measured.
[0020] Manufacturing equipment is used to produce products. For example, substrate processing equipment is used to produce substrates (e.g., wafers, semiconductors). The manufactured substrate has measured properties (e.g., property data). Various applications benefit from the measurement of different images (e.g., manufactured substrates) that share similar structures. Generally, multiple images of similar appearance can be measured to identify and analyze specific properties of different instances of the same structure (e.g., the same region or area of a semiconductor device), such as critical dimensions (CD), feature heights, hole depths, etc. For example, in the context of a cross-sectional image of a device, an engineer can measure the depth of a hole or groove to determine whether the device being manufactured meets the specification. Subsequently, the same property (e.g., depth) may be measured in similar holes or grooves of the same device in other images. Other images may be different instances of the same device on the same substrate (e.g., different devices on the same wafer) or different instances of the same device on different substrates (e.g., different devices on different wafers). For example, in a cross-sectional image, an engineer may be interested in measuring the height of a hole or the width of a hole at a certain height. For multiple other holes of the same type of structure, the height or width may be measured again. To achieve such measurements in an embodiment, automatic image registration and precise point placement are determined, and automatic measurements are generated based on the automatic precise point placement.
[0021] In image registration, each pixel of the first image can be transferred to the second image such that the image structure and coherence are preserved based on the similarity of the images. Rigid registration is perhaps the simplest way to transfer each pixel and involves a rotation and translation of the first image. Affine registration allows for shearing and scaling of the first image. Non-rigid registration enables images to be aligned with local deformations where each pixel is transferred independently of other pixels, thereby preserving more complex structural deformations and variability.
[0022] Performing measurements in several images presents challenges that typically result in manual measurements being performed in many images. For example, due to the similarities and dissimilarities of structures, and the goal of accurate and consistent results, measurements of the same point cannot typically be performed across different images of the same structure. Consistency and accuracy of measurements are goals to ensure reliable analysis and decision making in semiconductor manufacturing processes. However, manually measuring each instance of a specific attribute in different images is time consuming, laborious, and prone to error.
[0023] Existing solutions to this problem typically involve manual measurement techniques that rely on a human operator to visually identify and measure the target attribute or structure in each image. This manual approach not only increases the risk of inconsistency and inaccuracy across images, but also greatly limits the efficiency and throughput of the measurement process. In addition, the manual nature of these techniques makes them susceptible to operator fatigue, variations in measurement technology, and limitations of human perception.
[0024] Aspects and embodiments of the present disclosure address these and other shortcomings of the prior art by performing methods for measurement replication between structurally similar images, such as through image registration and precise point placement across images for repeatable measurements, thereby overcoming limitations of prior art measurement techniques.
[0025] In an embodiment, a processing device trains a machine learning model using a training data set including historical image data of a structure. The historical image data includes multiple images of the structure. The training includes marking regions of a source image of the multiple images to create a source mask for the source image. The training further includes providing the source image and the target image of the multiple images as training inputs to the machine learning model. The training further includes receiving a deformation field (also called a registration flow) between the source image and the target image as an output from the machine learning model. The training further includes applying the deformation field to the source mask to create a distorted source mask. The training further includes calculating a loss associated with the distorted source mask. The training further includes updating the weights of the machine learning model based on the loss. In some embodiments, the machine learning model may include at least one of a convolutional neural network (CNN) or an encoder-decoder network.
[0026] In some embodiments, the training may further include applying the deformation field to the source image to create a distorted source image. The training may further include calculating one or more similarity metric losses of the target image and the distorted source image. The training may further include updating the weights of the machine learning model based on the one or more similarity metric losses. In some embodiments, the loss associated with the distorted source mask may include a structural integrity loss, and the one or more similarity metric losses may include a structural similarity index (SSI) loss and a Pearson correlation coefficient (PCC) loss.
[0027] In some embodiments, the method may further include receiving an input comprising a template image and a test image, wherein the template image comprises a marker at a first position and a second position associated with a first measurement of the template image. The method may further include providing the template image and the test image as input to a trained machine learning model, wherein the trained machine learning model outputs a deformation field based on a non-rigid registration of the template image to the test image. The method may further include receiving an output from the trained machine learning model comprising a third position and a fourth position on the test image, the third position and the fourth position being usable to determine the first measurement of the test image. In some embodiments, after non-rigidly registering the test image to the template image, the third position of the test image may correspond to the first position of the template image.
[0028] In some embodiments, the first and second locations on the template image may be located at a vertical distance from the first reference point, and the third and fourth locations on the test image may be located at a vertical distance from the second reference point.
[0029] In some embodiments, the method may further include extracting a template signal along a first line that intersects a first position and a second position on the template image. The method may further include extracting a test signal along a second line that intersects a third position and a fourth position on the test image. The method may further include adjusting at least one of the third position or the fourth position using a dynamic time warping (DTW) transform of the test signal to the template signal. In some embodiments, the template signal and the test signal each include an average of pixel values of pixels along the first line and the second line, respectively.
[0030] In some embodiments, the method may further include identifying a local maximum of the template signal, wherein the local maximum of the template signal is a local maximum value that is closest to a point on the template signal corresponding to the first position on the template image. In some embodiments, the method may further include identifying a local minimum of the template signal, wherein the local minimum of the template signal is a local minimum value that is closest to a point on the template signal corresponding to the first position on the template image.
[0031] The method may further include identifying a test signal local maximum, wherein the test signal local maximum is the local maximum closest to a point on the test signal corresponding to the third position on the test image. The method may further include identifying a test signal local minimum, wherein the test signal local minimum is the local minimum closest to the point on the test signal corresponding to the third position on the test image.
[0032] The method may further include calculating a first ratio representing a relationship between a proximity of a point on the template signal to a maximum value of the template signal and a proximity of a point on the template signal to a minimum value of the template signal. The method may further include adjusting the third position so that a second ratio representing a relationship between a proximity of a point on the test signal to a maximum value of the test signal and a proximity of a point on the test signal to a minimum value of the test signal is approximately equal to the first ratio.
[0033] Aspects of the present disclosure result in technical advantages. Aspects of the present disclosure avoid the time-consuming, laborious and error-prone process of manually measuring each instance of a specific attribute of a structure in different images (e.g., the same critical dimension for different instances of the same device). Aspects of the present disclosure do not rely on a human operator to visually identify and measure the desired attribute in each image. Aspects of the present disclosure reduce the risk of inconsistency and inaccuracy in critical dimension (CD) measurements, and also improve the efficiency and throughput of the measurement process. Aspects of the present disclosure are not affected by operator fatigue, changes in measurement technology, or limitations of human perception. Aspects of the present disclosure can avoid delays in manufacturing and corresponding throughput losses. Aspects of the present disclosure allow for more accurate and consistent measurement of similar structures (e.g., of a manufactured substrate) in different images.
[0034] Although some embodiments of the present disclosure describe measurement duplication between structurally similar images associated with manufacturing systems (e.g., substrate processing equipment, manufacturing equipment, etc.), the present disclosure can be used for measurement duplication between structurally similar images associated with other fields (e.g., medical imaging, remote sensing, computer vision and robotics, geographic information systems, industrial inspection and quality control, astronomy and astrophysics, security and surveillance, forensics, agriculture and environmental monitoring, etc.).
[0035] Figure 1 1 is a block diagram illustrating an exemplary system 100 (exemplary system architecture) according to some embodiments. The system 100 (eg, via the automatic measurement component 114) may perform one or more of the methods described herein (eg, 7A to 7C The system 100 may include a client device 120, a manufacturing device 124, a sensor 126, a metrology device 128, an automatic measurement server 112, and / or a data store 140. In some embodiments, the automatic measurement server 112 is part of the automatic measurement system 110. In some embodiments, the automatic measurement system 110 further includes a server machine 180.
[0036] In some embodiments, one or more of the client devices 120, the manufacturing equipment 124, the sensors 126, the metrology equipment 128, the automated measurement server 112, the data store 140, and / or the server machine 180 are coupled to each other via the network 130. In at least one embodiment, the automated measurement system 110 trains one or more machine learning models to perform measurement iterations between structurally similar images. In at least one embodiment, the automated measurement system 110 implements one or more trained machine learning models (e.g., in the location determiner 172) to determine locations in an image from which to make measurements. In embodiments, the images to be measured may be generated during and / or after substrate fabrication.
[0037] In some embodiments, the network 130 is a public network that provides access to the automatic measurement server 112, the data store 140, and / or other publicly available computing devices to the client device 120. In some embodiments, the network 130 is a private network that provides access to the manufacturing equipment 124, the sensor 126, the metrology equipment 128, the data store 140, the automatic measurement server 112, and / or other dedicated computing devices to the client device 120. In some embodiments, the network 130 includes one or more wide area networks (WANs), local area networks (LANs), wired networks (e.g., Ethernet), wireless networks (e.g., 802.11 networks or Wi-Fi networks), cellular networks (e.g., long term evolution (LTE) networks), routers, hubs, switches, server computers, cloud computing networks, and / or combinations thereof.
[0038] In some embodiments, the client device 120 includes a computing device, such as a personal computer (PC), a laptop, a mobile phone, a smart phone, a tablet computer, a netbook computer, etc. In some embodiments, the client device 120 includes a local automatic measurement component 114, which may correspond to the automatic measurement component 114 of the automatic measurement server 112. The client device 120 includes an operating system that allows a user to merge, generate, view, and / or edit one or more of data (e.g., image data), provide instructions to the automatic measurement system 110 (e.g., a machine learning processing system), provide information to the automatic measurement system 112 (e.g., including an image to be measured, a template image with selected measurement points to be used to determine other target image measurement points), etc.
[0039] The manufacturing equipment 124 can produce a product, such as a substrate, a wafer, a semiconductor, an electronic device, etc., after performing one or more recipes and / or processes on the substrate over a period of time. The manufacturing equipment 124 may include one or more sensors configured to capture data (e.g., property data, such as image data) of the substrate before, during, and / or after the substrate processing operation. For example, the one or more sensors may be configured to capture image data and / or substrate images (e.g., cross-sectional images, scanning electron microscope (SEM) images, optical images, atomic force microscope images, superimposed images, infrared images, reflection images, transmission images, etc.), spectral data, non-spectral data, etc. for a portion of the substrate before, during, or after the substrate processing operation.
[0040] The manufacturing equipment 124 can perform processes on a substrate (e.g., a wafer, etc.) at a processing chamber. Examples of substrate processes include deposition processes that deposit one or more films on the surface of a substrate, etching processes that form patterns on the surface of a substrate, and the like. The manufacturing equipment 124 can perform each process according to a process recipe. The process recipe defines a set of specific operations to be performed on the substrate during the process, and may include one or more settings associated with each operation. One or more processes performed on the substrate form a structure on the substrate. The structure may have one or more regions, which may have one or more critical dimensions (e.g., which may be marked). For example, an image of a structure may have markings specifying the location of a coating, a gate, a drain, a mask, and / or other structures.
[0041] In some embodiments, fabrication equipment 124 includes sensors 126 configured to generate data (e.g., image data) associated with a substrate processed at fabrication system 100. For example, fabrication equipment 124 may include one or more sensors configured to generate image data (e.g., scanning electron microscope (SEM) images, optical images, atomic force microscope images, overlay images, infrared images, reflection images, transmission images, etc.) associated with a substrate and / or structure of a substrate before, during, and / or after a process (e.g., a deposition process).
[0042] The metrology equipment 128 may provide metrology data (e.g., image data) associated with a substrate processed by the manufacturing equipment 124. The metrology data may include one or more images of the processed substrate, such as a cross-sectional image, a top-down image, a SEM image, an optical image, an atomic force microscope image, an overlay image, an infrared image, a reflection image, a transmission image, and the like. In some embodiments, the metrology data may further include values of one or more types of surface profile property data (e.g., critical dimensions of one or more features included on the substrate surface, critical dimension uniformity across the substrate surface, edge placement error, and the like). The metrology data may be metrology data of a finished product or a semi-finished product. The metrology data for each substrate may be different. For example, the metrology data may be generated using a reflection measurement technique, an ellipsometry technique, a transmission electron microscope (TEM) technique, and the like. In some embodiments, the metrology data includes a cross-sectional image, a scanning electron microscope (SEM) image, an optical image, an atomic force microscope image, an overlay image, an infrared image, a reflection image, a transmission image, and the like. In some embodiments, the metrology equipment 128 includes an instance of an automatic measurement component 114, and may generate automatic measurements of the generated image during and / or after the image is generated.
[0043] In some embodiments, the automatic measurement server 112 and the server machine 180 each include one or more computing devices, such as a rack server, a router computer, a server computer, a personal computer, a mainframe computer, a laptop computer, a tablet computer, a desktop computer, a graphics processing unit (GPU), an accelerator application-specific integrated circuit (ASIC) (e.g., a tensor processing unit (TPU)), etc.
[0044] The automatic measurement server 112 may include an automatic measurement component 114, which may be additionally or alternatively included in the client device 120 and / or the metrology device 128. In some embodiments, the automatic measurement component 114 identifies the image data 142 to be processed (e.g., the historical image data 144 and / or the current image data 146). In some embodiments, the automatic measurement component 114 automatically determines the location in the image data 142 to be measured, and automatically generates a measurement based on the determined location. In some embodiments, the automatic measurement component 114 uses one or more trained machine learning models to determine the location to be measured. In some embodiments, the location determiner 172 is a trained machine learning model that is trained to receive a template image with a selected measurement point and one or more target images, and outputs a location on the one or more target images to be measured. In some embodiments, the location determiner 172 is trained using the historical image data 144. In some embodiments, the nodes of the location determiner 172 are updated using back propagation. For example, a distorted source image and a distorted source mask can be compared with a target image to determine the differences therebetween. These differences can then be used to update the nodes of the location determiner 172.
[0045] In some embodiments, the location determiner 172 generates an indication of the location to be measured. In embodiments, the location determiner is trained using supervised or semi-supervised machine learning (e.g., a supervised dataset, historical image data 144 tagged with the location to be measured, etc.). In some embodiments, a first subset of the historical image data 144 is labeled data, and a second subset of the historical image data 144 is unlabeled. In some embodiments, a dataset generator 176 of the server machine 180 generates a training dataset for training the location determiner 172.
[0046] In some embodiments, the manufacturing equipment 124 (e.g., a deposition chamber, a cluster tool, a wafer back grinding system, a wafer saw equipment, a die attach machine, a wire bonder, a die coating system, a molding equipment, a sealing equipment, a metal can welding machine, a deburring / trim / form / cut (DTFS) machine, a marking equipment, a wire handling equipment, etc.) is part of a substrate processing system (e.g., an integrated processing system). The manufacturing equipment 124 includes one or more of a controller, a housing system (e.g., a substrate carrier, a front opening wafer transfer box (FOUP), an automatic teaching FOUP, a process kit housing system, a substrate housing system, a cassette, etc.), a side storage box (SSP), an aligner device (e.g., an aligner chamber), a factory interface (e.g., an equipment front end module (EFEM)), a load lock, a transfer chamber, one or more processing chambers, a robot arm (e.g., disposed in the transfer chamber, disposed in the front interface, etc.), and the like. The housing system, SSP, and load lock mounted to the factory interface and the robotic arm provided in the factory interface are used to transfer contents (e.g., substrates, process kit rings, carriers, verification wafers, etc.) between the housing system, SSP, load lock, and factory interface. An aligner device can be provided in the factory interface to align the substrate. The load lock and process chamber can be mounted to the transfer chamber, and the robotic arm provided in the transfer chamber is used to transfer contents (e.g., substrates, process kit rings, carriers, verification wafers, etc.) between the load lock, the process chamber, and the transfer chamber. In some embodiments, the manufacturing equipment 124 includes components of the substrate processing system. In some embodiments, the image data 142 of the substrate depicts the results of the substrate undergoing one or more processes (e.g., deposition, etching, heating, cooling, transfer, processing, flow, etc.) performed by the components of the manufacturing equipment 124.
[0047] In some embodiments, sensor 126 provides image data 142 (e.g., image data, such as historical image data and current image data) of a substrate processed by manufacturing equipment 124 and / or one or more structures of the substrate (e.g., transistors, interconnects, gates, contacts, memory cells, etc.).
[0048] In some embodiments, the sensors 126 include one or more metrology tools, such as optical and imaging systems, imaging stations, etc. In some embodiments, the sensors 126 include one or more metrology tools, such as ellipsometers (for determining the properties and surfaces of thin films by measuring material properties such as layer thickness, optical constants, surface roughness, composition, and optical anisotropy), ion mills (for preparing heterogeneous host materials when wide areas of the material are uniformly thin), capacitance versus voltage (CV) systems (for measuring CV and capacitance versus time (Ct) characteristics of semiconductor devices), interferometers (for measuring distances in wavelengths and determining the wavelength of a specific light source), source measurement units (SMEs) magnetometers, profilometers, wafer probers (for measuring the distances of semiconductor wafers when the semiconductor wafers are being processed). The invention also includes a multi-functional semiconductor wafer testing system (for testing the semiconductor wafer before separation into individual dies or chips), a critical dimension scanning electron microscope (CD-SEM, used to ensure the stability of the manufacturing process by measuring the critical dimensions of the substrate), a reflectometer (for measuring the reflectivity and emissivity of the surface), a resistance probe (for measuring the resistivity of a thin film), a resistive high energy electron diffraction (RHEED) system (for measuring or monitoring the crystal structure or crystal orientation of epitaxial thin films of silicon or other materials), an X-ray diffractometer (for unambiguously determining the crystal structure, crystal orientation, film thickness and residual stress in silicon wafers, epitaxial films, or other substrates), and the like.
[0049] In some embodiments, the image data 142 is used to assess equipment health and / or product health (e.g., product quality). For example, the image data 142 may be image data used to measure a product (e.g., a manufactured wafer). In some embodiments, the image data 142 is received over a period of time.
[0050] In some embodiments, the sensor 126 and / or the metrology device 128 provides image data 142, which may include one or more of scanning electron microscope (SEM) images, energy dispersive x-ray (EDX) images, spatial position data, chip layer data, chip layout data, edge data, grayscale data, signal-to-noise data, spacing data, optical image data, and the like.
[0051] In some embodiments, the image data 142 includes data describing the visual attributes of the image, which may be organized into a matrix or a multi-dimensional array. In some embodiments, the image data allows for digital representation, processing, and analysis of pixel-level information of the image. In some embodiments, measurement data will be generated from the image data 142. The measurement data may relate to measurements of the substrate, such as critical dimensions. In some embodiments, the measurements of the image data 142 include size attribute data (e.g., data describing the size of the attributes of the substrate). In some embodiments, the measurements of the image data 142 include size attribute data (e.g., data describing the size of the attributes of the substrate). In some embodiments, the image data 142 includes SEM images (e.g., images captured by a scanning electron microscope using a focused electron beam to scan the surface of a substrate to create a high-resolution image). In some embodiments, the image data 142 includes EDX images (e.g., images generated from data collected using x-ray technology to identify the elemental composition of a material). In some embodiments, the measurements of the image data 142 include defect distribution data (e.g., data describing the distribution of defects on a substrate, such as space, time, etc.). In some embodiments, the measurements of the image data 142 include spatial position data (e.g., data describing the spatial position of the attributes, defects, elements, etc. of the substrate). In some embodiments, the image data 142 includes grayscale data (eg, data describing pixel brightness of an image of a substrate) and signal-to-noise data (eg, data describing a signal-to-noise ratio of a substrate measured using, for example, a spectroscopic measurement device).
[0052] In some embodiments, the image data 142 (e.g., historical image data 144, current image data 146, etc.) is processed (e.g., by the client device 120 and / or by the automatic measurement server 112). In some embodiments, the processing of the image data 142 includes generating features (e.g., measurements, critical dimensions, etc.). In some embodiments, the features are measurements or patterns (e.g., slopes, widths, heights, peaks, etc.) in the image data 142 or combinations of values from the image data 142.
[0053] In some embodiments, the metrology device 128 can be included as part of the manufacturing equipment 124. For example, the metrology device 128 can be included within or coupled to a processing chamber and configured to generate metrology data such as image data 142 of a substrate before, during, and / or after a process (e.g., a deposition process, an etching process, etc.) while the substrate remains in the processing chamber. In some cases, the metrology device 128 can be referred to as an in-situ metrology device. In another example, the metrology device 128 can be coupled to another station of the manufacturing equipment 124. For example, the metrology device can be coupled to a transfer chamber, a load lock, or a factory interface.
[0054] In some embodiments, the sensor 126 can be included as part of the manufacturing equipment 124. For example, the sensor 126 can be included inside or coupled to the processing chamber and configured to generate sensor data of the substrate before, during, and / or after a process (e.g., a deposition process, an etching process, etc.) while the substrate remains in the processing chamber. In some cases, the sensor 126 can be referred to as an in-situ sensor. In another example, the sensor 126 can be coupled to another station of the manufacturing equipment 124. For example, the sensor can be coupled to a transfer chamber, a load lock, or a factory interface.
[0055] In some embodiments, metrology equipment 128 (e.g., ellipsometric equipment, imaging equipment, spectroscopic equipment, etc.) is used to determine metrology data (e.g., image data, inspection data, spectroscopic data, ellipsometric data, material composition, optical, or structural data, etc.) corresponding to a substrate produced by manufacturing equipment 124 (e.g., substrate processing equipment). In some examples, metrology equipment 128 is used to inspect a portion (e.g., a structure, a region, a layer, etc.) of a substrate after the substrate is processed by manufacturing equipment 124. In some embodiments, metrology equipment 128 performs scanning acoustic microscopy (SAM), ultrasonic inspection, x-ray inspection, and / or computed tomography (CT) inspection.
[0056] In some examples, after fabrication equipment 124 deposits one or more layers on a substrate, metrology equipment 128 is used to determine the quality of the processed substrate (e.g., the size of features and / or structures, the uniformity of features and / or structures, etc.). In some embodiments, metrology equipment 128 includes an imaging device (e.g., a SAM device, an ultrasonic device, an x-ray device, a CT device, etc.). In some embodiments, image data 142 includes sensor data from sensor 126 and / or metrology data from metrology equipment 128 located in situ (inside a processing chamber).
[0057] In some embodiments, performance data 152 is measured by sensors 126 and / or metrology equipment 128 and may be associated with measurement consistency and accuracy of manufactured substrates.
[0058] In some embodiments, measurement data 160 may be generated for one or more images of the image data 142. The measurement data may be measurements of structures, critical dimensions, and / or other properties of features present in the image data 142. In embodiments, the measurement data 160 may be automatically generated for the one or more target images by the automatic measurement component 114 in response to the automatic measurement component 114 receiving and processing the labeled template image and the one or more target images. In some embodiments, the measurement data 160 may be included in the image data 142 (e.g., by annotating the image data 142).
[0059] In some embodiments, data store 140 is a memory (e.g., random access memory), a drive (e.g., a hard drive, a flash drive), a database system, or another type of component or device capable of storing data. In some embodiments, data store 140 includes multiple storage components (e.g., multiple drives or multiple databases) across multiple computing devices (e.g., multiple server computers). In some embodiments, data store 140 stores one or more of image data 142, performance data 152, and / or measurement data 160.
[0060] Measurement consistency and accuracy are beneficial in ensuring reliable analysis and decision making in semiconductor manufacturing processes. Manually measuring each instance of a particular attribute in different images is time consuming, laborious, and prone to error. By providing image data 142 to the automatic measurement component 114 and receiving automatic measurement data 160 from the automatic measurement component 114, the system 100 has a technical advantage of avoiding the time consuming, laborious, and error prone process of manually measuring each instance of a particular attribute in different images.
[0061] In some embodiments, the automatic measurement system 110 includes a server machine 180. The server machine 180 may include a data set generator 176 capable of generating a data set (e.g., a data input set and a target output set) to train, validate, and / or test one or more machine learning models (e.g., such as the position determiner 172). The data set generator 176 has the functionality of data collection, compilation, reduction, and / or partitioning to put the data in a form for machine learning. In some embodiments (e.g., for small data sets), partitioning (e.g., explicit partitioning) for post-training validation is not used. Repeated cross-validation (e.g., 5-fold cross-validation, leave-one-out cross-validation) can be used during training, where a given data set is effectively repeatedly partitioned into different training sets and validation sets during training.
[0062] In some embodiments, the dataset generator 176 may explicitly divide the historical data (e.g., the historical image data 144) into a training set (e.g., 60% of the historical data), a validation set (e.g., 20% of the historical data), and a test set (e.g., 20% of the historical data). Figure 2 Some operations of the data set generator 176 are described in detail.
[0063] The server machine 180 includes a training engine 182, a verification engine 184, and / or a testing engine 186. In some embodiments, an engine (e.g., training engine 182, verification engine 184, and testing engine 186) refers to hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, processing device, etc.), software (such as instructions running on a processing device, a general-purpose computer system, or a dedicated machine), firmware, microcode, or a combination thereof. The training engine 182 is capable of training a machine learning model (e.g., location determiner 172) using one or more feature sets associated with a training set from the dataset generator 176. In some embodiments, the training engine 182 generates multiple trained machine learning models, where each trained machine learning model (e.g., of the location determiner 172) corresponds to a different set of parameters for the training set (e.g., image data 142). In some embodiments, multiple models are trained using different objectives. For example, different models can be trained to automatically select locations to be measured for different structures and / or different devices. About Figures 2 to 4A Embodiments of training are described in more detail.
[0064] In some embodiments, performance data 152 may be associated with measurement consistency and accuracy of fabricated substrates (e.g., measurement repetition between structurally similar images). For example, performance data 152 may indicate critical dimension measurements. In some embodiments, at least a portion of performance data 152 is associated with measurement quality of substrates produced by fabrication equipment 124. Performance data 152 may indicate whether measurements were made correctly and / or accurately. For example, two images may be structurally similar, but not identical. Performance data 152 may indicate that measurements made on a first image were similarly measured on a second structurally similar, but not identical, image (e.g., through image-to-image registration, where each pixel is transferred from a first image to a second image such that the structure and coherence of the images is maintained).
[0065] In some embodiments, at least a portion of the performance data 152 is based on metrology data (e.g., image data) from the metrology device 128 (e.g., historical performance data 154 includes metrology data indicating a properly processed substrate, image data of a substrate, images of structures and / or features of a manufactured substrate, etc.) or sensor 126 (e.g., historical performance data 154 includes sensor data indicating a properly processed substrate, image data of a substrate, images of structures and / or features of a manufactured substrate, etc.). In some embodiments, at least a portion of the performance data 152 is based on an inspection of a substrate (e.g., current performance data 156 is based on an actual inspection). In some embodiments, the performance data 152 includes user input (e.g., via the client device 120) indicating substrate quality. In some embodiments, the performance data 152 includes an indication of an absolute value (e.g., inspection data of a substrate indicates that it differs from threshold data by a calculated value, a drift value differs from a threshold drift value by a calculated value) or a relative value (e.g., inspection data of the measurement data 160 indicates that it differs from threshold data by 5%, a drift value differs from a threshold drift value by 5%). In some embodiments, the performance data 152 indicates that a threshold error amount is met (eg, at most a 5% error in the measurement data after repeated measurements, at most a 5% error in production, etc.).
[0066] In some embodiments, image data 142 may be associated with a target image and / or a test image. In some embodiments, measurement data 160 may be associated with a deformation field (e.g., generated by position determiner 172). In some embodiments, measurement data 160 may be associated with applying the deformation field to a source mask to create a distortion.
[0067] In some embodiments, training the machine learning model may include computing a structural integrity loss (e.g., a structural similarity index (SSI) loss and / or a Pearson correlation coefficient (PCC) loss) of the warped source mask. In some embodiments, training the machine learning model may include applying a deformation field to a source image to create a warped source image, and computing one or more similarity metric losses of the target image and the warped source image. For example, measurement data 160 may be associated with a deformation field between a source image and a target image that has been provided as input to the machine learning model.
[0068] In some embodiments, the trained machine learning model (e.g., included in the position determiner 172) can be an unsupervised or semi-supervised machine learning model. In some embodiments, unsupervised training can include inputting the image data 142 as input to the machine learning model. In some embodiments, the machine learning model can be an encoder-decoder network, and the model can be trained by reconstructing the input image to map the spatial transformation (e.g., affine transformation) between the source image and the target image. In some embodiments, the machine learning model can be trained based on a training set of source images and target images. In some embodiments, the training set can further include a source mask image (e.g., a labeled source mask).
[0069] The source image can be transformed into a target image. In some embodiments, the mapping of the spatial transformation is referred to as a deformation field. The deformation field can be applied to the source image and / or the source mask. In some embodiments, the measurement generator 174 can apply the deformation field to the source image (e.g., with marked measurement locations) to determine the location of the measurement location on the target image. In some embodiments, precise point placement techniques can be used to more accurately place the measurement location (see Figure 7C ). In some embodiments, the source mask can be an annotated (labeled) version of the source image. In some embodiments, the source mask is an annotation of the structural composition of the source image. In some embodiments, a loss function can be used to measure the similarity between the target image and the distorted source image or the distorted source mask. The loss calculated using the loss function can be used to update the weights of a machine learning model (e.g., an encoder-decoder network), for example, by backpropagation. For a more detailed discussion of the training process, see Fig. 7A Description found in .
[0070] In some embodiments, training a machine learning model (e.g., where the machine learning model is part of the location determiner 172) may include updating weights of the machine learning model based on a structural integrity loss and / or one or more similarity metric losses (e.g., a structural similarity index (SSI) loss, a Pearson correlation coefficient (PCC) loss, etc.) (e.g., to improve measurement duplication between structurally similar images). In some embodiments, training a machine learning model (e.g., where the machine learning model is part of the location determiner 172) may include taking measurements of a substrate (e.g., structures and / or features of the substrate), adjusting the measurements of the substrate (e.g., structures and / or features of the substrate) for consistency and accuracy, applying a deformation field to a source mask to create a warped source mask, calculating a loss associated with the warped source mask, updating weights of the machine learning model based on the loss, applying the deformation field to a source to create a warped source image, calculating one or more similarity metric losses for the target image and the warped source image, updating weights of the machine learning model based on the one or more similarity metric losses, and the like. In some embodiments,
[0071] In some embodiments, the measurement data 160 is associated with an output (e.g., a deformation field) of the position determiner 172. In some embodiments, calculating a loss using the measurement data 160 and the image data 142 is associated with updating weights of a machine learning model (e.g., based on a structural integrity loss and / or a similarity metric loss) to improve measurement duplication between structurally similar images.
[0072] The performance data 152 includes historical performance data 154 and current performance data 156. The performance data 152 may indicate measurement quality, such as measurement duplication between structurally similar images, substrate throughput, substrate defects, etc. The performance data 152 may indicate whether measurements were made consistently and / or accurately. For example, two images may be structurally similar, but not identical. The performance data 152 may indicate that measurements made on a first image are similarly measured on a second image that is structurally similar, but not identical (e.g., through image-to-image registration, where each pixel is transferred from the first image to the second image such that the structure and coherence of the images is maintained).
[0073] The validation engine 184 can validate the trained machine learning model (e.g., the location determiner 172) using the corresponding feature set of the validation set from the dataset generator 176. For example, a first trained machine learning model trained using the first feature set of the training set is validated using the first feature set of the validation set. The validation engine 184 determines the accuracy of each of the trained machine learning models based on the corresponding feature set of the validation set.
[0074] The testing engine 186 can test the trained machine learning model using the corresponding feature set of the test set from the dataset generator 176. For example, a first trained machine learning model trained using the first feature set of the training set is tested using the first feature set of the test set. The testing engine 186 can determine whether the trained machine learning model (e.g., the trained machine learning model of the position determiner 172) has sufficient accuracy for production.
[0075] In some embodiments, a machine learning model refers to a model product created by a training engine 182 using a training set including data inputs and corresponding target outputs (e.g., conditions or ordinal levels that correctly classify corresponding training inputs). Patterns that map data inputs to target outputs can be found in the data set, and mappings that capture these patterns are provided for the machine learning model. In some embodiments, the machine learning model uses one or more of Gaussian process regression (GPR), Gaussian process classification (GPC), Bayesian neural network, neural network Gaussian process, deep belief network, Gaussian mixture model, or other probabilistic learning methods. Non-probabilistic methods can also be used, including one or more of support vector machines (SVM), radial basis functions (RBF), clustering, nearest neighbor algorithms (k-NN), linear regression, random forests, neural networks (e.g., artificial neural networks), etc. In some embodiments, the machine learning model is a multivariate analysis (MVA) regression model.
[0076] The automatic measurement component 114 provides the current image data 146 (e.g., as an input) to a trained machine learning model (e.g., the position determiner 172) and processes the current image data 146 using the trained machine learning model (e.g., processes on the input to obtain one or more outputs).
[0077] For purposes of illustration and not limitation, aspects of the present disclosure describe training one or more machine learning models using historical data (i.e., previous data, historical image data 144), and providing current image data 146 to one or more trained probabilistic machine learning models to determine measurement data 160. In other embodiments, heuristic models or rule-based models are used to determine measurement data 160 (e.g., without using trained machine learning models). In other embodiments, non-probabilistic machine learning models may be used.
[0078] In some embodiments, the functionality of the client device 120, the automatic measurement server 112, and / or the server machine 180 will be provided by a smaller number of machines. For example, in some embodiments, the server machine 180 and the automatic measurement server 112 are integrated into a single machine. In some embodiments, the data set generator 176, the training engine 182, the validation engine 184, the test engine 186, and / or the automatic measurement component 114 are distributed across more devices than shown.
[0079] In some embodiments, a "user" is represented as a single individual. However, other embodiments of the present disclosure encompass an entity that is controlled by multiple users and / or automated sources. In some examples, a group of individual users united as a group of administrators is considered a "user."
[0080] Although embodiments of the present disclosure are discussed in the context of determining measurement data 160 for measuring duplication between structurally similar images associated with a manufacturing system (e.g., substrate processing equipment, manufacturing equipment, etc.), in some embodiments, the present disclosure may also be generally applicable to measuring duplication between structurally similar images in various other fields involving images and image processing (e.g., medical imaging, remote sensing, computer vision and robotics, geographic information systems, industrial inspection and quality control, astronomy and astrophysics, security and surveillance, forensics, agriculture and environmental monitoring, etc.).
[0081] Figure 2 276 (e.g., associated with measurement duplication between structurally similar images, methods 700A-C, etc.) for creating a dataset for a machine learning model (e.g., associated with measurement duplication between structurally similar images, methods 700A-C, etc.) according to some embodiments. Figure 1 In some embodiments, the dataset generator 276 is Figure 1 The server machine 180 is part of Figure 2 The dataset generated by the dataset generator 276 can be used to train a machine learning model (e.g., see Figure 7B ) to enable the machine learning model to automatically determine locations to be measured and / or automatically generate measurements based on such determined locations (see, for example, Figure 7C ).
[0082] Dataset Generator 276 Dataset Generator 176 is a machine learning model (e.g., Figure 1 The data set generator 276 uses the historical image data 244 to create a data set. Figure 2 The system 200 shows a data set generator 276, a data input 210, and a target output 220 (eg, target data, such as a location to be measured).
[0083] In some embodiments, the data set generator 276 generates a data set (e.g., a training set, a validation set, a test set) including one or more data inputs 210 (e.g., a training input, a validation input, a test input). In some embodiments, the data set generator 276 does not generate a target output (e.g., for unsupervised learning). In some embodiments, the data set generator generates one or more target outputs 220 corresponding to the data inputs 210 (e.g., for supervised learning). The data set may also include mapping data that maps the data inputs 210 to the target outputs 220. The data inputs 210 are also referred to as "features," "attributes," or "information." In some embodiments, the data set generator 276 provides the data set to the training engine 182, the validation engine 184, or the testing engine 186, where the data set is used to train, validate, or test a machine learning model (e.g., associated with measurement repetitions between structurally similar images, methods 700A-C, etc.).
[0084] In some embodiments, the data set generator 276 generates the data input 210 and the target output 220. In some embodiments, the data input 210 includes one or more historical image data sets 244 (e.g., source images, target images, template images, test images, source masks, etc.) (e.g., associated with measurement iterations between structurally similar images, methods 700A-C, etc.). In some embodiments, the historical image data 244 includes image data from one or more types of sensors and / or metrology devices.
[0085] In some embodiments, the dataset generator 276 generates a first data input corresponding to the first historical image dataset 244A to train, validate, or test a first machine learning model, and the dataset generator 276 generates a second data input corresponding to the second historical image dataset 244B to train, validate, or test a second machine learning model (e.g., associated with measurement duplication between structurally similar images, methods 700A-C, etc.). The first data input may correspond to a first device and / or a first structure of the first device. The second data input may correspond to a second device or a second structure of the first device.
[0086] In some embodiments, dataset generator 276 generates historical performance dataset 254 to train, validate, or test a machine learning model (e.g., associated with measurement repetitions between structurally similar images, methods 700A-C, etc.).
[0087] The data input 210 and target output 220 used to train, validate, or test a machine learning model include information for a specific device and / or for a specific structure of a specific device.
[0088] In some embodiments, data items from a training data set are input into a machine learning model one at a time or in groups to train the machine learning model. The machine learning model processes the input to generate an output, for example, a location from which measurements will be taken. The artificial neural network includes an input layer consisting of values in the data points. The next layer is called a hidden layer, and the nodes at the hidden layer each receive one or more input values. Each node contains parameters (e.g., weights) applied to the input values. Thus, each node basically inputs the input value into a multivariate function (e.g., a nonlinear mathematical transformation) to generate an output value. The next layer can be another hidden layer or an output layer. In either case, the nodes at the next layer receive output values from the nodes at the previous layer, and each node applies a weight to the value, and then generates its own output value. This can be performed at each layer. The final layer is an output layer, in which there is a node for each category, prediction, and / or output that the machine learning model can generate.
[0089] Thus, the output may include one or more predictions or inferences (e.g., associated with the determination of a location to be used for measurement). For example, the output predictions or inferences may include one or more predictions of a deformation field, measurements based on the deformation field, and the like. The processing logic determines an error (i.e., a classification error) based on the difference between the output (e.g., prediction or inference) of the machine learning model and the target label associated with the input training data. The processing logic adjusts the weights of one or more nodes in the machine learning model based on the error. An error term or Δ may be determined for each node in the artificial neural network. Based on such an error, the artificial neural network adjusts one or more parameters (the weights of one or more inputs of the node) for one or more of its nodes. The parameters may be updated in a back-propagation manner, so that the nodes at the highest layer are updated first, followed by the nodes at the next layer, and so on. The artificial neural network contains multiple layers of "neurons", each of which receives input values from neurons in the previous layer. The parameters of each neuron include weights associated with the values received from each neuron at the previous layer. Thus, adjusting the parameters may include adjusting the weights assigned to each input of one or more neurons at one or more layers in the artificial neural network.
[0090] After one or more rounds of training, the processing logic can determine whether a stopping criterion has been met. The stopping criterion can be a target level of accuracy, a target number of processed images from the training data set, a target change in a parameter over one or more previous data points, a combination thereof, and / or other criteria. In some embodiments, the stopping criterion is met when at least a minimum number of data points have been processed and at least a threshold accuracy has been achieved. The threshold accuracy can be, for example, 70%, 80%, or 90% accuracy. In some embodiments, the stopping criterion is met if the accuracy of the machine learning model has stopped improving. If the stopping criterion has not been met, further training is performed. If the stopping criterion has been met, the training can be completed. Once the machine learning model has been trained, a retained portion of the training data set can be used to test the model.
[0091] Figure 3 3 is a block diagram illustrating a system 300 for generating an output including a measured position (position data 360) according to some embodiments. The system 300 is used to associate a measured position with a trained machine learning model (e.g., with measurement repetitions between structurally similar images, methods 700A-C, etc.) (e.g., Figure 1 The position data 360 is determined by a position determiner 172 ).
[0092] At block 310, the system 300 (e.g., Figure 1 The automatic measurement system 110 of the embodiment of the present invention performs data partitioning of the historical image data 354 (e.g., by Figure 1 The system 300 may generate a plurality of feature sets for each of the training set, validation set, and test set.
[0093] At block 312, system 300 performs model training using training set 302 (e.g., by Figure 1 In some embodiments, the system 300 trains multiple models using multiple feature sets of the training set 302 (e.g., a first feature set of the training set 302, a second feature set of the training set 302, etc.). For example, the system 300 trains the machine learning model to generate a first trained machine learning model using a first feature set in the training set, and to generate a second trained machine learning model using a second feature set in the training set. In some embodiments, the training is performed according to FIG. 4 .
[0094] At block 314, system 300 performs model validation using validation set 304 (e.g., by Figure 1700A-C). The system 300 uses a corresponding feature set of the validation set 304 to validate each trained model (e.g., associated with measurement duplication between structurally similar images, methods 700A-C, etc.). For example, the system 300 uses a first feature set in the validation set to validate a first trained machine learning model, and uses a second feature set in the validation set to validate a second trained machine learning model. At box 314, the system 300 determines the accuracy of each of the one or more trained models (e.g., through model validation), and determines whether one or more of the trained models have an accuracy that satisfies a threshold accuracy. In response to determining that one or more of the trained models have an accuracy that satisfies the threshold accuracy, the process returns to box 312, where the system 300 further performs model training on the one or more models. In response to determining that one or more of the trained models have an accuracy that satisfies the threshold accuracy, the process continues to box 316 for the one or more models.
[0095] In one embodiment, at block 316, the system 300 performs model selection to select one or more of the trained machine learning models for use in production. In some embodiments, block 316 is omitted.
[0096] At block 318, system 300 performs model testing using test set 306 (e.g., by Figure 1 The system 300 uses the first feature set in the test set to test the first trained machine learning model 308. The system 300 tests the first trained machine learning model using the first feature set in the test set to determine that the first trained machine learning model meets the threshold accuracy (e.g., based on the first feature set of the test set 306). In response to the accuracy of the selected model 308 not meeting the threshold accuracy (e.g., the selected model 308 overfits the training set 302 and / or the validation set 304 and is not applicable to other data sets, such as the test set 306), the process continues to box 312, where the system 300 further performs model training (e.g., retraining) using further training data. In response to determining that the selected model 308 has an accuracy that meets the threshold accuracy based on the test set 306, the process continues to box 320. In at least box 312, the model learns patterns in the historical image data to enable the model to receive a template image with a selected measurement location and one or more target images lacking an indication of the measurement location, and output the measurement location of the one or more target images. At block 318 , the system 300 applies the model to the remaining data (eg, the test set 306 ) to test the predictions (eg, the measured locations of the target images).
[0097] At block 320, the system 300 uses a trained model (e.g., the selected model 308) to receive current image data 346 (e.g., Figure 1The trained model may be used to determine (e.g., extract) position data 360 from the current image data 146 of the image processing unit 100. The position data 360 may indicate a measurement location from which an automatic measurement may be performed. In some embodiments, the trained model may further output a measurement using the determined measurement location. In an embodiment, the trained machine learning model may enable measurement repetition between structurally similar images.
[0098] In some embodiments, current image data 346 is received. In some embodiments, a model is retrained based on the current image data 346. In some embodiments, a new model is trained based on the current image data 346.
[0099] In some embodiments, one or more of blocks 310-320 occur in various orders and / or with other operations not presented and described herein. In some embodiments, one or more of blocks 310-320 are not performed. For example, in some embodiments, one or more of the data partitioning of block 310, the model validation of block 314, the model selection of block 316, and / or the model testing of block 318 are not performed.
[0100] Figure 4A is a block diagram illustrating training a machine learning model to perform image registration and / or determination of measurement locations in accordance with some embodiments.
[0101] In some embodiments, FIG. 4 depicts training and using a machine learning model 430. In some embodiments, training includes using a training data set including historical image data of a structure, the historical image data including multiple images of the structure. In some embodiments, for each iteration of training the machine learning model, the training data input includes a source image 401 and a target image 402, which are given to the machine learning model as training inputs. In some embodiments, the source image 401 is a labeled image including a source mask 405, the source mask containing labels for classifications of different regions of the source image 401. In some embodiments, a region 422 of the source image 401 and the source mask 405 may be labeled as a coating, a region 424 of the source image 401 and the source mask 405 may be labeled as a gate, and a region 426 of the source image 401 and the source mask 405 may be labeled as a mask. In some embodiments, the machine learning model 430 includes at least one of a CNN or an encoder-decoder network.
[0102] In some embodiments, CNN is a deep learning model that can be used for visual data (e.g., image data). In some embodiments, CNN can use convolutional layers to automatically learn basic features from input. Such layers can use filters to extract patterns and pass them through activation functions to achieve nonlinearity. In some embodiments, pooling layers can reduce spatial dimensions, and fully connected layers can perform classification.
[0103] In some embodiments, the encoder-decoder network can be used for sequence-to-sequence tasks where the input and output have different lengths or structures. In some embodiments, the encoder processes the input sequence (e.g., using a recurrent neural network (RNN) or a transformer) to create a compressed representation (e.g., called a context vector). In some embodiments, the decoder generates an output sequence based on the context vector (e.g., using an RNN or transformer network).
[0104] In some embodiments, source image 401 is marked to create source mask 405 for the source image. In some embodiments, source mask 405 is an annotation of the structural composition of the source image. For example, a coating, a gate, a drain, a mask, etc. can be marked to create a source mask. In some embodiments, region 422 of source image 401 and source mask 405 can be marked as a coating, region 424 of source image 401 and source mask 405 can be marked as a gate, and region 426 of source image 401 and source mask 405 can be marked as a mask. In some embodiments, source image 401 is manually marked. In some embodiments, source image 401 is automatically marked (e.g., using another trained machine learning model that has been trained to perform pixel-level classification and / or segmentation of images of a device).
[0105] In some embodiments, the source image 401 and the target image 402 may be provided as training inputs to the machine learning model 430. In some embodiments, the source mask 405 may also be provided as training inputs to the machine learning model 430. In some embodiments, the training data set may include multiple images. Multiple inputs may be assembled from multiple images. Each input may include a selection of a first image serving as a source image, a mask or other mark or annotation associated with the source image, and a second image serving as a target image. In an embodiment, each selected source image should be associated with a mask or other pixel-level annotation. The target image may or may not be associated with a mask. In the case where the target image is associated with a mask, the mask associated with the target image may not be used for training purposes. For the training of the machine learning model 430, many training data inputs may be assembled and may be input into the machine learning model 430 one at a time to perform an iterative training process.
[0106] In some embodiments, the number of unlabeled images can be substantially greater than the number of labeled (and usable as source images) images. However, this is not a problem because each different combination of source and target images can be used as a different training data input. For example, the same source image 401 can be paired with multiple different target images to form different data inputs. Thus, even a small number of labeled images can be used to form a large number of data inputs for training the machine learning model 430. In an example, the labeled multiple images can be 4 images out of a total of 100 images. Even this small number of images with an even smaller number of labeled images can result in a very large number of training data inputs. For example, 4 labeled images and 100 total images may result in approximately 400 different data inputs.
[0107] In some embodiments, at each iteration of training the machine learning model 430, the source image 401 and the target image 402 of the data input are input into the machine learning model 430. The machine learning model performs a non-rigid registration of the source image 401 to the target image 402, and outputs a deformation field 410 (also referred to as a registration flow) indicating the transformation applied to the source image 401 to achieve the target image 402.
[0108] In the field of image processing and computer vision, a deformation field (also called a displacement field or registration field or registration flow) is a fundamental concept used in various tasks such as image registration and image warping. It is a mathematical representation of the spatial correspondence between two images or image volumes. A deformation field comprises a set of vectors representing the spatial displacement of each pixel or voxel in one image to its corresponding position in the other image. Each vector in the field indicates how much a pixel should be moved in the x, y (and in 3D applications, z) directions to align with its corresponding pixel in the reference image. Depending on the application, the deformation field can be represented in different ways, but one common representation is a vector grid or a dense matrix. Each element in the grid corresponds to a specific pixel or voxel position in the source image, and its value represents the displacement required to map the position to the target image.
[0109] In registration, each pixel in the source image is transferred or registered to a pixel in the target image in a way that maintains structure and coherence. Registration between the source image and the target image can be achieved based on the similarity of the two images. Rigid registration is the simplest way to transfer each point and includes rotation and translation of the original image. However, rigid registration does not take into account the different shapes between the source image and the target image, which is common in semiconductor manufacturing. Affine registration allows the shearing and scaling of the source image to register the source image to the target image, but does not take into account the change in shape between the images. Non-rigid registration can align images with slightly different shapes. For non-rigid registration, local deformation can be performed in the source image to register the source image to the target image. Local deformation can include moving one or more points independently of one or more other points so that the points of the source image are aligned with the corresponding points of the target image. Non-rigid transformations can retain more complex structural deformations and variability between images, which rigid registration or affine registration may not retain. Thus, in an embodiment, the machine learning model 430 performs a non-rigid registration between the source image 401 and the target image 402 .
[0110] In an embodiment, the machine learning model 430 outputs a deformation field 410 (e.g., a registration field) between the source image 401 and the target image 402 based on the result of performing a registration (e.g., a non-rigid registration) between the source image 401 and the target image 402. In some embodiments, the deformation field 410 may represent a spatial mapping between corresponding points in the source image 401 and the target image 402, thereby indicating a local displacement or deformation that results in alignment between the two images. In an embodiment, the deformation field 410 is a two-dimensional deformation field and includes a magnitude and direction associated with each pixel in the source image 401 along the x-axis (dx) and along the y-axis (dy).
[0111] The deformation field 410 can be applied to the source image 401 to generate a distorted source image 411. In embodiments, the deformation field 410 can be applied to the source image 401 to generate a distorted version of the source image 401, which has been distorted or otherwise modified to resemble the target image as a result of the non-rigid registration. In some embodiments, the deformation field 410 can also be applied to the source mask 405 to create a distorted source mask 415.
[0112] Once fully trained, the deformation field 410 can be applied to the source image 401 to reproduce the target image 402. However, before the machine learning model 430 is fully trained, the deformation field 410 may not be accurate, and applying the deformation field 410 to the source image 401 may not accurately reproduce the target image 402. Thus, one or more errors can be determined based at least in part on the deformation field 410. In one embodiment, the error is determined based on applying the deformation field 410 to the source image 401 (to result in a distorted source image 411) and the source mask 405 (to result in a distorted source mask 415). The distorted source image 411 and the distorted source mask 415 can be compared to the target image 402 to determine the differences therebetween. These differences can then be used to update the nodes of the machine learning model 430 using back propagation. In some embodiments, one or more similarity metrics are calculated based on the comparison of the distorted source image 411 to the target image 402. In some embodiments, multiple (e.g., two) similarity metrics are calculated based on the comparison of the distorted source image 411 to the target image 402. In some embodiments, additional similarity metrics are calculated based on a comparison of the warped source mask 415 to the target image 402. Each of the similarity metrics can be used to update weights associated with nodes of the machine learning model 430 (e.g., via back-propagation) to improve the ability of the machine learning model 430 to register the source image to the target image (e.g., to improve the accuracy of the deformation field 410 output by the machine learning model 430(), which can be applied to the source image to cause the source image to match or approximately match the target image).
[0113] In the context of comparing a target image to a source image, a similarity metric loss function is a mathematical function used to quantify the similarity or dissimilarity between the two images. The purpose of using such a loss function is to train the model to minimize the difference between the target image and the source image, thereby improving the model's ability to perform tasks such as image comparison, image registration, image reconstruction, or image retrieval. There are various commonly used similarity metric loss functions, depending on the specific task and the characteristics of the image. Some of the popular ones include:
[0114] 1) Mean Squared Error (MSE) Loss: This is the basic and widely used loss function, which calculates the average squared difference between corresponding pixels of the target image and the source image. It penalizes larger differences more heavily;
[0115] 2) L1 loss (mean absolute error): Similar to MSE, but it calculates the mean absolute difference between pixels. L1 loss is less sensitive to outliers than MSE;
[0116] 3) Structural Similarity Index (SSIM) loss; SSIM is a perceptual-based metric that measures the similarity between two images by comparing brightness, contrast, and structure. It takes into account both local and global image features;
[0117] 4) Cosine similarity loss: This measures the cosine of the angle between the target image and the source image vector, treating them as high-dimensional feature vectors;
[0118] 5) Triplet loss: Triplet loss is commonly used in tasks such as image retrieval and face recognition. It aims to ensure that the distance between the target image and the source image is smaller than the distance between the target image and other non-matching images;
[0119] 6) Contrastive loss: This loss function encourages similar images to be closer in feature space and dissimilar images to be farther apart; and
[0120] 7) Pearson Correlation Coefficient (PCC): A statistical measure of the linear relationship between two continuous variables. It is used to quantify how well the relationship between two variables can be described by a straight line.
[0121] In some embodiments, the similarity metric calculated based on the comparison of the warped source image 411 and the target image includes a structural similarity index (SSI) loss and / or a Pearson correlation coefficient (PCC) loss. Other similarity metrics may be used additionally or alternatively.
[0122] In the context of comparing a target image and a source image, the Pearson correlation coefficient can be used as a similarity metric loss. To do this, the image should first be converted into a one-dimensional array or vector. Each element of the vector represents a pixel intensity value or feature value of the image. Then, the Pearson correlation coefficient between the target image vector and the source image vector can be calculated. The Pearson correlation coefficient ranges from -1 to +1. A value of +1 indicates a perfect positive linear relationship between the two images, which means that they change directly and proportionally. A value of -1 indicates a perfect negative linear relationship, implying that one image increases while the other decreases. Values close to 0 suggest that there is little or no linear relationship between the two images. In an embodiment, in order to use the Pearson correlation coefficient as a similarity metric loss, the goal can be to maximize the PCC during the training process. In this way, the machine learning model 430 learns to find patterns and similarities between the target image and the source image that can be described by a linear relationship.
[0123] In some embodiments, the PCC loss measures the linear correlation between two sets of numerical data. For example, an image can be represented as numerical data, such as pixel values or extracted features. The image can be converted into a feature vector or matrix. The vector can then be flattened or reshaped for alignment, and the PCC between the two vectors can be calculated. In some embodiments, a higher positive PCC indicates more similarity / correlation, while a negative value indicates dissimilarity / irrelevance. In some embodiments, the calculated SSI loss can have values ranging from -1 to 1, where values from -1 to 0 indicate negative correlation, values of 0 indicate irrelevance, and values from 0 to 1 indicate positive correlation. In some embodiments, the weights of the machine learning model can be updated based on one or more similarity metric losses.
[0124] The Structural Similarity Index (SSIM) is a perceptual metric that evaluates the similarity between two images. It is designed to capture the structural information and visual quality of an image by measuring the perceptual change in structural information between a reference (original or source) image and a distorted (target) image. The SSIM index takes into account three main components of an image: brightness, contrast, and structure. It is based on the understanding that the human visual system is highly sensitive to these aspects of an image. The SSIM index ranges from -1 to 1, where 1 indicates a perfect match between the two images, 0 indicates no similarity, and -1 implies a perfect mismatch. The formula for calculating the SSIM index is as follows:
[0125] SSIM(x,y)=l(x,y)·C(x,y)·s(x,y)
[0126] Where l(x,y) is the brightness similarity, c(x,y) is the contrast similarity, and s(x,y) is the structural similarity. Brightness similarity measures the similarity of overall brightness and intensity between two images. It is calculated using the mean and variance of the pixel intensities in the two images. Contrast similarity quantifies the difference in contrast, or local variations in pixel intensity. It is calculated based on the covariance of pixel intensities in the reference and distorted images. Structural similarity takes into account patterns and structures in the images. It is determined by the covariance of pixel intensities after taking into account brightness and contrast information. Higher SSIM values indicate greater structural similarity between images.
[0127] In some embodiments, the structural integrity loss of the distorted source mask 415 can be calculated. In some embodiments, calculating the structural integrity loss can refer to a calculated value that quantifies the degradation or distortion of important visual features that define the structure of an image (e.g., source image 401). The structural integrity loss can be used to measure the deviation between the original image (e.g., source mask 405) and the processed image (e.g., distorted source mask), thereby indicating the degree of structural information loss. In some embodiments, the structural integrity loss is calculated on the distorted source mask between marked pixels (source mask). In some embodiments, if the adjacent pixels are different, the difference between the adjacent pixels is calculated as 1, and if the adjacent pixels are the same, the difference is calculated as 0. In some embodiments, the structural integrity loss can be represented by the equation:
[0128]
[0129] In some embodiments, the weights of the machine learning model may be updated based on the loss (e.g., via backpropagation).
[0130] In some embodiments, the weights of the machine learning model can be updated (e.g., via back propagation) based on the loss and one or more similarity metric losses. In some embodiments, the structural integrity loss and one or more similarity metric losses can be combined, and the weights of the machine learning model can be updated based on the combined structural integrity loss and one or more similarity metric losses. In some embodiments, each of the calculated similarity metric losses is used to determine, via back propagation, adjustments to be made to weights associated with nodes of the machine learning model 430. By calculating multiple different similarity loss functions and updating the machine learning model 430 based on each of the similarity loss functions, training of the machine learning model 430 can be achieved more quickly and with a smaller set of training data than would otherwise be achievable.
[0131] Figure 4B4 is a block diagram illustrating the use of a machine learning model to perform image registration between input source and target images according to some embodiments. In some embodiments, the inference phase of the machine learning model 430 can be performed after the machine learning model 430 is trained. The inference phase includes inputting a source image 401 and a target image 402 into the machine learning model 430, and receiving a deformation field 410 as an output. In some embodiments, an input including a template image and a test image may be accepted, wherein the template image includes markings of a first position and a second position associated with a first measurement of the template image. In some embodiments, the first and second positions are marked on the source image. In some embodiments, a separate input indicating the coordinates of the first and second positions is provided with the source image. The output deformation field can be applied to the source image to create the target image again, together with an indication of the distorted coordinates of the first and second positions in the target image from the source image. In some embodiments, the template image (e.g., source image) and the test image (e.g., target image) can be provided as input to the trained machine learning model. In some embodiments, an output including a third position and a fourth position on the test image can be received from the trained machine learning model, and the third position and the fourth position can be used to determine the first measurement of the test image.
[0132] In an embodiment, in order to accurately replicate measurements from a template image or source image (where the locations for one or more measurements have been marked) to other images, at least two points can be registered between the source image and the target image for the measurement. Using the registration flow, the two points of each measurement can be registered to determine the location on the target image from which the measurement is generated. The determined points can then be used to automatically generate measurements (e.g., by measuring the distance between points in the target image or distorted source image).
[0133] The output position output by the trained machine learning model to be used for the measurement in the target image may be close to correct, but may not be completely correct. Thus, in some embodiments, one or more additional operations are performed on the source image, the distorted source image, and / or the target image to improve the accuracy of the position. In some embodiments, using image registration in a trained machine can relieve a user (e.g., an engineer) from performing many manual steps to align structures.
[0134] FIG. 5A to FIG. 5C is an example of precise point placement according to some embodiments.
[0135] refer to Figure 5AIn some embodiments, precise point placement may include receiving inputs including a template image 511A and a test image, wherein the template image 511A includes markings of a first location 501A and a second location 502A on the template image 511A associated with a first measurement 531A of the template image 511A (also referred to as a source image). In some embodiments, the first location 501A may be the coordinates of a particular pixel on the template image 511A. In some embodiments, the template image 511A and the test image (also referred to as a target image) may be provided as inputs to a trained machine learning model.
[0136] In some embodiments, the template signal may be extracted along a first line 541A intersecting a first position 501A and a second position 502A on the template image 511A.
[0137] refer to Figure 5B , in some embodiments, includes a third location 503B having a marking (e.g., Figure 5A ) and the fourth position 504B (e.g., corresponding to the first position 501A of Figure 5A The output of a test image 512B (e.g., a distorted source image, a distorted template image, etc.) of the training machine learning model may be output by the trained machine learning model and may be associated with a second measurement 531B of the test image 512B (e.g., also referred to as a distorted source image, a distorted template image, etc.). In some embodiments, the trained machine learning model outputs a deformation field, and the processing logic applies the deformation field to Figure 5A The template image 511A is used to generate a test image 512B (e.g., a distorted template image) corresponding to the target image (e.g., distorted using a deformation field to look like the target image). Figure 5B As shown, on the test image 512B (eg, the distorted template image) associated with the second measurement 531B of the test image 512B, the test image 512B (eg, the distorted template image) may include a portion corresponding to the third position 503B. Figure 5A The first position 501A of the distortion and the fourth position 504B corresponding to Figure 5A The distorted second position 502A.
[0138] In some embodiments, the test signal may be extracted along a second line 551B that intersects a third location 503B and a fourth location 504B on a test image 512B (eg, a distorted template image, a distorted source image, etc.).
[0139] In some embodiments, the distorted source image or the distorted template image includes an adjusted third position and an adjusted fourth position. In some embodiments, the adjusted third position and the adjusted fourth position are placed using a precise point. Figure 5B For more details, see 7A to 7C Instructions in .
[0140] In an example of precise point placement using DTW, such as FIG. 5A to FIG. 5C As can be seen in FIG. 1 , the third position 503B and the fourth position 504B can be adjusted in the test image (e.g., the distorted source image) to match the positions of the points placed in the template image (e.g., the first position 501A, the second position 502A). In some embodiments, DTW can be used to match the positions of the automatically generated points on the test image with the points placed on the test image (in terms of signal similarity).
[0141] In some embodiments, the third location (e.g., Figure 5C The third position 503C) or Figure 5C At least one of the fourth positions 504C of the test signal can be adjusted using a DTW transform of the test signal to the template signal to produce a test signal after DTW. In some embodiments, the third position can represent a position of a measurement point before DTW, and the measurement point can be adjusted based on the DTW transform. In some embodiments, the test signal can be a query sequence and the template signal can be a reference sequence. In some embodiments, the template signal and the test signal each include a first line (e.g., Figure 5A 541A) and a second line (eg, Figure 5B In some embodiments, the average value of the pixel values along the first line and the second line may be an average grayscale value of the pixels, etc.
[0142] In some embodiments, after the DTW transformation of the test signal to the template signal, a test signal after DTW is generated. In some embodiments, the third position after DTW (eg, corresponding to the point after DTW) corresponds to a more accurate point placement on the test signal after DTW.
[0143] In some examples, according to some embodiments, precise point placement includes using local ratio adjustment. In some embodiments, for more finely adjusted measurement iterations on the template image, the ratio of the closest local minimum and maximum on the signal for each location on the signal can be found.
[0144] The ratio can be calculated as (point - local minimum) / (local maximum - local minimum) and after calculation on the template image, the point in the test image can be fine-tuned.
[0145] In some embodiments, a template signal local maximum may be identified, wherein the template signal local maximum is a region that is consistent with the template image (eg, Figure 5A The first position (eg, Figure 5A The local maximum value closest to the first point on the template signal corresponding to the first position 501A).
[0146] In some embodiments, a template signal local minimum may be identified, where the template signal local minimum is a local minimum that is consistent with the template image (e.g., Figure 5A The first position (eg, Figure 5A The local minimum value closest to the first point on the template signal corresponding to the first position 501A).
[0147] In some embodiments, a test signal local maximum may be identified, where the test signal local maximum is a signal that is consistent with a test image (eg, a distorted source image, Figure 5B A third position (eg, Figure 5B The point after DTW on the test signal corresponding to the third position 503B) is closest to the local maximum.
[0148] In some embodiments, a test signal local minimum may be identified, where the test signal local minimum is a local minimum that is consistent with a test image (eg, a distorted source image, Figure 5B A third position (eg, Figure 5B The point after DTW on the test signal corresponding to the third position 503B) is closest to the local minimum.
[0149] In some embodiments, the points represent points after DTW before applying the DTW transform.
[0150] In some embodiments, a first ratio may be calculated that represents a relationship between a proximity of the first point on the template signal to a local maximum of the template signal and a proximity of the first point on the template signal to a local minimum of the template signal.
[0151] In some embodiments, the third position (eg, Figure 5B In some embodiments, after non-rigid registration of the test image to the template image 511A, the third position 503B of the test image 512B (e.g., the distorted template image) may correspond to the first position 501A of the template image 511A.
[0152] In some embodiments, after precise point placement (e.g., using DTW and / or local ratio adjustment by adjusting the positions to cause the ratios to be approximately equal), Figure 5C The third position 503C can be more accurately placed on the distorted source image 512C and achieve a more accurate measurement of the test image. Similarly, the fourth position 504B can be adjusted (e.g., to obtain the fourth position 504C), thereby achieving a more precise and accurate measurement of the test image.
[0153] In some embodiments, the process described and performed on a single location (e.g., the location of the measurement) can be applied to a second location, a third location, a fourth location, etc. In some embodiments, the method can also accommodate multiple template images when each signal at each specified location of the template image is passed through DTW along with the test image, and the final location of the point on the test image is the average of each approximation (e.g., from each template image), giving a more robust approximation of the location on the test image (e.g., measurement repetitions).
[0154] In some embodiments, by predicting measurement locations on the template image based on a DTW approach rather than simple edge detection, precise point placement can be robust to noise and other artifacts.
[0155] FIG. 6A to FIG. 6B is an example of precise point placement using reference points according to some embodiments. In some embodiments, to restrict the movement of a point, a reference such as a reference point and / or a reference line may be used.
[0156] refer to Fig. 6A In some embodiments, the reference point may be a reference line (e.g., a first reference line 630A). In some embodiments, the first position 601A and the second position 602A of the template image 611A are located at a vertical distance 620A from the first reference line 630A. Figure 6B , the third position 603B and the fourth position 604B of the test image 612B are located at a vertical distance 620B (approximately equal to) from the second reference line 632B. Fig. 6A In some embodiments, the first reference line 630A and the second reference line 632B may be positioned at features of the template image 511A and the test image 512B, respectively (e.g., the top of a feature or the bottom of a feature, such as a coating, a drain, a source, etc.).
[0157] 7A to 7Cis a flow chart of methods 700A-C for associating measurement duplication between structurally similar images according to certain embodiments. In some embodiments, methods 700A-C are performed by processing logic, which includes hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, processing device, etc.), software (such as instructions running on a processing device, a general purpose computer system, or a dedicated machine), firmware, microcode, or a combination thereof. In one embodiment, method 700A can be performed by a computer system, such as Figure 1 Computer system architecture 100. In other or similar embodiments, one or more operations of method 700A may be performed by one or more other machines not depicted in the figures. In some embodiments, methods 700A-C are performed at least in part by automatic measurement system 110. In some embodiments, method 700A is performed by client device 120 (e.g., automatic measurement component 114) and / or automatic measurement system 110 (e.g., automatic measurement component 114). In some embodiments, method 700B is performed by server machine 180 (e.g., training engine 182, etc.). In some embodiments, method 700C is performed by prediction server 112 (e.g., automatic measurement component 114) and / or client device 120 (e.g., automatic measurement component 114). In some embodiments, a non-transitory storage medium stores instructions that, when executed by a processing device (e.g., a processing device of automatic measurement system 110, server machine 180, prediction server 112, client device 120, etc.), cause the processing device to perform one or more of methods 700A-C.
[0158] To simplify explanation, methods 700A-C are depicted and described as a series of operations. However, operations according to the present disclosure may occur in various orders and / or simultaneously, and with other operations not presented and described herein. Furthermore, in some embodiments, not all of the operations shown are performed to implement methods 700A-C according to the disclosed subject matter. Furthermore, those skilled in the art will understand and appreciate that methods 700A-C may alternatively be represented as a series of interrelated states via state diagrams or events.
[0159] Fig. 7A is a flow chart of a method of correlating measurement repetitions between structurally similar images in accordance with aspects of the present disclosure.
[0160] See also Fig. 7AIn some embodiments, processing logic implementing method 700A trains a machine learning model using a training data set including historical image data of a structure, the historical image data including multiple images of the structure. In some embodiments, training includes blocks 701 through 706. In some embodiments, at block 701, processing logic implementing method 700A processes a source image and a target image of the same type of structure using the machine learning model to generate a deformation field between the source image and the target image.
[0161] At block 702, processing logic applies a deformation field to a source mask associated with a source image to create a distorted source mask, wherein the source mask includes markings of one or more regions of the source image. In some embodiments, processing logic generates the source mask based on the markings of the one or more regions of the source image. In some embodiments, processing logic applies the deformation field to the source image to create the distorted source image.
[0162] At block 703, processing logic calculates a first loss based on a comparison of the warped source mask and the target image. In some embodiments, the first loss based on a comparison of the warped source mask and the target image may include a structural integrity loss.
[0163] At block 704, processing logic calculates a second loss based on a comparison of the target image and the warped source image. In some embodiments, processing logic calculates one or more similarity metric losses based on a comparison of the target image and the warped source image. In some embodiments, the one or more similarity metric losses based on a comparison of the target image and the warped source image may include a structural similarity index (SSI) loss and a Pearson correlation coefficient (PCC) loss.
[0164] As described above, in some examples, more than one loss may be calculated. For example, processing logic may calculate a first loss (e.g., a similarity metric loss) based on a comparison of a target image and a distorted source image. Processing logic may also calculate a second loss based on a comparison of a distorted source mask and the target image. In some examples, both the first loss (e.g., based on a comparison of the target image and the distorted source image) and the second loss (e.g., based on a comparison of the distorted source mask and the target image) may be used to train a machine learning model, as described in block 705. In some embodiments, any number of losses (e.g., a first loss, a second loss, a third loss, etc.) may be calculated and may be used to train a machine learning model, as described in block 705.
[0165] At box 705, processing logic trains the machine learning model based at least in part on the first or second loss. In some embodiments, processing logic trains the machine learning model based at least in part on the first and second losses. In some embodiments, processing logic trains the machine learning model based at least in part on one or more similarity metric losses. In some embodiments, processing logic trains the machine learning model based on both a first loss (e.g., based on a comparison of a warped source mask and a target image) and a second loss (e.g., based on a comparison of a target image and a warped source image). In some embodiments, the second loss may include multiple similarity metric losses, including a structural similarity index (SSI) loss, a Pearson correlation coefficient (PCC) loss, and the like. In some embodiments, the machine learning model may include at least one of a CNN or an encoder-decoder network.
[0166] In some embodiments, the processing logic may further receive an input comprising a template image and a test image, wherein the template image comprises a marker of a first position and a second position associated with a first measurement of the template image. In some embodiments, the processing logic may further provide the template image and the test image as input to a trained machine learning model, wherein the trained machine learning model outputs a deformation field based on a non-rigid registration of the template image to the test image. In some embodiments, the processing logic may receive an output from the trained machine learning model comprising a third position and a fourth position on the test image, the third position and the fourth position being usable to determine the first measurement of the test image. In some embodiments, after non-rigidly registering the test image to the template image, the third position of the test image may correspond to the first position of the template image. In some embodiments, during registration, each position from the first image (e.g., the source image) is translated to the second image (e.g., the target image) by a vector describing the motion of the source image to the target image (e.g., represented as a registration flow of the deformation field).
[0167] In some embodiments, in order to accurately copy the first measurement from the template image to the test image, the two locations (e.g., the first location and the second location of the test image) should be registered. Using the registration flow, the location of the measurement on the template image is registered and the corresponding measurement is repeated on the test image.
[0168] In the template image, the placed measurements are done manually. On the test image, the measurements are repeated using registration. In some embodiments, the repeated measurements can be performed using precise point placement techniques (e.g., FIG. 5A to FIG. 5C , Figure 6, and Figure 7C ) for further adjustments.
[0169] In some embodiments, the first position and the second position of the template image are located at a vertical distance from the first reference point, and the third position and the fourth position of the test image are located at a vertical distance from the second reference point. In some embodiments, the reference point can be a reference line, or the reference line can include a reference point.
[0170] In some embodiments, in order to copy the measurement from the template image to other images (e.g., test images), the signal (e.g., template signal and / or test signal) needs to be extracted near each placement point (e.g., position, measurement point, etc.). In some embodiments, the processing logic may further extract the template signal along a first line that intersects the first position and the second position on the template image. In some embodiments, the processing logic may further extract the test signal along a second line that intersects the third position and the fourth position on the test image. In some embodiments, the processing logic may further use a DTW transform of the test signal to the template signal to adjust at least one of the third position or the fourth position. In some embodiments, the template signal and the test signal each include an average of the pixel values of the pixels along the first line and the second line, respectively.
[0171] In some embodiments, the template signal and the test signal can be extracted in the measurement direction (e.g., horizontally). In some embodiments, the signal can be extracted parallel to a first line intersecting the first and second positions on the template image and a second line intersecting the third and fourth positions on the test image.
[0172] In some embodiments, the processing logic may further identify a template signal local maximum, wherein the template signal local maximum is a local maximum value that is closest to a point on the template signal corresponding to the first position on the template image and a point on the template signal corresponding to the first position on the template image. In some embodiments, the processing logic may further identify a template signal local minimum value, wherein the template signal local minimum value is a local minimum value that is closest to a point on the template signal corresponding to the first position on the template image.
[0173] In some embodiments, the processing logic may further identify a test signal local maximum, wherein the test signal local maximum is a local maximum that is closest to a point on the test signal corresponding to a third position on the test image.
[0174] In some embodiments, the processing logic may further identify a test signal local minimum, wherein the test signal local minimum is a local minimum that is closest to a point on the test signal corresponding to a third position on the test image.
[0175] In some embodiments, the processing logic may further calculate a first ratio representing a relationship between a proximity of a point on the template signal to a maximum value of the template signal and a proximity of a point on the template signal to a minimum value of the template signal. In some embodiments, the processing logic may further adjust the third position so that a second ratio representing a relationship between a proximity of a point on the test signal to a maximum value of the test signal and a proximity of a point on the test signal to a minimum value of the test signal is approximately equal to the first ratio.
[0176] In some embodiments, the first and second ratios may be calculated by subtracting the y-value of the signal local minimum from the y-value of the first position and dividing by the y-value of the signal local maximum and the difference between the y-value of the signal local maximum.
[0177] In some embodiments, any number of locations can be identified on the template image and the locations can be registered to the test image.Such locations can correspond to multiple measurements.
[0178] In some embodiments, multiple template images may have locations identified and may be transformed through DTW along with the test image. Repeated measurements (eg, locations corresponding to the measurements) may be based on an average of all approximations of the template image DTW.
[0179] Figure 7B is a method for training a machine learning model (e.g., Figure 1 The position determiner 172 of FIG. 170 ) determines prediction data associated with measured repetitions between structurally similar images (e.g., Figure 1 Flow chart of a method for obtaining measurement data 160).
[0180] See also Figure 7B At block 710 of method 700B, processing logic identifies historical image data (eg, historical image data 144, historical images, etc.) The historical image data may include images from historical substrates, historical features of substrates, historical structures of substrates, and the like.
[0181] In some embodiments, at block 712, processing logic uses a training dataset including historical image data and a training dataset including historical performance data (e.g., deformation fields, warped source masks, losses associated with the warped source masks, one or more similarity metric losses of the target image and the warped source image, Figure 1The machine learning model may be trained based on a target output of historical performance data 154 of the substrate to generate a trained machine learning model. The historical performance data may include historical measurements of the substrate (e.g., features of the substrate, structure of the substrate, etc.) and / or data from historical images, such as values of one or more of measurement accuracy, measurement consistency, etc. The performance data including the historical performance data may include calculation of a loss associated with a distorted source mask after image registration for measurement repetitions between structurally similar images (e.g., structural integrity loss), calculation of a similarity metric loss after image registration for measurement repetitions between structurally similar images (e.g., structural similarity index (SSI) loss, Pearson correlation coefficient (PCC) loss, etc.), etc.
[0182] In some embodiments, at least a portion of performance data 152 is associated with a measured quality of substrates produced by manufacturing equipment 124. Performance data 152 can indicate whether measurements were made correctly and / or accurately. For example, two images may be structurally similar, but not identical. Performance data 152 can indicate that measurements made on a first image are similarly measured on a second structurally similar, but not identical, image (e.g., through image-to-image registration, where each pixel is transferred from the first image to the second image such that the structure and coherence of the images is maintained).
[0183] Historical performance data can be associated with the quality of the image registration, such as by computing a loss associated with a distorted source mask (e.g., a structural integrity loss) and / or a similarity metric loss (e.g., a structural similarity index (SSI) loss, a Pearson correlation coefficient (PCC) loss, etc.) after measuring repeated image registrations between structurally similar images (e.g., structurally similar images of historical substrates).
[0184] Figure 7C is a machine learning model trained to use repeated associations of measurements between structurally similar images (e.g., Figure 1 method 700C of the position determiner 172) and causing corrective action to be performed.
[0185] See also Figure 7C At block 720 of method 700C, input is received including a template image and a test image, wherein the template image includes markings at first and second locations associated with a first measurement of the template image.
[0186] At block 722, processing logic provides the template image and the test image as input to a trained machine learning model (e.g., via Figure 7BThe machine learning model is trained in block 712 of the present invention, wherein the trained machine learning model outputs a deformation field based on a non-rigid registration of the template image to the test image. In some embodiments, the trained machine learning model can be associated with measured repetitions between structurally similar images (e.g., image data). In some embodiments, the machine learning model includes at least one of a CNN or an encoder-decoder network.
[0187] At block 724 , processing logic determines, based on applying the deformation field to the template image, third and fourth locations on the test image corresponding to the first and second locations on the template image, respectively, wherein the third and fourth locations may be used to determine a first measurement of the test image.
[0188] At block 726 , processing logic extracts a template signal along a first line that intersects the first location and the second location on the template image.
[0189] At block 728 , processing logic extracts a test signal along a second line that intersects the third location and the fourth location on the test image.
[0190] In some embodiments, the template signal and the test signal each comprise an average of pixel values of pixels along the first line and the second line, respectively.
[0191] At block 730 , processing logic adjusts at least one of the third position or the fourth position using a DTW transform of the test signal to the template signal.
[0192] At block 732, processing logic identifies a template signal local maximum, wherein the template signal local maximum is a local maximum that is closest to a point on the template signal corresponding to the first location on the template image.
[0193] At block 734, processing logic identifies a template signal local minimum, wherein the template signal local minimum is a local minimum that is closest to a point on the template signal corresponding to the first location on the template image.
[0194] At block 736, processing logic identifies a test signal local maximum, wherein the test signal local maximum is a local maximum that is closest to a point on the test signal corresponding to the third location on the test image.
[0195] At block 738, processing logic identifies a test signal local minimum, wherein the test signal local minimum is the local minimum that is closest to a point on the test signal corresponding to the third position on the test image.
[0196] At block 740 , processing logic calculates a first ratio that represents a relationship between a proximity of a point on the template signal to a maximum value of the template signal and a proximity of a point on the template signal to a minimum value of the template signal.
[0197] At block 742, processing logic adjusts the third position so that a second ratio representing the relationship between the proximity of a point on the test signal to a test signal maximum and the proximity of a point on the test signal to a test signal minimum is approximately equal to the first ratio.
[0198] In some embodiments, after non-rigid registration of the test image to the template image, the third position of the test image corresponds to the first position of the template image.
[0199] In some embodiments, the first position and the second position of the template image are located at a vertical distance from the first reference point, and the third position and the fourth position of the test image are located at a vertical distance from the second reference point. In some embodiments, the reference point can be a reference line, or the reference line can include a reference point.
[0200] In some embodiments, Fig. 7A Blocks 701-708 of include training a machine learning model to repeat measurements between structurally similar images based on a source image, a target image, and / or a source mask. In some embodiments, the trained machine learning model can be used to repeat measurements between structurally similar images based on a source image, a target image, and / or a source mask.
[0201] In some embodiments, the image data 142 may include template images, test images, source images, target images, source masks, etc., and the trained machine learning model of box 722 is trained using data inputs including historical template images, test images, source images, target images, and / or source masks and target outputs including historical performance data 154 (e.g., deformation fields, measurements, etc.).
[0202] In some embodiments, the image data (e.g., for example, template images, test images, source images, target images, source masks, etc.) and the trained machine learning model of block 722 are trained using data inputs including historical template images, test images, source images, target images, and / or source masks and target outputs including historical performance data 154 (e.g., deformation field measurements, etc.). The measurement data 160 of block 724 may be associated with performance data predicted based on the image data (e.g., repeatedly measured performance data).
[0203] Figure 8 is a block diagram illustrating a computer system 800 according to some embodiments. In some embodiments, the computer system 800 is one or more of the client device 120, the automated measurement system 110, the server machine 180, the prediction server 112, and the like.
[0204] In some embodiments, the computer system 800 is connected (e.g., via a network, such as a local area network (LAN), an intranet, an extranet, or the Internet) to other computer systems. In some embodiments, the computer system 800 operates in a client-server environment with the capabilities of a server or client computer, or operates as a peer computer in a peer-to-peer or distributed network environment. In some embodiments, the computer system 800 is provided by a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a network device, a server, a network router, a switch or a bridge, or any device capable of executing an instruction set (continuously or otherwise) for the action taken by the device. In addition, the term "computer" shall include any computer collection that executes an instruction set (or multiple instruction sets) alone or in combination to perform any one or more methods described herein.
[0205] In a further aspect, the computer system 800 includes a processing device 802, a volatile memory 804 (e.g., a random access memory (RAM)), a non-volatile memory 806 (e.g., a read-only memory (ROM) or an electrically erasable programmable ROM (EEPROM)), and a data storage device 818, which communicate with each other via a bus 808.
[0206] In some embodiments, the processing device 802 is provided by one or more processors, such as a general-purpose processor (such as, for example, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor that implements other types of instruction sets, or a microprocessor that implements a combination of types of instruction sets) or a special-purpose processor (such as, for example, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), or a network processor).
[0207] In some embodiments, the computer system 800 further includes a network interface device 822 (e.g., coupled to the network 874). In some embodiments, the computer system 800 also includes a video display unit 810 (e.g., a liquid crystal display (LCD)), an alphanumeric input device 812 (e.g., a keyboard), a cursor control device 814 (e.g., a mouse), and a signal generating device 820.
[0208] In some embodiments, the data storage device 818 includes a non-transitory computer-readable storage medium 824 on which are stored instructions 826 encoding any one or more of the methods or functions described herein, including instructions for Figure 1 The components (eg, the automatic measurement component 114, etc.) of the present invention are encoded with instructions for implementing the methods described herein (eg, one or more of the methods 700A-C).
[0209] In some embodiments, the instructions 826 also reside, in whole or in part, within the volatile memory 804 and / or within the processing device 802 during execution by the computer system 800, and thus, in some embodiments, the volatile memory 804 and the processing device 802 also constitute machine-readable storage media.
[0210] Although the computer-readable storage medium 824 is illustrated as a single medium in the illustrative example, the term "computer-readable storage medium" shall include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more executable instruction sets. The term "computer-readable storage medium" shall also include any tangible medium that can store or encode a set of instructions for execution by a computer, the set of instructions causing the computer to perform any one or more of the methodologies described herein. The term "computer-readable storage medium" shall include, but is not limited to, solid-state memories, optical media, and magnetic media.
[0211] The methods, components, and features described herein may be implemented by independent hardware components, or may be integrated into the functionality of other hardware components such as application specific integrated circuits (ASICs), FPGAs, DSPs, or similar devices. In addition, the methods, components, and features may be implemented by firmware modules or functional circuitry within a hardware device. In addition, the methods, components, and features may be implemented in any combination of hardware devices and computer program components or in a computer program.
[0212] Unless expressly stated otherwise, terms such as "training", "labeling", "providing", "receiving", "applying", "computing", "updating", "extracting", "adjusting", "calculating", "identifying", "determining", "causing", "executing", "obtaining", "accessing", "adding", "using", etc. refer to actions and processes performed or implemented by a computer system that manipulates and transforms data represented as physical (electronic) quantities within computer system registers and memories into other data similarly represented as physical quantities within computer system memories or registers or other such information storage, transmission or display devices. In addition, the terms "first", "second", "third", "fourth", etc. as used herein mean labels for distinguishing between different elements and may not have ordinal meanings according to their numerical labels.
[0213] The examples described herein also relate to devices for performing the methods described herein. Such devices may be specially constructed for performing the methods described herein, or may include a general-purpose computer system selectively programmed by a computer program stored in the computer system. Such a computer program may be stored in a computer-readable tangible storage medium.
[0214] The methods and illustrative examples described herein do not inherently relate to any specific computer or other device. Various general-purpose systems may be used in accordance with the teachings described herein, or it may prove convenient to construct more specialized equipment to perform each of the methods described herein and / or its individual functions, routines, subroutines, or operations. Examples of the structures of various these systems are set forth in the description above.
[0215] The above description is intended to be illustrative rather than limiting. Although the present disclosure has been described with reference to specific illustrative examples and implementations, it will be appreciated that the present disclosure is not limited to the described examples and implementations. The scope of the present disclosure should be determined with reference to the following claims and the full scope of equivalents authorized by the claims.
Claims
1. A non-transitory computer-readable storage medium storing instructions, wherein when the instructions are executed, the processing device performs the following operations, the operations comprising: receiving an input comprising a template image and a test image, wherein the template image comprises a marker at a first position and a second position associated with a first measurement of the template image; providing the template image and the test image as input to a trained machine learning model, wherein the trained machine learning model outputs a deformation field based on a non-rigid registration of the template image to the test image; and Based on applying the deformation field to the template image, third and fourth positions on the test image corresponding to the first and second positions on the template image, respectively, are determined, wherein the third and fourth positions can be used to determine the first measurement of the test image.
2. The non-transitory computer-readable storage medium of claim 1, the operations further comprising: extracting a template signal along a first line intersecting the first position and the second position on the template image; extracting a test signal along a second line intersecting the third position and the fourth position on the test image; and At least one of the third position or the fourth position is adjusted using a dynamic time warping (DTW) transform of the test signal to the template signal. 3 . The non-transitory computer-readable storage medium of claim 2 , wherein the template signal and the test signal each comprise an average of pixel values of pixels along the first line and the second line, respectively.
4. The non-transitory computer-readable storage medium of claim 2, the operations further comprising: identifying a local maximum of a template signal, wherein the local maximum of the template signal is a local maximum closest to a point on the template signal corresponding to the first position on the template image; identifying a local minimum of a template signal, wherein the local minimum of the template signal is a local minimum that is closest to a point on the template signal corresponding to the first position on the template image; identifying a test signal local maximum, wherein the test signal local maximum is a local maximum closest to a point on the test signal corresponding to the third position on the test image; identifying a test signal local minimum, wherein the test signal local minimum is a local minimum that is closest to a point on the test signal corresponding to the third position on the test image; calculating a first ratio, the first ratio representing a relationship between a proximity of the point on the template signal to a maximum value of the template signal and a proximity of the point on the template signal to a minimum value of the template signal; and The third position is adjusted so that a second ratio representing the relationship between the proximity of the point on the test signal to the test signal maximum and the proximity of the point on the test signal to the test signal minimum is approximately equal to the first ratio. 5 . The non-transitory computer-readable storage medium of claim 1 , wherein the third position of the test image corresponds to the first position of the template image after non-rigid registration of the test image to the template image.
6. The non-transitory computer-readable storage medium of claim 2, wherein the first position and the second position of the template image are positioned at a vertical distance from a first reference point, and the third position and the fourth position of the test image are positioned at the vertical distance from a second reference point.
7. The non-transitory computer-readable storage medium of claim 1, wherein the machine learning model comprises at least one of a convolutional neural network (CNN) or an encoder-decoder network.
8. A system, comprising: Memory; and A processing device, coupled to the memory, the processing device being configured to: receiving an input comprising a template image and a test image, wherein the template image comprises a marker at a first position and a second position associated with a first measurement of the template image; providing the template image and the test image as input to a trained machine learning model, wherein the trained machine learning model outputs a deformation field based on a non-rigid registration of the template image to the test image; and Based on applying the deformation field to the template image, third and fourth positions on the test image corresponding to the first and second positions on the template image, respectively, are determined, wherein the third and fourth positions can be used to determine the first measurement of the test image.
9. The system according to claim 8, wherein the processing device is further configured to: extracting a template signal along a first line intersecting the first position and the second position on the template image; extracting a test signal along a second line intersecting the third position and the fourth position on the test image; and At least one of the third position or the fourth position is adjusted using a dynamic time warping (DTW) transform of the test signal to the template signal.
10. The system of claim 9, wherein the template signal and the test signal each comprise an average of pixel values of pixels along the first line and the second line, respectively.
11. The system according to claim 9, wherein the processing device is further configured to: identifying a local maximum of a template signal, wherein the local maximum of the template signal is a local maximum closest to a point on the template signal corresponding to the first position on the template image; identifying a local minimum of a template signal, wherein the local minimum of the template signal is a local minimum that is closest to a point on the template signal corresponding to the first position on the template image; identifying a test signal local maximum, wherein the test signal local maximum is a local maximum closest to a point on the test signal corresponding to the third position on the test image; identifying a test signal local minimum, wherein the test signal local minimum is a local minimum that is closest to a point on the test signal corresponding to the third position on the test image; calculating a first ratio, the first ratio representing a relationship between a proximity of the point on the template signal to a maximum value of the template signal and a proximity of the point on the template signal to a minimum value of the template signal; and The third position is adjusted so that a second ratio representing the relationship between the proximity of the point on the test signal to the test signal maximum and the proximity of the point on the test signal to the test signal minimum is approximately equal to the first ratio.
12. The system of claim 8, wherein the third position of the test image corresponds to the first position of the template image after non-rigid registration of the test image to the template image.
13. The system of claim 9, wherein the first position and the second position of the template image are positioned at a vertical distance from a first reference point, and the third position and the fourth position of the test image are positioned at the vertical distance from a second reference point.
14. The system of claim 8, wherein the machine learning model comprises at least one of a convolutional neural network (CNN) or an encoder-decoder network.
15. A method comprising: Processing a source image and a target image of the same type of structure using a machine learning model to generate a deformation field between the source image and the target image; applying the deformation field to a source mask associated with the source image to create a warped source mask, wherein the source mask contains indicia of one or more regions of the source image; computing a loss based on a comparison of the warped source mask and the target image; and The machine learning model is trained based at least in part on the loss. 16 . The method of claim 15 , further comprising generating the source mask based on markings of one or more regions of the source image.
17. The method according to claim 15, wherein the training further comprises: applying the deformation field to the source image to create a distorted source image; computing one or more similarity metric losses based on a comparison of the target image and the warped source image; and The machine learning model is trained based at least in part on the one or more similarity metric losses.
18. The method of claim 17, wherein the loss based on the comparison of the warped source mask and the target image comprises a structural integrity loss, and wherein the one or more similarity metric losses based on the comparison of the target image and the warped source image comprise a structural similarity index (SSI) loss and a Pearson correlation coefficient (PCC) loss.
19. The method of claim 15, wherein the machine learning model comprises at least one of a convolutional neural network (CNN) or an encoder-decoder network.
20. The method of claim 15, further comprising: receiving an input comprising a template image and a test image, wherein the template image comprises a marker at a first position and a second position associated with a first measurement of the template image; providing the template image and the test image as input to the trained machine learning model, wherein the trained machine learning model outputs a deformation field based on a non-rigid registration of the template image to the test image; and Based on applying the deformation field to the template image, determining a third position and a fourth position on the test image corresponding to the first position and the second position, Wherein the third position and the fourth position may be used to determine the first measurement of the test image.