Method and apparatus for classifying images, and defect inspection system for semiconductor manufacturing process
By generating an attention mask and updating the image classification model, the classification performance degradation of the artificial intelligence classifier when domain shifted image input is solved, and efficient image classification under domain changes is achieved.
Patent Information
- Application Number
- CN202510010132.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-04
- Filing Date
- 2025-01-03
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the classification performance of artificial intelligence classifiers decreases when domain shifted image input, making it difficult to adapt to the modification of image characteristics of domain changes.
By generating an attention mask, perform domain adaptation to the image based on the attention mask, update the image classification model, and use the loss function to optimize the model parameters to improve the accuracy of image classification.
Even if the image domain changes, images can still be effectively classified, improving the accuracy and robustness of image classification.
Smart Images

Figure CN120259712A_ABST
Abstract
Description
[0001] This application claims the priority and benefit of Korean Patent Application No. 10-2024-0001795, filed with the Korean Intellectual Property Office on January 4, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present disclosure relates to a method and apparatus for classifying images. Background Art
[0003] Test-time adaptation is a technique for improving the classification performance of images in a domain different from the domain of images used for training an artificial intelligence classifier (domain-shifted images) when the domain-shifted images are input to the trained artificial intelligence classifier. Generally, an adaptation process for adapting the trained artificial intelligence classifier to the domain-shifted images may be performed such that the parameters of the encoder of the artificial intelligence classifier are updated, and then the classification of the domain-shifted images is performed.
[0004] If the domain of an image is shifted (or changed), the characteristics of the image may be modified, and an artificial intelligence classifier trained with images in a different domain may exhibit low classification performance for the domain-shifted images. Summary of the Invention
[0005] In one general aspect, a method for classifying an image includes: generating an attention mask based on features of the image; updating an image classification model by performing domain adaptation on the image based on the attention mask; and determining a class of the image using the updated image classification model.
[0006] The step of generating the attention mask may include: generating spatial features by embedding the image into a latent space; and generating an attention mask based on the spatial features.
[0007] The step of updating the image classification model may include: calculating a loss function by masking a difference between the image and a reconstructed image using the attention mask, the reconstructed image being generated by decoding the spatial features; and updating the image classification model based on a calculation result of the loss function.
[0008] The step of generating an attention mask based on the spatial features may include: generating a set of attention maps based on the spatial features; merging a plurality of attention maps included in the set of attention maps; and generating an attention mask by sampling the merged attention maps.
[0009] The step of merging the plurality of attention maps included in the set of attention maps may include: merging the plurality of attention maps by performing tiling or layer averaging.
[0010] The step of generating an attention mask by performing sampling on the merged attention map may include: generating an attention mask by performing Bernoulli sampling or thresholding on each patch of the merged attention map.
[0011] The step of generating an attention mask based on spatial features may include: generating an attention vector based on spatial features; and generating an attention mask by performing sampling on each of a plurality of elements in the attention vector.
[0012] Each of the plurality of elements included in the attention vector may correspond to a respective patch belonging to the spatial features, and each of the plurality of elements may be determined based on a respective similarity between a respective feature corresponding to the respective patch and a reference feature of the spatial features.
[0013] The respective similarity may be a cosine similarity between the respective feature corresponding to the respective patch and the reference feature.
[0014] The step of generating spatial features by embedding the image into a latent space may include: transforming the image; and generating spatial features by embedding the transformed image into a latent space.
[0015] In another general aspect, an apparatus for classifying an image includes: one or more processors and a memory, wherein the memory stores instructions that are configured to cause the one or more processors to perform processing that includes: generating an attention mask based on features of the image; using the attention mask to update an artificial intelligence (AI) model through domain adaptation of a domain of the image; and using the updated AI model to determine a class of the image.
[0016] The AI model may be trained based on images of a source domain, and the domain of the image may be different from the source domain.
[0017] The step of generating an attention mask may include: generating spatial features by embedding the image into a latent space; and generating an attention mask based on the spatial features.
[0018] The step of updating the AI model may include: calculating a loss function by masking a difference between the image and a reconstructed image using the attention mask, the reconstructed image being generated based on decoding the spatial features; and updating the AI model based on the calculated loss.
[0019] The step of generating an attention mask based on spatial features may include: generating a set of attention maps based on spatial features; merging a plurality of attention maps included in the set of attention maps; and generating an attention mask by performing sampling on the merged attention map.
[0020] The steps of generating an attention mask based on spatial features may include: generating an attention vector based on the spatial features; and generating an attention mask by performing sampling on multiple elements in the attention vector.
[0021] The steps of generating spatial features by embedding the image into a latent space may include: transforming the image; and generating spatial features by embedding the transformed image into the latent space.
[0022] In another general aspect, a defect inspection system for a semiconductor manufacturing process includes: a domain adaptation device configured to generate an attention mask based on an image acquired from an inspection device in a manufacturing environment, and update an artificial intelligence (AI) model trained based on a source image in a source domain by performing domain adaptation using the attention mask, and an artificial intelligence model configured to be updated by domain adaptation and determine a category of the image after domain adaptation.
[0023] When the domain adaptation device performs domain adaptation using the attention mask, the domain adaptation device may further be configured to: decode spatial features of the image output by the artificial intelligence model to reconstruct the image; and update the artificial intelligence model by calculating a loss function using the image, the reconstructed image, and the attention mask.
[0024] When the domain adaptation device updates the artificial intelligence model by calculating a loss function using the image, the reconstructed image, and the attention mask, the domain adaptation device may further be configured to: calculate the loss function by masking a difference between the image and the reconstructed image using the attention mask, and update the artificial intelligence model based on the calculation of the loss function.
[0025] Other features and aspects will be apparent from the following detailed description, the drawings, and the claims. Description of the Drawings
[0026] Figure 1 Illustrates a system for classifying an image according to one or more embodiments.
[0027] Figure 2 Illustrates a method for performing attention-based domain adaptation by a system for classifying an image according to one or more embodiments.
[0028] Figure 3 Illustrates a mask generator according to one or more embodiments.
[0029] Figure 4 Illustrates a method for generating a mask according to one or more embodiments.
[0030] Figure 5Illustrates an updated device for classifying images according to one or more embodiments.
[0031] Figure 6 Illustrates a system for classifying images according to another embodiment.
[0032] Figure 7 Illustrates a method for performing attention-based domain adaptation by a system for classifying images according to another embodiment.
[0033] Figure 8 Illustrates a mask generator according to another embodiment.
[0034] Figure 9 Illustrates a method for generating a mask according to another embodiment.
[0035] Figure 10 Illustrates a system for classifying images of a semiconductor manufacturing process according to one or more embodiments.
[0036] Figure 11 Illustrates input images of different domains according to one or more embodiments.
[0037] Figure 12 Illustrates a neural network according to one or more embodiments.
[0038] Figure 13 Illustrates a system for classifying images according to one or more embodiments.
[0039] Throughout the drawings and the detailed description, unless otherwise described or provided, the same or similar reference numerals will be understood to represent the same or similar elements, features, and structures. The drawings may not be drawn to scale, and for clarity, illustration, and convenience, the relative sizes, proportions, and depictions of elements in the drawings may be exaggerated. Detailed Description
[0040] The following detailed description is provided to assist the reader in obtaining a thorough understanding of the methods, devices, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will be apparent after understanding the disclosure of this application. For example, the order of operations described herein is merely an example and is not limited to the order of operations set forth herein. Rather, the order of operations may be changed as will be apparent after understanding the disclosure of this application, except for operations that must occur in a particular order. In addition, descriptions of features known after understanding the disclosure of this application may be omitted for greater clarity and conciseness.
[0041] The features described herein can be implemented in various forms and should not be construed as limited to the examples described herein. Instead, the examples described herein are provided only to illustrate some of the many possible ways of implementing the methods, devices, and / or systems described herein that will be apparent after understanding the disclosure of the present application.
[0042] The terms used herein are for the purpose of describing various examples only and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are intended to include the plural forms as well. As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more of them. By way of non-limiting example, the terms "comprises," "comprising," and "having" specify the presence of the stated features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.
[0043] Throughout the specification, when a component or element is described as "connected to," "coupled to," or "joined to" another component or element, it can be directly "connected to," "coupled to," or "joined to" the other component or element, or there can reasonably be one or more other components or elements intervening between them. When a component or element is described as "directly connected to," "directly coupled to," or "directly joined to" another component or element, there may be no other elements intervening between them. Similarly, expressions such as "between" and "immediately between" and "adjacent to" and "immediately adjacent to" can also be interpreted as described above.
[0044] Although terms such as "first," "second," and "third" or A, B, (a), (b), etc. may be used herein to describe various members, components, regions, layers, or parts, these members, components, regions, layers, or parts are not limited by these terms. Each of these terms is not used to define, for example, the nature, order, or sequence of the corresponding member, component, region, layer, or part, but is only used to distinguish the corresponding member, component, region, layer, or part from other members, components, regions, layers, or parts. Thus, the first member, first component, first region, first layer, or first part mentioned in the examples described herein can also be referred to as the second member, second component, second region, second layer, or second part without departing from the teachings of the examples.
[0045] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs, based on an understanding of the disclosure of this application. Unless explicitly defined as such herein, terms (such as those defined in a common dictionary) shall be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the disclosure of this application, and shall not be interpreted in an idealized or overly formal sense. The term "may" as used herein with respect to an example or embodiment (e.g., what may be included or implemented with respect to an example or embodiment) means that there is at least one example or embodiment that includes or implements such a feature, without limiting all examples thereto.
[0046] The artificial intelligence (AI) model of the present disclosure may be a machine learning model that learns at least one task and may be implemented as a computer program in the form of instructions executed by a processor. The task learned by the AI model represents a task to be solved or a task to be executed by machine learning. The AI model may be implemented as a computer program executed on a computing device, may be downloaded via a network, or may be sold as a product. Optionally, the AI model may be networked with various devices. Figure 12 An example of a neural network of the AI model is shown.
[0047] Figure 1 A system for classifying an image according to one or more embodiments is shown. Figure 2 A method for performing attention-based domain adaptation by a system for classifying an image according to one or more embodiments is shown. Figure 3 A mask generator according to one or more embodiments is shown, and Figure 4 A method for generating a mask according to one or more embodiments is shown.
[0048] In some embodiments, the system 100 for classifying an image may perform attention-based domain adaptation on an input image to determine the class of the input image, update the image classification device 110 (e.g., the artificial intelligence model in the image classification device 110) according to the result of the attention-based domain adaptation, and may classify the input image into a class using the updated image classification device 110.
[0049] In some embodiments, before performing attention-based domain adaptation, an artificial intelligence (AI) model in the image classification device 110 may be trained based on images of the source domain. That is, the AI model trained based on the source domain may be adapted to a new target domain different from the source domain through "attention-based domain adaptation using images in the target domain", and then the image classification device 110 may classify the category of the input image by using the inference of the domain-adapted AI model. Through attention-based domain adaptation, the domain-adapted AI model in the image classification device 110 can perform the classification task well not only for images in the source domain but also for images in the target domain.
[0050] Referring Figure 1 , the image classification system 100 according to one or more embodiments may include an image classification device 110, an image transformer 120, a decoder 130, a mask generator 140, and a domain adapter 150. The image classification device 110 may include an encoder 111 and a classifier 112.
[0051] In some embodiments, the image classification device 110 may be pre-trained with images in the source domain. When an image is input to the pre-trained image classification system 100, the image classification system 100 may perform attention-based domain adaptation for the image classification device 110 by using the input image, and then the image classification device 110 updated through attention-based domain adaptation may determine the category of the input image. Even if the domain of the input image is a domain different from the source domain, the image classification device 110 may accurately classify the category of the input image through attention-based domain adaptation.
[0052] On the other hand, the image classification system 100 may further include a domain determiner (not shown) for determining whether the domain of the input image is the same as the training domain (or source domain) of the image classification device 110. The domain determiner may determine whether domain adaptation needs to be performed based on determining whether the domains are the same.
[0053] For example, when the domain determiner determines that the input image belongs to the source domain, the image classification device 110 may classify the category of the input image without domain adaptation. On the other hand, when the domain determiner determines that the input image belongs to another domain different from the source domain, the image classification system 100 may perform domain adaptation on the input image and may perform category classification on the input image by using the image classification device 110 updated through domain adaptation.
[0054] Next, the image classification system 100 can perform attention-based domain adaptation on the input image (e.g., perform attention-based domain adaptation using the input image), and can use the domain-adapted image classification device 110 "updated by attention-based domain adaptation" to determine the category of the input image. Since the image classification system 100 performs attention-based domain adaptation on the input image, the parameters of the image classification device 110 can be updated, and the updated image classification device 110 can infer the category of the input image.
[0055] Referring to Figure 1 and Figure 2 , the encoder 111 of the image classification device 110 can generate spatial features by embedding the input image x into the latent space ( v )(S110). The encoder 111 can output spatial features according to the input image x , regardless of the domain of the input image x . If the image transformed by the image transformer 120 is input to the encoder 111, the encoder 111 can generate spatial features according to the transformed input image . The spatial features generated by the encoder 111 can be transmitted to the classifier 112, the decoder 130, and the mask generator 140.
[0056] The classifier 112 of the image classification device 110 can receive spatial features from the encoder 111, and can generate a set of attention maps A based on the received spatial features (S120). In some embodiments, the set of attention maps may include the attention maps generated at each layer of the classifier 112 . When there are n layers in the classifier 112, the set of attention maps A can be expressed as Equation 1 below.
[0057] (Equation 1)
[0058] In some embodiments, the image transformer 120 can use a predetermined method to transform the input image. For example, the image transformer 120 can perform image enhancement on the input image, or can perform a masking operation on the input image. Optionally, the image transformer 120 can transmit the input image to the image classification device 110 as it is without transformation (identity function). Referring to Figure 1 , the image transformer 120 is represented by a dashed line to illustrate the case where the image transformer 120 transmits the input image to the image classification device 110 as it is without transforming the input image.
[0059] In some embodiments, the decoder 130 may process the spatial features received from the encoder 111. Decode to reconstruct the input image x (S130). In other words, the decoder 130 may Reconstruct the input image x To output the reconstructed image .
[0060] In some embodiments, the mask generator 140 may use the attention map set A received from the classifier 112 to generate a mask that provides attention. m (or attention mask) (S140). In other words, providing an attention mask m Can be generated based on the input image. Mask m It can be used to select image patches that are robust to domain shifts (or changes).
[0061] Reference Figure 3 , the mask generator 140 according to one or more embodiments may include an attention map merger 141 and a sampler 142 .
[0062] Reference Figure 3 and Figure 4 The attention map merger 141 may merge the multiple attention maps included in the attention map set A to generate a merged attention map ( S141 ) For example, the attention map merger 141 may merge a plurality of attention maps by performing rollout or layer average.
[0063] In some embodiments, the sampler 142 may Perform sampling to generate the mask m (S142). The sampler 142 may be based on the merged attention map For example, the sampler 142 may generate a mask by performing Bernoulli sampling on each image block of the merged attention map or performing thresholding on each image block of the merged attention map according to a predetermined value.
[0064] Return to reference Figure 1 and Figure 2 , the domain adapter 150 may use the mask generated by the mask generator 140 m For the input image x and the reconstructed image received from the decoder 130 The domain adapter 150 may calculate the loss function shown in Equation 2 by masking the difference (or distance) between them (S150).
[0065] (Equation 2)
[0066] In Equation 2, N represents the total number of image patches, and i represents the index of each image patch. In Equation 2, the loss function can be calculated by multiplying the L2 norm value between the th i image patch of the reconstructed image x and the i th m image patch of the input image i by the value of the mask . In some embodiments, can be 0 or 1. According to Equation 2, after the difference between the input image and the reconstructed image is masked by the mask, the masked differences can be summed up so that the sum value is output as the result of the loss function. In other words, the domain adapter 150 can perform domain adaptation on the regions selected by the mask in the input image.
[0067] Thereafter, the image classification system 100 can update the image classification device 110 (e.g., the artificial intelligence model in the image classification device 110) based on the result of the loss function determined according to the input image, the reconstructed image, and the mask (S160).
[0068] Since the image classification device 110 is updated based on the result of the loss function determined according to the input image, the reconstructed image, and the mask, the attention-based domain adaptation can be completed, and then the updated image classification device 110 can classify the input image used for domain adaptation. That is to say, the attention can represent the difference between the input image and the reconstructed image, and can be achieved by the AI learning scheme through the mask.
[0069] In some embodiments, the parameters of the encoder 111 of the image classification device 110 can be updated by attention-based domain adaptation, and the parameters of the classifier 112 can not be updated. In other cases, the image classification system 100 can update the parameters of the classifier 112 by performing few-shot adaptation using different images in the same target domain as the target domain of the input image.
[0070] An image classification system 100 according to one or more embodiments may update an encoder 111 of an image classification device 110 by performing attention-based domain adaptation (first domain adaptation) on an input image, and may update a classifier 112 of the image classification device 110 by performing few-shot adaptation (second domain adaptation) using different images in the same domain as the domain of the input image. Thereafter, the image classification system 100 may classify the category of the input image using the updated image classification device 110.
[0071] In some embodiments, in the attention-based domain adaptation for the input image, the image classification system 100 may generate a mask using a set of attention maps generated by the encoder 111 and the classifier 112, and may update the parameters of the encoder 111 based on a loss function calculated by using the generated mask.
[0072] In addition, in the few-shot adaptation for the input image, the image classification system 100 may generate features of labeled images in the same domain as the domain of the input image, may perform classification based on the generated features, and may update the classifier 112 based on a loss function calculated by using the classification result and the label of the input image.
[0073] Figure 5 is a block diagram showing an updated device for classifying images according to one or more embodiments.
[0074] Referring to Figure 5 , the updated image classification device 210 may classify the category of the input image for attention-based domain adaptation. The input image for attention-based domain adaptation may belong to the source domain or the target domain.
[0075] In some embodiments, the encoder 211 updated by attention-based domain adaptation may generate spatial features from the input image, and the classifier 212 may infer the category y of the input image based on the spatial features generated by the encoder 211. The classifier 212 may not be updated by attention-based domain adaptation.
[0076] Optionally, the classifier 212 may be updated by attention-based domain adaptation and / or few-shot adaptation, and the classifier 212 updated by attention-based domain adaptation and / or few-shot adaptation may be used to infer the category of the input image.
[0077] In some embodiments, the image classification system 100 may perform attention-based domain adaptation on an input image using the image classification device 110, the image transformer 120, the decoder 130, the mask generator 140, and the domain adapter 150, and may determine the class of the input image using the image classification device 110 updated by the attention-based domain adaptation. When the class of the input image is inferred, the image transformer 120, the decoder 130, the mask generator 140, and the domain adapter 150 are not used.
[0078] As described above, the image classification system 100 may perform attention-based domain adaptation on an input image using a mask generated based on the attention map of the input image, so that even if the domain of the input image is different from the source domain or the domain of the input image is unknown, the class of the input image is successfully classified.
[0079] Figure 6 A system for classifying an image according to another embodiment is shown Figure 7 A method for performing attention-based domain adaptation by a system for classifying an image according to another embodiment is shown Figure 8 A mask generator according to another embodiment is shown, and Figure 9 A method for generating a mask according to another embodiment is shown.
[0080] Referring to Figure 6 , the image classification system 300 according to one or more embodiments may include an image classification device 310, an image transformer 320, a decoder 330, a mask generator 340, and a domain adapter 350. The image classification device 310 may include an encoder 311 and a classifier 312.
[0081] In another embodiment, the image classification system 300 may perform attention-based domain adaptation on an input image, and may determine the class of the input image using the image classification device 310 updated by the attention-based domain adaptation. Since the image classification system 300 performs attention-based domain adaptation on the input image, the parameters of the encoder 311 of the image classification device 310 may be updated, and then the class of the input image may be determined by the updated image classification device 310.
[0082] Referring to Figure 6 and Figure 7 , the encoder 311 of the image classification device 310 may generate spatial features ( x ) (S210) by embedding the input image v into a latent space. The encoder 311 may output spatial features according to the input image x regardless of the domain of the input image x . When the image transformed by the image transformer 320When the image is input to the encoder 311, the encoder 311 may Generate spatial features . Spatial characteristics The spatial features generated by the encoder 311 may include N image blocks. It can be transmitted to the decoder 330 and the mask generator 340.
[0083] In another embodiment, the image converter 320 may convert the input image using a predetermined method. For example, the image converter 320 may perform image enhancement on the input image, or may perform a masking operation on the input image. Alternatively, the image converter 320 may transmit the input image as is to the image classification device 310 without converting the input image.
[0084] In another embodiment, the decoder 330 may receive the spatial features received from the encoder 311. Decode to reconstruct the input image x (S220). In other words, the decoder 330 may Reconstruct the input image x To output the reconstructed image .
[0085] In another embodiment, the mask generator 340 may be based on the spatial features received from the encoder 311. To generate an attention-providing mask (or attention mask) m (S230) Mask m Can be used to select image patches that are robust to domain shifts (or changes). Figure 8 , the mask generator 340 according to one or more embodiments may include an attention vector generator 341 and a sampler 342.
[0086] Reference Figure 8 and Figure 9 , the attention vector generator 341 can be based on the spatial features generated by the encoder 311 To generate the attention vector a (S231). Equation 3 can be expressed according to the spatial characteristics Generated attention vector a .
[0087] (Equation 3)
[0088] Referring to Equation 3, the attention vector a Can include multiple elements ( s k )(in, k is 1 andn a random integer between), and each of the plurality of elements included in the attention vector may correspond to an image patch belonging to the spatial feature. The attention vector may be determined based on the similarity between the reference feature of the spatial feature and the "feature corresponding to the image patch belonging to the spatial feature .
[0089] In another embodiment, the encoder 311 or the mask generator 340 may use various methods to determine the reference feature of the spatial feature . For example, the reference feature of the spatial feature may be determined as the average value of the features of each image patch of the spatial feature , and the method for determining the reference feature of the spatial feature is not limited thereto.
[0090] In another embodiment, the similarity ( k ) between the k-th feature ( f ) corresponding to the k-th image patch of the spatial feature and the reference feature ( f cls ) may be determined as in Equation 4 below. s k ).
[0091] (Equation 4)
[0092] In Equation 4, s k may represent the cosine similarity between the k -th feature ( f k ) and the reference feature ( f cls ). However, the method for calculating the similarity between the feature corresponding to the image patch of the spatial feature and the reference feature is not limited thereto.
[0093] In another embodiment, the sampler 342 may generate a mask ( a ) by sampling the attention vector m (S232). The sampler 342 may perform sampling based on the magnitude of each element of the attention vector a . For example, the sampler 342 may generate a mask by performing Bernoulli sampling on each element of the attention vector or thresholding each element of the attention vector according to a predetermined value. In another embodiment, a binary mask may be generated by assigning 1 to the sampled elements of the attention vector and 0 to the non-sampled elements.
[0094] Refer toFigure 6 and Figure 7 the domain adapter 350 can calculate a loss function (S240) by masking "the difference (or distance) between the input image m and the reconstructed image received from the decoder 330" using the mask generated by the mask generator 340. The domain adapter 350 can calculate the result of the loss function of Equation 2 to perform domain adaptation on the regions of the input image selected by the mask. x Thereafter, the image classification system 300 can update the image classification device 310 (S250) based on the loss function determined according to the input image, the reconstructed image, and the mask. Based on the loss function calculated according to the input image, the reconstructed image, and the mask, the encoder 311 of the image classification device 310 can be updated to complete attention-based domain adaptation, and then the updated image classification device 310 can classify the input image for domain adaptation.
[0095] In another embodiment, when performing attention-based domain adaptation on an input image, the classifier 312 of the image classification device 310 may not be updated. Even if a mask is generated based on the spatial features generated by the encoder 311 and the encoder 311 of the image classification device 310 is updated based on the result of domain adaptation performed on a part (e.g., image patches of an image) selected by the mask, the classifier 312 of the image classification device 310 may not be updated.
[0096] In this case, in order to improve the classification performance for the input image even if the input image does not belong to the source domain of the image classification device 310, the image classification system 300 can update the classifier 312 of the image classification device 310 by performing few-shot adaptation using at least one of the same-domain images within the same domain as the domain of the input image. The image classification device 310 including the classifier 312 updated by few-shot adaptation and the encoder 311 updated by attention-based domain adaptation can successfully classify the input image that does not belong to the source domain.
[0097] As described above, the image classification system 300 can perform attention-based domain adaptation on the input image using the mask generated according to the attention vector of the input image, so that even if the domain of the input image changes, the category of the input image can be successfully classified.
[0098]
[0099] Figure 10 A system for classifying images of a semiconductor manufacturing process (e.g., a defect inspection system for a semiconductor manufacturing process) according to one or more embodiments is shown, and Figure 11 input images of different domains according to one or more embodiments are shown.
[0100] In some embodiments, in an in-fabrication environment of a semiconductor manufacturing process, the domain adaptation device 400 can update the image classification device 500 by performing adaptation at test time on an image transmitted from an inspection device. When the input image is an image of a manufactured product generated in real time, the adaptation at test time can be performed in real time, so that the image classification model can be adapted at test time. However, the test timing of the adaptation here is not limited to this. The image classification device 500 updated by the adaptation at test time can infer the category of the image. x When performing a test, the adaptation at test time is used to update the image classification device 500. When the input image is an image of a manufactured product generated in real time, the adaptation at test time can be performed in real time, so that the image classification model can be adapted at test time. However, the test timing of the adaptation here is not limited to this. The image classification device 500 updated by the adaptation at test time can infer the category of the image. x In some embodiments, the domain adaptation device 400 can perform attention-based domain adaptation as a mechanism for adaptation at test time for the input image.
[0101] In some embodiments, the domain adaptation device 400 can perform adaptation at test time when an image is transmitted from an inspection device of a semiconductor, and can allow the image classification device 500 to be adapted to the domain of the transmitted image, even if the domain of the image is not always the same or the domain of the image is unknown.
[0102] When the image classification device 500 is trained based on images of a source domain, the image input from the inspection device can be an image of the source domain or an image of a target domain different from the source domain. Even if an image of any domain is input, the image classification device 500 can accurately classify the category of the input image during the semiconductor manufacturing process through adaptation at test time.
[0103] A change in the domain of the input image (e.g., a change from the source domain to the target domain) can occur due to a change in the generation of the product, an addition of a manufacturing step of the product, etc. Optionally, the change in the domain can occur due to a change in the design of the semiconductor, an addition / change of an inspection device, an addition / change of an inspection method, etc. If a change in the domain of the input image occurs during the semiconductor manufacturing process, the domain adaptation device 400 according to one or more embodiments can update the image classification device 500 through attention-based domain adaptation, and the updated image classification device 500 can accurately classify the category of the input image.
[0104] Referring to Figure 10 and Figure 11 , when an image is input from an inspection device, the domain adaptation device 400 according to one or more embodiments can update the image classification device 500 by performing attention-based domain adaptation (S310). Since the effect of domain adaptation can be reduced when the difference between the domain of the input image and the "source domain to which the images used for training the image classification device 500 belong" is large, the domain adaptation device 400 according to one or more embodiments can update the image classification device 500 by performing attention-based domain adaptation on regions of the input image that are robust to changes in the domain.
[0105] In some embodiments, the mask generator of the domain adaptation device 400 may generate an attention mask based on an input image. The mask generated by the mask generator according to the attention to the input image may allow selective execution of domain adaptation on "regions that are robust to domain changes".
[0106] In some embodiments, the classifier of the image classification device 500 trained based on images of a source domain may generate a set of attention maps according to an input image, and the set of attention maps generated by the classifier may be used as the attention to the input image. The attention maps may represent regions that are robust to domain changes and may represent regions that can improve the classification performance of the image. The mask generator may generate such a mask: the attention may be provided by combining multiple attention maps within the set of attention maps generated by the classifier and performing sampling on the combined attention maps.
[0107] Optionally, the encoder of the image classification device 500 trained based on images of a source domain may generate spatial features according to an input image, and the spatial features generated by the encoder may be used to generate the attention to the input image. The mask generator may generate an attention vector according to the spatial features generated by the encoder, and may generate an attention mask by performing sampling on the elements of the attention vector (e.g., each of the multiple elements).
[0108] In some embodiments, the domain adapter of the domain adaptation device 400 may perform domain adaptation on the regions selected by the mask in the input image.
[0109] Referring to Figure 10 and Figure 11 , the domain adaptation device 400 may use a decoder, a mask generator, and a domain adapter to perform attention-based domain adaptation and may update the image classification device 500. The domain adaptation device 400 may also use an image transformer for attention-based domain adaptation.
[0110] Thereafter, the image classification device 500 updated by attention-based domain adaptation may determine the x category y of the input image (S320). That is, in the step of inferring the category of the input image, the configuration of the domain adaptation device 400 for attention-based domain adaptation (e.g., an image transformer, a decoder, a mask generator, and a domain adapter) may not be used.
[0111] As described above, in the fab environment of the manufacturing process, the domain adaptation device 400 may update an artificial intelligence (AI) model through attention-based domain adaptation without additional training of the classification model, and the updated AI model may classify the category of the input image.
[0112] Figure 12Shows a neural network according to one or more embodiments.
[0113] Referring Figure 12 , a neural network 1200 according to one or more embodiments may include an input layer 1210, a hidden layer 1220, and an output layer 1230. Each of the input layer 1210, the hidden layer 1220, and the output layer 1230 may include a corresponding set of nodes, and the connection strength between the nodes may correspond to weight values (e.g., Figure 12 W in 1,1 W 1,2 W i,j etc.). This may be referred to as connection weights. The set of nodes included in each of the input layer 1210, the hidden layer 1220, and the output layer 1230 may be fully connected to each other or not fully connected. In some embodiments, the number of parameters (weight values and bias values) may be equal to the number of connections within the neural network 1200.
[0114] The input layer 1210 may include a set of input nodes x1 to x i , and the number of input nodes x1 to x i may correspond to the number of independent input variables. To train the neural network 1200, a training set may be input to the input layer 1210, and if a test data set is input to the input layer 1210 of the trained neural network 1200, an inference result (e.g., a category) may be output from the output layer 1230 of the trained artificial neural network 1200. In some embodiments, the input layer 1210 may have a structure suitable for processing large-scale inputs.
[0115] The hidden layer 1220 may be provided between the input layer 1210 and the output layer 1230, and may include at least one hidden layer or 12201 to 1220 n hidden layers. The output layer 1230 may include at least one output node or y1 to y j output nodes. An activation function may be used in the hidden layer 1220 and the output layer 1230. In some embodiments, the neural network 1200 may be learned by adjusting the weight values of the hidden nodes included in the hidden layer 1220.
[0116] Figure 13 Shows a system for classifying images according to one or more embodiments.
[0117] An image classification system according to an embodiment may be implemented as a computer system (e.g., a computer-readable medium). Referring Figure 13, the computer system 1300 includes a processor 1310 and a memory 1320 (the processor 1310 represents any single processor or any combination of processors (e.g., CPU, GPU, accelerator, etc.)). The memory 1320 can be connected to the processor 1310 to store various information for driving the processor 1310 or at least one program executed by the processor 1310.
[0118] The processor 1310 can implement the functions, processes, or methods proposed in the embodiments. The operation of the computer system 1300 according to some embodiments can be implemented by the processor 1310.
[0119] The memory 1320 can be provided inside or outside the processor, and the memory can be connected to the processor through various known means. The memory can be various forms of volatile or non-volatile storage media. For example, the memory can include a read-only memory (ROM) or a random access memory (RAM).
[0120] The embodiments can be implemented by a program (in the form of source code, executable instructions, etc.) that implements the functions corresponding to the configurations of the embodiments or a recording medium (not the signal itself) recording the program. This program can be easily implemented by those of ordinary skill in the art to which the present disclosure pertains according to the descriptions of the above embodiments. That is, using the above descriptions, engineers, etc. can easily, for example, formulate source code corresponding to the descriptions, compile the source code into instructions, and the instructions, when executed by the processor 1310, will cause the processor to perform physical operations similar to the above descriptions. Specifically, the methods according to some embodiments (e.g., image preprocessing methods, etc.) can be implemented in the form of program instructions, which can be executed by various computer devices and recorded on a computer-readable medium. The computer-readable medium can independently or combinatorially include program instructions, data files, data structures, etc. The program instructions recorded on the computer-readable medium can be specifically designed and configured for the embodiments or can be known to those skilled in the art of computer software for use. The computer-readable recording medium can include hardware devices configured to store and execute program instructions. For example, the computer-readable recording medium can be a hard disk, magnetic media (such as floppy disks and magnetic tapes), optical media (such as CD-ROMs and DVDs), magneto-optical media (such as floppy optical disks), ROM, RAM, flash memory, etc. The program instructions can include high-level language codes executable by a computer using an interpreter, etc., and machine language codes generated by a compiler.
[0121] Herein, regarding Figures 1 to 13The described computing devices, electronic devices, processors, memories, displays, information output systems and hardware, storage devices, and other devices, apparatuses, units, modules, and components are implemented by or represent hardware components. Examples of hardware components that can be used, where appropriate, to perform the operations described in this application include controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtractors, multipliers, dividers, integrators, and any other electronic components configured to perform the operations described in this application. In other examples, one or more of the hardware components that perform the operations described in this application are implemented by computing hardware (e.g., by one or more processors or computers). A processor or computer can be implemented by one or more processing elements (such as, logic gate arrays, controllers, and arithmetic logic units, digital signal processors, microcomputers, programmable logic controllers, field programmable gate arrays, programmable logic arrays, microprocessors, or any other device or combination of devices configured to respond and execute instructions in a defined manner to achieve a desired result). In one example, a processor or computer includes or is connected to one or more memories that store instructions or software executed by the processor or computer. The hardware components implemented by the processor or computer can execute instructions or software for performing the operations described in this application (such as, an operating system (OS) and one or more software applications running on the OS). The hardware components can also access, manipulate, process, create, and store data in response to the execution of the instructions or software. For simplicity, the singular terms "processor" or "computer" can be used in the description of the examples described in this application, but in other examples, multiple processors or computers can be used, or a processor or computer can include multiple processing elements, or multiple types of processing elements, or both. For example, a single hardware component or two or more hardware components can be implemented by a single processor, or two or more processors, or a processor and a controller. One or more hardware components can be implemented by one or more processors or a processor and a controller, and one or more other hardware components can be implemented by one or more other processors or additional processors and additional controllers. One or more processors or a processor and a controller can implement a single hardware component or two or more hardware components. The hardware components can have any one or more of different processing configurations, examples of any one or more of different processing configurations include single processors, independent processors, parallel processors, single instruction single data (SISD) multiprocessing, single instruction multiple data (SIMD) multiprocessing, multiple instruction single data (MISD) multiprocessing, and multiple instruction multiple data (MIMD) multiprocessing.
[0122] Figures 1 to 13The method of performing the operations described in this application, as shown, is performed by computing hardware (e.g., by one or more processors or computers), which is implemented as described above to execute instructions or software to perform the operations performed by the method described in this application. For example, a single operation or two or more operations may be performed by a single processor, or two or more processors, or a processor and a controller. One or more operations may be performed by one or more processors or a processor and a controller, and one or more other operations may be performed by one or more other processors or additional processors and additional controllers. One or more processors or a processor and a controller may perform a single operation or two or more operations.
[0123] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement the hardware components and perform the method described above may be written as a computer program, code segment, instruction, or any combination thereof, for individually or jointly instructing or configuring one or more processors or computers to operate as a machine or special-purpose computer to perform the operations performed by the hardware components and method described above. In one example, the instructions or software include machine code (such as machine code generated by a compiler) that is directly executed by one or more processors or computers. In another example, the instructions or software include high-level code that is executed by one or more processors or computers using an interpreter. The instructions or software may be written in any programming language based on the block diagrams and flowcharts shown in the figures and the corresponding descriptions herein, which disclose algorithms for performing the operations performed by the hardware components and the method described above.
[0124] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement the hardware components and execute the methods described above, along with any associated data, data files, and data structures, can be recorded, stored, or fixed in one or more non-transitory computer-readable storage media, or recorded, stored, or fixed on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk memory, hard disk drive (HDD), solid state drive (SSD), flash memory, card-type memory (such as, multimedia card or micro card (e.g., Secure Digital (SD) or Extreme Digital (XD))), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store instructions or software and any associated data, data files, and data structures in a non-transitory manner and provide the instructions or software and any associated data, data files, and data structures to one or more processors or computers such that the one or more processors or computers can execute the instructions. In one example, the instructions or software and any associated data, data files, and data structures are distributed across a networked computer system such that the instructions and software and any associated data, data files, and data structures are stored, accessed, and executed by one or more processors or computers in a distributed manner.
[0125] Although the present disclosure includes specific examples, it will be apparent after understanding the disclosure of this application that various changes in form and detail can be made in these examples without departing from the spirit and scope of the claims and their equivalents. The examples described herein will be considered only as descriptive and not for purposes of limitation. The description of a feature or aspect in each example will be considered applicable to similar features or aspects in other examples. Appropriate results can be achieved if the described techniques are performed in a different order, and / or if the components in the described system, architecture, device, or circuit are combined in a different manner, and / or replaced or supplemented by other components or their equivalents.
[0126] Accordingly, in addition to the above disclosure, the scope of the disclosure may be defined by the claims and their equivalents, and all variations within the scope of the claims and their equivalents should be construed as being included in the disclosure.
Claims
1. A method for classifying an image, the method comprising: Generating an attention mask based on features of the image; Updating an image classification model by performing domain adaptation on the image based on the attention mask; And Using the updated image classification model to determine the category of the image.
2. The method according to claim 1, wherein, The step of generating the attention mask includes: Generating spatial features by embedding the image into a latent space; and Generating an attention mask based on the spatial features.
3. The method according to claim 2, wherein The step of updating the image classification model includes: Calculating a loss function by masking the difference between the image and the reconstructed image using the attention mask, the reconstructed image being generated by decoding the spatial features; and Updating the image classification model based on the calculation result of the loss function.
4. The method according to claim 2, wherein The step of generating an attention mask based on the spatial features includes: Generating a set of attention maps based on the spatial features; Merging a plurality of attention maps included in the set of attention maps; and Generating an attention mask by sampling the merged attention maps.
5. The method according to claim 4, wherein The step of merging the plurality of attention maps included in the set of attention maps includes: Merging the plurality of attention maps by performing tiling or layer averaging.
6. The method according to claim 4, wherein, The step of generating an attention mask by sampling the merged attention maps includes: Generating an attention mask by performing Bernoulli sampling or thresholding on each patch of the merged attention maps.
7. The method according to claim 2, wherein, The step of generating an attention mask based on the spatial features includes: Generating an attention vector based on the spatial features; and Generating an attention mask by sampling each of a plurality of elements in the attention vector.
8. The method according to claim 7, wherein Each of the plurality of elements included in the attention vector corresponds to a corresponding patch belonging to the spatial features, and Each of the plurality of elements is determined based on a respective similarity between a respective feature corresponding to the corresponding patch and a reference feature of the spatial features.
9. The method according to claim 8, wherein The respective similarity is a cosine similarity between the respective feature corresponding to the corresponding patch and the reference feature.
10. The method according to any one of claims 2 to 9, wherein, The step of generating spatial features by embedding the image into a latent space includes: Transforming the image; and Generating spatial features by embedding the transformed image into a latent space.
11. An apparatus for classifying an image, the apparatus comprising: One or more processors and a memory, Wherein the memory stores instructions configured to cause the one or more processors to perform processing, the processing including: Generating an attention mask based on features of the image; Updating an artificial intelligence model by performing domain adaptation on the domain of the image using the attention mask; and Using the updated artificial intelligence model to determine the category of the image.
12. The device according to claim 11, wherein, The artificial intelligence model is trained based on images of a source domain, and the domain of the image is different from the source domain.
13. The device according to claim 11, wherein, The step of generating the attention mask includes: Generating spatial features by embedding the image into a latent space; and Generating an attention mask based on the spatial features.
14. The device according to claim 13, wherein, The step of updating the artificial intelligence model includes: Calculating a loss function by masking the difference between the image and the reconstructed image with an attention mask, where the reconstructed image is generated based on decoding spatial features; and Updating the artificial intelligence model based on the calculated loss.
15. The device according to claim 13, wherein, The step of generating an attention mask based on spatial features includes:[[]] Generating a set of attention maps based on spatial features; Merging a plurality of attention maps included in the set of attention maps; and Generating an attention mask by performing sampling on the merged attention maps.
16. The device according to claim 13, wherein, The step of generating an attention mask based on spatial features includes:[[]] Generating an attention vector based on spatial features; and Generating an attention mask by performing sampling on a plurality of elements in the attention vector.
17. The device according to claim 13, wherein, The step of generating spatial features by embedding the image into a latent space includes:[[]] Transforming the image; and Generating spatial features by embedding the transformed image into a latent space.
18. A defect inspection system for a semiconductor manufacturing process, the defect inspection system comprising:[[]] A domain adaptation device configured to: generate an attention mask based on an image obtained from an inspection device in a manufacturing environment, and perform domain adaptation by using the attention mask to update an artificial intelligence model trained based on a source image in a source domain, and An artificial intelligence model configured to: be updated by domain adaptation and determine the category of the image after domain adaptation.
19. The defect inspection system according to claim 18, wherein,[[]] When the domain adaptation device performs domain adaptation by using the attention mask, the domain adaptation device is further configured to:[[]] Decode the spatial features of the image output by the artificial intelligence model to reconstruct the image; and Update the artificial intelligence model by calculating a loss function by using the image, the reconstructed image, and the attention mask.
20. The defect inspection system according to claim 19, wherein,[[]] When the domain adaptation device updates the artificial intelligence model by calculating a loss function by using the image, the reconstructed image, and the attention mask, the domain adaptation device is further configured to: calculate the loss function by masking the difference between the image and the reconstructed image with the attention mask, and update the artificial intelligence model based on the calculation of the loss function.
Citation Information
Patent Citations
Display panel and electric apparatus
KR1020240001795A