Method and apparatus for training a pre-trained ai model

By calculating the gradient similarity of the loss function between the pre-trained domain and the target domain, and optimizing the parameter update of the diffusion probability model, the problem of long training time of the diffusion probability model is solved, and more efficient domain adaptation is achieved.

CN120278293APending Publication Date: 2025-07-08SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510009501.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-05
Filing Date
2025-01-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The diffusion probability model is computationally expensive and time-consuming during training, making it difficult to efficiently adapt to the new target domain.

Method used

By calculating the gradient similarity of the loss function between the pretrained domain and the target domain, using low-rank adaptation and bias term fine-tuning methods, the parameter update rate of the pretrained AI model is optimized and the transfer learning time is shortened.

Benefits of technology

The training efficiency of the diffusion probability model in the target domain is improved, the computing resource requirements are reduced, and the training time is shortened.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278293A_ABST
    Figure CN120278293A_ABST
Patent Text Reader

Abstract

A method and apparatus for training a pre-trained AI model are provided. The method for training a pre-trained AI model includes calculating a first gradient of a loss function of a first image generated by the pre-trained AI model and a second gradient of a loss function of a second image generated by the pre-trained AI model, calculating a similarity between the first gradient and the second gradient, and updating a pre-trained AI model based on the similarity.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority and the benefit of Korean Patent Application No. 10-2024-0002289, filed with the Korean Intellectual Property Office on January 5, 2024, the entire content of which is incorporated herein by reference. Technical Field

[0002] The present disclosure relates to a method and an apparatus for training a pre-trained AI model using similarities between domains. Background Art

[0003] Diffusion probabilistic models are used as generative artificial intelligence (generative AI) models in the field of image generation. Compared with existing generative adversarial networks (GANs), diffusion probabilistic models have the advantage of being able to finely adjust image generation through various conditions such as text and region specification. However, since diffusion probabilistic models gradually generate images from noise during training, they require a large amount of computation and the training may take a long time. Summary of the Invention

[0004] The present invention is provided to introduce a selection of concepts that will be further described in the detailed description below in a simplified form. The present invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter.

[0005] In one general aspect, a method for training a pre-trained artificial intelligence (AI) model includes: receiving a first image and a second image from the pre-trained AI model, determining a first gradient of a first loss function of the first image generated by the pre-trained AI model and a second gradient of a second loss function of the second image generated by the pre-trained AI model; determining a similarity between the first gradient and the second gradient; and updating the pre-trained AI model based on the similarity, wherein the first image and the second image respectively correspond to different domains.

[0006] The first image may correspond to a pre-training domain of the pre-trained AI model, and the second image may correspond to a target domain different from the pre-training domain.

[0007] The step of determining a first gradient of a first loss function of the first image generated by the pre-trained AI model and a second gradient of a second loss function of the second image generated by the pre-trained AI model may include: determining the first gradient from a partial derivative of the first loss function with respect to parameters of a layer in the pre-trained AI model; and determining the second gradient from a partial derivative of the second loss function with respect to parameters of a layer in the pre-trained AI model.

[0008] The step of determining the similarity between the first gradient and the second gradient may include: determining the similarity for each layer between the first gradient and the second gradient.

[0009] The step of updating the pre-trained AI model based on the similarity may include: determining a reweighted gradient based on the similarity; and using the reweighted gradient to update the parameters of the layers in the pre-trained AI model.

[0010] The reweighted gradient may be determined as the product of the similarity for each layer and the second gradient.

[0011] In another general aspect, an apparatus for training a pre-trained artificial intelligence (AI) model includes: one or more processors and a memory, wherein the memory stores instructions configured to cause the one or more processors to perform processing including: receiving, from each of a plurality of pre-trained AI models, respective image pairs each including a first image corresponding to a pre-trained domain and a second image corresponding to a target domain; determining a similarity between the pre-trained domain and the target domain based on the respective image pairs; and selecting at least one pre-trained AI model from the plurality of pre-trained AI models by comparing the similarities respectively corresponding to each of the plurality of pre-trained AI models, to update the selected at least one pre-trained AI model.

[0012] The processing of determining a similarity between the pre-trained domain and the target domain based on the respective image pairs may include: determining a first gradient of a first loss function of the first image and a second gradient of a second loss function of the second image; and determining a similarity between the first gradient and the second gradient.

[0013] The processing of selecting at least one pre-trained AI model from the plurality of pre-trained AI models by comparing the similarities respectively corresponding to each of the plurality of pre-trained AI models may include: selecting the pre-trained AI model corresponding to the maximum similarity.

[0014] The processing may further include: updating the selected pre-trained AI model based on the maximum similarity.

[0015] The processing of determining a first gradient of a loss function of the first image and a second gradient of a loss function of the second image may include: determining the first gradient from the partial derivative of the first loss function with respect to the parameters of the layers in the pre-trained AI model; and determining the second gradient from the partial derivative of the second loss function with respect to the parameters of the layers in the pre-trained AI model.

[0016] The processing of determining a similarity between the pre-trained domain and the target domain based on the respective image pairs may further include: calculating a cosine similarity for each layer between the first gradient and the second gradient.

[0017] The process of updating a selected pre-trained AI model based on maximum similarity may include: determining a reweighted gradient based on the cosine similarity for each layer between a first gradient and a second gradient; and using the reweighted gradient to update the parameters of the layers in the selected pre-trained AI model.

[0018] The process may further include: determining a reweighted gradient based on the product of the cosine similarity for each layer and the second gradient.

[0019] In another general aspect, a system for training an artificial intelligence (AI) model that has been pre-trained in a pre-training domain to learn a target domain different from the pre-training domain, the system includes: one or more processors; and a memory, where the memory stores instructions that are configured to cause the one or more processors to perform a process, the process including: receiving a first image in the pre-training domain and a second image in the target domain from the AI model; determining a correlation value indicating the correlation between the pre-training domain and the target domain based on the first image corresponding to the pre-training domain and the second image corresponding to the target domain; and performing an adjustment based on the correlation value to adapt the AI model to the target domain.

[0020] The process of determining the correlation between the pre-training domain and the target domain based on the first image and the second image may include: determining the similarity between a first gradient of a first loss function determined by the first image and a second gradient of a second loss function determined by the second image.

[0021] The process of determining the similarity may include: determining the first gradient from the partial derivative of the first loss function with respect to the parameters of the layers in the AI model; and determining the second gradient from the partial derivative of the second loss function with respect to the parameters of the layers in the AI model.

[0022] The process of determining the similarity may further include: determining the similarity for each layer between the first gradient and the second gradient.

[0023] The process of performing an adjustment based on the correlation value to adapt the AI model to the target domain may include: determining a reweighted gradient based on the similarity for each layer; and using the reweighted gradient to update the parameters of the layers in the AI model.

[0024] The reweighted gradient may be determined based on the product of the similarity for each layer and the second gradient. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 Illustrates a system for training a pre-trained AI model according to one or more embodiments.

[0026] Figure 2Illustrates reverse denoising processing according to one or more embodiments.

[0027] Figure 3 Illustrates a method for training a pre-trained AI model according to one or more embodiments.

[0028] Figure 4 Illustrates a system for training a pre-trained AI model according to one or more embodiments.

[0029] Figure 5 Illustrates a method for training a pre-trained AI model according to one or more embodiments.

[0030] Figure 6 Illustrates the cosine similarity for each layer between gradients of loss functions according to one or more embodiments.

[0031] Figure 7 Illustrates a defect inspection system in semiconductor manufacturing processes according to one or more embodiments.

[0032] Figure 8 Illustrates a neural network structure of an encoder and a decoder according to one or more embodiments.

[0033] Figure 9 Illustrates a training device according to one or more embodiments.

[0034] Throughout the drawings and the detailed description, unless otherwise described or provided, the same or similar reference numerals will be understood to represent the same or similar elements, features, and structures. The drawings may not be drawn to scale, and for clarity, illustration, and convenience, the relative dimensions, proportions, and depictions of elements in the drawings may be exaggerated. Detailed Description

[0035] The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, apparatuses, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will be apparent after understanding the disclosure of this application. For example, the order of operations described herein is merely an example and is not limited to those set forth herein. Rather, the order of operations may be changed as will be apparent after understanding the disclosure of this application, except for operations that must occur in a specific order. Additionally, descriptions of features known after understanding the disclosure of this application may be omitted for greater clarity and conciseness.

[0036] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Instead, the examples described herein are provided only to illustrate some of the many possible ways of implementing the methods, apparatuses, and / or systems described herein that will be apparent after understanding the disclosure of this application.

[0037] The terms used herein are for the purpose of describing various examples only and are not intended to limit the present disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more of them. As a non-limiting example, the terms "comprise," "include," and "have" specify the presence of the stated features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.

[0038] Throughout the specification, when a component or element is described as "connected to," "coupled to," or "joined to" another component or element, it may be directly "connected to," "coupled to," or "joined to" the other component or element, or one or more other components or elements may reasonably be present therebetween. When a component or element is described as "directly connected to," "directly coupled to," or "directly joined to" another component or element, no other elements may be present therebetween. Similarly, expressions such as "between" and "immediately between" and "adjacent to" and "immediately adjacent to" may also be interpreted as described above.

[0039] Although terms such as "first," "second," and "third" or A, B, (a), (b), etc. may be used herein to describe various components, assemblies, regions, layers, or portions, these components, assemblies, regions, layers, or portions are not limited by these terms. Each of these terms is not used to define, for example, the nature, order, or sequence of the corresponding component, assembly, region, layer, or portion, but is only used to distinguish the corresponding component, assembly, region, layer, or portion from other components, assemblies, regions, layers, or portions. Thus, the first component, the first assembly, the first region, the first layer, or the first portion mentioned in the examples described herein may also be referred to as the second component, the second assembly, the second region, the second layer, or the second portion without departing from the teachings of the examples.

[0040] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains based on an understanding of the present application. Unless explicitly defined as such herein, terms (such as those defined in a general dictionary) will be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the disclosure of the present application, and will not be interpreted in an idealized or overly formal sense. The use of the term "may" with respect to an example or embodiment herein (e.g., with respect to what an example or embodiment may include or may achieve) means that there is at least one example or embodiment that includes or achieves such a feature, although not all examples are limited thereto.

[0041] Therefore, in order to shorten the training time of the diffusion model, a method of adapting a pre-trained model that utilizes a large amount of data to a new target domain through fine-tuning can be used. Low-rank adaptation (LoRA) and bias-terms fine-tuning (BitFit) are methods for fine-tuning existing large-scale language models based on transformers. Compared with existing models, the above methods can adapt an artificial intelligence (AI) model to the target domain by learning with less than 1% of the parameters.

[0042] The artificial intelligence model (AI model) of the present disclosure is a machine learning model that learns at least one task (multiple tasks) and can be implemented as a computer program executed by a processor. The task learned by the AI model can represent a problem to be solved by machine learning or a job to be performed by machine learning. The AI model can be implemented as a computer program running on a computing device, a computer program downloaded through a network, or a computer program sold in the form of a product. Optionally, the AI model can be connected to various devices through a network. In addition, the AI model can interoperate with various devices through a network.

[0043] Figure 1 A system for training a pre-trained AI model according to one or more embodiments is shown, Figure 2 A reverse denoising process according to one or more embodiments is shown, and Figure 3 A method for training a pre-trained AI model according to one or more embodiments is shown.

[0044] In some embodiments, the training device 100 may perform transfer learning based on the relationship or similarity between the target domain and the pre-training domain of the pre-trained AI model 200. The target domain is a domain different from the pre-training domain, and the training device 100 may perform transfer learning on the pre-trained AI model 200 to adapt the pre-trained AI model 200 to the target domain. For example, the training device 100 may fine-tune the pre-trained AI model 200 based on the similarity between "the gradient of the loss function determined by the image corresponding to the pre-training domain" and "the gradient of the loss function determined by the image corresponding to the target domain".

[0045] In some embodiments, the pre-trained AI model 200 may be a generative diffusion model that generates images of the training domain from noise images. The AI model 200 can be pre-trained through forward diffusion processing and reverse denoising processing.

[0046] In the forward diffusion process, the AI model 200 can generate a noise image by adding random noise to the input image in the pre-trained domain to generate a noise image . In the t-th step (1 ≤ t ≤ T) of the forward diffusion process, noise sampled from a fixed normal distribution can be added to the image.

[0047] In the reverse denoising process, the AI model 200 can generate a result image with a probability distribution similar to that of the input image by removing noise with a normal distribution from the noise image to generate a result image with a probability distribution similar to that of the input image . In the reverse denoising process, noise sampled from the learned normal distribution can be subtracted from the image at each step. The AI model 200 can be pre-trained by updating the parameters (e.g., the mean m and standard deviation ) representing the probability distribution from which the noise to be subtracted from the image is sampled.

[0048] In some embodiments, as shown in Equation 1 below, the training device 100 can determine the loss function based on the image at step t with noise randomly sampled from a normal distribution (mean = 0, standard deviation = 1) to determine the loss function .

[0049] Equation 1

[0050] In Equation 1, is the noise added to the image at step t , and is the image generated by the AI model 200 when the image is input to the AI model 200 at step t of the noise. That is, through the reverse denoising process, the training device 100 can learn the normal distribution of the noise added to the image in the forward diffusion process from the image to the image based on the difference between the image output from the AI model 200 when the image at any step t is input to the AI model 200 and the image to the image .

[0051] is the image sampled at any step t and is shown by Equation 2 below.

[0052] Equation 2

[0053] In Equation 2, is a value determined by a parameter indicating the magnitude of noise to be added to an image and and .

[0054] Referring to Figure 1 , the training device 100 may include: a gradient calculator 110, a similarity calculator 120, and a parameter updater 130.

[0055] In some embodiments, the gradient calculator 110 may calculate the gradient of the loss function of the first image corresponding to the pre-training domain and the gradient of the loss function of the second image corresponding to the target domain, respectively. The first image and the second image are generated by the pre-trained AI model 200. The pre-trained AI model 200 may generate the first image corresponding to the pre-training domain and the second image corresponding to the target domain through reverse denoising, and send the first image corresponding to the pre-training domain and the second image corresponding to the target domain to the gradient calculator 110 of the training device 100.

[0056] In some embodiments, the first image corresponding to the pre-training domain may be an image having the same semantics as the images belonging to the pre-training domain. In addition, the second image corresponding to the target domain may be an image having the same semantics as the images belonging to the target domain. For example, the pre-training domain may be a natural domain including natural images, and the target domain may be a semiconductor domain including semiconductor images.

[0057] Referring to Figure 2 , the first image corresponding to the pre-training domain and the second image corresponding to the target domain are shown, and the first image and the second image are generated by the pre-trained AI model 200 at any step of the reverse denoising process performed by the training device 100.

[0058] In Figure 2 , when the image with the noise added to the image at step t belonging to the pre-training domain is input to the pre-trained AI model 200, the pre-trained AI model 200 may generate an image corresponding to the pre-training domain .

[0059] In addition, when the image with the noise added to the image at step t belonging to the target domain is input to the pre-trained AI model 200, the pre-trained AI model 200 may generate an image corresponding to the target domain .

[0060] In Figure 2 , it is emphasized that, compared with and In contrast, in and respectively, one step of noise is removed.

[0061] In some embodiments, the similarity calculator 120 may calculate the similarity between "the gradient of the loss function calculated from the first image corresponding to the pre-training domain" and "the gradient of the loss function calculated from the second image corresponding to the target domain".

[0062] In some embodiments, the parameter updater 130 may calculate a reweighted gradient based on the similarity between the gradient of the loss function corresponding to the pre-training domain and the gradient of the loss function corresponding to the target domain, and use the reweighted gradient to update the parameters of the pre-trained AI model 200.

[0063] Referring to Figure 3 , the gradient calculator 110 of the training device 100 may calculate the loss functions of the image corresponding to the pre-training domain generated by the pre-trained AI model 200 and the image corresponding to the target domain respectively, and calculate the gradients of the respective loss functions (S110).

[0064] The following Equation 3 respectively represents the loss function calculated from the image corresponding to the pre-training domain and the loss function calculated from the image corresponding to the target domain . .

[0065] Equation 3

[0066] Referring to Equation 3, the loss function may include a term , which represents that noise is added to the image at the t-th step of the image corresponding to the pre-training domain . In addition, the loss function may include a term , which represents that noise is added to the image at the t-th step of the image corresponding to the target domain . The following Equation 4 and Equation 5 represent the gradient of the loss function corresponding to the pre-training domain and the gradient of the loss function corresponding to the target domain.

[0067] Equation 4 Gradient of the loss function corresponding to the pre-training domain: Gradient of the loss function corresponding to the target domain: ​In some embodiments, the gradient calculator 110 may calculate gradients corresponding to the images of the pre-training domain and the images of the target domain from the partial derivatives of the loss function with respect to (or regarding) the parameters of the layers in the AI model 200. Equation 5 below represents the gradient of the loss function determined by the gradient calculator 110 with respect to the parameters of the layer of the parameters.

[0068] Equation 5 The gradient of the loss function corresponding to the pre-training domain with respect to the parameters of each layer : The gradient of the loss function corresponding to the target domain with respect to the parameters of each layer : Referring to Figure 3 , the similarity calculator 120 may calculate the similarity (S120) between "the gradient of the loss function calculated from the first image corresponding to the pre-training domain" and "the gradient of the loss function calculated from the second image corresponding to the target domain".

[0069] In some embodiments, the similarity between "the gradient of the loss function calculated from the first image corresponding to the pre-training domain" and "the gradient of the loss function calculated from the second image corresponding to the target domain" may be used by the training device 100 as a numerical indicator representing the relationship between the pre-training domain and the target domain. As a non-limiting example, the similarity between "the gradient of the loss function calculated from the first image corresponding to the pre-training domain" and "the gradient of the loss function calculated from the second image corresponding to the target domain" may be determined by cosine similarity. Other similarities (e.g., Euclidean distance) may be used.

[0070] In some embodiments, the similarity calculator 120 may determine the similarity for each layer based on the gradients of the loss function related to the parameters of the layers included in the pre-trained AI model 200. The update rate of each layer in the pre-trained AI model 200 may be determined according to the similarity for each layer between the gradient of the loss function corresponding to the pre-training domain and the gradient of the loss function corresponding to the target domain determined by the similarity calculator 120.

[0071] Equation 6 below represents the similarity between the gradients of the loss function with respect to the parameters of the layer of the parameters. :

[0072] Equation 6

[0073] In response to the similarity between the gradients of the loss function corresponding to the pre-training domain with respect to each layer and the gradients of the loss function corresponding to the target domain with respect to each layer being relatively large, the similarity between the gradients of the loss function with respect to the parameters of layer can be determined to be a relatively small value by function . Optionally, in response to the similarity between the gradients of the loss function corresponding to the pre-training domain with respect to each layer and the gradients of the loss function corresponding to the target domain with respect to each layer being relatively small, the similarity between the gradients of the loss function with respect to the parameters of layer can be determined to be a relatively large value by function . Referring to

[0074] , the parameter updater 130 can calculate a reweighted gradient based on the similarity between the gradients of the loss function corresponding to the pre-training domain and the gradients of the loss function corresponding to the target domain (S130), and use the reweighted gradient to update the parameters of the pre-trained AI model 200 (S140). Figure 3

[0075] In some embodiments, the parameter updater 130 can calculate the reweighted gradients of the respective layers included in the pre-trained AI model 200 based on the similarity for each layer between the gradients of the loss functions of the respective domains. As shown in Equation 7 below, the reweighted gradient of the th layer included in the pre-trained AI model 200 can be determined as: the similarity for each layer multiplied by "the gradient of the loss function with respect to the parameters of this layer calculated from the images corresponding to the target domain".

[0076] Equation 7

[0077] Referring to Equation 7, the parameter updater 130 can use the reweighted gradients to update the parameters of the respective layers in the pre-trained AI model 200.

[0078] In some embodiments, when the similarity for each layer between the gradients of the loss functions of the respective domains is relatively high, the corresponding layer will be updated relatively less. Additionally, when the similarity for each layer between the gradients of the loss functions of the respective domains is relatively low, the corresponding layer will be updated relatively more.

[0079] For example, when the high-frequency information between the pre-training domain and the target domain is relatively similar and the semantic information is relatively different, according to the similarity for each layer between the gradients of the loss function, the layers corresponding to the high-frequency information can be updated relatively less, and the layers corresponding to the semantic information can be updated relatively more.​​​​

[0080] As described above, the training device 100 according to one or more embodiments performs updates on each layer in the pre-trained AI model based on the similarity between the pre-training domain used in the pre-training of the AI model and the target domain, thereby optimizing the update rate of each layer in the pre-trained AI model and shortening the time required for transfer learning regarding the pre-trained AI model.

[0081] Figure 4 A system for training a pre-trained AI model according to another embodiment is shown, and Figure 5 A method for training a pre-trained AI model according to another embodiment is shown.

[0082] In another embodiment, the training device 100 may receive image pairs of a first image corresponding to the pre-training domain and a second image corresponding to the target domain from each of the pre-trained AI models 2001 to 200 n and may determine the similarity between the pre-training domain and the target domain based on the image pairs. Thereafter, the training device 100 may compare the multiple similarities respectively corresponding to each of the pre-trained AI models 2001 to 200 n to determine at least one suitable or optimal AI model for transfer learning adapted to the target domain.

[0083] Referring to Figure 4 , the training device 100 may include a gradient calculator 110, a similarity calculator 120, and a parameter updater 130, and may further include a model determiner 140.

[0084] In another embodiment, the gradient calculator 110 may calculate the gradients of the loss functions corresponding to the pre-training domain and the target domain based on the image pairs generated by each of the pre-trained AI models 2001 to 200 n

[0085] In another embodiment, the similarity calculator 120 may calculate the similarity between the gradients of the loss functions of the image pairs of each of the pre-trained AI models 2001 to 200 n . Based on the image pairs received from each of the pre-trained AI models 2001 to 200 n , the similarity calculator 120 may calculate the similarity between the pre-training domain and the target domain for each of the pre-trained AI models 2001 to 200 n

[0086] In another embodiment, the model determiner 140 may compare the multiple similarities corresponding to each of the pre-trained AI models 2001 to 200 n such that the model determiner 140 among the multiple pre-trained AI models 2001 to 200 n ​​Select at least one best pre-trained AI model to be domain-adapted by transfer learning from among them. Thereafter, the parameter updater 130 may update the parameters of the selected pre-trained AI model.

[0087] Referring to Figure 5 , the gradient calculator 110 of the training device 100 may calculate the gradients of the loss function of the images of the pre-trained domain and the images of the target domain generated by the multiple pre-trained AI models 2001 to 200 n (S210).

[0088] For example, after the first AI model 2001 is pre-trained in a domain including conventional images, the gradient calculator 110 may calculate the gradient of the loss function of the first image corresponding to the pre-trained domain generated by the first AI model 2001 and the gradient of the loss function of the second image corresponding to the target domain (e.g., the domain of scanning electron microscope (SEM) images of semiconductors). In addition, after the nth AI model 200 n is pre-trained in a domain including semiconductor images, the gradient calculator 110 may calculate the gradient of the loss function of the first image corresponding to the pre-trained domain generated by the nth AI model 200 n and the gradient of the loss function of the second image corresponding to the target domain. That is, the gradient calculator 110 may receive image pairs of images corresponding to the pre-trained domain and the target domain respectively from the multiple AI models 2001 to 200 n , where the multiple AI models are pre-trained in different domains respectively. Then, the gradient calculator 110 may calculate the gradient of the loss function based on the received image pairs.

[0089] At least one of the multiple pre-trained AI models 2001 to 200 n may be pre-trained in the domain of semiconductor images. At this time, the semiconductor images belonging to the pre-trained domain and the images belonging to the target domain including semiconductor images may have different types respectively. For example, images of different types of semiconductors may belong to different domains, and even for images of the same type of semiconductor, images obtained by different measurement devices (SEM, etc.) may also belong to different domains.

[0090] Referring to Figure 5 , the similarity calculator 120 of the training device 100 may calculate the similarity between the gradients of the loss function of the image pairs generated by the multiple pre-trained AI models 2001 to 200 n respectively (S220). In some embodiments, the similarity calculator 120 of the training device 100 may calculate the similarity between the gradients of the loss function of the image pairs generated by the multiple pre-trained AI models 2001 to 200 nThe similarity between the gradients of the loss functions of the generated image pairs is determined as a numerical metric representing the relationship between the pre-training domains of multiple pre-trained AI models 2001 to 200 n and the target domain.

[0091] Referring to Figure 5 , the model determiner 140 of the training device 100 can be based on the similarity between the gradients of the loss functions of the image pairs generated by multiple pre-trained AI models 2001 to 200 n select at least one pre-trained AI model from multiple pre-trained AI models 2001 to 200 n (S230). In some embodiments, the model determiner 140 of the training device 100 can select the pre-trained AI model corresponding to the maximum similarity among multiple pre-trained AI models 2001 to 200 n .

[0092] For example, when the similarity between the gradients of the loss functions of the image pairs generated by the first pre-trained AI model is relatively large, it can be determined that the first pre-trained AI model can be relatively easily transferred to the target domain. Optionally, when the similarity between the gradients of the loss functions of the image pairs generated by the second pre-trained AI model is relatively large, it can be determined that the second pre-trained AI model can be adapted to the target domain with relatively few parameter updates.

[0093] Conversely, when the similarity between the gradients of the loss functions of the image pairs generated by the third pre-trained AI model is relatively small, it can be determined that it is relatively difficult to adapt the third pre-trained AI model to the target domain. Optionally, when the similarity between the gradients of the loss functions of the image pairs generated by the fourth pre-trained AI model is relatively small, the fourth pre-trained AI model can be determined to require a relatively large number of parameter updates to be able to adapt to the target domain.

[0094] After that, the parameter updater 130 of the training device 100 can calculate the reweighted gradient based on the similarity between the gradients of the loss functions of the image pairs generated by the selected pre-trained AI model (S240), and use the reweighted gradient to update the parameters of the selected pre-trained AI model (S250).

[0095] As described above, the training device 100 according to another embodiment can select a pre-trained AI model optimized for the target domain based on the similarity between the target domain and the pre-training domain from among multiple pre-trained AI models. Therefore, the fine-tuning speed of the pre-trained AI model can be accelerated.

[0096] Figure 6 Shows the cosine similarity for each layer between the gradients of the loss functions according to one or more embodiments.

[0097] In Figure 6 the curve graph, the x-axis may represent the index of the layers included in the pre-trained AI model, and the y-axis may represent the cosine similarity of each layer. Referring to Figure 6 , it can be seen that as the time step of the transfer learning performed by the training device 100 according to one or more embodiments increases, the cosine similarity between the gradients of the loss functions in all layers of the AI model 200 increases.

[0098] Referring to Figure 6 , since the similarity between the domains of the layers with smaller indices and the layers with larger indices is relatively large, the training device 100 can perform tuning on the layers with smaller indices and the layers with larger indices at a small update rate. Since the similarity between the domains of the layers with indices from 50 to 75 is relatively small, the training device 100 can perform tuning on the layers with indices from 50 to 75 at a large update rate.

[0099] Table 1 shows the results of transfer learning in which an AI model pre-trained using a high-quality face (Flickr-Faces-HQ, FFHQ) dataset is adapted to a high-quality animal face (Animal Faces-HQ, AFHQ) dataset.

[0100] Table 1

[0101] Referring to Table 1, compared with other learning methods in Table 1, the cosine-similarity-based reweighting method according to one or more embodiments shows the best clean Fréchet inception distance (FID) metric. The clean-FID metric can represent the quality of the images generated by the AI model. The lower the clean-FID metric, the better the performance of the AI model that generates the images. In addition, it can be seen that the cosine-similarity-based reweighting method according to one or more embodiments can achieve the performance corresponding to the clean-FID metric by only using 1500 time steps.

[0102] Figure 7 Shows a defect inspection system in a semiconductor manufacturing process according to one or more embodiments.

[0103] Referring to Figure 7 , the defect inspection system 10 for a semiconductor manufacturing process may include: a defect inspection device 300, a measurement device 400, and a training device 100.

[0104] The defect inspection device 300 can inspect for defects from various images obtained during semiconductor manufacturing processes based on inferences from one or more AI models. During semiconductor manufacturing processes, in-fab wafers pass through several apparatuses, chambers, etc. One or more AI models can include a classification AI model and a generative AI model. The classification AI model is configured to classify images obtained during the manufacturing process, and the generative AI model is configured to generate images for learning the classification AI model. The generative AI model can generate images of a specific domain and provide the images to the classification AI model, and the classification AI model can use the images generated by the generative AI model to perform training on image classification. In some embodiments, the generative AI model can be an image-to-image (im2im) conversion model.

[0105] The measurement device 400 can perform required measurements during semiconductor manufacturing processes and obtain images of semiconductors, wafers, etc. The images obtained by the measurement device 400 can be classified by the classification AI model of the defect inspection device 300 into normal images and / or defect images.

[0106] The training device 100 can train the classification AI model and the generative AI model. In addition, the training device 100 can adapt a pre-trained AI model to a domain including semiconductor images through transfer learning of the generative AI model regarding "generating images for training the classification AI model". The pre-trained AI model can be pre-trained based on images from a natural domain or can be pre-trained based on images from a semiconductor domain different from the domain that needs to be fine-tuned.

[0107] In some embodiments, the training device 100 can determine the correlation between the pre-training domain of the pre-trained AI model and the target domain of the transfer learning and can perform adjustment based on the quantified correlation to adapt the pre-trained AI model to the target domain.

[0108] For example, the training device 100 can allow the pre-trained AI model to generate images of the pre-training domain and images of the target domain and can calculate the gradients of the loss function based on the images of the pre-training domain and the images of the target domain, respectively. After that, the training device 100 can perform transfer learning on the trained AI model by using the similarity between "the gradient of the loss function determined from the image corresponding to the pre-training domain" and "the gradient of the loss function determined from the image corresponding to the target domain" as the quantified correlation between the pre-training domain and the target domain.

[0109] In some embodiments, the training device 100 may select at least one pre-trained AI model from multiple pre-trained AI models based on the correlation between the pre-training domain of the pre-trained AI model and the target domain of transfer learning, and perform transfer learning on the selected pre-trained AI model. The training device 100 may quantify the correlation between the pre-training domain of the pre-trained AI model and the target domain, and select at least one pre-trained AI model among the multiple pre-trained AI models based on the quantified correlation between the domains.

[0110] As described above, the training device 100 according to one or more embodiments determines the correlation between the pre-training domain of the pre-trained AI model and the target domain for transfer learning, and performs layer-by-layer update of the pre-trained AI model based on this correlation, so that the update rate for each layer of the pre-trained AI model can be optimized, and the time required for transfer learning of the pre-trained AI model can be shortened.

[0111] Figure 8 Shows the neural network structure of the encoder and decoder according to one or more embodiments.

[0112] Referring to Figure 8 , the encoder 810 and decoder 820 according to one or more embodiments may have a neural network (NN) structure including an input layer 8101 and 8201, hidden layers 8102 and 8202, and output layers 8103 and 8203, respectively. In some embodiments, the encoder 810 may have the encoder structure of the above-mentioned generative AI model.

[0113] In addition, the decoder 820 may have the decoder structure of the above-mentioned generative AI model. The input layers 8101 and 8201, hidden layers 8102 and 8202, and output layers 8103 and 8203 of the encoder 810 and decoder 820 may each include a corresponding set of nodes, and the connection strength between each node may correspond to a weight (connection weight). The nodes included in the input layers 8101 and 8201, hidden layers 8102 and 8202, and output layers 8103 and 8203 may be connected to each other in a fully connected architecture.

[0114] The number of parameters (weights and biases) may be equal to the number of connections in the neural network 800. The input layers 8101 and 8201 may include input nodes, and the number of input nodes may correspond to the number of independent input variables.

[0115] To train the encoder 810, an image pair may be input to the input layer 8101. When the image pair is input into the input layer 8101 of the encoder 810, a denoised image pair may be output as an inference result from the output layer 8203 of the trained decoder 820.

[0116] The hidden layers 8102 and 8202 may be located between the input layers 8101 and 8201 and the output layers 8103 and 8203, and the hidden layers 8102 and 8202 may include at least one hidden layer. The output layers 8103 and 8203 may include at least one output node. An activation function may be used in the hidden layers 8102 and 8202 and the output layers 8103 and 8203 to determine the node output / activation.

[0117] In some embodiments, the encoder 810 and the decoder 820 may be trained by updating the weights and / or parameters of the hidden nodes included in the hidden layers 8102 and 8202.

[0118] Figure 9 A training device according to one or more embodiments is shown.

[0119] The training device according to one or more embodiments may be implemented as a computer system (e.g., a computer-readable medium). Referring to Figure 9 , the computer system 900 includes one or more processors 910 and a memory 920. The one or more processors 910 represent any single processor or any combination of processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), an accelerator, etc.). The memory 920 may be connected to the one or more processors 910 and may store instructions or programs that are configured to cause the one or more processors 910 to perform processing including any of the methods described above.

[0120] The one or more processors 910 may implement the functions, steps, or methods proposed in the embodiments. The operation of the computer system 900 according to one or more embodiments may be implemented by the one or more processors 910. The one or more processors 910 may include a GPU, a CPU, and / or an NPU. When the operation of the computer system 900 is implemented by the one or more processors 910, each task may be divided among the one or more processors 910 according to the load. For example, when one processor is a CPU, the other processors may be a GPU, an NPU, a field-programmable gate array (FPGA), and / or a digital signal processor (DSP).

[0121] The memory 920 may be disposed inside / outside the processor and may be connected to the processor in various ways known to those skilled in the art. The memory represents various forms of volatile storage media or non-volatile storage media (other than the signal itself), and for example, the memory may include a read-only memory (ROM) and a random access memory (RAM). In another way, the memory may be a PIM (processing in memory) including a logic unit for performing self-contained operations.

[0122] In another way, some functions of a device or system for training a pre-trained AI model can be provided by a neuromorphic chip including neurons, synapses, and an inter-neuron connection module. A neuromorphic chip is a computer device that mimics the structure of a biological nervous system and can perform neural network operations.

[0123] Meanwhile, embodiments are implemented not only by the devices and / or methods described so far, but also by a program or a recording medium recording the program that implements functions corresponding to the configuration of the embodiments, and such implementation can be easily achieved by those skilled in the art to which the present disclosure pertains from the description provided above. Specifically, the method according to the present disclosure (for example, a method for training a pre-trained AI model, etc.) can be implemented in the form of program instructions executable by various computer devices. A computer-readable medium may include program instructions, data files, data structures, etc. alone or in combination. The program instructions recorded on the computer-readable medium may be specifically designed and configured for the embodiments. The computer-readable recording medium may include a hardware device configured to store and execute the program instructions. For example, the computer-readable recording medium includes: magnetic media (such as hard disks, floppy disks, and magnetic tapes), optical recording media (such as CD-ROMs and DVDs), and optical disks (such as floppy disks). It can be magneto-optical media, ROM, RAM, flash memory, etc. The program instructions may include not only machine language codes generated by a compiler, but also high-level language codes executable by a computer through an interpreter, etc.

[0124] Herein, regarding Figures 1 to 9The described computing devices, electronic devices, processors, memories, displays, information output systems and hardware, storage devices, and other devices, apparatuses, units, modules, and components are implemented by or represent hardware components. Examples of hardware components that can be used to perform the operations described in this application include, where appropriate: controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtracters, multipliers, dividers, integrators, and any other electronic components configured to perform the operations described in this application. In other examples, one or more of the hardware components that perform the operations described in this application are implemented by computing hardware (e.g., by one or more processors or computers). A processor or computer can be implemented by one or more processing elements, such as logic gate arrays, controllers, and arithmetic logic units, digital signal processors, microcomputers, programmable logic controllers, field programmable gate arrays, programmable logic arrays, microprocessors, or any other device or combination of devices configured to respond and execute instructions in a defined manner to achieve a desired result. In one example, the processor or computer includes or is connected to one or more memories that store instructions or software executed by the processor or computer. The hardware components implemented by the processor or computer can execute instructions or software (such as an operating system (OS) and one or more software applications running on the OS) for performing the operations described in this application. The hardware components can also access, manipulate, process, create, and store data in response to the execution of the instructions or software. For the sake of brevity, the singular terms "processor" or "computer" can be used in the description of the examples described in this application, but in other examples, multiple processors or computers can be used, or the processor or computer can include multiple processing elements or multiple types of processing elements or both. For example, a single hardware component or two or more hardware components can be implemented by a single processor, or two or more processors, or a processor and a controller. One or more hardware components can be implemented by one or more processors, or a processor and a controller, and one or more other hardware components can be implemented by one or more other processors, or additional processors and additional controllers. One or more processors, or a processor and a controller, can implement a single hardware component or two or more hardware components. The hardware components can have any one or more of different processing configurations, where examples of different processing configurations include: single processor, independent processors, parallel processors, single instruction single data (SISD) multiprocessing, single instruction multiple data (SIMD) multiprocessing, multiple instruction single data (MISD) multiprocessing, and multiple instruction multiple data (MIMD) multiprocessing.

[0125] Figures 1 to 9The method of performing the operations described in this application, as shown, is performed by computing hardware (e.g., by one or more processors or computers), which is implemented to execute instructions or software for performing the operations described in this application that are performed by the method. For example, a single operation or two or more operations can be performed by a single processor, or two or more processors, or a processor and a controller. One or more operations can be performed by one or more processors or a processor and a controller, and one or more other operations can be performed by one or more other processors, or additional processors and additional controllers. One or more processors, or a processor and a controller, can perform a single operation or two or more operations.

[0126] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement the hardware components and perform the method as described above can be written as a computer program, code segment, instruction, or any combination thereof to individually or jointly direct or configure one or more processors or computers to operate as a machine or special-purpose computer to perform the operations performed by the hardware components and method as described above. In one example, the instructions or software include machine code (such as machine code generated by a compiler) that is directly executed by one or more processors or computers. In another example, the instructions or software include higher-level code that is executed by one or more processors or computers using an interpreter. The instructions or software can be written in any programming language based on the block diagrams and flowcharts shown in the figures and the corresponding descriptions herein, which disclose algorithms for performing the operations performed by the hardware components and method as described above.

[0127] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement the hardware components and perform the methods described above, and any associated data, data files, and data structures may be recorded, stored, or fixed in one or more non-transitory computer-readable storage media, or recorded, stored, or fixed on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage devices, hard disk drives (HDD), solid state drives (SSD), flash memory, card memory (such as, micro multimedia cards or cards (e.g., Secure Digital (SD) or Extreme Digital (XD))), magnetic tape, floppy disks, magneto-optical data storage devices, optical data storage devices, hard disks, solid state disks, and any other device configured to store instructions or software and any associated data, data files, and data structures in a non-transitory manner and provide the instructions or software and any associated data, data files, and data structures to one or more processors or computers such that the one or more processors or computers can execute the instructions. In one example, the instructions or software and any associated data, data files, and data structures are distributed across a networked computer system such that the instructions and software and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0128] Although this disclosure includes specific examples, it will be apparent after understanding the disclosure of this application that various changes in form and detail may be made in these examples without departing from the spirit and scope of the claims and their equivalents. The examples described herein are considered to be illustrative only and not for purposes of limitation. The description of a feature or aspect in each example is considered to be applicable to similar features or aspects in other examples. Appropriate results may be achieved if the described techniques are performed in a different order, and / or if the components in the described systems, architectures, devices, or circuits are combined in a different manner, and / or replaced or supplemented by other components or their equivalents.

[0129] Accordingly, in addition to the foregoing disclosure, the scope of the present disclosure may also be defined by the claims and their equivalents, and all variations within the scope of the claims and their equivalents should be construed as being included in the present disclosure.

Claims

1. A method for training a pre-trained artificial intelligence model, the method comprising: Receiving a first image and a second image from the pre-trained artificial intelligence model; Determining a first gradient of a first loss function of the first image generated by the pre-trained artificial intelligence model and a second gradient of a second loss function of the second image generated by the pre-trained artificial intelligence model; Determining a similarity between the first gradient and the second gradient; And Updating the pre-trained artificial intelligence model based on the similarity, Wherein the first image and the second image respectively correspond to different domains.

2. The method according to claim 1, wherein, The first image corresponds to the pre-training domain of the pre-trained artificial intelligence model, and the second image corresponds to a target domain different from the pre-training domain.

3. The method according to claim 1, wherein, The step of determining a first gradient of a first loss function of the first image generated by the pre-trained artificial intelligence model and a second gradient of a second loss function of the second image generated by the pre-trained artificial intelligence model includes: Determining the first gradient from the partial derivative of the first loss function with respect to the parameters of the layers in the pre-trained artificial intelligence model; and Determining the second gradient from the partial derivative of the second loss function with respect to the parameters of the layers in the pre-trained artificial intelligence model.

4. The method according to claim 3, wherein, The step of determining a similarity between the first gradient and the second gradient includes: Determining the similarity for each layer between the first gradient and the second gradient.

5. The method according to claim 4, wherein, The step of updating the pre-trained artificial intelligence model based on the similarity includes: Determining a reweighted gradient based on the similarity; and Using the reweighted gradient to update the parameters of the layers in the pre-trained artificial intelligence model.

6. The method according to claim 5, wherein, The reweighted gradient is determined as the product of the similarity for each layer and the second gradient.

7. A device for training a pre-trained artificial intelligence model, the device comprising: One or more processors and a memory, wherein the memory stores instructions configured to cause the one or more processors to perform processing, the processing including: Receiving, from each of a plurality of pre-trained artificial intelligence models, an image pair each including a first image corresponding to a pre-training domain and a second image corresponding to a target domain; Determining a similarity between the pre-training domain and the target domain based on each image pair; and Selecting at least one pre-trained artificial intelligence model from the plurality of pre-trained artificial intelligence models by comparing the similarities respectively corresponding to each of the plurality of pre-trained artificial intelligence models, to update the selected at least one pre-trained artificial intelligence model.

8. The device according to claim 7, wherein, The processing of determining a similarity between the pre-training domain and the target domain based on each image pair includes: Determining a first gradient of a first loss function of the first image and a second gradient of a second loss function of the second image; and Determining a similarity between the first gradient and the second gradient.

9. The device according to claim 8, wherein, The process of selecting at least one pre-trained artificial intelligence model from the multiple pre-trained artificial intelligence models by comparing the similarities respectively corresponding to each of the multiple pre-trained artificial intelligence models includes: Selecting the pre-trained artificial intelligence model corresponding to the maximum similarity.

10. The device according to claim 9, wherein, The process further includes: Updating the selected pre-trained artificial intelligence model based on the maximum similarity.

11. The device according to claim 8, wherein, The process of determining the first gradient of the loss function of the first image and the second gradient of the loss function of the second image includes: Determining the first gradient from the partial derivative of the first loss function with respect to the parameters of the layer in the pre-trained artificial intelligence model; and Determining the second gradient from the partial derivative of the second loss function with respect to the parameters of the layer in the pre-trained artificial intelligence model.

12. The device according to claim 8, wherein, The process of determining the similarity between the pre-trained domain and the target domain based on each image pair further includes: Calculating the cosine similarity for each layer between the first gradient and the second gradient.

13. The device according to claim 10, wherein, The process of updating the selected pre-trained artificial intelligence model based on the maximum similarity includes: Determining the reweighted gradient based on the cosine similarity for each layer between the first gradient and the second gradient; and Using the reweighted gradient to update the parameters of the layer in the selected pre-trained artificial intelligence model.

14. The device according to claim 13, wherein, The process further includes: Determining the reweighted gradient based on the product of the cosine similarity for each layer and the second gradient.

15. A system for training an artificial intelligence model that has been pre-trained in a pre-trained domain to learn a target domain different from the pre-trained domain, the system includes: One or more processors; And a memory, Wherein, the memory stores instructions configured to cause the one or more processors to execute a process, the process includes: Receiving a first image in the pre-trained domain and a second image in the target domain from the artificial intelligence model; Determining a correlation value indicating the correlation between the pre-trained domain and the target domain based on the first image corresponding to the pre-trained domain and the second image corresponding to the target domain; and Performing an adjustment based on the correlation value to adapt the artificial intelligence model to the target domain.

16. The system according to claim 15, wherein, The process of determining the correlation between the pre-trained domain and the target domain based on the first image and the second image includes: Determining the similarity between the first gradient of the first loss function determined by the first image and the second gradient of the second loss function determined by the second image.

17. The system according to claim 16, wherein, The process of determining the similarity includes: Determining the first gradient from the partial derivative of the first loss function with respect to the parameters of the layer in the artificial intelligence model; and Determining the second gradient from the partial derivative of the second loss function with respect to the parameters of the layer in the artificial intelligence model.

18. The system according to claim 17, wherein, The process of determining the similarity further includes: Determine the similarity for each layer between the first gradient and the second gradient.

19. The system according to claim 18, wherein The process of performing an adjustment based on the correlation value to adapt the artificial intelligence model to the target domain includes: Determine a reweighted gradient based on the similarity for each layer; and Use the reweighted gradient to update the parameters of the layers in the artificial intelligence model.

20. The system according to claim 19, wherein Determine the reweighted gradient based on the product of the similarity for each layer and the second gradient.

Citation Information

Patent Citations

  • Heteroaryl derivative compounds, and uses thereof

    KR1020240002289A