Apparatus and method for measuring performance of ai model and apparatus for training ai model
By adding perturbation-based noise to the AI model and updating the perturbation, inferred uncertainty is generated, the problem of high uncertainty of the AI model is solved, the robustness and reliability of the model are improved, verification costs are saved, and classification performance of specific categories is enhanced.
Patent Information
- Application Number
- CN202510007343.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-30
- Filing Date
- 2025-01-03
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to effectively measure and reduce the uncertainty of artificial intelligence models, affecting their reliability and performance.
By adding noise based on perturbation determination in the linear classifier of the AI model, inferred uncertainty is generated and the perturbation is updated during training to improve model robustness, the noise is independently sampled using a standard normal distribution and adjusting the perturbation size according to the category importance.
Improves the robustness and reliability of AI models, reduces sensitivity to noise, saves test set verification time and cost, and improves classification performance for specific categories.
Smart Images

Figure CN120256887A_ABST
Abstract
Description
[0001] This application claims the priority and benefits of Korean Patent Application No. 10-2024-0001164, filed with the Korean Intellectual Property Office on January 3, 2024, and Korean Patent Application No. 10-2024-0200292, filed with the Korean Intellectual Property Office on December 30, 2024. The entire contents of the Korean patent applications are incorporated herein by reference. Technical Field
[0002] The present disclosure relates to a method for measuring the performance of an artificial intelligence (AI) model using perturbation and a device using the method. Background Art
[0003] Techniques for classifying input data using an AI model are used in various fields such as image recognition, autonomous driving, and telemedicine. In addition, efforts are being made to learn and study how to reduce the uncertainty of the AI model and improve the reliability of the AI model. Summary of the Invention
[0004] The present Summary of the Invention is provided to introduce, in a brief form, a selection of concepts that are further described in the Detailed Description below. The present Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to assist in determining the scope of the claimed subject matter.
[0005] In one general aspect, a device for measuring the performance of an artificial intelligence (AI) model includes: one or more processors and a memory; and the memory stores instructions that are configured to cause the one or more processors to perform processing, the processing including: determining noise based on perturbation; adding the noise to representative vectors corresponding to respective classes of an input image determined by a linear classifier of the AI model; and generating an inference uncertainty of the AI model with respect to the input image based on a plurality of inference results output by the linear classifier of the AI model using the input data and the representative vectors to which the noise is added, wherein the input data is multimedia data.
[0006] The step of determining noise based on perturbation may include: sampling noise from a noise distribution that follows a standard normal distribution and scaling the sampled noise by a magnitude of the perturbation.
[0007] The step of sampling noise from a noise distribution that follows a standard normal distribution may include: sampling noise independently for each class in the noise distribution.
[0008] The step of adding the noise to each representative vector may include: adding the noise independently to each of the representative vectors.
[0009] The processing may further include: determining perturbations proportionally to the corresponding importance of each category.
[0010] The multiple inference results may correspond to the category classification results of the feature vectors of the images generated by the encoder of the AI model, and the category classification results may be determined by a linear classifier classifying the feature vectors into categories of each category using the representative vectors added with noise.
[0011] The processing may further include: receiving a predetermined number of category classification results from the linear classifier as inference results.
[0012] The processing may further include: determining the inference uncertainty of the AI model based on the standard deviation of the multiple inference results.
[0013] Based on determining that the standard deviation of the multiple inference results satisfies a predetermined deviation threshold condition, the multiple inference results of the AI model may not be used as the final results, or based on determining that the standard deviation of the multiple inference results does not satisfy a predetermined deviation threshold condition, the multiple inference results of the AI model may be used as the final results.
[0014] The standard deviation may be statistically determined from the multiple score vectors output as the multiple inference results.
[0015] The maximum element in each score vector may be in the same position within each score vector.
[0016] In another general aspect, an apparatus for training an artificial intelligence (AI) model, the apparatus includes: one or more processors and a memory; and the memory stores instructions configured to cause the one or more processors to perform processing, the processing including: adding noise determined based on each perturbation to a representative vector corresponding to each category; updating each perturbation in response to training the AI model using input data and a weight matrix including the representative vectors added with noise, where the input data is multimedia data; and determining whether to terminate the training of the AI model based on the updated each perturbation.
[0017] The step of adding noise determined based on each perturbation to a representative vector corresponding to each category may include: sampling noise from a noise distribution that follows a normal distribution; scaling the sampled noise by the magnitude of each perturbation; and adding the scaled noise to the representative vector corresponding to each category.
[0018] The step of sampling noise from a noise distribution that follows a normal distribution may include: independently sampling noise for each category in the noise distribution.
[0019] The step of adding the noise determined based on each perturbation to the representative vectors corresponding to each category may include: sampling first noise for a first representative vector to be added to the representative vector from a noise distribution that follows a normal distribution, and sampling second noise for a second representative vector to be added to the representative vector from the noise distribution; and adding the first noise to the first representative vector and adding the second noise to the second representative vector.
[0020] The magnitude of each updated perturbation may be greater than the magnitude of each perturbation before updating each perturbation.
[0021] The step of determining whether to terminate the training of the AI model based on the updated perturbations may include: comparing the magnitude of the updated perturbations with a predetermined perturbation reference value.
[0022] The processing may further include: generating category uncertainties for each category based on the updated perturbations; and determining whether to terminate the training of the AI model based on the category uncertainties for each category.
[0023] In another general aspect, a method for measuring the performance of an artificial intelligence (AI) model for classifying input data, the method being executed by one or more processors, the method including: determining perturbations based on respective category importances; adding noise determined based on the perturbations to representative vectors corresponding to respective categories, wherein the AI model has been trained to infer a category for input data of the AI model; receiving, using the representative vectors with added noise, multiple inference results for the input output from the AI model; and generating an inference uncertainty of the AI model based on the multiple inference results, wherein the input data is multimedia data.
[0024] The input may be an image of a semiconductor product generated during a semiconductor manufacturing process.
[0025] Other features and aspects will be apparent from the following detailed description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 Illustrates a training apparatus for an AI model according to one or more embodiments.
[0027] Figure 2 Illustrates a training method for an AI model according to one or more embodiments.
[0028] Figure 3 Illustrates a training system for an AI model according to one or more embodiments.
[0029] Figure 4Shows an inference device using an AI model according to one or more embodiments.
[0030] Figure 5 Shows an inference method using an AI model according to one or more embodiments.
[0031] Figure 6 Shows a neural network according to one or more embodiments.
[0032] Figure 7 Shows a system for measuring the performance of an AI model according to one or more embodiments.
[0033] Throughout the drawings and the detailed description, unless otherwise described or provided, the same or similar reference numerals will be understood to refer to the same or similar elements, features, and structures. The drawings may not be to scale, and for clarity, illustration, and convenience, the relative dimensions, proportions, and depictions of the elements in the drawings may be exaggerated. Detailed Description
[0034] The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, devices, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will be apparent after understanding the disclosure of this application. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of this application, except for operations that must occur in a particular order. Additionally, descriptions of features known after understanding the disclosure of this application may be omitted for greater clarity and conciseness.
[0035] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Instead, the examples described herein are provided only to illustrate some of the many possible ways of implementing the methods, devices, and / or systems described herein that will be apparent after understanding the disclosure of this application.
[0036] The terms used herein are for describing various examples only and will not be used to limit the disclosure. Unless the context clearly indicates otherwise, the singular is intended to also include the plural form. As used herein, the term "and / or" includes any one of the related listed items and any combination of any two or more. As a non-limiting example, the terms "comprises," "comprising," and "having" indicate the presence of the stated features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.
[0037] Throughout the specification, when a component or element is described as being "connected to", "coupled to", or "joined to" another component or element, it can be directly "connected to", "coupled to", or "joined to" the other component or element, or one or more other components or elements can reasonably be present therebetween. When a component or element is described as being "directly connected to", "directly coupled to", or "directly joined to" another component or element, no other elements can be present therebetween. Similarly, expressions such as "between" and "immediately between" and "adjacent to" and "immediately adjacent to" can also be interpreted as described above.
[0038] Although terms such as "first", "second", and "third" or A, B, (a), (b), etc. may be used herein to describe various members, components, regions, layers, or parts, these members, components, regions, layers, or parts should not be limited by these terms. Each of these terms is not used to define, for example, the nature, order, or sequence of the corresponding member, component, region, layer, or part, but is only used to distinguish the corresponding member, component, region, layer, or part from other members, components, regions, layers, or parts. Thus, without departing from the teachings of the examples, the first member, first component, first region, first layer, or first part referred to in the examples described herein can also be referred to as the second member, second component, second region, second layer, or second part.
[0039] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains and based on an understanding of the disclosure of this application. Unless explicitly defined as such herein, terms (such as those defined in a general dictionary) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the disclosure of this application, and should not be interpreted in an idealized or overly formal sense. The use of the term "can" herein with respect to an example or embodiment (e.g., with respect to what an example or embodiment can include or implement) means that there is at least one example or embodiment that includes or implements such a feature, while all examples are not limited thereto.
[0040] An artificial intelligence (AI) model according to the present disclosure is a machine learning model that learns at least one task and can be implemented as a computer program executed by a processor. The task learned by the AI model can refer to a task to be solved by machine learning or a task to be performed by machine learning. The AI model can be implemented as a computer program that runs on a computing device, is downloaded via a network, or is sold in the form of a product. Optionally, the AI model can be linked to various devices via a network.
[0041] To quantify the uncertainty of an AI model, during the inference of the AI model, some parameters of the AI model can be randomly discarded, and the uncertainty of the AI model can be determined based on the change in the prediction distribution of the AI model. In addition, mix-up of input data can be used to measure the uncertainty of the AI model.
[0042] Figure 1 Discloses a training device for an AI model according to one or more embodiments, and Figure 2 Discloses a method for training an AI model according to one or more embodiments.
[0043] Referring to Figure 1 , according to some embodiments, the training device 100 may include an AI model 110, a perturbation adder 120 (or "perturbator"), and a training determiner 130.
[0044] In some embodiments, the AI model 110 may include an encoder 111 and a classifier 112 (e.g., a linear classifier). The encoder 111 may extract features of the data input to the AI model 110 (or referred to as input data) and generate a feature vector. The classifier 112 may determine the class of the input data through an operation between a weight matrix configured for the classification of the input data and the feature vector of the input data. In some embodiments, noise based on perturbations of each class may be added to the weight matrix of the classifier 112.
[0045] In some embodiments, the AI model 110 may be trained to classify input data into classes. The input data may be multimedia data (e.g., at least one of an image (static image or video), text, speech, etc.), and the class may be any type of category to which the input data belongs and may be the purpose of classification by the AI model 110. For example, the input data may be an image of a road and its background acquired during the autonomous driving of a vehicle. In addition, the class / category may have the type of objects included in the road image (person, vehicle, road structure, traffic signal device, lane, crosswalk, etc.).
[0046] In some embodiments, the input data may be a product image of a product captured during the process of manufacturing the product (e.g., an image of a semiconductor wafer). In an embodiment for an image of a product, the class may represent various types of defects that occur during the manufacturing process of the product. For example, the class may be contamination, foreign matter, fracture, the state of solder marks (solder balls), and / or character abnormality, etc.
[0047] In some embodiments, the classifier 112 of the AI model 110 may use a weight matrix to convert the feature vector of the data output by the encoder 111 into a score vector, the dimension of which is the number of classes (each element of the score vector is the score or probability that the input data belongs to the corresponding class). Equation 1 below represents the linear classification process performed in the classifier 112.
[0048] Equation 1
[0049] In Equation 1, represents the weight matrix of the classifier 112 ( represents the weight matrix is the transpose matrix of), represents the feature vector of the data output from the encoder 111, and represents the bias vector that affects / biases the scores without interacting with the feature vector. When the feature vector input to the classifier 112 has dimension n and when the number of classes is m, the weight matrix may have dimension .
[0050] In some embodiments, the classifier 112 may convert the feature vector of the input data into a score vector representing the scores for each class through operations such as Equation 1.
[0051] In some embodiments, the weight matrix of the classifier 112 may be a set of class representative vectors corresponding to the classes respectively. Equation 2 represents the weight matrix of the classifier 112 .
[0052] Equation 2
[0053] Referring to Equation 2, the weight matrix includes representative vectors corresponding to the classes respectively (i = 1 to m). That is, the weight matrix may be composed of m representative vectors of m corresponding classes. Equation 3 represents the representative vector of the i-th class .
[0054] Equation 3
[0055] In Equation 3, the representative vector of the i-th class may include n weight values (w 1,i to w n,i ). The feature vector can be multiplied by the representative vector of the i-th class a matrix multiplication operation therebetween to determine the score of the eigenvector for the i-th category.
[0056] In some embodiments, the elements of the score vector representing the result of the linear classification of the input terms may include the scores of the vectors of the input terms for each of the respective categories. Additionally, a given score in the score vector may represent the probability that the input term of the eigenvector belongs to the category corresponding to the given score. The dimension of the score vector is equal to the number of categories (the number of categories is m). The i-th element of the score vector corresponding to the i-th category can be determined by a matrix multiplication operation between the eigenvector and the representative vector of the i-th category.
[0057] In some embodiments, the perturbation adder 120 may add noise determined based on the perturbation to the weight matrix of the classifier 112 and update the perturbation through the training of the AI model 110. In some embodiments, during the training of the AI model 110, the perturbation adder 120 may update the perturbation (e.g., the value / size of the perturbation) in the direction of increasing the perturbation for each category.
[0058] In some embodiments, when the training cycle (epoch) of the AI model 110 terminates, the training determiner 130 may determine whether to terminate the training of the AI model 110 based on the updated perturbations for the respective categories. In some embodiments, the training determiner 130 may terminate the training of the AI model 110 in response to the sizes of the updated perturbations for the respective categories each being greater than the perturbation reference value.
[0059] Optionally, the training determiner 130 may calculate the category uncertainty for each category based on the updated perturbations and terminate the training of the AI model 110 in response to the category uncertainties for the respective categories each being less than the uncertainty reference value. When the size of the updated perturbation for one of the categories is less than the perturbation reference value or the category uncertainty for one of the categories is greater than the uncertainty reference value, the training determiner 130 may determine that additional training of the AI model 110 is required for the corresponding category. A large category uncertainty for a particular category may indicate an insufficient amount of data belonging to that category or the existence of data with incorrect labels.
[0060] Referring to Figure 2 , the perturbation adder 120 may determine the perturbations for the respective categories (S110). In some embodiments, the perturbation adder 120 may use the standard deviation (or variance 2 ) of the noise distribution from which the noise is sampled as the perturbation for each category. The perturbation may be determined as a scaler value. For example, the noise distribution may follow a normal distribution (i.e., a Gaussian distribution) or a standard normal distribution.
[0061] In some embodiments, the perturbation adder 120 may independently determine perturbations for each category. For example, for i and j (where i ≠ j), the perturbation for the i-th category may be , and the perturbation for the j-th category may be . The perturbations independently determined for each category may be updated to various magnitudes / values according to the training of the AI model 110.
[0062] In some embodiments, the perturbation adder 120 may determine the noise for each category based on the perturbations of the categories, and then the determined noise may be added to the representative vector of the category (S120) (which may be done for each category). For example, the perturbation adder 120 may randomly sample the noise for the i-th category from a noise distribution representing the standard deviation (mean = 0, standard deviation = normal distribution), and then the sampled noise may be added to the representative vector of the i-th category (which may be done for each category). Optionally, the perturbation adder 120 may randomly sample the noise for the i-th category from the standard normal distribution (mean = 0, standard deviation = 1), multiply the sampled noise by the perturbation , and then add the noise multiplied by the perturbation to the representative vector of the i-th category. The noise sampled from the normal distribution with standard deviation and the noise sampled from the standard normal distribution and then multiplied by the perturbation may be the same. That is, the noise may be scaled by the perturbation.
[0063] Equation 4 represents the representative vector (of the i-th category) to which the noise according to the corresponding perturbation is added .
[0064] Equation 4
[0065] In Equation 4, N represents the noise vector randomly sampled from the noise distribution, and includes N1 to N as its elements n . The magnitude of the perturbation may indicate / control the noise level to be added to the representative vector of the category. The larger the magnitude of the perturbation , the greater the noise added to the representative vector of the category, and the smaller the magnitude of the perturbation , the smaller the noise added to the representative vector of the category.
[0066] In some embodiments, the noise to be added to the representative vector of the i-th category and the noise to be added to the representative vector of the j-th category may be sampled independently from a noise distribution. For example, the perturbation adder 120 may independently perform the sampling of the noise to be added to the representative vector of the i-th category and independently perform the sampling of the noise to be added to the representative vector of the j-th category. Optionally, the perturbation adder 120 may sample the noise from the noise distribution, multiply the sampled noise by the perturbation for each category, and add the multiplication result to the representative vector of each category.
[0067] Referring Figure 2 , the AI model 110 may be trained using the training set by the perturbation adder 120 based on the operation between (i) the feature vector (of the input data) output from the encoder 111 and (ii) the weight matrix added with noise (S130).
[0068] In some embodiments, a score vector may be determined through the operation between the feature vector and the weight matrix added with noise. The i-th element of the score vector may represent a score corresponding to (or being) the probability that the input data (corresponding to the feature vector) will be determined to be in the i-th category. The score vector may be converted into a one-hot vector through one-hot encoding (the element "1" represents the most likely category).
[0069] In some embodiments, the input data input to the AI model 110 may be determined to be in the category corresponding to the maximum value among the elements of the score vector (or the category corresponding to the "1" of the one-hot vector). The loss function of the AI model 110 (e.g., ) may output the difference between the category of the input data output by the AI model 110 and the predetermined label (ground truth) of the input data.
[0070] In some embodiments, the AI model 110 may update the weights and parameters within the AI model 110 to minimize the difference of the loss function. In addition, the perturbation for each category determined by the perturbation adder 120 may also be updated during the training of the AI model 110 (e.g., as part of the training pass).
[0071] When the weights and parameters of the AI model 110 are updated in the training of the AI model 110 to minimize the difference of the loss function, the perturbation adder 120 may gradually increase the perturbation for each category by updating. In some embodiments, a normalized loss function may be added to the loss function of the AI model 110 to increase the magnitude of the perturbation. Equation 5 represents an example of the normalized loss function, and Equation 6 represents the loss function of the AI model 110 added with the normalized loss function.
[0072] Equation 5
[0073] Equation 6
[0074] P in Equation 5 and Equation 6 th is a perturbation threshold for preventing the perturbation from increasing too much, and for example, the number of classes m × 0.01 (0.01m) can be used as the perturbation threshold. For example, in Equation 6, can be a predefined coefficient.
[0075] In some embodiments, since the perturbation for each class increases through the training of the AI model 110, due to the magnitude of the increased perturbation, the AI model 110 can be trained to be robust to noise (and by extension, general noise).
[0076] The perturbation adder 120 can add noise to the weight matrix of the classifier 112 based on the updated perturbation, and the classifier 112 can classify the next input data in the input data set (training set) using the weight matrix with added noise based on the updated perturbation.
[0077] In some embodiments, when the training of the AI model 110 is completed for the input data set (e.g., when a predetermined number of epochs have passed), the training determiner 130 can determine when (or whether) to terminate the training of the AI model 110 based on the magnitude of the updated perturbation. For example, the training determiner 130 can determine whether to terminate the training of the AI model 110 by comparing the magnitude of the updated perturbation with a predetermined perturbation reference value. Alternatively, the training determiner 130 can calculate the class uncertainty for each class based on the updated perturbation and compare the calculated class uncertainty with an uncertainty reference value to determine whether to terminate the training of the AI model 110.
[0078] In some embodiments, the training determiner 130 can determine the class uncertainty (S140) for each class based on the updated perturbation. For example, Equation 7 represents the class uncertainty calculated according to the magnitude of the perturbation.
[0079] Equation 7
[0080] In Equation 7, the class uncertainty U of the i-th class i can be determined by a monotonically decreasing function (such as an exponential function with a base value between 0 and 1 (e.g., 1 / e 100 ). In some embodiments, the training determiner 130 can make the difference in the class uncertainty for each class larger by using a monotonically decreasing function such as Equation 7.
[0081] In some embodiments, the training determiner 130 may compare the class uncertainty of each class with an uncertainty reference value (S150), and when the class uncertainty of at least one class is greater than the uncertainty reference value, the training determiner 130 may determine to additionally collect data of the corresponding class (S160). After that, the AI model 110 may be retrained using the training set with the added data (e.g., by using data augmentation techniques).
[0082] In some embodiments, when there is a class with "updated perturbation less than the perturbation reference value" after a predetermined number of rounds are completed or when there is a class with "class uncertainty determined based on the updated perturbation greater than the uncertainty reference value", it may be determined that the training of the AI model 110 for the corresponding class is insufficient.
[0083] When the magnitude of the updated perturbation of a class is greater than the perturbation reference value, this may indicate that the AI model 110 has been trained to be robust to perturbations even when a large level of noise is added to the classifier 112. In other words, the generalization performance of the AI model 110 for the corresponding class may be determined to be high. Optionally, when the class uncertainty of a class is less than the uncertainty reference value, the classification performance of the AI model 110 for the corresponding class may be determined to be high.
[0084] For additionally collecting data of a specific class, the domain in which the data of the corresponding class can be collected may be determined. For example, when the data is image data captured during the manufacturing process of a semiconductor and the class that needs to be additionally collected is the fracture of solder balls on the front side of a semiconductor wafer, the domain in which the image data on the front side of the wafer can be collected is determined, and then the image data included in the corresponding class (solder ball fracture) in the corresponding domain may be collected.
[0085] When data of a class with a perturbation magnitude less than the perturbation reference value or a class with a class uncertainty greater than the uncertainty reference value is additionally collected, the AI model 110 may perform additional training using the additionally collected data set. The AI model 110 may perform training and update the perturbation until the perturbation magnitude for the class for which the data is additionally collected becomes greater than the perturbation reference value or until the class uncertainty of the class becomes less than the uncertainty reference value.
[0086] When the perturbation magnitudes of all classes are greater than the perturbation reference value or the class uncertainties of all classes are less than the uncertainty reference value, the training determiner 130 may terminate the training of the AI model 110 (S170).
[0087] In some embodiments, the AI model 110 for a category's category uncertainty can be used as an indicator of the performance of the trained AI model 110 for that category. By measuring the category uncertainty of each category of the AI model 110 trained by the training determiner 130, the performance of the AI model 110 can be verified using only the training set (i.e., no separate test set is required).
[0088] According to the apparatus and method for AI model training, the AI model 110 trained using the training set does not need to be verified by a separate test set, and thus the time and cost required to generate the test set can be saved. In other words, a test set generated by strict requirements (e.g., not overlapping with the training set, data being strictly and accurately labeled, a sufficiently large amount of data being included, etc.) is not required, and thus the cost and time saved can be quite significant.
[0089] Furthermore, in one exemplary embodiment, categories that require additional data collection can be identified, and thus the time and cost required to collect data for all categories indiscriminately can be saved, and the problem of insufficient data for specific categories can be prevented.
[0090] Figure 3 Shows a training system for an AI model according to one or more embodiments.
[0091] Referring to Figure 3 , the training system for an AI model according to one or more embodiments may include a training device 100 for training the AI model, an inspection device 10, and a data collection device 20.
[0092] In some embodiments, the inspection device 10 may store product images acquired for product inspection (such as non-destructive inspection and destructive inspection of products). The inspection device 10 may acquire product images during the process of manufacturing a product, and may use the product images to perform non-destructive inspection. The inspection device 10 may sample the product during the manufacturing process and perform destructive inspection on the sampled product.
[0093] In some embodiments, the data collection device 20 may collect product images for training and testing the AI model from the inspection device 10. Optionally, the data collection device 20 may acquire images for autonomous driving of a vehicle (e.g., images of roads and backgrounds). In this case, the images may include various types of objects (people, vehicles, road structures, buildings, traffic signal devices, lanes, crosswalks, etc.).
[0094] In some embodiments, the training device 100 may determine the class uncertainty of each class of the AI model based on the perturbations updated through the training of the AI model. For example, when the magnitude of the updated perturbation of at least one class is large, the training device 100 determines that the class uncertainty of the corresponding class is high, and may retrain the AI model for that class. Optionally, when the class uncertainty determined based on the magnitude of the updated perturbation of at least one class is high, the training device 100 may retrain the AI model for that class.
[0095] In some embodiments, the training device 100 may request the data collection device 20 to additionally collect data of the class determined to have a relatively high class uncertainty for additional training of the AI model. The data collection device 20 may determine the domain in which data of the class with high class uncertainty can be collected in order to additionally collect data of that class.
[0096] For example, when the data is image data collected during the manufacturing process of a semiconductor and the class for which additional collection is requested is the class having cracks on the back surface of the wafer, the data collection device 20 may determine that the image of the back surface of the semiconductor wafer is the domain for additional image data collection (corresponding to the crack class) and obtain additional images of that domain from one or more specific corresponding inspection devices 10 that acquire images of that domain.
[0097] As another example, when the data is image data of a semiconductor and the class for which additional collection is requested is the class that does not meet the threshold dimension (CD) standard of the semiconductor profile, the data collection device 20 may collect additional image data belonging to the corresponding class (CD non-compliance) from a destructive inspection device (e.g., a scanning electron microscope) that provides semiconductor images of the corresponding domain.
[0098] In this way, the training device 100 according to one or more embodiments can determine the class that requires additional training of the AI model, can obtain the class that requires additional training of the AI model, and thereby can save data collection time and cost.
[0099] Figure 4 An inference device using an AI model according to one or more embodiments is shown, and Figure 5 An inference method using an AI model according to one or more embodiments is shown.
[0100] Referring to Figure 4 , the inference device 200 according to one or more embodiments may include an AI model 210, a perturbation adder 220, and a performance measurement device 230.
[0101] In some embodiments, the AI model 210 is a trained model and may include an encoder 211 and a classifier 212. The encoder 211 may extract features from the input data input to the AI model 210 and generate a feature vector of the input data. The classifier 212 may classify (e.g., infer) the class of the input data through an operation between the weight matrix for classification of the input data and the feature vector of the input data. In some embodiments, noise based on respective perturbations for respective classes may be added to the weight matrix of the classifier 212.
[0102] In some embodiments, the AI model 210 may perform inference to classify the input data into one of the classes. The input data may be an image, text, speech, etc., and the classes may be categories to which the input data belongs; such classification may be the purpose of the AI model 210. For example, if the input data is an image acquired during autonomous driving of a vehicle, the classes into which the image may be classified may include a person, a vehicle, a road structure, a building, a traffic signal device, a lane, a crosswalk, etc.
[0103] In some embodiments, the input data may be a product image acquired during a process of manufacturing a product (e.g., a semiconductor), and each class may represent a type of defect that occurs during the manufacturing process of the product. For example, the classes may be contamination, foreign matter, breakage, the state of solder marks (solder balls), character anomalies, etc.
[0104] In some embodiments, the classifier 212 of the AI model 210 may use a weight matrix to convert the feature vector output by the encoder 211 into a score vector having the same dimension as the number of classes. The linear classification process performed in the classifier 212 may be implemented as represented by Equation 1.
[0105] In some embodiments, the classifier 212 may convert the feature vector of the input data into a score vector representing the scores for each class through an operation such as Equation 1.
[0106] In some embodiments, the weight matrix of the classifier 212 may be a set of representative vectors for respective classes. The weight matrix of the classifier 212 and the representative vector for the i-th class may be represented as Equation 2 and Equation 3.
[0107] In some embodiments, the elements of the score vector representing the result of linear classification may represent the scores of the feature vector for each class. Each score may represent the probability that the input data of the feature vector belongs to the class corresponding to the score. The dimension of the score vector may be equal to the number of classes (the number of classes is m). Among the elements of the score vector, the element corresponding to the i-th class of the score vector may be determined through a matrix multiplication operation between the feature vector and the representative vector for the i-th class.
[0108] In some embodiments, the perturbation adder 220 may determine the perturbations for each category based on the importance of each category, determine the noise for the category according to the respective perturbations for each category, and add the determined noise to each representative vector of the category included in the weight matrix of the classifier 212. Table 1 shows the category importance for determining defects in semiconductor products (such as wafers).
[0109] (Table 1)
[0110] In Table 1, the numbers of the category importance may indicate the priorities of the categories. In the example of Table 1, the category of solder ball loss is the highest priority category, and the category of observation error is the lowest priority category. In other words, since categories such as solder ball loss, crack, and short circuit are the categories that the AI model 210 needs to accurately infer, the inference accuracy of this category may have a significant impact on the overall performance of the AI model 210. On the other hand, if the inference accuracy of the AI model 210 for categories such as observation error is relatively low, the overall performance of the AI model 210 may not be significantly affected. In some embodiments, the categories classified by the AI model 210 may have different category importances, and the inference accuracy and / or uncertainty may be measured differently for each category according to the corresponding category importance. The perturbation adder 220 may determine a relatively large perturbation for a category with a relatively high category importance (e.g., a category importance greater than the importance threshold), and determine a relatively low perturbation for a category with a relatively low category importance (e.g., a category importance less than or equal to the importance threshold). In one example, the perturbation adder 220 may determine the perturbations in proportion to the category importance of each category. For example, when there are m categories, the perturbation adder 220 may use the largest size of the perturbation =m / 100 to add noise to the representative vector of the category with the highest category importance. Optionally, the perturbation adder 220 may use the smallest size of the perturbation =1 / 100 to add noise to the representative vector of the category with the lowest category importance.
[0111] In some embodiments, the performance measurement device 230 may measure the performance of the AI model 210 based on the inference results of the AI model 210 for the same input data, where the AI model 210 includes the classifier 212 with noise added according to the category importance. The performance measurement device 230 may determine the inference uncertainty of the AI model 210 based on the inference results of the AI model 210, and measure the performance of the AI model 210 according to the inference uncertainty of the AI model 210. The inference uncertainty of the AI model 210 may be determined by the standard deviation or variance of the inference results for the same input data.
[0112] According to one or more embodiments, the perturbation adder 220 may add noise according to "perturbation based on class importance" to the representative vector of a class, and then the performance measurement device 230 may measure the performance of the inference device 200 according to class importance with additional references (e.g., different criteria). For example, the performance measurement device 230 may measure the inference performance of the AI model 210 more strictly for classes with relatively high class importance, and measure the inference performance of the AI model 210 more leniently for classes with relatively low class importance.
[0113] In some embodiments, for a class with relatively high class importance, when the perturbation adder 220 adds a corresponding relatively high noise (determined based on a relatively large perturbation due to the high class importance) to the representative vector of the class, the possibility of deterioration of the inference performance of the AI model 210 for the class is high. Even when the noise added to the representative vector is relatively large (based on a relatively large perturbation), when the inference uncertainty of the AI model 210 is measured to be low, the performance measurement device 230 may strictly evaluate the reliability of the AI model 210.
[0114] On the other hand, for a class with relatively low class importance, when the perturbation adder 220 adds a corresponding noise (determined based on a relatively small perturbation due to the small class importance) to the representative vector of the class, the possibility of deterioration of the inference performance of the AI model 210 for the class is low. When the class importance is low, the performance of the AI model 210 for the class is not very important, so it may not be necessary to strictly evaluate the inference performance of the AI model 210.
[0115] Referring to Figure 4 and Figure 5 , in some embodiments, the perturbation adder 220 may determine the perturbation of each class based on the respective class importance (S210), and add each noise (determined based on each perturbation) to the respective representative vector of the class (S220). The representative vector w' of the i-th class to which the noise according to the perturbation is added i may be given by Equation 4.
[0116] The perturbation adder 220 may randomly sample the noise from a predetermined noise distribution (e.g., a predetermined noise distribution that follows a standard normal distribution) (e.g., sample the noise in the predetermined noise distribution), and may scale the randomly sampled noise by respective predetermined perturbations according to the respective class importance. The perturbation adder 220 may add the noise scaled by the perturbation to the representative vector of the class corresponding to the perturbation . When the noise scaled by a relatively large perturbation is added to the corresponding representative vector of the corresponding class, the inference uncertainty of the class may increase.
[0117] In some embodiments, the performance measurement device 230 may obtain the inference result for the input data from the trained AI model 210 (S230). The trained AI model 210 may perform inference on the input data using the weight matrix including the representative vectors with added noise, and output the score vector of the input data as the inference result.
[0118] In some embodiments, the score vector output from the trained AI model 210 may include the scores of the input data for each of the respective categories for which the AI model 210 has been trained to recognize. The i-th element of the score vector may represent the score of the input data with respect to the i-th category. In this case, the inference result of the trained AI model 210 may be the category corresponding to the maximum element of the score vector. For example, when the maximum element in the score vector is the i-th element, the inference result of the trained AI model 210 may be determined as: "The input data is the i-th category".
[0119] In some embodiments, the performance measurement device 230 may repeatedly perform noise addition and collect the inference results of the trained AI model 210 for the same input data, and then determine the inference uncertainty of the AI model 210 when the number of repetitions reaches a predetermined number I (S240 and S250).
[0120] As just pointed out, steps S220 and S230 may be repeated until the number of collected inference results reaches the predetermined number I. Before I inference results are collected, the perturbation adder 220 may sample the noise again from the noise distribution and add the newly sampled noise scaled by "predetermined perturbations according to category importance" to the representative vectors of each category. The trained AI model 210 may re-output the inference result for the same input data using the weight matrix with the repeatedly added noise according to the repeated new sampling.
[0121] In some embodiments, the performance measurement device 230 may measure the performance of the trained AI model 210 based on the inference results of the trained AI model 210 for the same input data.
[0122] In some embodiments, the performance measurement device 230 may calculate the standard deviation of the inference results of the trained AI model 210 Determine the inference uncertainty of the trained AI model 210. For example, when the standard deviation of the inference results of the trained AI model 210 satisfies (e.g., is greater than) a predetermined deviation threshold, the performance measurement device 230 may determine that the inference uncertainty of the trained AI model 210 is high, and may not use the inference results of the trained AI model 210. Optionally, when the standard deviation of the inference results of the trained AI model 210 does not satisfy (e.g., is less than or equal to) the predetermined deviation threshold, the performance measurement device 230 may determine the inference uncertainty of the trained AI model 210 as a reliable level and use the inference results.
[0123] For example, the standard deviation of the I inference results output by the trained AI model 210 can be statistically determined from the I score vectors. In some embodiments, the standard deviation of the I inference results output by the trained AI model 210 can be statistically determined from the elements with the maximum scores in each of the I score vectors. In this case, the indices of the elements with the maximum scores in each of the I score vectors inferred from the same input data (element instances of different score vectors) may all be the same (which indicates that the trained AI model 210 has low inference uncertainty and reliable inference results).
[0124] In some embodiments, when the indices of the elements with the maximum scores in each of the I score vectors inferred from the same input data are not exactly the same, the performance measurement device 230 may determine that the inference uncertainty of the trained AI model 210 is the highest. In other words, when all i score vectors are not encoded as the same one-hot vector, it is determined that (i) different inference results are output for the same input data, and (ii) thus the performance measurement device 230 may distrust the inference results of the trained AI model 210 without calculating the standard deviation of the inference results (i.e., the standard deviation can be calculated to determine reliability / uncertainty).
[0125] In some embodiments, the performance measurement device 230 may use a predetermined function to determine the inference uncertainty of the trained AI model 210 from the standard deviation of the inference results of the trained AI model 210 , and determine the inference uncertainty U of the trained AI model 210. Equation 8 represents the inference uncertainty U of the trained AI model 210 determined using a predetermined function from the standard deviation of the inference results . infer .
[0126] Equation 8
[0127] Referring to Equation 8, an exponential function can be used to determine the inference uncertainty U infer. In some embodiments, a monotonically increasing function such as an exponential function can be used as a predetermined function for determining inference uncertainty.
[0128] For example, when the inference uncertainty of the trained AI model 210 is higher than a predetermined uncertainty threshold, the performance measurement device 230 may not use the inference result of the trained AI model 210 as a conclusive / inferential determination regarding the corresponding input data. Optionally, when the inference uncertainty of the trained AI model 210 is lower than a predetermined uncertainty threshold, the performance measurement device 230 may use the inference result of the trained AI model 210 as the final inference regarding the corresponding input data.
[0129] In some embodiments, when the inference uncertainty of the trained AI model 210 is determined to be high (e.g., higher than a predetermined uncertainty threshold), it may be determined that retraining of the trained AI model 210 is required. In this case, the performance measurement device 230 may determine to perform retraining of the AI model 210 only for those classes with high class importance. When it is determined that retraining of the trained AI model 210 is required, data for only one or more classes indicated as having high inference uncertainty (by the inference result) may be additionally collected.
[0130] When it is determined that the inference result of the trained AI model 210 is reliable, the inference result of the trained AI model 210 may be used (S260).
[0131] As described above, the inference device 200 according to one or more embodiments adds noise based on perturbations for each class to the weight matrix of the trained AI model 210 to more strictly measure the performance of the trained AI model 210 for classes with high class importance. That is, although noise based on large perturbations is added to the representative vector corresponding to the class with high class importance, when the trained AI model 210 outputs an inference result with high reliability for that class, the inference device 200 may use the inference result of the trained AI model 210. Therefore, the performance of the trained AI model 210 can be easily measured according to the importance of the class, and the training objectives of the AI model 210 can be accurately located and improved.
[0132] In addition, the performance of the trained AI model 210 can be verified in the inference step of the AI model through the inference uncertainty of the trained AI model 210, thereby improving the classification performance for classes with high class importance. For example, in the case of product defect detection by the trained AI model 210, the defect detection performance of the trained AI model 210 can be improved, and depending on the application, an improvement in the product yield rate can be expected, and autonomous driving can be improved, etc.
[0133] Figure 6 Shows an artificial neural network according to one or more embodiments.
[0134] Referring Figure 6 , a neural network (NN) 600 according to one or more embodiments may include an input layer 610, a hidden layer section 620, and an output layer 630. The input layer 610, the layers in the hidden layer section 620 (e.g., layers 6201 to 620 n ), and the output layer 630 may include respective sets of nodes, and the connection strengths between the nodes of the layers (usually, adjacent layers, but not necessarily) may be represented as weights (weighted connections). The nodes included in the input layer 610, the layers in the hidden layer section 620, and the output layer 630 may be fully connected to each other. In some embodiments, the number of parameters (e.g., the number of weights and the number of biases) may be equal to the number of weighted connections in the neural network 600.
[0135] The input layer 610 may include input nodes x1 to x i , and the number of input nodes x1 to x i may correspond to the number of independent variables of the input data. A training set may be input to the input layer 610 for training the neural network 600. When test data is input to the input layer 610 of the trained neural network 600, an inference result may be output from the output layer 630 of the trained neural network 600. In some embodiments, the input layer 610 may have a structure suitable for processing large-scale inputs. In a non-limiting example, the neural network 600 may include a convolutional neural network combined with one or more fully connected layers.
[0136] The layers of the hidden layer section 620 may be located between the input layer 610 and the output layer 630, and may include at least one of the hidden layers 6201 to 620 n . The output layer 630 may include at least one of the output nodes y1 to y j . Activation functions may be used for the layers of the hidden layer section 620 and the output layer 630. In some embodiments, the neural network 600 may be learned by adjusting the weights of the nodes included in the hidden layer section 620.
[0137] Figure 7 Shows a system for measuring the performance of an AI model according to one or more embodiments.
[0138] A system for measuring the performance of an AI model according to one or more embodiments may be implemented as a computer system.
[0139] Referring Figure 7, the computer system 700 may include one or more processors 710 and a memory 720. The memory 720 may be connected to one or more processors 710 and may store instructions or programs configured to cause the one or more processors 710 to perform processing including any of the methods described above.
[0140] One or more processors 710 may implement the functions, stages, or methods described for various embodiments. The operation of the computer system 700 according to one or more embodiments may be implemented by one or more processors 710. One or more processors 710 may include a graphics processing unit (GPU), a central processing unit (CPU), and / or a neural processing unit (NPU). When the operation of the computer system 700 is executed by one or more processors 710, each task may be divided among the one or more processors 710 according to the load. For example, when one processor is a CPU, the other processors may be GPUs, NPUs, field-programmable gate arrays (FPGAs), and / or digital signal processors (DSPs).
[0141] The memory 720 may be disposed inside / outside the processor and may be connected to the processor by various means known to those skilled in the art. The memory represents various forms of volatile or non-volatile storage media (but not the signal itself), and for example, the memory may include a read-only memory (ROM) and a random access memory (RAM). In another way, the memory may be a PIM (processing in memory) including logic units for performing independent operations (for example, a bit cell may act as a persistent bit storage device and may have both circuit elements for performing operations on the stored bit data).
[0142] In another way, some functions of the yield prediction device (for example, training the yield prediction model and / or the path generation model, inference of the yield prediction model and / or the path generation model) may be provided by a neuromorphic chip including neurons, synapses, and inter-neuron connection modules. A neuromorphic chip is a computer device that mimics the structure of a biological nervous system and may perform neural network operations.
[0143] Meanwhile, the embodiments can be implemented not only by the apparatuses and / or methods described so far, but also by a program (instructions) that implements functions corresponding to the configurations of the embodiments or a recording medium on which the program is recorded, and such an implementation can be easily achieved by any person skilled in the art to which this description pertains from the description provided above. Specifically, the methods according to the present disclosure (e.g., the yield prediction method, etc.) can be implemented in the form of program instructions executable by various computer devices. The computer-readable medium may include program instructions, data files, data structures, etc. alone or in combination. The program instructions recorded on the computer-readable medium can be specifically designed and configured for the embodiments. The computer-readable recording medium may include a hardware device configured to store and execute program instructions. For example, the computer-readable recording medium includes magnetic media (such as hard disks, floppy disks, and magnetic tapes), optical recording media (such as CD-ROMs and DVDs), and magneto-optical media (such as floppy disks). It can be ROM, RAM, flash memory, etc. The program instructions may include not only machine language codes generated by a compiler but also high-level language codes executable by a computer through an interpreter, etc.
[0144] Herein, with respect to Figures 1 to 7The described computing devices, electronic devices, processors, memories, displays, information output systems and hardware, storage devices, and other devices, apparatuses, units, modules, and components are implemented by or represent hardware components. Examples of hardware components that can be used to perform the operations described in this application include, where appropriate, controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtractors, multipliers, dividers, integrators, and any other electronic components configured to perform the operations described in this application. In other examples, one or more of the hardware components that perform the operations described in this application are implemented by computing hardware (e.g., by one or more processors or computers). A processor or computer can be implemented by one or more processing elements (such as logic gate arrays, controllers, and arithmetic logic units, digital signal processors, microcomputers, programmable logic controllers, field programmable gate arrays, programmable logic arrays, microprocessors, or any other device or combination of devices configured to respond and execute instructions in a defined manner to achieve a desired result). In one example, a processor or computer includes or is connected to one or more memories that store instructions or software executed by the processor or computer. The hardware components implemented by the processor or computer can execute instructions or software (such as an operating system (OS) and one or more software applications running on the OS) for performing the operations described in this application. The hardware components can also access, manipulate, process, create, and store data in response to the execution of the instructions or software. For brevity, the singular terms "processor" or "computer" can be used in the description of the examples described in this application, but in other examples, multiple processors or computers can be used, or a processor or computer can include multiple processing elements, or multiple types of processing elements, or both. For example, a single hardware component or two or more hardware components can be implemented by a single processor, or two or more processors, or a processor and a controller. One or more hardware components can be implemented by one or more processors, or a processor and a controller, and one or more other hardware components can be implemented by one or more other processors, or additional processors and additional controllers. One or more processors or a processor and a controller can implement a single hardware component, or two or more hardware components. The hardware components can have any one or more of different processing configurations, examples of different processing configurations including single processors, independent processors, parallel processors, single instruction single data (SISD) multiprocessing, single instruction multiple data (SIMD) multiprocessing, multiple instruction single data (MISD) multiprocessing, and multiple instruction multiple data (MIMD) multiprocessing.
[0145] Figures 1 to 7The method of performing the operations described in this application, shown in [description], is performed by computing hardware (e.g., by one or more processors or computers), which is implemented to execute instructions or software as described above to perform the operations performed by the method described in this application. For example, a single operation or two or more operations can be performed by a single processor, or two or more processors, or a processor and a controller. One or more operations can be performed by one or more processors, or a processor and a controller, and one or more other operations can be performed by one or more other processors, or additional processors and additional controllers. One or more processors or a processor and a controller can perform a single operation, or two or more operations.
[0146] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement the hardware components and perform the method as described above can be written as a computer program, code segment, instruction, or any combination thereof, for individually or jointly instructing or configuring one or more processors or computers to operate as a machine or special-purpose computer to perform the operations performed by the hardware components and method as described above. In one example, the instructions or software include machine code (such as machine code generated by a compiler) that is directly executed by one or more processors or computers. In another example, the instructions or software include higher-level code that is executed by one or more processors or computers using an interpreter. The instructions or software can be written in any programming language based on the block diagrams and flowcharts shown in the figures and the corresponding descriptions herein, which disclose algorithms for performing the operations performed by the hardware components and method as described above.
[0147] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement the hardware components and execute the methods as described above, as well as any associated data, data files, and data structures, can be recorded, stored, or fixed in one or more non-transitory computer-readable storage media, or can be recorded, stored, or fixed on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage devices, hard disk drives (HDD), solid state drives (SSD), flash memory, card memory (such as, multimedia card or micro card (e.g., Secure Digital (SD) or Extreme Digital (XD))), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device that is configured to store instructions or software and any associated data, data files, and data structures in a non-transitory manner and provide the instructions or software and any associated data, data files, and data structures to one or more processors or computers such that the one or more processors or computers can execute the instructions. In one example, the instructions or software and any associated data, data files, and data structures are distributed across a networked computer system such that the instructions and software and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0148] Although this disclosure includes specific examples, it will be apparent after understanding the disclosure of this application that various changes in form and detail can be made in these examples without departing from the spirit and scope of the claims and their equivalents. The examples described herein are considered to be illustrative only and not for purposes of limitation. The description of a feature or aspect in each example should be considered applicable to similar features or aspects in other examples. Appropriate results can be achieved if the described techniques are performed in a different order, and / or if the components in the described systems, architectures, devices, or circuits are combined in a different manner, and / or replaced or supplemented by other components or their equivalents.
[0149] Accordingly, in addition to the foregoing disclosure, the scope of the present disclosure is also defined by the claims and their equivalents, and all variations within the scope of the claims and their equivalents should be construed as being included in the present disclosure.
Claims
1. An apparatus for measuring the performance of an artificial intelligence model, the apparatus comprising: one or more processors and a memory; and the memory stores instructions that are configured to cause the one or more processors to perform processing, the processing including: determining noise based on perturbations; adding the noise to representative vectors corresponding to respective categories of an input image determined by a linear classifier of the artificial intelligence model; and generating an inference uncertainty of the artificial intelligence model for the input image based on a plurality of inference results output by the linear classifier of the artificial intelligence model using input data and the representative vectors added with noise, wherein the input data is multimedia data.
2. The apparatus according to claim 1, wherein the step of determining noise based on perturbations includes: sampling noise from a noise distribution that follows a standard normal distribution, scaling the sampled noise by a magnitude of the perturbation.
3. The apparatus according to claim 2, wherein, The step of sampling noise from a noise distribution that follows a standard normal distribution includes: independently sampling noise for each category in the noise distribution.
4. The apparatus according to claim 1, wherein, The step of adding the noise to respective representative vectors includes: independently adding the noise to each of the respective representative vectors.
5. The device according to claim 4, wherein, The processing further includes: determining perturbations in proportion to the importance corresponding to respective categories.
6. The apparatus according to claim 1, wherein the plurality of inference results correspond to a category classification result of a feature vector of an image generated by an encoder of the artificial intelligence model, and the category classification result is determined by classifying the feature vector into respective categories by the linear classifier using the representative vectors added with noise.
7. The apparatus according to claim 6, wherein the processing further includes: receiving a predetermined number of category classification results from the linear classifier as inference results.
8. The apparatus according to claim 1, wherein the processing further includes: determining the inference uncertainty of the artificial intelligence model based on a standard deviation of the plurality of inference results.
9. The apparatus according to claim 6, wherein based on determining that the standard deviation of the plurality of inference results satisfies a predetermined deviation threshold condition, the plurality of inference results of the artificial intelligence model are not used as final results, or based on determining that the standard deviation of the plurality of inference results does not satisfy a predetermined deviation threshold condition, the plurality of inference results of the artificial intelligence model are used as final results.
10. The apparatus according to claim 6, wherein the standard deviation is statistically determined from a plurality of score vectors output as the plurality of inference results.
11. The apparatus according to claim 10, wherein the maximum element in each score vector is in the same position within each score vector.
12. An apparatus for training an artificial intelligence model, the apparatus comprising: one or more processors and a memory; and the memory stores instructions that are configured to cause the one or more processors to perform processing, the processing including: adding noise determined based on respective perturbations to representative vectors corresponding to respective categories; Updating respective perturbations in response to training an artificial intelligence model using input data and a weight matrix including representative vectors with added noise, where the input data is multimedia data; and Determining whether to terminate the training of the artificial intelligence model based on the updated respective perturbations.
13. The apparatus according to claim 12, wherein The step of adding noise determined based on respective perturbations to representative vectors corresponding to respective categories includes: Sampling noise from a noise distribution subject to a normal distribution; Scaling the sampled noise by the magnitude of respective perturbations; and Adding the scaled noise to representative vectors corresponding to respective categories.
14. The apparatus according to claim 13, wherein The step of sampling noise from a noise distribution subject to a normal distribution includes: Sampling noise independently for each category in the noise distribution.
15. The apparatus according to claim 12, wherein The step of adding noise determined based on respective perturbations to representative vectors corresponding to respective categories includes: Sampling first noise for a first representative vector to be added to the representative vector from a noise distribution subject to a normal distribution, and sampling second noise for a second representative vector to be added to the representative vector from the noise distribution; and Adding the first noise to the first representative vector and adding the second noise to the second representative vector.
16. The apparatus according to claim 12, wherein The magnitude of each of the updated perturbations is greater than the magnitude of each of the perturbations before updating the respective perturbations.
17. The apparatus according to claim 12, wherein The step of determining whether to terminate the training of the artificial intelligence model based on the updated respective perturbations includes: Comparing the magnitude of each of the updated perturbations with a predetermined perturbation reference value.
18. The apparatus according to claim 12, wherein, The processing further includes: Generating category uncertainties for respective categories based on the updated respective perturbations; and Determining whether to terminate the training of the artificial intelligence model based on the category uncertainties for respective categories.
19. A method for measuring the performance of an artificial intelligence model for classifying input data, the method being executed by one or more processors, the method including: Determining perturbations based on respective category importances; Adding noise determined based on the perturbations to representative vectors corresponding to respective categories, where the artificial intelligence model has been trained to infer a category for input data of the artificial intelligence model; Receiving multiple inference results of the input output from the artificial intelligence model using the representative vectors with added noise; and Generating an inference uncertainty of the artificial intelligence model based on the multiple inference results, where the input data is multimedia data.
20. The method according to claim 19, wherein The input is an image of a semiconductor product generated during a semiconductor manufacturing process.
Citation Information
Patent Citations
Method for applying a resin composition without overspray and a resin composition for use in the method
KR1020240001164A