Method and apparatus for measuring performance of artificial intelligence model using perturbation
The method adds class-specific perturbations to AI models' weight matrices for precise performance measurement and training, reducing costs and improving data collection efficiency by eliminating the need for separate test sets.
Patent Information
- Application Number
- JP2025001075
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-30
- Filing Date
- 2025-01-06
- Publication Date
- 2025-07-15
AI Technical Summary
Existing AI models lack efficient methods to quantify uncertainty and measure performance accurately, particularly for classes with varying importance, necessitating separate test sets that increase time and cost.
A method and apparatus that add class-specific perturbations to the weight matrix of AI models during training and inference, using a perturbation adder and performance measurer to calculate inference uncertainty and measure performance based on class importance, eliminating the need for separate test sets.
Enables precise measurement of AI model performance for classes with high importance, reduces data collection and verification costs, and improves training targeting by identifying classes needing additional data.
Smart Images

Figure 2025106230000001_ABST
Abstract
Description
Technical Field
[0001] This description relates to a method and apparatus for measuring the performance of an AI model using perturbations.
Background Art
[0002] Techniques for classifying the class of input data using an AI model are widely used in various fields such as image recognition, autonomous driving, and telemedicine. And research is being conducted to reduce the uncertainty of such AI models and improve their reliability.
[0003] In order to quantify the uncertainty of an AI model, by randomly dropping out some parameters of the model during inference of the AI model, the uncertainty of the AI model is determined based on the change in the prediction distribution of the AI model. Also, mix-up of input data can be used to measure the uncertainty of an AI model.
Summary of the Invention
Problems to be Solved by the Invention
[0004] One embodiment provides an apparatus for measuring the performance of an AI model using perturbations.
[0005] Another embodiment provides an apparatus for training an AI model using perturbations.
[0006] Another embodiment provides a method for measuring the performance of an AI model using perturbations.
Means for Solving the Problems
[0007] According to one embodiment, an apparatus for measuring the performance of an AI model is provided. The apparatus includes a perturbation adder that determines a perturbation for each of a plurality of classes based on class importance and adds noise determined based on the perturbation to representative vectors of the plurality of classes, and a performance measurer that calculates inference uncertainty of the AI model from a plurality of inference results output by the AI model using a weighted value matrix including the representative vectors to which the noise has been added.
[0008] According to another embodiment, an apparatus for training an AI model is provided. The apparatus includes a perturbation adder that adds noise determined based on a perturbation to representative vectors of a plurality of classes and updates the magnitude of the perturbation when the AI model is trained using a weighted value matrix including the representative vectors to which the noise has been added, and a training determiner that determines the end of training of the AI model based on the updated perturbation.
[0009] According to still another embodiment, a method for measuring the performance of an AI model is provided. The method includes determining a perturbation for each of a plurality of classes based on class importance, determining noise based on the perturbation and adding the noise to representative vectors of the plurality of classes, and calculating inference uncertainty of the AI model from a plurality of inference results output by the AI model using a weighted value matrix including the representative vectors to which the noise has been added.
Advantages of the Invention
[0010] By adding noise based on class-specific perturbations to the weight matrix of the trained AI model 210, the performance of the trained AI model 210 for classes with high importance can be strictly measured. As a result, the performance of the trained AI model can be easily measured according to the importance of the classes, and the training target of the AI model can be accurately targeted. In addition, since it is not necessary to verify the AI model trained using the training set with a test set generated separately from the training set, the time and cost required for generating the test set can be saved.
Brief Description of Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Mode for Carrying Out the Invention
[0012] Hereinafter, with reference to the accompanying drawings, embodiments of the present disclosure will be described in detail so that those having ordinary knowledge in the technical field to which the present invention pertains can easily implement them. However, the present disclosure can be implemented in various different forms and is not limited to the embodiments described herein. And, in order to clearly explain the present disclosure in the drawings, parts that are unnecessary for the explanation are omitted, and similar parts throughout the specification are denoted by similar reference numerals.
[0013] Throughout the specification, when a part states that a certain component "includes", this means that other components can be further included, rather than excluding other components unless there is a particularly contrary statement.
[0014] Expressions described as singular in this specification can be interpreted as singular or plural unless an explicit expression such as "one" or "single" is used.
[0015] In this specification, "and / or" includes each and all combinations of one or more of the recited components.
[0016] In this specification, terms including ordinal numbers such as first, second, etc. are used to describe various components, but the components are not limited by the terms. The terms are only used for the purpose of distinguishing one component from another. For example, without departing from the scope of the present disclosure, the first component can be named the second component, and similarly, the second component can also be named the first component.
[0017] In the flowcharts described with reference to the drawings in this specification, the order of operations can be changed, various operations can be combined, or a certain operation can be divided, and a specific operation may not be performed.
[0018] The artificial intelligence model (AI model) of the present disclosure is a machine learning model that learns at least one task and can be implemented as a computer program executed by a processor. The task learned by the artificial intelligence model can refer to the problem to be solved through machine learning or the operation to be performed through machine learning. The artificial intelligence model can be implemented as a computer program executed on a computing device, downloaded through a network, or sold in the form of a product. Or the artificial intelligence model can interact with various devices through a network.
[0019] FIG. 1 is a block diagram showing a training device for an AI model according to an embodiment, and FIG. 2 is a flowchart showing a training method for an AI model according to an embodiment.
[0020] Referring to FIG. 1, a training device 100 according to an embodiment may include an artificial intelligence (AI) model 110, a perturbation adder 120, and a training determiner 130.
[0021] In one embodiment, the AI model 110 may include an encoder 111 and a classifier 112. The encoder 111 can extract the features of the data input to the AI model 110 and generate a feature vector. The classifier 112 can classify the class of the input data through the operation between the weight value matrix for classifying the input data and the feature vector of the input data. In one embodiment, noise based on perturbations for each class may be added to the weight value matrix of the classifier 112.
[0022] In one embodiment, the AI model 110 is trained to be able to classify input data into one of a plurality of classes. The input data may be an image, text, voice, etc., and the class may be a category to which the input data belongs and may be the purpose of classification by the AI model 110. For example, the input data is an image of a road and its surroundings acquired during the autonomous driving of a vehicle, and the class may be the type of object (person, vehicle, road structure, traffic signal device, lane, crosswalk, etc.) included in the image of the road and its surroundings.
[0023] In one embodiment, the input data may be a product image (e.g., semiconductor) during the manufacturing process of a product, and each class may be able to indicate the type of defect that occurs during the manufacturing process of the product. For example, the class may be contamination, foreign matter, breakage, state of solder ball marks, character abnormality, etc.
[0024] In one embodiment, the classifier 112 of the AI model 110 can use a weight matrix to convert the feature vector of the data output by the encoder 111 into a score vector having the same dimension as the number of classes. The following Equation 1 can show the linear classification process performed by the classifier 112.
[0025]
Equation
[0026] In one embodiment, the classifier 112 can convert the feature vector of the input data into a score vector indicating the score for each class through operations such as Equation 1.
[0027] In one embodiment, the weight matrix of the classifier 112 may be a set of class-representative vectors. Equation 2 can represent the weight matrix W of the classifier 112.
[0028]
Equation
[0029]
Equation
[0030] In one embodiment, the elements of the score vector indicating the result of linear classification can indicate the scores of the feature vectors for each class, and the scores can indicate the probability that the input data of the feature vector belongs to the class corresponding to the score. The dimension of the score vector is the same as the number of classes m. The element of the score vector corresponding to the i-th class is determined by the matrix product operation between the feature vector and the representative vector of the i-th class.
[0031] In one embodiment, the perturbation adder 120 can add the noise determined based on the perturbation to the weighted value matrix of the classifier 112 and update the magnitude of the perturbation through the training of the AI model 110. In one embodiment, when the AI model 110 is trained, the perturbation adder 120 can update the magnitude of the perturbation in the direction of increasing the magnitude of the perturbation for each class.
[0032] In one embodiment, the training determiner 130 can determine whether to end the training of the AI model 110 based on the updated perturbation of each class when the training cycle (epoch) of the AI model 110 ends. In one embodiment, the training determiner 130 can end the training of the AI model 110 when the magnitude of the updated perturbation of all classes is greater than the perturbation reference value.
[0033] Alternatively, the training determiner 130 can calculate the class uncertainty of multiple classes based on the updated perturbation and end the training of the AI model 110 when the class uncertainty of all classes is less than the uncertainty reference value. When the magnitude of the updated perturbation of at least one class is less than the perturbation reference value or the class uncertainty of at least one class is greater than the uncertainty reference value, the training determiner 130 can determine that the AI model 110 needs to be additionally trained for the class. A large class uncertainty of a specific class can mean that the amount of data belonging to the class is insufficient or there is data with incorrect labels.
[0034] Referring to FIG. 2, the perturbation adder 120 can determine the perturbation for each class (S110). In one embodiment, the perturbation adder 120 is the standard deviation σ (or variance σ) of the noise distribution from which the noise is sampled. 2) can be used as a perturbation (scalar value) for each class. The noise distribution may follow a normal distribution (Gaussian distribution) or a standard normal distribution.
[0035] In one embodiment, the perturbation adder 120 can independently determine the perturbation for each class. For example, for different i and j from each other (i≠j), the perturbation for the i-th class is σ i and the perturbation for the j-th class may be σ j . The perturbations independently determined for each class can be updated with different magnitudes from each other by training the AI model 110.
[0036] In one embodiment, the perturbation adder 120 can determine noise based on the perturbation and add the noise to the representative vector of each class (S120). For example, the perturbation adder 120 can sample noise randomly from a noise distribution (mean = 0, standard deviation = σ i ), which is a normal distribution with a standard deviation of σ i , and add the sampled noise to the representative vector of the i-th class. Or the perturbation adder 120 can sample noise randomly from a standard normal distribution (mean = 0, standard deviation = 1), multiply the sampled noise by the perturbation σ i , and then add the noise multiplied by the perturbation to the representative vector of the i-th class. The noise sampled from the normal distribution with a standard deviation JPEG2025106230000005.jpg22 and the noise after being multiplied by the perturbation JPEG2025106230000006.jpg22 (the noise scaled only by the perturbation JPEG2025106230000007.jpg22) may be the same.
[0037] The following Equation 4 can show the representative vector W i ’ of the i-th class to which noise is added by the perturbation.
[0038] [Number] In Equation 4, N can represent a noise vector randomly sampled from a noise distribution and can include N1 to N2 as elements. The magnitude of the perturbation can indicate the level of noise added to the representative vector of the class. The larger the magnitude σ of the perturbation, the higher the level of noise added to the representative vector of the class, and the smaller the magnitude σ of the perturbation, the lower the level of noise added to the representative vector of the class.
[0039] In one embodiment, the noise added to the representative vector of the i-th class and the noise added to the representative vector of the j-th class can be sampled independently from the noise distribution. For example, the perturbation adder 120 can independently sample the noise added to the representative vector of the i-th class and the noise added to the representative vector of the j-th class. Alternatively, after sampling noise from the noise distribution, the perturbation adder 120 can multiply the sampled noise by different magnitudes of perturbations corresponding to each class, and then add the sampled noise multiplied by the perturbation to the representative vector of each class.
[0040] Referring to FIG. 2, the AI model 110 is trained on the training set (S130) based on the operation between the feature vector of the data output from the encoder 111 and the weighted value matrix to which noise is added by the perturbation adder 120.
[0041] In one embodiment, a score vector is determined by the operation between the feature vector and the weighted value matrix to which noise is added, and the i-th element of the score vector can indicate the score at which the data corresponding to the feature vector is determined to be in the i-th class. The score vector can be converted into a one-hot vector by one-hot encoding.
[0042] In one embodiment, the class of the data input to the AI model 110 can be determined as the class corresponding to the largest value among the elements of the score vector (or the class corresponding to 1 in the one-hot vector). The loss function of the AI model 110 can output the difference between the class of the input data output by the AI model 110 and the predetermined label of the input data.
[0043] In one embodiment, the AI model 110 can update the weighting values and parameters within the AI model 110 such that the difference in the loss function is minimized, and the perturbations for each class determined by the perturbation adder 120 can also be updated together during the training of the AI model 110.
[0044] When the weighting values and parameters are updated so that the difference in the loss function is minimized in the training of the AI model 110, the perturbation adder 120 can cause the perturbations for each class to gradually increase with the update. In one embodiment, a regularization loss function for increasing the magnitude of the perturbation is added to the loss function of the AI model 110. Equation 5 can illustrate an example of the regularization loss function, and Equation 6 can illustrate the loss function of the AI model 110 to which the regularization loss function is added.
[0045]
Equation
[0046]
Equation
[0047] In one embodiment, since the perturbations for each class increase through the training of the AI model 110, the AI model 110 is trained to be robust to noise.
[0048] The perturbation adder 120 adds noise to the weight matrix of the classifier 112 based on the updated perturbation, and the classifier 112 can classify the next input data in the input data set (training set) using the weight matrix with the added noise based on the updated perturbation.
[0049] In one embodiment, when the training of the AI model 110 for the input data set is completed (for example, when a predetermined number of epochs have elapsed), the training determiner 130 can determine the end of the training of the AI model 110 based on the magnitude of the updated perturbation. For example, the training determiner 130 can compare the magnitude of the updated perturbation with a predetermined perturbation reference value to determine whether to end the training of the AI model 110. Alternatively, the training determiner 130 can calculate the class-by-class class uncertainty based on the updated perturbation and compare the calculated class uncertainty with an uncertainty reference value to determine whether to end the training of the AI model 110.
[0050] In one embodiment, the training determiner 130 can determine the class-by-class class uncertainty based on the updated perturbation (S140). For example, the class uncertainty calculated by the magnitude of the perturbation may be as shown in Equation 7.
[0051]
Equation
[0052] In one embodiment, the training determiner 130 compares the class uncertainty of each class with an uncertainty reference value (S150), and when the class uncertainty of at least one class is greater than the uncertainty reference value, it can determine to additionally collect the data of the class (S160). Then, the AI model 110 is trained again using the training set with the additional data.
[0053] In one embodiment, after a predetermined number of epochs are completed, if there is a class where the magnitude of the updated perturbation is smaller than the perturbation reference value, or if there is a class where the class uncertainty determined based on the updated perturbation is greater than the uncertainty reference value, it is determined that the training of the AI model 110 for the class is insufficient.
[0054] In one embodiment, if the magnitude of the updated perturbation of a certain class is greater than the perturbation reference value, it can mean that the AI model 110 has been robustly trained against the perturbation despite a large level of noise being added to the classifier 112. That is, it can be determined that the generalization performance of the AI model 110 for the class is high. Or if the class uncertainty of a certain class is smaller than the uncertainty reference value, it can be determined that the classification performance of the AI model 110 for the class is high.
[0055] For additional collection of data for a specific class, the domain where the data of the class is collected is determined. For example, when the data is an image collected in the semiconductor manufacturing process and the class that requires additional collection is ball breakage on the front of the semiconductor wafer, the domain where the image of the front of the wafer is collected is determined, and the images belonging to the class (ball breakage) within the domain are collected.
[0056] When data of a class having a perturbation smaller than the perturbation reference value or a class having a class uncertainty larger than the uncertainty reference value is additionally collected, the AI model 110 can perform additional training on the additionally collected dataset. The AI model 110 can perform training to update the perturbation until the perturbation magnitude of the class for which data is additionally collected becomes larger than the perturbation reference value or the class uncertainty of the class becomes smaller than the uncertainty reference value.
[0057] After that, when the perturbation magnitude of all classes is larger than the perturbation reference value or the class uncertainty of all classes is smaller than the uncertainty reference value, the training decision maker 130 can end the training of the AI model 110 (S170).
[0058] In one embodiment, the class uncertainty of the AI model 110 for a specific class can be used as an index indicating the performance of the AI model 110 for the class. The performance of the AI model 110 is verified only with the dataset for training without a separate test set by measuring the class uncertainty for each class of the trained AI model 110 by the training decision maker 130.
[0059] According to the AI model training apparatus and method according to one embodiment, since the AI model 110 trained using the training set does not need to be verified by a test set generated separately from the training set, the time and cost required for generating the test set can be saved. That is, since a test set generated by strict requirements (for example, not overlapping with the training set, data being strictly and accurately labeled, containing a sufficiently large amount of data, etc.) is not required, the cost and time can be saved very much.
[0060] Also, according to the training apparatus and method of the AI model according to one embodiment, since it is possible to identify classes that require additional data collection, it is possible to save the time and cost required for collecting data indiscriminately for all classes, and it is possible to prevent the problem of insufficient data for specific classes.
[0061] FIG. 3 is a diagram showing a training system of an AI model according to one embodiment.
[0062] Referring to FIG. 3, the training system of the AI model according to one embodiment may include a training apparatus 100 for training the AI model, a plurality of inspection apparatuses 10, and a data collection apparatus 20.
[0063] In one embodiment, the plurality of inspection apparatuses 10 can store product images obtained for product inspection such as non-destructive inspection and destructive inspection of products. The plurality of inspection apparatuses 10 can acquire product images during the manufacturing process of the product and perform non-destructive inspection using the product images. The plurality of inspection apparatuses 10 can sample products during the manufacturing process and perform destructive inspection on the sampled products.
[0064] In one embodiment, the data collection apparatus 20 can collect product images for training and testing of the AI model from the plurality of inspection apparatuses 10. Alternatively, the data collection apparatus 20 can acquire roads and road surrounding images necessary for autonomous driving of the vehicle. At this time, the roads and road surrounding images include various types of objects (such as people, vehicles, road structures, traffic signal devices, lanes, crosswalks, etc.).
[0065] In one embodiment, the training device 100 can determine the class-by-class uncertainty of the AI model based on perturbations updated through the training of the AI model. For example, when the magnitude of the updated perturbation for at least one class is large, the training device 100 can determine that the class uncertainty for that class is high and retrain the AI model for that class. Or, when the class uncertainty determined based on the magnitude of the updated perturbation for at least one class is high, the training device 100 can retrain the AI model for that class.
[0066] In one embodiment, the training device 100 can request the data collection device 20 to additionally collect data of classes determined to have high class uncertainty for additional training of the AI model. The data collection device 20 can determine the domain in which the data of the class with high class uncertainty is collected in order to additionally collect the data of the class with high class uncertainty.
[0067] For example, when the data is image data during the semiconductor manufacturing process and the class for which additional collection is requested is a crack on the back of the wafer, the data collection device 20 can determine the back image of the semiconductor wafer as the domain and additionally collect the image data belonging to that class (crack) from the inspection device 10 that acquires the images of that domain.
[0068] Or, when the data is semiconductor image data and the class for which additional collection is requested does not meet the critical dimension (CD) standard of the semiconductor cross-section, the data collection device 20 can additionally collect the image data belonging to that class (CD non-compliance) from a semiconductor destructive inspection device (for example, a scanning electron microscope (SEM)) that acquires the images of that domain.
[0069] In this way, by the training device 100 according to one embodiment determining the classes that require additional training of the AI model, the time and cost of the data collection device 10 for collecting data of the classes can be saved.
[0070] FIG. 4 is a block diagram showing an inference device using an AI model according to one embodiment, and FIG. 5 is a flowchart showing an inference method using an AI model according to one embodiment.
[0071] Referring to FIG. 4, an inference device 200 according to one embodiment can include an AI model 210, a perturbation adder 220, and a performance measurer 230.
[0072] In one embodiment, the AI model 210 is a trained model and can include an encoder 211 and a classifier 212. The encoder 211 can extract features from the data input to the AI model 210 to generate a feature vector. The classifier 212 can classify the class of the input data through an operation between a weight matrix for classification of the input data and the feature vector of the input data. In one embodiment, noise based on perturbations for each class is added to the weight matrix of the classifier 212.
[0073] In one embodiment, the AI model 210 can perform an inference to classify input data into one of a plurality of classes. The input data may be an image, text, voice, etc., and the class may be a category to which the input data belongs and is the purpose of classification by the AI model 210. For example, the input data may be a surrounding image acquired during autonomous driving of a vehicle, and the class may be the type of object (person, vehicle, road structure, traffic signal device, lane, crosswalk, etc.) included in the image acquired during driving of the vehicle.
[0074] In one embodiment, the input data may be a product image (e.g., semiconductor) during the manufacturing process of a product, and each class can indicate the type of defect that occurs during the manufacturing process of the product. For example, the class may be contamination, foreign matter, breakage, state of solder joint (ball), character abnormality, etc.
[0075] In one embodiment, the classifier 212 of the AI model 210 can convert the feature vector output by the encoder 211 into a score vector having the same dimension as the number of classes using a weight value matrix. The linear classification process performed by the classifier 212 can be shown as in Equation 1.
[0076] In one embodiment, the classifier 212 can convert the feature vector of the input data into a score vector indicating the score for each class through an operation such as Equation 1.
[0077] In one embodiment, the weight value matrix of the classifier 212 may be a set of representative vectors for each class. The weight value matrix of the classifier 212 and the representative vector of the i-th class can be expressed as in Equation 2 and Equation 3.
[0078] In one embodiment, the elements of the score vector indicating the result of linear classification can indicate the scores of the feature vectors for each class, and the scores can indicate the probability that the input data of the feature vector belongs to the class corresponding to the score. The dimension of the score vector is the same as the number of classes m. The element of the score vector corresponding to the i-th class among the elements of the score vector can be determined by the matrix product operation between the feature vector and the representative vector of the i-th class.
[0079] In one embodiment, the perturbation adder 220 determines perturbations of different magnitudes for each class based on class importance, determines noise by class-specific perturbations, and can add the determined noise to the representative vectors of the classes included in the weight value matrix of the classifier 212. Table 1 below shows the class importance for defect discrimination of semiconductor products (such as wafers).
[0080] Class importance for defect determination
[0081] [Table 1] In Table 1, the numbers of class importance can indicate the priority order of classes. For example, the class of ball dropout may be the class with the highest priority, and the class of observation error may be the class with the lowest priority. That is, classes such as ball dropout, crack, and ball short circuit are classes that the AI model 210 needs to accurately infer. Therefore, the inference accuracy for these classes may have a great impact on the performance of the AI model 210. On the contrary, even if the inference accuracy of the AI model 210 for classes such as observation error is relatively low, it may not have much impact on the performance of the AI model 210. In one embodiment, the multiple classes classified by the AI model 210 have different class importances, and the inference accuracy and / or uncertainty need to be measured differently for each class according to the class importance. The perturbation adder 220 can determine relatively large perturbations for classes with relatively high class importance and relatively low perturbations for classes with relatively low class importance. For example, when there are m classes, the perturbation adder 220 can add noise using the perturbation σ = m / 100 with the largest magnitude to the representative vector of the class with the highest class importance. Or the perturbation adder 220 can add noise using the perturbation σ = 1 / 100 with the smallest magnitude to the representative vector of the class with the lowest class importance.
[0082] In one embodiment, the performance measurement device 230 can measure the performance of the AI model 210 based on a plurality of inference results for the same input data of the classifier 212 having noise according to class importance. The performance measurement device 230 can determine the inference uncertainty of the AI model 210 based on a plurality of inference results of the AI model 210, and can measure the performance of the AI model 210 according to the inference uncertainty of the AI model 210. The inference uncertainty of the AI model 210 can be determined by the standard deviation or variance of a plurality of inference results for the same input data.
[0083] According to one embodiment, by adding the perturbation adder 220 to add noise due to perturbation based on class importance to the representative vector of the class, the performance measurement device 230 can measure the performance of the inference device 200 by other criteria according to class importance. For example, the performance measurement device 230 can strictly measure the inference performance of the AI model 210 for classes with relatively high class importance, and can leniently measure the inference performance of the AI model 210 for classes with relatively low class importance.
[0084] In one embodiment, for a class having a relatively high class importance, if the perturbation adder 220 adds noise determined based on a relatively large perturbation to the representative vector of the class, the inference performance of the AI model 210 for the class is likely to decrease. That is, when the inference uncertainty of the AI model 210 is measured to be low even if noise determined based on a relatively large perturbation is added to the representative vector, the performance measurement device 230 can highly evaluate the reliability of the AI model 210.
[0085] Or, for a class having a relatively low class importance, if the perturbation adder 220 adds noise determined based on a relatively small perturbation to the representative vector of the class, the inference performance of the AI model 210 for the class is less likely to decrease. If the class importance is low, since the performance of the AI model 210 for the class is not so important, it may not be necessary to strictly measure the inference performance of the AI model 210.
[0086] Referring to FIG. 4, in one embodiment, the perturbation adder 220 determines perturbations for each of a plurality of classes based on class importance (S210), and can add noise determined based on the perturbations to the representative vectors of each class (S220). The representative vector W of the i-th class to which noise due to perturbation is added i ’ may be as shown in Equation 4.
[0087] The perturbation adder 220 randomly samples noise from a predetermined noise distribution (following a standard normal distribution), and scales the randomly sampled noise by a perturbation predetermined according to class importance JPEG2025106230000013.jpg22 can only be scaled. The perturbation adder 220 scales the perturbation JPEG2025106230000014.jpg22 only-scaled noise by the perturbation JPEG2025106230000015.jpg22 and can add it to the representative vector of the class corresponding thereto. When noise scaled by a relatively large perturbation is added to the representative vector of a specific class, the inference uncertainty of that class may increase.
[0088] In one embodiment, the performance measurer 230 can obtain an inference result for input data from the trained AI model 210 (S230). The trained AI model 210 can perform inference on the input data using a weighted value matrix including the representative vectors to which noise is added, and output a score vector of the input data as an inference result.
[0089] In one embodiment, the score vector output from the trained AI model 210 can include scores for each class of the input data. The i-th element of the score vector can indicate the score for the i-th class of the input data. At this time, the inference result of the trained AI model 210 can be determined as the class corresponding to the largest element in the score vector. For example, when the largest element in the score vector is the i-th element, the inference result of the trained AI model 210 can be "the input data is the i-th class".
[0090] In one embodiment, the performance measurer 230 can collect a plurality of inference results of the trained AI model 210 for the same input data, and determine the inference uncertainty of the AI model 210 when the number of collected inference results reaches a predetermined number (I) (S240, S250).
[0091] In one embodiment, steps S220 and S230 are repeated until I inference results of the trained AI model 210 are collected. Until I inference results are collected, the perturbation adder 220 can resample noise from the noise distribution and add the newly sampled noise scaled by a perturbation predetermined according to class importance to the representative vector of each class. The trained AI model 210 can output the inference result for the same input data again using the weighted value matrix with noise added by the new sampling.
[0092] In one embodiment, the performance measurer 230 can measure the performance of the trained AI model 210 based on a plurality of inference results of the trained AI model 210 for the same input data.
[0093] In one embodiment, the performance measurer 230 is the standard deviation σ of a plurality of inference results of the trained AI model 210 infercan be determined as the inference uncertainty of the trained AI model 210. For example, when the standard deviation of a plurality of inference results of the trained AI model 210 is greater than a predetermined deviation threshold, the performance measuring device 230 may determine that the inference uncertainty of the trained AI model 210 is high and may not use the inference result of the trained AI model 210. Or, when the standard deviation of a plurality of inference results of the trained AI model 210 is less than a predetermined deviation threshold, the performance measuring device 230 may determine that the inference uncertainty of the trained AI model 210 is at a reliable level and may use the inference result of the trained AI model 210.
[0094] In one embodiment, the standard deviation σ of I inference results output by the trained AI model 210 infer may be a statistical value of an element having a maximum value in each of the I score vectors. At this time, the indexes of the elements having the maximum value in the I score vectors inferred from the same input data may all be the same.
[0095] In one embodiment, if the indexes of the elements having the maximum value in the I score vectors inferred from the same input data are not all the same, the performance measuring device 230 may determine that the inference uncertainty of the trained AI model 210 is maximum. That is, if the I score vectors are not all encoded as the same one-hot vector, it is determined that different inference results are output for the same input data. Therefore, the performance measuring device 230 does not trust the inference result of the trained AI model 210 without having to calculate the standard deviation of the plurality of inference results.
[0096] In one embodiment, the performance measuring device 230 can determine the inference uncertainty of the trained AI model 210 using a predetermined function from the standard deviation σ of a plurality of inference results of the trained AI model 210 infer The following Equation 8 shows the standard deviation σ of a plurality of inference results inferThe inference uncertainty U of the trained AI model 210 determined using a function pre-determined from infer is shown.
[0097]
Number
[0098] For example, when the inference uncertainty of the trained AI model 210 is higher than a pre-determined uncertainty threshold, the performance measurer 230 may not use the inference result of the trained AI model 210. Or, when the inference uncertainty of the trained AI model 210 is lower than a pre-determined uncertainty threshold, the performance measurer 230 can use the inference result of the trained AI model 210.
[0099] In one embodiment, when it is determined that the inference uncertainty of the trained AI model 210 is high, re-training of the trained AI model 210 can be determined. At this time, the performance measurer 230 can determine re-training of the trained AI model 210 only for classes with high class importance. When re-training of the trained AI model 210 is determined, data of the class indicated by the inference result showing high inference uncertainty is additionally collected.
[0100] In one embodiment, when the inference result of the trained AI model 210 is determined to be at a reliable level, the inference result of the trained AI model 210 is used (S260).
[0101] As described above, the inference device 200 according to one embodiment can add noise based on class-specific perturbations to the weight value matrix of the trained AI model 210 to more strictly measure the performance of the trained AI model 210 for classes with high importance. That is, when the trained AI model 210 outputs a highly reliable inference result for a class even if noise based on a large perturbation is added to the representative vector corresponding to the class with high importance, the inference device 200 can use the inference result of the trained AI model 210. Thereby, the performance of the trained AI model 210 can be easily measured according to the importance of the class, and the training target of the AI model 210 can be accurately targeted.
[0102] In addition, by verifying the performance of the trained AI model 210 through the inference uncertainty of the trained AI model 210 at the inference stage of the AI model, the classification performance for classes with high importance can be improved. For example, in the case of detecting product defects through the trained AI model 210, the defect detection performance of the trained AI model 210 can be improved, and thus an improvement in the product yield can be expected.
[0103] FIG. 6 is a diagram showing an artificial neural network according to one embodiment.
[0104] Referring to FIG. 6, an artificial neural network (ANN) 600 according to one embodiment can include an input layer 610, a hidden layer 620, and an output layer 630. The input layer 610, the hidden layer 620, and the output layer 630 each include a plurality of nodes, and the strength of the connection between each node can correspond to a weight value (weight connection). The plurality of nodes included in the input layer 610, the hidden layer 620, and the output layer 630 can be connected to each other in a fully connected manner. In one embodiment, the number of parameters (weight values and biases) may be the same as the number of weight connections in the artificial neural network 600.
[0105] The input layer 610 can include a plurality of input nodes (x1 to x i ), and the number of input nodes (x1 to x i ) can correspond to the number of independent variables of the input data. For the training of the artificial neural network 600, a training set is input to the input layer 610, and when the inference target data is input to the input layer 610 of the trained artificial neural network 600, the inference result is output at the output layer 630 of the trained artificial neural network 600. In one embodiment, the input layer 610 can have a structure suitable for processing large-scale inputs.
[0106] The hidden layer 620 is located between the input layer 610 and the output layer 630 and can include at least one hidden layer 6201 to 620 n . The output layer 630 can include at least one output node y1 to y j . Activation functions can be used for the hidden layer 620 and the output layer 630. In one embodiment, the artificial neural network 600 is trained in a manner that adjusts the weight values of the hidden nodes included in the hidden layer 620.
[0107] FIG. 7 is a block diagram showing a performance measurement system of an AI model according to one embodiment.
[0108] A performance measurement system of an AI model according to one embodiment can be realized as a computer system, for example, a computer-readable medium. Referring to FIG. 7, the computer system 700 includes at least one processor 710 and a memory 720. The memory 720 can be connected to the processor 710 and store various information for driving the processor 710 or at least one program executed by the processor 710.
[0109] The processor 710 can implement the functions, processes, or methods proposed in the embodiments. The operations of the computer system 700 according to the embodiments can be realized by the processor 710. At least one processor 710 can include at least one of a GPU, a CPU, and an NPU. When the operations of the computer system 700 are realized by at least one processor 710, each task is divided among the at least one processor 710 according to the load. For example, when one processor is a CPU, the other processor can be any one of a GPU, an NPU, an FPGA, and a DSP.
[0110] In the embodiments described herein, the memory 720 may be disposed inside or outside the processor, and the memory can be connected to the processor through various known means. The memory is various forms of volatile or non-volatile storage media. For example, the memory can include a read-only memory (ROM) or a random access memory (RAM).
[0111] On the one hand, an embodiment is not only realized through the aforementioned device and / or method, but can also be realized through a program that realizes the functions corresponding to the configuration of the embodiment or a recording medium on which the program is recorded. Such realization can be easily achieved by an ordinary technician in the technical field to which this description pertains from the description of the aforementioned embodiment. Specifically, the method according to the embodiment (for example, an image preprocessing method, etc.) is realized in the form of program instructions that can be executed through various computer means and can be recorded on a computer-readable medium. A computer-readable medium can include program instructions, data files, data structures, etc. alone or in combination. The program instructions recorded on the computer-readable medium can be those specifically designed and configured for the embodiment or those known and usable by an ordinary technician in the field of computer software. The computer-readable recording medium can include a hardware device configured to store and execute program instructions. For example, the computer-readable recording medium can be a magnetic medium such as a hard disk, a floppy disk, and a magnetic tape, an optical medium such as a CD-ROM and a DVD, a magneto-optical medium such as a floptical disk, a ROM, a RAM, a flash memory, etc. The program instructions can include not only machine language code generated by a compiler but also high-level language code that can be executed by a computer through an interpreter, etc.
[0112] Although the embodiments have been described in detail above, the scope of the rights described herein is not limited thereto, and various modifications and improvements made by those skilled in the art using the basic concepts defined in the following claims also fall within the scope of the rights described herein.
Explanation of Reference Numerals
[0113] 10 Inspection device 20 Data collection device 100 Training Device 110, 210 Artificial Intelligence (AI) Model 111, 211 Encoder 112, 212 Classifier 120, 220 Perturbation Adder 130 Training Decision Maker 200 Inference Device 230 Performance Measurer 700 Computer System 710 Processor 720 Memory
Claims
1. An apparatus for measuring the performance of an AI model, comprising a processor and a memory, the memory including instructions configured for the processor to execute a process, the process determining noise based on perturbation; adding the noise to each representative vector corresponding to a plurality of classes of a linear classifier of the AI model; determining the inference uncertainty of the AI model for the image based on a plurality of inference results for one image output by the linear classifier of the AI model using the representative vectors added with the noise The apparatus includes.
2. The process sampling noise from a noise distribution following a standard normal distribution; scaling the sampled noise by the magnitude of the perturbation further includes The step of adding the noise to each representative vector corresponding to a plurality of classes of a linear classifier of the AI model is the step of adding the scaled noise to each representative vector The apparatus according to claim 1, including.
3. The step of sampling the noise is The apparatus according to claim 2, wherein noise for each of the plurality of classes is independently sampled from the noise distribution.
4. The step of adding the noise to each representative vector corresponding to a plurality of classes of a linear classifier of the AI model is the step of independently adding the noise to each of the plurality of representative vectors corresponding to the classes The apparatus according to claim 1, including.
5. The process further includes determining perturbation for each of the plurality of classes in proportion to the class importance of the plurality of classes The apparatus according to claim 4, including.
6. The plurality of inference results correspond to class classification results of the feature vector of the image generated by an encoder of the AI model, The class classification result is determined by classifying the feature vector into one of the plurality of classes by the linear classifier using the plurality of representative vectors added with the noise. The apparatus according to claim 1.
7. The process further includes receiving a predetermined number of class classification results from the linear classifier as the plurality of inference results The apparatus according to claim 6, including.
8. The step of determining the inference uncertainty of the AI model for the image is Determining the standard deviation of the plurality of inference results as the inference uncertainty of the AI model The apparatus according to any one of claims 1 to 7, comprising:
9. The inference result of the AI model is not used as a final result based on a determination that the standard deviation of the inference result satisfies a condition of a predetermined deviation threshold, or The apparatus according to claim 8, wherein the inference result of the AI model is used as a final result based on a determination that the standard deviation of the inference result does not satisfy a condition of a predetermined deviation threshold.
10. The apparatus according to claim 8, wherein the standard deviation is statistically determined based on a plurality of score vectors output as the inference result.
11. The apparatus according to claim 10, wherein the indices of elements having the maximum score within the plurality of score vectors are the same.
12. A method for training an AI model, comprising: Adding noise determined based on perturbation to a plurality of representative vectors respectively corresponding to a plurality of classes; Updating the magnitude of the perturbation by training the AI model performed using a weighted value matrix including the plurality of representative vectors to which the noise is added; Determining the end of the training of the AI model based on the updated perturbation And a method.
13. The step of adding noise determined based on perturbation to a plurality of representative vectors respectively corresponding to a plurality of classes includes: Sampling noise from a noise distribution following a normal distribution; Scaling the sampled noise by the magnitude of the perturbation; Adding the scaled noise to the plurality of representative vectors The method according to claim 12, comprising:
14. The step of sampling noise from a noise distribution following a normal distribution includes: Sampling noise independently for each of the plurality of classes from the noise distribution The method according to claim 13, comprising:
15. The step of adding noise determined based on perturbation to a plurality of representative vectors respectively corresponding to a plurality of classes is Sampling the noise added to the first representative vector among the plurality of representative vectors from a noise distribution following a normal distribution, and sampling the noise added to a second representative vector different from the first representative vector among the plurality of representative vectors from the noise distribution; Adding the noise added to the first representative vector to the first representative vector, and adding the noise added to the second representative vector to the second representative vector; The method according to claim 12, comprising:
16. The method according to claim 12, wherein the magnitude of the updated perturbation is larger than the magnitude of the perturbation before being updated.
17. The step of determining the end of the training of the AI model based on the updated perturbation: Comparing the magnitude of the updated perturbation with a predetermined perturbation reference value to determine the end of the training of the AI model The method according to any one of claims 12 to 16, comprising:
18. The step of determining the end of the training of the AI model based on the updated perturbation: Calculating the class uncertainty of the plurality of classes based on the updated perturbation; Determining the end of the training of the AI model based on the class uncertainty of the plurality of classes The method according to any one of claims 12 to 16, comprising:
19. A method for measuring the performance of an image classification model, comprising: Adding, to each of a plurality of representative vectors corresponding to a plurality of classes, noise determined based on a predetermined perturbation for each of the plurality of classes; Receiving a predetermined number of inference results for one image performed by the image classification model using the plurality of representative vectors to which the noise is added; Calculating the inference uncertainty of the image classification model based on the predetermined number of inference results The method comprising:
20. The method according to claim 19, wherein the image is an image of a semiconductor product collected in a semiconductor manufacturing process.
Citation Information
Cited By
Camera module capable of acquiring an accurate image
US12621552B2