Hybrid generative model device and out-of-distribution determination method using same

The hybrid generative model addresses the challenge of distinguishing in-distribution and out-of-distribution data using Wasserstein distance and mutual information, ensuring robust image classification by quantifying damage difficulty, thus improving model performance in applications like autonomous vehicles and deepfakes.

WO2026101076A1PCT designated stage Publication Date: 2026-05-15IOPS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
IOPS CO LTD
Filing Date
2025-10-24
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing image classification models struggle to effectively distinguish between in-distribution and out-of-distribution data, particularly in applications like autonomous vehicles and deepfakes, where new input data not seen during training can lead to misclassification.

Method used

A hybrid generative model using Wasserstein distance, mutual information, and minimum description length to determine out-of-distribution data, incorporating a normalizing flow model and global mean pooling for improved classification performance.

Benefits of technology

Effectively distinguishes between in-distribution and out-of-distribution data, ensuring image integrity by measuring damage difficulty through covariate and semantic changes, enhancing model robustness against unseen data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025017018_15052026_PF_FP_ABST
    Figure KR2025017018_15052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a hybrid generative model device and an out-of-distribution determination method using same, wherein the hybrid generative model device comprises: a first input unit for inputting training data; a second input unit for inputting test data; a learning model unit for inputting the training data to a hybrid generative model to train the hybrid generative model, and inputting the test data to the hybrid generative model and allowing the hybrid generative model to produce an output; and an out-of-distribution (OOD) determination unit for determining OOD through a Wasserstein distance, which is measured according to the output of the hybrid generative model, between the training data and the test data and mutual information about the training data and the test data, and a minimal description length of the test data. Thereby, it is possible to effectively secure the integrity of an image as well as effectively determine in-distribution data and out-of-distribution data by using the hybrid generative model.
Need to check novelty before this filing date? Find Prior Art

Description

Hybrid generative model device and method for determining out-of-distribution using the same

[0001] The present invention relates to a hybrid generative model device and a method for determining out-of-distribution using the same, which can effectively distinguish between distribution data and out-of-distribution data using a hybrid generative model and effectively ensure image integrity by determining out-of-distribution (OOD) through the Wasserstein distance and mutual information of training data and test data measured according to the output of the hybrid generative model and the minimum description length of the test data.

[0002]

[0003] As is well known, the field of image classification using deep learning has recently been actively researched. However, general image classification models output the closest class based on the assumption that the image belongs to a target distribution, and out-of-distribution (OOD) detection, which detects abnormal data outside the target distribution, is emerging as an important problem in deep learning image classification.

[0004] OOD detection in such image classification is a technique that allows a model to determine whether an image to be classified belongs to a target distribution. Techniques are being proposed such as training the model with OOD detection as a task during the training process, or analyzing parameters calculated when an image is input into the model.

[0005] As described above, ODD detection in image classification models is to determine that OOD data is input to the model when it does not correspond to the target distribution. To address this, the softmax score output from the model can be used to distinguish between In Distribution (ID) data and OOD data. For example, techniques such as input processing on images and techniques that consider OOD data from the training stage of the model have been proposed.

[0006] Meanwhile, in image classifiers for autonomous vehicles, new input data that was not seen during training (e.g., wild animals passing on a highway) may be input, and to address this, ODD detection is essential. In broad classification systems that process images, such as in the field of deepfakes, anomaly detection, new object detection, and inductive and transfer learning tasks are emerging, and various techniques are being proposed.

[0007]

[0008] [Prior Art Literature]

[0009] (Patent Document) Korean Published Patent No. 10-2022-0157835 (Published Nov. 29, 2022)

[0010]

[0011] The present invention aims to provide a hybrid generative model device and a method for determining out-of-distribution using the same, which can effectively distinguish between distribution data and out-of-distribution data using a hybrid generative model and effectively ensure image integrity by determining OOD through Wasserstein distance and mutual information of training data and test data measured according to the output of the hybrid generative model and the minimum explanation length of the test data.

[0012]

[0013] The purposes of the embodiments of the present invention are not limited to those mentioned above, and other unmentioned purposes will be clearly understood by those skilled in the art from the description below.

[0014]

[0015] According to one aspect of the present invention, a hybrid generative model device may be provided, comprising: a first input unit for inputting training data; a second input unit for inputting test data; a learning model unit for inputting the training data into a hybrid generative model to train it, and inputting and outputting the test data into the hybrid generative model; and an OOD determination unit for determining Out-Of-Distribution (OOD) through the Wasserstein Distance and Mutual Information of the training data and test data measured according to the output of the hybrid generative model, and the Minimal Description Length of the test data.

[0016] In addition, according to one aspect of the present invention, a hybrid generative model device may be provided in which the training data utilizes in-distribution data and the test data includes out-of-distribution data.

[0017] In addition, according to one aspect of the present invention, the hybrid generative model may be provided as a hybrid generative model device including a normalizing flow model.

[0018] In addition, according to one aspect of the present invention, the OOD determination unit may be provided with a hybrid generative model device that determines the OOD by deriving a damage difficulty ranking using the Wasserstein distance, mutual information, and minimum explanation length.

[0019] In addition, according to one aspect of the present invention, the OOD discrimination unit may be provided with a hybrid generative model device that derives the damage difficulty ranking based on changes in covariates and semantic changes for the training data and test data.

[0020]

[0021] According to another aspect of the present invention, a method for determining out-of-distribution using a hybrid generative model device may be provided, comprising: a step of inputting training data through a first input unit; a step of inputting the training data into a hybrid generative model in a learning model unit to train it; a step of inputting test data through a second input unit; a step of inputting and outputting the test data to the hybrid generative model in the learning model unit; and a step of determining out-of-distribution (OOD) in an OOD determination unit through the Wasserstein distance and mutual information of the training data and test data measured according to the output of the hybrid generative model, and the minimal description length of the test data.

[0022] In addition, according to another aspect of the present invention, a method for determining out-of-distribution using a hybrid generative model device may be provided, wherein the training data utilizes in-distribution data and the test data includes out-of-distribution data.

[0023] In addition, according to another aspect of the present invention, the hybrid generative model may provide a method for determining out-of-distribution using a hybrid generative model device including a normalizing flow model.

[0024] In addition, according to another aspect of the present invention, the step of determining the OOD may be provided as an out-of-distribution determination method using a hybrid generative model device that determines the OOD by deriving a damage difficulty ranking using the Wasserstein distance, mutual information, and minimum explanation length in the OOD determination unit.

[0025] In addition, according to another aspect of the present invention, the step of determining the OOD may be provided as an out-of-distribution determination method using a hybrid generative model device that derives the damage difficulty ranking based on the change in covariates and the change in semantics for the training data and test data in the OOD determination unit.

[0026]

[0027] The present invention determines OOD through Wasserstein distance and mutual information of training data and test data measured according to the output of a hybrid generative model, and the minimum explanation length of the test data, thereby enabling effective determination of attribution data and out-of-distribution data using a hybrid generative model, as well as effectively ensuring image integrity.

[0028]

[0029] FIG. 1 is a block diagram of a hybrid generative model device according to one embodiment of the present invention, and

[0030] FIGS. 2 to 8 are drawings for explaining the detailed configuration of a hybrid generative model device according to an embodiment of the present invention, and

[0031] FIG. 9 is a flowchart illustrating the process of determining out-of-distribution using a hybrid generative model device according to another embodiment of the present invention.

[0032]

[0033] The advantages and features of the embodiments of the present invention, and the methods for achieving them, will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below but may be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims. Throughout the specification, the same reference numerals refer to the same components.

[0034] In describing the embodiments of the present invention, specific descriptions of known functions or configurations will be omitted if it is determined that such detailed descriptions could unnecessarily obscure the essence of the invention. Furthermore, the terms described below are defined in consideration of their functions in the embodiments of the present invention, and these definitions may vary depending on the intentions or practices of the user or operator. Therefore, such definitions should be based on the content throughout this specification.

[0035] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.

[0036]

[0037] FIG. 1 is a block diagram of a hybrid generative model device according to one embodiment of the present invention, and FIGS. 2 to 8 are drawings for explaining the detailed configuration of a hybrid generative model device according to one embodiment of the present invention.

[0038]

[0039] Referring to FIGS. 1 to 8, a hybrid generative model device according to one embodiment of the present invention may include a first input unit (110), a second input unit (120), a learning model unit (130), an OOD discrimination unit (140), etc.

[0040]

[0041] The first input unit (110) is a component that inputs training data, and the training data can be, for example, In Distribution data.

[0042] Training data using such attribution data (ID) can be input into the learning model unit (130) for training a hybrid generative model.

[0043]

[0044] The second input unit (120) is a component that inputs test data, and the test data may include, for example, Out-Of-Distribution data.

[0045] Test data containing such out-of-distribution data (OOD) can be input into a learning model unit (130) to be input into a learned hybrid generative model, and separately, can be transmitted to an OOD discrimination unit (140) to obtain the Minimal Description Length of the test data.

[0046] For example, Figure 2 shows a schematic diagram of ID (including augmented internal distribution) and OOD (covariate change and semantic change) and samples for each distribution. It can be seen that the deep generative model trained on MNIST can be generalized to the augmented distribution and overlaps with the covariate change, and that the semantic change is furthest from the internal distribution, but the relative distance may vary depending on whether there is the fewest or no intersection with the internal distribution for the most difficult type of covariate change.

[0047]

[0048] The learning model unit (130) is a component that inputs training data into a hybrid generative model to train it, and inputs and outputs test data into a hybrid generative model. The hybrid generative model may include, for example, a Normalizing Flow model.

[0049] Hybrid generative models including such normalizing flow models can utilize classifiers and Glow models that include latent feature flattening, two linear layers, and an output layer, or that include two convolutional layers without latent feature flattening, global mean pooling, and an output layer. Based on hybrid generative models, for data classification, the distance between training data and test data can be measured and detected using cosine similarity, Wasserstein distance, log distribution (log(p(z))), etc., derived from latent information by utilizing global mean pooling in the latent space.

[0050] As described above, classification performance can be improved by directly using latent features through a hybrid generative model composed of latent features, global mean pooling, and an output layer.

[0051] For example, FIG. 3 shows the architecture of a hybrid generative model according to one embodiment of the present invention, where ID represents training data as in-distribution data and ODD represents test data including out-of-distribution data, and ID and OOD are each input to f(x), which is a feature encoding function of a normalizing flow model, and each latent feature (z) can be output as the output of the function.

[0052] In addition, the latent features output corresponding to the ID are input into a Global Pooling Linear Model (GAPLM) to learn a classification task, through which features usable for classification and image generation can be obtained.

[0053] In addition, the features obtained corresponding to the ID can be input into f-1(z), which is the feature decoding function (i.e., image reconstruction function) of the normalizing flow model, to be output as a pseudo-ID(x'), and the features obtained corresponding to the ID and the latent features corresponding to the OOD are each generated densities ( , It can be output as ).

[0054] Using the output values ​​described above, the distance between training data and test data can be measured and detected using cosine similarity, Wasserstein distance, log distribution (log(p(z))), etc.

[0055]

[0056] The OOD determination unit (140) is a component that determines OOD (Out-Of-Distribution) through the Wasserstein distance and mutual information of the training data and test data measured according to the output of the hybrid generative model, and the minimum explanation length of the test data, and can determine OOD by deriving a damage difficulty ranking using the Wasserstein distance, mutual information, and minimum explanation length.

[0057] This OOD discrimination unit (140) can derive a damage difficulty ranking based on changes in covariates and semantic changes in training data and test data, and the distance index (i.e., similarity index) that compares the model accuracy index (F1-score) and the distance (i.e., similarity) represents the damage difficulty, and the log distribution plays the same role as the distance index, and accordingly, it can be confirmed that when the damage difficulty is low, the area where the ID and the damage data (OOD) overlap is larger, and through this, it can be seen that the classification of damage data can be effectively performed even if the training of the hybrid generative model is performed using ID without using damage data.

[0058] Here, the log distribution of the semantic change data can be separated from the ID and covariate changes.

[0059] The OOD discrimination unit (140) described above can measure the difficulty of damage according to the covariate change attribute using distribution indicators (e.g., cosine similarity, Wasserstein distance, etc.) and model accuracy indicators (e.g., F1, etc.) for training data and test data, and can quantify the covariate change and semantic change based on a hybrid generative model. Here, the log density distribution of ID, covariate change and semantic change can move according to the covariate attribute.

[0060] Additionally, the OOD discrimination unit (140) measures the distance between training data and test data based on a hybrid generative model in which global average pooling in the latent space is used for data classification, and the measurement method may include cosine similarity, Wasserstein distance, log distribution, etc.

[0061] Meanwhile, in Figure 4, ID and OOD (covariate change and semantic change) can be distinguished for MNIST, MNIST-C, and Fashion-MNIST through a hybrid generative model, and it can be seen that this distribution is arranged in a generative density in a similar manner as shown in Figure 2.

[0062] To explain the OOD, covariate shift, and semantic shift described above, when a training dataset is derived from a source distribution (p(x,y)), this data is used in a prediction model (p(y|x)). However, if the test dataset has a target distribution (q(x,y)) that is different from the source distribution (i.e., p≠q), this is called a distribution shift or dataset shift, and the hybrid generative model can detect covariate shift and semantic shift among the major types of distribution shifts.

[0063] And, covariance shift (or domain shift) means that the characteristic distribution (p(x)) changes and the prediction model (p(y|x)) remains fixed. For example, the same tree may have different visual images in satellite images depending on the season, Class 1 numbers of different colored IDs (e.g., MNIST data) may have the same label, and the background of a photograph of a cow may be a variety of grassy fields.

[0064] In addition, the covariate distribution (q(x)) is different from the training distribution (p(x)), meaning that the image and label are both changed from source to target when the source domain is p(x)p(y|x) while the target domain is q(x)q(y|x), which is the semantic shift that out-of-distribution primarily aims to solve.

[0065] Meanwhile, regarding the training joint distribution, the test risk can be derived by first decomposing the covariate variation into q(x,y)=q(x)q(y|x) and from q(y|x)=p(y|x) as shown in Equation 1 below.

[0066] [Mathematical Formula 1]

[0067]

[0068] Here, min w??∫dxq(x) means optimizing the parameter w, which is used to minimize the objective function by adjusting the parameter w in a given function or model f, ∫dxq(x) is an integral over the input variable x, where q(x) is a function representing the distribution of x and represents the prior distribution of the data, and is used to evaluate the average over x through weighted integration over q(x), and ∫dyp(y|x) is an integral over the distribution p(y|x) conditioned on x with respect to y, which represents calculating the expected value of y after sampling it according to the conditional probability distribution (y|x).

[0069] In addition, in l(f(f(x,w),y)), f(x,w) is a function calculated using input x and parameter w, which is evaluated again with y as f, and the loss function l is applied to the result, and the loss function l represents a function that indicates the difference between the output of this model and the target value (or actual value).

[0070] Equation 1 above represents an optimization problem in which a model f finds the optimal parameter w for given inputs x and y. The model f generates an output value through the input x and parameter w, and the result can be further combined with y and evaluated by a loss function l, and finally, w that minimizes this loss can be found.

[0071] Meanwhile, regarding the damage dataset (OOD), it can be usefully applied in applications that can prevent image fraud or forgery (e.g., fake face recognition) by measuring vulnerability to damage datasets (e.g., impulse noise, etc.) and can also be used in satellite imagery. Damage datasets such as IMAGENET-C and MNIST-C can be used to measure the robustness of a model. The IMAGENET-C and MNIST-C datasets consist of 19 and 16 damage types, respectively, and MNIST-C, which is derived from IMAGENET-C and CIFAR10-C, is data adapted to MNIST to measure OOD in terms of model performance.

[0072] Here, MNIST-C was selected based on four impairment principles (e.g., (1) non-triviality, (2) semantic invariance, (3) realism, (4) breadth), and impairment was designed to degrade the accuracy of the model. Since image impairment occurs in real environments, the dataset includes attributes from sensors, environments, and physical factors, and was designed with redundancy with other impairment datasets in mind.

[0073] To elaborate on the hybrid generative model described above, the hybrid generative model according to one embodiment of the present invention can learn distribution data using a normalizing flow based on the GlOW model according to the architecture shown in FIG. 3, and can be performed by calculating the damage difficulty for out-of-distribution data to measure changes in covariates.

[0074] Specifically, the hybrid generative model is composed of a linear model through a normalizing flow based on a GLOW model and global mean pooling according to the architecture shown in Fig. 3, where the GLOW model comprises a 1×1 convolution as shown in Equation 2 below and for It may include a logarithmic determination formula such as. Here, It represents the function in the model's latent space, and represents the weight of each layer, and is the determinant of the weight matrix Wl of each layer.

[0075] [Mathematical Formula 2]

[0076]

[0077] Equation 2 above defines a joint distribution p(x,y) model consisting of feature and label pairs (xn,yn) using a global mean pooling linear model (GAPLM) for the output of the normalizing flow model.

[0078] In this global mean pooling linear model (GAPLM), derived from latent features (z) This can serve as the final layer of a selective classifier using global mean pooling, and the classification probability can be expressed as Equation 3 below.

[0079] [Mathematical Formula 3]

[0080]

[0081] Here, g -1 silver Represents a link function such as

[0082]

[0083] Next, regarding the optional classification for OOD detection in hybrid generative models, during the prediction or OOD detection step, the user can selectively proceed with classification based on a threshold rejection rule; for example, the observed a threshold If it is smaller, the data can be rejected as OOD, and a generator element regarding whether the given test data belongs to ID or OOD. The model can be determined based on.

[0084] To explain the estimation of such test data damage, IMAGENET-C can evaluate the robustness and damage difficulty of a classifier based on five damage levels. Corruption Error (CE) is a standardized performance metric that measures the performance degradation of a model on a damaged dataset after it has been trained on a clean dataset. Unlike CE, Corruption Difficulty Ranking (CDR) can directly measure the distance between distributions through a hybrid generative model according to an embodiment of the present invention without additional testing at each damage level, and can simultaneously evaluate the complexity of the input data in terms of model and data uncertainty.

[0085]

[0086] Meanwhile, mutual information (or mutual information, MI) is an indicator for calculating uncertainty between two random variables X and Y, and can be defined as shown in Equation 4 below.

[0087] [Mathematical Formula 4]

[0088]

[0089] The above mathematical formula 4 describes how similar the test data is to the training data and can be interpreted as a measure of the level of image damage. Wasserstein distance can be used to measure how far the OOD is from the ID, and this can be used to calculate the difference between the change in covariates and the change in semantics.

[0090]

[0091] This Wasserstein distance can be defined as shown in Equation 5 below.

[0092] [Mathematical Formula 5]

[0093]

[0094] Here, J(P,Q) represents all joint distributions J of (X,Y) with marginal probabilities P and Q, and each random variable is class It represents, and this metric tends to capture noise types in damage data.

[0095] The mutual information (MI) described above is a measure of dependency between two random variables, which is 0 when the two distributions are independent and becomes infinity when the two distributions are identical. From the perspective of representation learning, the Wasserstein Dependency Measure (WDM), a modified version of mutual information in which KL divergence is changed to Wasserstein distance, can be provided as the posterior probability.

[0096] Here, the combination of mutual information and Wasserstein distance in the latent space of a hybrid generative model according to one embodiment of the present invention can serve as an indicator that captures important features for measuring model uncertainty, damage difficulty, and OOD detection, and can be used as a prior distribution in terms of KL divergence of WDM as shown in Equation 6 below.

[0097] [Mathematical Formula 6]

[0098]

[0099] Here, each distribution represents the uncertainty of the model in the normalizing flow of Equation 2 above. as, It can include both discriminative and generative elements of the model.

[0100] Additionally, to measure aleatoric uncertainty along with image complexity, a Minimum Description Length (MDL) score can be introduced, which indicates how difficult it is for the input to reach the desired quality; MDL is an implicit inductive bias for measuring complexity.

[0101]

[0102] Meanwhile, the damage dataset can be quantified through the inherent data complexity and the model's ODD capability. By integrating mutual information and Wasserstein distance, the Corruption Difficulty Ranking (CDR), a distance-based damage measurement technique, can be expressed as shown in Equation 7 below.

[0103] [Mathematical Formula 7]

[0104]

[0105] Here, WDrank, MIrank, and MDLrank represent the respective rank values ​​of Wasserstein distance, mutual information, and MDL, respectively, and α, β, and γ represent weighting coefficients.

[0106] The CDR (Damage Difficulty Ranking) as in Equation 7 above can calculate the correlation or inverse correlation for the damage dataset tested by the hybrid generative model. As shown in Table 1 below, the most damaged images correspond to the top ranks in the CDR, and for example, the identity damage attribute may be in the subgroup, and the scale attribute may be in the subgroup in relation to mutual information.

[0107] [Table 1]

[0108]

[0109] [Table 2]

[0110]

[0111] Here, IQ (Image quality) represents image quality, RN (Random Noise) represents random noise, Con (Content) represents content, Geo (Geometry) represents geometry, and the arrow direction (↑, ↓) indicates which deformation is more difficult in each indicator.

[0112] Tables 1 and 2 above represent the CDR of damage data for MNIST and CIFAR10, respectively. Each damage difficulty metric ranks different damage characteristics, with WD ranking mainly noise types, MI ranking noise and quality, and MDL ranking quality. For CIFAR10-C, WD ranks noise, MI ranking quality, and MDL ranking quality. Overall, CDR represents an integrated damage metric that quantifies image quality, noise, and content (or blur, noise, digital, and weather).

[0113] From Tables 1 and 2 above, it can be seen that CDR is inversely proportional to Mutual Information (MI) and proportional to Wasserstein Distance (WD) and MDL. The overall ranking can be defined according to the ranking calculated using the rankings of WD, MI, and MDL. For example, in Table 1 above, fog ranks 1st, 4th, and 1st in WD, MI, and MDL, respectively. Finally, the score for fog is 5+4+1=10, and the lowest score indicates the most difficult damage. Here, α, β, and γ are set to 1.

[0114]

[0115] To describe a simulation using a hybrid generative model device according to one embodiment of the present invention as described above, first, the training dataset for learning the hybrid generative model uses only ID data, and the OOD dataset, which is the test data, may include a dataset having covariate changes (damage data) and semantic changes.

[0116] For example, datasets such as MNIST, MNIST-C, Fashion-MNIST, CIFAR10, and CIFAR10-C can be used. MNIST is used as training data and contains digit images (0–9) consisting of single-channel images of size 28×28. MNIST-C is used as covariate change (damage) data and contains images of damaged digits, including various damage types such as brightness, Canny edge, fog, dotted line, and glass blur. Fashion-MNIST is used as semantic change data and contains grayscale images of 10 fashion items (e.g., T-shirts and tops, pants, pullovers, dresses, coats, etc.) and has the same size as MNIST and MNIST-C.

[0117] In addition, CIFAR10 is used as training data and includes small-sized (32×32) RGB images of various classes (e.g., birds, cats, airplanes, cars, deer, etc.), CIFAR10-C is used as covariate change (damage) data and is a damaged version of CIFAR10 and includes various damage types such as brightness, contrast, elastic transformation, impulse noise, snow, zoom blur, etc., and SVHN is used as semantic change data and is a digit image dataset similar to MNIST, which includes images of house numbers taken on the street and includes RGB channels.

[0118] Next, to demonstrate the OOD detection and damage estimation capabilities of the hybrid generative model, the dataset groups are described as follows: the first group consists of MNIST, MNIST-C, and Fashion-MNIST, where MNIST is used as in-domain data, MNIST-C is used as covariate change (damage) data, and Fashion-MNIST is used as semantic change data.

[0119] And, the second group consists of CIFAR10, CIFAR10-C, and SVHN, where CIFAR10 is used as in-domain data, CIFAR10-C is used as covariate change (damage) data, and SVHN is used as semantic change data.

[0120]

[0121] Figures 5 and 6 show the relationship between the Wasserstein distance (WD) and the F1 score of CIFAR10-C, based on the results of performing a hybrid generative model using the training and test data described above. Figure 5 shows a plot between the WD and the F1 score, and Figure 6 shows a plot between the WD and the mutual information (MI). These visual representations can visually explain how damage difficulty affects the performance of the model.

[0122] Meanwhile, the log p(x) graph derived from the output of the hybrid generative model is shown in Figures 7 and 8, which are aligned according to the CDRs described in Tables 1 and 2, and show that, from the perspective of the potential characteristics of the hybrid generative model, the covariate distributions close to ID tend to be relatively easily damaged. That is, difficult damage is far from the ID distribution, and some damages further from ID can be interpreted as very difficult damages even more so than semantic change. It can be seen that these damages require a lot of effort to generalize when learning covariate change, and that semantic change has a distribution distinct from the ID distribution as OOD.

[0123] Specifically, Figure 7 shows the log distribution (log p(x)) values ​​in the MNIST-C dataset, with the order in accordance with Table 1 above, where the most difficult variation is located at the top left and the easiest variation is located at the bottom right. In this graph, the ID, covariate variation, and semantic variation are represented starting from the left waveform, respectively, and the ID and semantic variation data (FAN-MNIST) are used identically in all graphs.

[0124] In addition, Figure 8 shows the log distribution (log p(x)) values ​​of CIFAR10-C, with the most difficult variation located in the top left and the easiest variation in the bottom right according to the order presented in Table 2 above. In this graph, the ID, covariate variation, and semantic variation are represented starting from the left waveform, respectively, and the same ID data and semantic variation data (SVHN) are used identically in all graphs.

[0125]

[0126] Finally, regarding the damage difficulty ranking, Table 1 and Table 2 above show the damage difficulty ranking (CDR) of damage data in MNIST and CIFAR10, respectively. It can be seen that high-ranking damages such as fog and impulse noise in MNIST-C also have high rankings in WD (Wassestein distance) and MI (mutual information). Since these two indicators are derived from latent features, their behavior appears similar, but it can be confirmed that MDL (minimum explanation length) shows a slightly different pattern from WD and MI.

[0127] For example, by referring to the stripe of MNIST-C and the top-ranked damage of CIFAR10-C, the damage difficulty ranking (CDR) described above can determine OOD by simultaneously considering latent features (WD, MI) and data complexity (MDL).

[0128] When compared to the performance of conventional technology, it can be seen that the Conv3(GAN) model shows similar test accuracy for low CDRs, such as impulse noise in MNIST-C data.

[0129] Meanwhile, to further explain the Global Average Pooling Linear Model (GAPLM) of the hybrid generative model, global average pooling applied through the convolutional layers of an MLP (Multilayer Perceptron) can approximate the confidence map more accurately than a GLM (Generalized Linear Model), and the architecture of the hybrid generative model has the advantage of being able to accommodate other network structures.

[0130] In addition, when comparing the linear layer and GAPLM structures that use flattened latent features, as shown in Table 3 below, it can be seen that while both models show similar performance on the MNIST dataset, the linear model shows better performance on the CIFAR10 dataset. However, since GAPLM summarizes latent feature z generated from the flow, it can be seen that the number of parameters is small, and the number of parameters in the classification layer has decreased from 15,360 to 240.

[0131] [Table 3]

[0132]

[0133] In this way, since GAPLM summarizes z with minimal loss of classification performance, it can be confirmed that z is a suitable feature for classification.

[0134]

[0135] Accordingly, according to one embodiment of the present invention, by determining OOD through Wasserstein distance and mutual information of training data and test data measured according to the output of a hybrid generative model and the minimum explanation length of the test data, it is possible to effectively distinguish between attribution data and out-of-distribution data using a hybrid generative model, as well as effectively ensure the integrity of the image.

[0136]

[0137] FIG. 9 is a flowchart illustrating the process of determining out-of-distribution using a hybrid generative model device according to another embodiment of the present invention. Here, since the specific details regarding the process of determining out-of-distribution using a hybrid generative model device have been described in detail in one embodiment of the present invention, the process will be described in a general manner below.

[0138]

[0139] Referring to FIG. 9, training data can be input through the first input unit (110) (step 210).

[0140] Here, training data can use, for example, in-distribution data.

[0141]

[0142] And, training data can be input into the hybrid generative model in the learning model section (130) to train it (step 220).

[0143] Here, the hybrid generative model may include, for example, a Normalizing Flow model.

[0144]

[0145] Meanwhile, test data can be input through the second input unit (120) (step 230).

[0146]

[0147] Here, the test data may include, for example, out-of-distribution data.

[0148]

[0149] Next, test data can be input and output to the hybrid generative model in the learning model section (130) (step 240).

[0150]

[0151] Next, the OOD (Out-Of-Distribution) can be determined through the Wasserstein Distance and Mutual Information of the training data and test data measured according to the output of the hybrid generative model in the OOD determination unit (140), and the Minimal Description Length of the test data (step 250).

[0152] In the step (250) of determining the above OOD, the OOD determination unit (140) can determine the OOD by deriving a damage difficulty ranking using Wasserstein distance, mutual information, and minimum explanation length.

[0153] In addition, in the step (250) of determining the above OOD, the OOD determination unit (140) can derive a damage difficulty ranking based on changes in covariates and semantic changes for training data and test data.

[0154]

[0155] Accordingly, according to another embodiment of the present invention, by determining OOD through Wasserstein distance and mutual information of training data and test data measured according to the output of a hybrid generative model and the minimum explanation length of the test data, it is possible to effectively distinguish between attribution data and out-of-distribution data using a hybrid generative model, as well as effectively ensure the integrity of the image.

[0156]

[0157] Although various embodiments of the present invention have been presented and described in the above description, the present invention is not necessarily limited thereto, and those skilled in the art will readily understand that various substitutions, modifications, and changes are possible within the scope of the technical concept of the present invention.

[0158]

[0159] [Explanation of the symbol]

[0160] 110 : First input section

[0161] 120 : Second input section

[0162] 130 : Learning Model Section

[0163] 140 : OOD Discriminator

Claims

1. A first input unit for inputting training data; A second input unit for inputting test data; A learning model unit that inputs the above training data into a hybrid generative model for training, and inputs and outputs the above test data into the hybrid generative model; and An OOD determination unit that determines Out-Of-Distribution (OOD) through the Wasserstein Distance and Mutual Information of the training data and test data measured according to the output of the hybrid generative model, and the Minimal Description Length of the test data; A hybrid generative model device including 2. In Claim 1, The above training data is, Using In Distribution data, The above test data is, Includes Out-Of-Distribution data Hybrid generative model device.

3. In Claim 2, The above hybrid generative model is, Includes a Normalizing Flow model Hybrid generative model device.

4. In Claim 3, The above OOD discrimination unit is, Determining the above OOD by deriving a damage difficulty ranking using the above Wasserstein distance, mutual information, and minimum explanation length Hybrid generative model device.

5. In Claim 4, The above OOD discrimination unit is, Deriving the above injury difficulty ranking based on changes in covariates and semantic changes for the above training data and test data Hybrid generative model device.

6. A step of inputting training data through the first input unit; A step of inputting the above training data into a hybrid generative model in the learning model section to train it; A step of inputting test data through the second input unit; A step of inputting and outputting the test data to the hybrid generative model in the above learning model unit; and A step of determining Out-Of-Distribution (OOD) through the Wasserstein Distance and Mutual Information of the training data and test data measured according to the output of the hybrid generative model in the OOD determination unit, and the Minimal Description Length of the test data; A method for determining out-of-distribution using a hybrid generative model device including 7. In Claim 6, The above training data is, Using In Distribution data, The above test data is, Using Out-Of-Distribution data Method for determining out-of-distribution using a hybrid generative model device.

8. In Claim 7, The above hybrid generative model is, Includes a Normalizing Flow model Method for determining out-of-distribution using a hybrid generative model device.

9. In Claim 8, The step of determining the above OOD is, The above OOD discrimination unit derives a damage difficulty ranking using the above Wasserstein distance, mutual information, and minimum explanation length to identify the above OOD. Method for determining out-of-distribution using a hybrid generative model device.

10. In Claim 9, The step of determining the above OOD is, The above OOD discrimination unit derives the above injury difficulty ranking based on changes in covariates and semantic changes for the above training data and test data. Method for determining out-of-distribution using a hybrid generative model device.