Method for detecting out-of-domain and computing device for executing the same

KR103026208B1Active Publication Date: 2026-09-29AJOU UNIV IND ACADEMIC COOP FOUND
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
KR1020230014067
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-02-02
Publication Date
2026-09-29
Estimated Expiration
2043-02-02

Smart Images

  • Figure R1020230014067_ABST
    Figure R1020230014067_ABST
Patent Text Reader

Abstract

A method for detecting out-of-distribution samples and a computing device for performing the same are disclosed. A method for detecting out-of-distribution samples according to one disclosed embodiment is performed in a computing device having one or more processors and a memory for storing one or more programs executed by one or more processors, and comprises a first learning step and a second learning step based on an artificial neural network, wherein the first learning step comprises the step of inputting learning data belonging to a normal domain into a first artificial neural network and training the first artificial neural network to generate an embedding vector from the input learning data, and the step of inputting the embedding vector into a second artificial neural network and a third artificial neural network, respectively, and training the second artificial neural network and the third artificial neural network to assign pseudo-labels to the embedding vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] An embodiment of the present invention relates to an out-of-distribution sample detection technique. Background Technology

[0003] Out-of-domain (OOD) detection is the task of distinguishing whether a given sample, without labels, belongs within a normal domain. In other words, OOD detection is the task of classifying a given sample as an anomalous domain if it falls into a domain other than the domain defined as normal (i.e., the normal domain). However, in reality, it is difficult to obtain anomalous domains, so machine learning models have been trained using only data belonging to the normal domain (i.e., normal data). Prior art literature

[0005] Korean Registered Patent Publication No. 10-2420994 (July 18, 2022) The problem to be solved

[0006] The disclosed embodiment is intended to provide an out-of-distribution sample detection method capable of accurately detecting out-of-distribution samples and a computing device for performing the same. means of solving the problem

[0008] A method for detecting out-of-distribution samples according to a disclosed embodiment is a method performed in a computing device having one or more processors and a memory storing one or more programs executed by said one or more processors, comprising a first learning step and a second learning step based on an artificial neural network, wherein the first learning step comprises: a step of inputting learning data belonging to a normal domain into a first artificial neural network and training said artificial neural network to generate an embedding vector from the input learning data; and a step of inputting said embedding vector into a second artificial neural network and a third artificial neural network, respectively, and training said artificial neural network and said third artificial neural network to assign a pseudo-label to said embedding vector.

[0009] A computing device according to one disclosed embodiment comprises one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and wherein the one or more programs include instructions for performing a first learning step and a second learning step based on an artificial neural network, and the instructions for performing the first learning step include: instructions for inputting learning data belonging to a normal domain into a first artificial neural network and training the first artificial neural network to generate an embedding vector from the input learning data; and instructions for inputting the embedding vector into a second artificial neural network and a third artificial neural network, respectively, and training the second artificial neural network and the third artificial neural network to assign a pseudo-label to the embedding vector. Effects of the invention

[0011] According to the disclosed embodiment, by performing unsupervised learning to assign pseudo-labels using normal data to an artificial neural network of an out-of-distribution sample detection device, and then performing supervised learning to predict pseudo-labels for the normal data to which pseudo-labels have been assigned, it is possible to train the artificial neural network using only normal data and to detect out-of-distribution samples more accurately. Brief explanation of the drawing

[0013] FIG. 1 is a drawing showing an out-of-distribution sample detection device for a first learning step according to an embodiment of the present invention. FIG. 2 is a drawing showing an out-of-distribution sample detection device for a second learning step according to an embodiment of the present invention. FIG. 3 is a diagram schematically showing a change in the embedding space in an out-of-distribution sample detection device according to an embodiment of the present invention. FIG. 4 is a schematic diagram illustrating an out-of-distribution sample detection device for a test step in one embodiment of the present invention. FIG. 5 is a flowchart illustrating a method for detecting out-of-distribution samples according to an embodiment of the present invention. FIG. 6 is a block diagram illustrating a computing environment including a computing device suitable for use in exemplary embodiments. Specific details for implementing the invention

[0014] Hereinafter, specific embodiments of the present invention will be described with reference to the drawings. The following detailed description is provided to facilitate a comprehensive understanding of the methods, apparatus, and / or systems described herein. However, this is merely illustrative and the present invention is not limited thereto.

[0015] In describing the embodiments of the present invention, detailed descriptions of known technologies related to the present invention are omitted if it is determined that such detailed descriptions may unnecessarily obscure the essence of the present invention. Furthermore, the terms described below are defined in consideration of their functions within the present invention, and these may vary depending on the intentions or practices of the user or operator. Therefore, such definitions should be based on the content throughout this specification. Terms used in the detailed description are intended merely to describe the embodiments of the present invention and should not be limiting in any way. Unless explicitly stated otherwise, expressions in the singular form include the meaning of the plural form. In this description, expressions such as "include" or "comprise" are intended to refer to certain characteristics, numbers, steps, actions, elements, parts thereof, or combinations thereof, and should not be interpreted to exclude the existence or possibility of one or more other characteristics, numbers, steps, actions, elements, parts thereof, or combinations thereof other than those described.

[0016] Additionally, terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. These terms may be used for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component.

[0018] FIG. 1 is a diagram showing an out-of-distribution sample detection device for a first learning step according to an embodiment of the present invention, FIG. 2 is a diagram showing an out-of-distribution sample detection device for a second learning step according to an embodiment of the present invention, and FIG. 3 is a diagram schematically showing a change in the embedding space in an out-of-distribution sample detection device according to an embodiment of the present invention.

[0019] Referring to FIGS. 1 to 3, the out-of-distribution sample detection device (100) may include an embedding module (102) and a pseudo-label module (104). The out-of-distribution sample detection device (100) may perform machine learning-based out-of-distribution sample detection tasks.

[0020] The learning stage of the artificial neural networks constituting the out-of-distribution sample detection device (100) can be performed in two stages. The first learning stage is for clustering the learning data and assigning pseudo-labels, and is a stage of learning the artificial neural networks of the embedding module (102) and the pseudo-label module (104). Since the first learning stage is performed without assigning labels to the learning data, unsupervised learning is performed.

[0021] In the second learning step, training data with pseudo-labels is input to the embedding module (102) that has completed the learning in the first learning step, so that the training data can be classified. Since the second learning step is performed with the training data having pseudo-labels, supervised learning is carried out.

[0022] Hereinafter, the first learning step for the embedding module (102) and the pseudo-label module (104) will be described with reference to FIGS. 1 and FIGS. 3.

[0023] The embedding module (102) receives training data and can generate an embedding vector from the training data. In one embodiment, the training data may be text data. The set of training data is It can be expressed as x i is the i-th training data, and M can represent the total number of training data. Here, the training data can be data belonging to the normal domain. Also, the training data is not assigned labels related to categories.

[0024] The embedding module (102) may include a first artificial neural network (102a). The first artificial neural network (102a) may be trained to generate an embedding vector from each input training data. The first artificial neural network (102a) may be an artificial neural network that tokenizes training data, which is text data, and then converts the token sequence into an embedding vector. In one embodiment, the first artificial neural network (102a) may be BERT (Bidirectional Encoder Representations from Transformer), but is not limited thereto.

[0025] The first artificial neural network (102a) is input training data (x) as in Equation 1. i embedding vector(e) from ) i Can generate ).

[0026] (Mathematical Formula 1)

[0027]

[0028] : Artificial neural network constituting the first artificial neural network (102a)

[0029] The pseudo-label module (104) may be configured to receive each embedding vector output from the embedding module (102) and to assign a pseudo-label to each embedding vector. Here, since the training data is not labeled, the pseudo-label module (104) assigns a pseudo-label to each embedding vector (i.e., training data) through unsupervised learning.

[0030] The pseudo-label module (104) can cluster embedding vectors and assign a pseudo-label to each embedding vector according to the cluster to which the embedding vector belongs. The pseudo-label module (104) may include a second artificial neural network (104a) and a third artificial neural network (104b).

[0031] The second artificial neural network (104a) can make embedding vectors located close to each other in the space where the embedding vectors are located (which may be referred to as the embedding space) become closer to each other. That is, the second artificial neural network (104a) can be trained so that embedding vectors located close to each other in the embedding space belong to the same cluster.

[0032] Specifically, the second artificial neural network (104a) can calculate a soft probability that each embedding vector belongs to cluster k among K pre-set clusters. The number of clusters K can be appropriately determined. The soft probability (q) that the i-th embedding vector belongs to cluster k ik ) can be calculated by the following mathematical formula 2.

[0033] (Mathematical Formula 2)

[0034]

[0035] K: Number of clusters

[0036] e i : i-th embedding vector

[0037] μ k : Center of the k-th cluster

[0038] α: Pre-set constant

[0039] The second artificial neural network (104a) is the soft probability (q) that the i-th embedding vector belongs to cluster k. ik An auxiliary target distribution can be calculated to refine ). Here, the auxiliary target distribution is the soft probability (q) that the i-th embedding vector belongs to cluster k. ik It may be a distribution that causes ) to increase. That is, the auxiliary target distribution sharpens the distribution of embedding vectors in the embedding space, thereby the soft probability (q ik It can be a distribution that helps increase the soft probability (q) that the i-th embedding vector belongs to cluster k. ik Auxiliary target distribution (p) to refine ) ik ) can be calculated by the following mathematical formula 3.

[0040] (Mathematical Formula 3)

[0041]

[0042] f k : Soft probability (q ik exponents for normalizing )

[0043] Here, It can be represented as follows. M can be the total number of embedding vectors.

[0044] The second artificial neural network (104a) can be trained such that the soft probability of each embedding vector located in the embedding space belonging to a specific cluster approaches an auxiliary target distribution that improves the said soft probability. In this case, the second artificial neural network (104a) can perform clustering by making embedding vectors located close to each other in the embedding space become closer to each other. The loss function (L) of the second artificial neural network (104a) cluster ) can be expressed by the following mathematical formula 4.

[0045] (Mathematical Formula 4)

[0046]

[0047] M: Total number of embedding vectors

[0048] l i C : Soft probability that the i-th embedding vector belongs to cluster k (q ik ) and soft probability(q ik The auxiliary target distribution (p) that improves ) ik A function that calculates the difference between ).

[0049] In one embodiment, the function (l i C ) can use KL divergence (Kullback-Leibler divergence) and can be expressed by the following mathematical formula 5.

[0050] (Mathematical Formula 5)

[0051]

[0052] The third artificial neural network (104b) can be configured to perform clustering while separating embedding vectors that are gathered in the embedding space. That is, the third artificial neural network (104b) can be trained so that embedding vectors that are gathered in one place or overlapping in the embedding space are separated from each other and clustered. At this time, the third artificial neural network (104b) can be trained so that embedding vectors having the same or similar attributes are densely clustered together.

[0053] When clustering embedding vectors located close to each other in the embedding space using the second artificial neural network (104a) to make them closer to each other, the clustering efficiency may be reduced if the embedding vectors are gathered in one place or overlap. Accordingly, the third artificial neural network (104b) can be used to separate and cluster embedding vectors that are gathered in one place or overlap in the embedding space.

[0054] Specifically, the out-of-distribution sample detection device (100) inputs the i-th training data twice into the first artificial neural network (102a) to obtain a pair of embedding vectors (e i o , e i 1 ) can be obtained. That is, the i-th training data is input into the first artificial neural network (102a) to obtain an embedding vector (e i o ) obtain ) and input the i-th training data back into the first artificial neural network (102a) to obtain an embedding vector (e i 1 You can obtain ).

[0055] Here, a pair of embedding vectors (e i o , e i 1 ) is obtained by inputting the same i-th training data into the first artificial neural network (102a) respectively, but if a different dropout mask is used in the first artificial neural network (102a), a pair of embedding vectors (e) of different values ​​are obtained. i o , e i 1 You will be able to obtain ).

[0056] The third artificial neural network (104b) is a pair of embedding vectors (e) of the i-th training data. i o , e i 1) are made to get closer to each other, and a pair of embedding vectors (e) of the i-th training data i o , e i 1 The embedding vectors of ) and other training data (i.e., the remaining training data excluding the i-th training data) can be trained to diverge from each other. The loss function (l) of the third artificial neural network (104b) for such i-th training data. i CL ) can be expressed as shown in the following mathematical formula 6.

[0057] (Mathematical Formula 6)

[0058]

[0059] M: Total number of training data

[0060] : Embedding vector(e i o Output value of the third artificial neural network (104b) for )

[0061] : Embedding vector(e i 1 Output value of the third artificial neural network (104b) for )

[0062] : Embedding vector(e j Output value of the third artificial neural network (104b) for )

[0063] j: The j-th training data among the training data excluding the i-th training data

[0064] sim( , ) : and liver similarity value

[0065] sim( , ) : and liver similarity value

[0066] In addition, the total loss function (L) of the third artificial neural network (104b) CL) can be expressed as shown in the following mathematical formula 7.

[0067] (Mathematical Formula 7)

[0068]

[0069] Here, the final loss function of the first training step (L stage1 ) can be expressed by the following mathematical formula 8.

[0070] (Mathematical Formula 8)

[0071]

[0072] λ : weight

[0073] The final loss function of the first training step (L stage1 The artificial neural networks of the embedding module (102) and the pseudo-label module (104) can be trained using ). As a result, each input training data (x i For ), pseudo-label(y i pseudo ) can be granted.

[0074] Hereinafter, the second learning step will be described with reference to FIGS. 2 and FIGS. 3. The second learning step can be performed on an embedding module (102) that has completed the learning of the first learning step. In the state where only the first learning step has been performed, as shown in FIG. 3, among the OOD samples, FAR OOD samples can be effectively detected as they are separated from the cluster of normal samples in the embedding space, but NEAR OOD samples may not be detected well as they are located at the boundary of the cluster of normal samples. Here, FAR OOD samples may be OOD samples that have attributes completely different from normal samples (IN DOMAIN). NEAR OOD samples may be OOD samples that are not normal samples but have attributes that are partially similar to normal samples. In the second learning step, supervised learning using pseudo-labels can be performed to further separate each cluster in the embedding space.

[0075] The embedding module (102) is a pseudo-label (y i pseudo Training data (x) assigned i ) can be received as input. The embedding module (102) receives training data (x i Receive ) as input, and the received training data (x i ) pseudo-label(y i pseudo ) can predict. That is, in the second learning step, the first artificial neural network (102a) can predict the learning data (x i Receive ) as input, and the received training data (x i ) pseudo-label(y i pseudo It can be trained to predict ). At this time, the first artificial neural network (102a) predicts the training data (x i ) pseudo-label(y i pseudo It can be trained such that the difference between ) and the actual pseudo-labels (i.e., correct values) of the corresponding training data is minimized. In the second training step, the loss function (L) of the first artificial neural network (102a) stage2 ) can be expressed by the following mathematical formula 9.

[0076] (Mathematical Formula 9)

[0077]

[0078] M: Total number of training data

[0079] p i : Prediction probability distribution for the pseudo-label of the i-th training data

[0080] y i pseudo : Pseudo-labels assigned to the i-th training data

[0081] After going through the second learning step, as shown in Figure 3, separation between each cluster is achieved, and it can be confirmed that NEAR OOD can also be detected more accurately.

[0082] According to the disclosed embodiment, unsupervised learning is performed on an artificial neural network of an out-of-distribution sample detection device (100) to assign pseudo-labels using normal data, and then supervised learning is performed to predict pseudo-labels on the normal data to which pseudo-labels have been assigned. This allows the artificial neural network to be trained using only normal data and enables more accurate detection of out-of-distribution samples.

[0083] In this specification, the term "module" may refer to a functional and structural combination of hardware for carrying out the technical concept of the present invention and software for driving said hardware. For example, the "module" may refer to a logical unit of a specific code and a hardware resource for executing said code, and does not necessarily refer to physically connected code or a single type of hardware.

[0085] When the learning of the out-of-distribution sample detection device (100) is completed (the first learning step and the second learning step are completed), the out-of-distribution sample detection device (100) can detect out-of-distribution samples using the embedding module (102). That is, in the test phase, out-of-distribution sample detection can be performed on the input sample using only the embedding module (102) that has completed the learning.

[0086] FIG. 4 is a schematic diagram illustrating an out-of-distribution sample detection device for a test step (inference step) in one embodiment of the present invention. Referring to FIG. 4, the out-of-distribution sample detection device (100) can detect whether an input sample is an out-of-distribution sample by using a distance-based scoring function. The out-of-distribution sample detection device (100) inputs the sample into an embedding module (102) to generate an embedding vector for the sample, and can calculate a score value for the sample based on the embedding vector and a preset scoring function. The out-of-distribution sample detection device (100) can detect the sample as a normal sample if the calculated score value is greater than or equal to a preset threshold, and detect the sample as an out-of-distribution sample if the calculated score value is less than the preset threshold. The out-of-distribution sample detection device (100) can calculate a score value (s) for the sample through the following mathematical formula 10.

[0087] (Mathematical Formula 10)

[0088]

[0089] h: Embedding vector for the sample

[0090] μ k : Center of the k-th cluster

[0091] T : Transpose matrix

[0092] : Covariance matrix of each cluster k

[0094] FIG. 5 is a flowchart illustrating a method for detecting out-of-distribution samples according to an embodiment of the present invention. In the illustrated flowchart, the method is described by dividing it into a plurality of steps, but at least some of the steps may be performed in a different order, combined with other steps, omitted, divided into detailed steps, or performed with one or more steps not illustrated added.

[0095] Referring to FIG. 5, the out-of-distribution sample detection device (100) can perform a first learning step (first learning step) (S 101). In the first learning step, the artificial neural networks of the embedding module (102) and the pseudo-label module (104) can be trained to generate embedding vectors from input learning data and to cluster the embedding vectors to assign pseudo-labels.

[0096] Next, the out-of-distribution sample detection device (100) can perform a second learning step (second learning step) (S 103). In the second learning step, training data with pseudo-labels is input, and the artificial neural network of the embedding module (102) can be trained to predict pseudo-labels from the input training data.

[0097] Next, the out-of-distribution sample detection device (100) can generate an embedding vector from an input sample using a learned embedding module (102) and detect whether the sample is an out-of-distribution sample based on this (S 105).

[0099] FIG. 6 is a block diagram illustrating a computing environment (10) including a computing device suitable for use in exemplary embodiments. In the illustrated embodiments, each component may have different functions and capabilities in addition to those described below, and may include additional components in addition to those described below.

[0100] The illustrated computing environment (10) includes a computing device (12). In one embodiment, the computing device (12) may be an out-of-distribution sample detection device (100).

[0101] The computing device (12) includes at least one processor (14), a computer-readable storage medium (16), and a communication bus (18). The processor (14) can cause the computing device (12) to operate according to the exemplary embodiment described above. For example, the processor (14) can execute one or more programs stored in the computer-readable storage medium (16). The one or more programs may include one or more computer-executable instructions, and the computer-executable instructions may be configured to cause the computing device (12) to perform operations according to the exemplary embodiment when executed by the processor (14).

[0102] A computer-readable storage medium (16) is configured to store computer-executable instructions or program code, program data and / or other suitable forms of information. A program (20) stored in the computer-readable storage medium (16) includes a set of instructions executable by a processor (14). In one embodiment, the computer-readable storage medium (16) may be memory (volatile memory such as random access memory, non-volatile memory, or a suitable combination thereof), one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other forms of storage media that are accessed by a computing device (12) and capable of storing desired information, or a suitable combination thereof.

[0103] The communication bus (18) interconnects various other components of the computing device (12), including the processor (14) and the computer-readable storage medium (16).

[0104] The computing device (12) may also include one or more input / output interfaces (22) and one or more network communication interfaces (26) that provide interfaces for one or more input / output devices (24). The input / output interfaces (22) and the network communication interfaces (26) are connected to a communication bus (18). The input / output devices (24) may be connected to other components of the computing device (12) through the input / output interfaces (22). An exemplary input / output device (24) may include an input device such as a pointing device (such as a mouse or trackpad), a keyboard, a touch input device (such as a touchpad or touchscreen), a voice or sound input device, various types of sensor devices and / or imaging devices, and / or an output device such as a display device, a printer, a speaker and / or a network card. An exemplary input / output device (24) may be included inside the computing device (12) as a component constituting the computing device (12), or it may be connected to the computing device (12) as a separate device distinct from the computing device (12).

[0106] Although representative embodiments of the present invention have been described in detail above, those skilled in the art will understand that various modifications can be made to the above-described embodiments without departing from the scope of the present invention. Therefore, the scope of the present invention should not be limited to the described embodiments, but should be defined by the claims set forth below as well as equivalents thereof. Explanation of the symbols

[0108] 10: Computing Environment 12: Computing device 14 : Processor 16: Computer-readable storage media 18: Communication bus 20 : Program 22 : Input / Output Interface 24 : Input / Output Devices 26: Network communication interface 100: Out-of-distribution sample detection device 102 : Embedding Module 102a: The first artificial neural network 104: Pseudo-label module 104a: Second Artificial Neural Network 104b: The Third Artificial Neural Network

Claims

Claim 1 A method performed in a computing device having one or more processors and a memory for storing one or more programs executed by said one or more processors, comprising an artificial neural network-based first learning step and a second learning step, wherein the first learning step comprises the step of inputting learning data belonging to a normal domain into a first artificial neural network and training the first artificial neural network to generate an embedding vector from the input learning data; The method includes the step of inputting the embedding vectors into a second artificial neural network and a third artificial neural network, respectively, and training the second artificial neural network and the third artificial neural network to assign pseudo-labels to the embedding vectors, wherein the step of training the second artificial neural network trains the second artificial neural network such that embedding vectors located close to each other in the embedding space where the embedding vectors are located belong to the same cluster, and the step of training the third artificial neural network trains the third artificial neural network such that embedding vectors that are gathered or overlapping in the embedding space are separated and clustered, wherein a pair of embedding vectors of the i-th training data are made closer to each other, and the third artificial neural network is trained such that a pair of embedding vectors of the i-th training data and the embedding vectors of the remaining training data excluding the i-th training data are made farther apart from each other, and the step of training the third artificial neural network trains the third artificial neural network such that the i-th training data is input twice into the first artificial neural network and a pair of embeddings of the i-th training data from the first artificial neural network A method for detecting out-of-distribution samples, wherein a vector is obtained, and a pair of embedding vectors of different values ​​are obtained using different dropout masks in the first artificial neural network. Claim 2 delete Claim 3 A method for detecting out-of-distribution samples according to claim 1, wherein the step of training the second artificial neural network comprises: a step of calculating a soft probability that the embedding vector belongs to a specific cluster among a preset number of clusters; a step of calculating an auxiliary target distribution to improve the soft probability that the embedding vector belongs to the specific cluster; and a step of training the second artificial neural network such that for each embedding vector located in the embedding space, the soft probability becomes closer to the auxiliary target distribution. Claim 4 In claim 3, the loss function (L) of the second artificial neural network cluster ) is an out-of-distribution sample detection method represented by the following mathematical formula. (Mathematical formula) M: Total number of embedding vectors i C : Soft probability that the i-th embedding vector belongs to cluster k (q ik ) and soft probability(q ik The auxiliary target distribution (p) that improves ) ik A function that calculates the difference between ). Claim 5 delete Claim 6 delete Claim 7 delete Claim 8 In claim 4, the loss function (L) of the third artificial neural network CL ) is an out-of-distribution sample detection method represented by the following mathematical formula. (Mathematical formula) M: Total number of embedding vectors i CL : A function that causes a pair of embedding vectors of the i-th training data to be close to each other, and causes the pair of embedding vectors of the i-th training data and the embedding vectors of the remaining training data (excluding the i-th training data) to be far apart from each other. Claim 9 In claim 8, the final loss function (L) of the first learning step stage1 ) is an out-of-distribution sample detection method represented by the following mathematical formula. (Mathematical formula) λ : weight Claim 10 A method for detecting out-of-distribution samples according to claim 1, wherein the second learning step inputs pseudo-labeled learning data into the first artificial neural network and trains the first artificial neural network to predict the pseudo-label of the input learning data. Claim 11 In claim 10, the loss function (L) of the second learning step stage2 ) is an out-of-distribution sample detection method represented by the following mathematical formula. (Mathematical formula) M: Total number of training data points p i : Predicted probability distribution y for the pseudo-label of the i-th training data i pseudo : Pseudo-labels assigned to the i-th training data Claim 12 The method for detecting out-of-distribution samples according to claim 10 further comprises the steps of: inputting a sample to be tested into the first artificial neural network and calculating a score value for the sample based on an embedding vector for the sample generated by the first artificial neural network and a preset scoring function; and detecting whether the sample is an out-of-distribution sample by comparing the score value with a preset threshold value. Claim 13 One or more processors; memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing a first learning step and a second learning step based on an artificial neural network, wherein the instructions for performing the first learning step include instructions for inputting learning data belonging to a normal domain into a first artificial neural network and training the first artificial neural network to generate an embedding vector from the input learning data; and input the embedding vectors into the second artificial neural network and the third artificial neural network, respectively, and include a command to train the second artificial neural network and the third artificial neural network to assign pseudo-labels to the embedding vectors; the command to train the second artificial neural network includes a command to train the second artificial neural network such that embedding vectors located close to each other in the embedding space where the embedding vectors are located belong to the same cluster; the command to train the third artificial neural network includes a command to train the third artificial neural network such that embedding vectors gathered or overlapping in the embedding space are separated and clustered; wherein the third artificial neural network is trained such that a pair of embedding vectors of the i-th training data become close to each other, and the pair of embedding vectors of the i-th training data and the embedding vectors of the remaining training data excluding the i-th training data become far apart from each other; and the command to train the third artificial neural network includes inputting the i-th training data into the first artificial neural network twice so that the first artificial A computing device that obtains a pair of embedding vectors of the i-th training data from a neural network, wherein the first artificial neural network uses different dropout masks to obtain a pair of embedding vectors of different values. Claim 14 delete Claim 15 A computing device according to claim 13, wherein a command for training the second artificial neural network comprises: a command for calculating a soft probability that the embedding vector belongs to a specific cluster among a preset number of clusters; a command for calculating an auxiliary target distribution to improve the soft probability that the embedding vector belongs to the specific cluster; and a command for training the second artificial neural network such that for each embedding vector located in the embedding space, the soft probability becomes closer to the auxiliary target distribution. Claim 16 delete Claim 17 delete Claim 18 A computing device according to claim 13, wherein the command for performing the second learning step comprises a command for inputting pseudo-labeled learning data into the first artificial neural network and training the first artificial neural network to predict the pseudo-label of the input learning data. Claim 19 A computer program stored on a non-transitory computer-readable storage medium, wherein the computer program comprises one or more instructions, and when the instructions are executed by a computing device having one or more processors, the computing device causes the computing device to perform a first learning step and a second learning step based on an artificial neural network, wherein the first learning step comprises the step of inputting learning data belonging to a normal domain into a first artificial neural network and training the first artificial neural network to generate an embedding vector from the input learning data; The method includes the step of inputting the embedding vectors into a second artificial neural network and a third artificial neural network, respectively, and training the second artificial neural network and the third artificial neural network to assign pseudo-labels to the embedding vectors, wherein the step of training the second artificial neural network trains the second artificial neural network such that embedding vectors located close to each other in the embedding space where the embedding vectors are located belong to the same cluster, and the step of training the third artificial neural network trains the third artificial neural network such that embedding vectors that are gathered or overlapping in the embedding space are separated and clustered, wherein a pair of embedding vectors of the i-th training data are made closer to each other, and the third artificial neural network is trained such that a pair of embedding vectors of the i-th training data and the embedding vectors of the remaining training data excluding the i-th training data are made farther apart from each other, and the step of training the third artificial neural network trains the third artificial neural network such that the i-th training data is input twice into the first artificial neural network and a pair of embeddings of the i-th training data from the first artificial neural network A computer program that acquires a vector, wherein a pair of embedding vectors of different values ​​are acquired using different dropout masks in the first artificial neural network.

Citation Information

Patent Citations

  • Techniques for retrieving document data

    KR102458457B1

  • Apparatus and method for detecting anomaly

    KR102363737B1