Method and system for operation of electronic device that performs mbti that optimizes distance between classes by using support embedding
Patent Information
- Application Number
- PCT/KR2025/009589
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2025-07-04
- Publication Date
- 2026-10-01
Smart Images

Figure KR2025009589_01102026_PF_FP_ABST
Abstract
Description
Method of operation and system of an electronic device performing MBTI that optimizes inter-class distances using support embeddings
[0001] The present disclosure relates to a method of operation of an electronic device, and more specifically, to a method of operation of an electronic device and a system that enables the generation of an image having clearer characteristics by utilizing a plurality of text embeddings optimized for different classes and optimizing an embedding for a specific class.
[0002] With the recent rapid advancement of AI-based generative models, technologies capable of learning new concepts using only a few images (few-shot learning) are garnering attention. Unlike traditional large-scale data learning methods, technologies that can effectively learn and apply new concepts from limited datasets are in demand across various industries.
[0003] In particular, Textual Inversion is a representative generative model learning method that meets this need and enables the learning of new concepts. Textual Inversion is utilized to learn new concepts in text-to-image and image-to-image conversion, and DA-Fusion is an augmentation model that utilizes this and can be used for data augmentation in few-shot image classification.
[0004] These methods have opened up the possibility of effectively learning new concepts without the vast datasets required by existing deep learning models, but they still have some limitations. For example, if the model is not optimized by considering the relationship with existing classes when learning new concepts, it may learn inaccurate concepts or undesirable mixing may occur. Additionally, if the generated embeddings become excessively similar to superclasses or similar but different classes, a problem may arise where the model fails to clearly distinguish the concept.
[0005] The present disclosure aims to provide a method and system for operating an electronic device that enables the generation of an image having clearer characteristics by utilizing multiple text embeddings optimized for different classes and optimizing the embedding for a specific class.
[0006] The purposes of the present disclosure are not limited to those mentioned above, and other purposes and advantages of the present disclosure not mentioned may be understood from the following description and will be more clearly understood from the embodiments of the present disclosure. Furthermore, it will be readily apparent that the purposes and advantages of the present disclosure can be realized by the means and combinations thereof set forth in the claims.
[0007] A method of operation of an electronic device according to one embodiment of the present disclosure comprises: acquiring a plurality of text embeddings optimized for each of different classes; acquiring a target embedding for a target class that does not match the different classes; identifying at least one of the plurality of text embeddings as a support embedding based on the cosine similarity of each of the plurality of text embeddings to the target embedding; and optimizing the target embedding for the target class by maximizing the distance between the target embedding and the support embedding based on a first image generated based on the target embedding and a second image generated based on the support embedding.
[0008] At this time, the step of obtaining a target embedding for the target class may involve obtaining text related to the target class, identifying a vector value in a latent space corresponding to the obtained text, and obtaining the target embedding corresponding to the identified vector value.
[0009] Meanwhile, the step of identifying at least one of the plurality of text embeddings as a support embedding may include identifying the rank of each of the plurality of text embeddings based on the cosine similarity of each of the plurality of text embeddings with respect to the target embedding, and setting the text embedding whose identified rank is greater than or equal to a threshold value as the support embedding.
[0010] Here, the method of operation of the electronic device may include: a step of acquiring a first noise that needs to be added to a reference noise at each time step during the process in which a diffusion model generates an image based on a text embedding; a step of acquiring a second noise predicted at each time step during the process in which the target embedding is input to the diffusion model and the diffusion model performs reverse diffusion on the reference noise to generate the first image based on the target embedding; and a step of acquiring a third noise predicted at each time step during the process in which the support embedding is input to the diffusion model and the diffusion model performs reverse diffusion on the reference noise to generate the second image based on the support embedding.
[0011] At this time, the method of operation of the electronic device may include the step of calculating a first loss according to the loss function of the diffusion model based on the first noise and the second noise for each time step, the step of calculating a second loss according to the loss function of the DMM (Diffusion-Based Distance Metric) based on the second noise and the third noise for each time step and a penalty set for each time step, and the step of calculating a third loss according to the time-adaptive normalized loss function based on the first noise and the second noise for each time step and a penalty set for each time step.
[0012] At this time, the step of optimizing the target embedding for the target class may include the step of calculating a final loss based on the first to third losses, and the step of adjusting the target embedding in a direction such that the calculated final loss decreases and the second loss increases.
[0013] At this time, the step of calculating the final loss may include the step of applying a first hyperparameter having a positive value to the second loss, the step of applying a second hyperparameter having a positive value to the third loss, and the step of calculating the final loss by subtracting the second loss to which the first hyperparameter has been applied from the value calculated by summing the first loss and the third loss to which the second hyperparameter has been applied.
[0014] Here, the step of adjusting the target embedding may adjust the target embedding such that the cosine similarity of the support embedding to the target embedding is reduced.
[0015] An electronic device according to one embodiment of the present disclosure includes a memory storing at least one instruction and a processor connected to said memory, wherein the processor acquires a plurality of text embeddings optimized for each of different classes and a target embedding for a target class, identifies at least one of said text embeddings as a support embedding based on the cosine similarity of each of said text embeddings to said target embedding, and optimizes said target embedding for said target class by maximizing the distance between said target embedding and said support embedding based on a first image generated based on said target embedding and a second image generated based on said support embedding.
[0016] A system according to one embodiment of the present disclosure includes a user terminal that transmits text related to a target class to an electronic device, and an electronic device that identifies a vector value in a latent space corresponding to the text received from the user terminal, obtains a target embedding corresponding to the identified vector value, obtains a plurality of text embeddings optimized for each of different classes, identifies at least one of the plurality of text embeddings as a support embedding based on the cosine similarity of each of the plurality of text embeddings to the target embedding, and optimizes the target embedding for the target class by maximizing the distance between the target embedding and the support embedding based on a first image generated based on the target embedding and a second image generated based on the support embedding.
[0017] Through the present disclosure, by utilizing multiple text embeddings optimized for different classes to maximize the distance between the support embedding and the target embedding, the embedding of the target class can be optimized, thereby enabling more sophisticated class representation and clear image generation.
[0018] In addition, through the present disclosure, by performing optimization by considering both DMM loss and time-adaptive normalization loss together, more precise embedding adjustment is possible and the quality of the generated image can be improved.
[0019] In addition, the MBTI technique according to the present disclosure can generate more detailed images than existing techniques and can contribute to improving image classification performance by utilizing it as a data augmentation technique.
[0020] FIG. 1 is a flowchart for explaining the operation of an electronic device according to one embodiment of the present disclosure.
[0021] FIG. 2 is a diagram illustrating the operation of an electronic device optimizing text embeddings according to one embodiment of the present disclosure.
[0022] FIG. 3 is a diagram illustrating the operation performed by an electronic device to train a diffusion model according to one embodiment of the present disclosure.
[0023] FIG. 4 is a drawing for explaining an image generated based on MBTI and an original image performed by an electronic device according to one embodiment of the present disclosure.
[0024] FIG. 5 is a graph for comparing performance when some of the elements constituting MBTI according to one embodiment of the present disclosure are removed.
[0025] FIG. 6 is a diagram illustrating the operation of an electronic device training a diffusion model based on an image of a target class according to an additional embodiment of the present disclosure.
[0026] FIG. 7 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure.
[0027] FIG. 8 is a drawing illustrating the configuration of a system according to one embodiment of the present disclosure.
[0028] The embodiments described herein are subject to various modifications and may have various forms; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope of specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In relation to the description of the drawings, similar reference numerals may be used for similar components.
[0029] In describing the present disclosure, if it is determined that a detailed description of related known functions or configurations could unnecessarily obscure the essence of the present disclosure, such detailed description is omitted.
[0030] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concept of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to make the present disclosure more faithful and complete and to fully convey the technical concept of the present disclosure to those skilled in the art.
[0031] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit the scope of the rights. The singular expression includes the plural expression unless the context clearly indicates otherwise.
[0032] In the present disclosure, expressions such as “have,” “may have,” “include,” or “may include” indicate the presence of such features (e.g., numerical values, functions, actions, or components such as parts) and do not exclude the presence of additional features.
[0033] In the present disclosure, expressions such as “A or B,” “at least one of A or / and B,” or “one or more of A or / and B” may include all possible combinations of items listed together. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” may refer to cases including (1) at least one A, (2) at least one B, or (3) both at least one A and at least one B.
[0034] Expressions such as "first," "second," "first," or "second" used in this disclosure may modify various components regardless of order and / or importance, and are used only to distinguish one component from another and do not limit said components.
[0035] Where it is stated that a component (e.g., Component 1) is "(operatively or communicatively) coupled with / to" or "connected to" another component (e.g., Component 2), it should be understood that the component may be directly connected to the other component or connected through the other component (e.g., Component 3).
[0036] On the other hand, when it is stated that a certain component (e.g., a first component) is "directly connected" or "directly coupled" to another component (e.g., a second component), it may be understood that no other component (e.g., a third component) exists between the certain component and the other component.
[0037] As used in this disclosure, the expression “configured to” may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean only “specifically designed to” in hardware.
[0038] Instead, in some situations, the expression “device configured to do something” may mean that the device is “capable of doing something” together with other devices or components. For example, the phrase “processor configured (or set) to perform A, B, and C” may mean a dedicated processor for performing those operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or application processor) capable of performing those operations by executing one or more software programs stored in a memory device.
[0039] In the embodiments, a 'module' or 'part' performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Additionally, a plurality of 'modules' or a plurality of 'parts' may be integrated into at least one module and implemented by at least one processor, except for the 'module' or 'part' that needs to be implemented in specific hardware.
[0040] Meanwhile, various elements and areas in the drawings are depicted schematically. Accordingly, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0041] Hereinafter, embodiments according to the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement them.
[0042] FIG. 1 is a flowchart for explaining the operation of an electronic device according to one embodiment of the present disclosure.
[0043] The electronic device (100) can perform training of a diffusion model that generates images based on embeddings.
[0044] In one embodiment, the electronic device (100) can perform learning of the diffusion model using the MBTI (Metric-Based Textual Inversion) technique.
[0045] The diffusion model can perform a forward diffusion process that adds Gaussian noise (: Equation 1) and a reverse diffusion process that removes noise through a Markov chain (: Equation 2).
[0046]
[0047]
[0048] In mathematical formula 1, represents the initial noise distribution, and refers to the target data sample. The denoising process is a parameter By learning, we enable the restoration of noise-added samples back to the original data. The objective function of the diffusion process (: Equation 3) is to minimize the difference in distribution between the forward and backward processes.
[0049]
[0050] DDPM (Ho, Jain, and Abbeel 2020) simplified this objective function and proposed a loss function as shown in Equation 4.
[0051]
[0052] In mathematical formula 4, means noise, and represents a denoising network. This process is performed in image space x and latent space z.
[0053] In addition, the diffusion model may correspond to a Latent Diffusion Model (LDM). An LDM consists of an autoencoder and a denoising U-Net. The autoencoder maps an image to a latent code z and may include various information such as class labels and segmentation masks. An LDM can add noise to the data z in the latent space at a specific time step t and then remove the noise using a denoising network. The loss function of this process is defined in the latent space as shown in Equation (5).
[0054]
[0055] Here, is a latent vector containing noise at time t, and c(y) represents condition information from various sources (e.g., text description, class label, segmentation mask, etc.). The decoder can convert these latent representations back into images.
[0056] Text inversion, which forms the basis of MBTI, is a method that generates contextually relevant images by utilizing a pre-trained LDM. Text inversion focuses on finding unique words s* in text sources y that match specific concepts, and to this end, it optimizes text embeddings w* that find s* using 3 to 5 images. The optimization of embeddings by text inversion is expressed as Equation 6.
[0057]
[0058] Here, w* enables various conditional image generation. For example, image generation may be possible based on text prompts such as “A photo of s*” or “a drawing of s*”. The w described below is a condition source It means.
[0059] The electronic device (100) can optimize text embeddings for a class by maximizing the distance of text embeddings for each different class through MBTI. Here, a class refers to a category, concept, type, etc. related to an image. For example, a class may include a dog, cat, car, aircraft, flower, person, etc., and may include a rose, sunflower, Siberian Husky, poodle, Persian, Siamese, passenger car, van, bus, helicopter, airplane, male, female, etc., but is not limited thereto.
[0060] MBTI introduces a mechanism to maximize inter-class distance in text embeddings while maintaining the integrity of class attributes in the generated images. Additionally, by optimizing text embeddings through the application of cross-class loss, it can improve class distinction and enhance the detail of the generated images.
[0061] The core innovation of MBTI is the utilization of support embeddings. This enables efficient similarity measurement and distance calculation during the back-end process of the diffusion model. Support embeddings are the k text embeddings from previously optimized classes that are most similar to the text embedding currently being optimized; they play a role in helping to progressively represent the target class more precisely at each noise removal step. MBTI enables precise and stable metric learning by applying time-adaptive penalties and regularization techniques when measuring distances between text embeddings.
[0062] Experimental results confirmed that MBTI generates images containing fine details more effectively than existing methods (e.g., DA-Fusion), thereby enabling the acquisition of images with enhanced clarity and specificity. Additionally, it was confirmed that overall image classification performance is significantly improved when MBTI is utilized as a data augmentation technique.
[0063] Referring to FIG. 1, the electronic device (100) can obtain a plurality of text embeddings optimized for each of the different classes (S110).
[0064] In one embodiment, the electronic device (100) may have a plurality of text embeddings optimized for each of different classes stored in advance by performing the operation to be described later to optimize the text embedding for at least one class for that class.
[0065] As an additional embodiment, the electronic device (100) may obtain multiple text embeddings optimized for each of different classes through user input.
[0066] The electronic device (100) can obtain a target embedding for a target class (S120). Here, the target class refers to a class that does not match the other classes described above, and may correspond to a class corresponding to an image that the user wants to additionally generate.
[0067] In one embodiment, the electronic device (100) can obtain text related to a target class through user input, identify a vector value in a latent space corresponding to the obtained text, and obtain a target embedding corresponding to the identified vector value.
[0068] The electronic device (100) can identify at least one of the multiple text embeddings as a supporting embedding based on the cosine similarity of each of the multiple text embeddings to the target embedding (S130).
[0069] In one embodiment, the electronic device (100) calculates the cosine similarity of each of a plurality of text embeddings for a target embedding, identifies the rank of each of the plurality of text embeddings according to the calculated cosine similarity, and can identify a text embedding whose identified rank is greater than or equal to a threshold value (: K to be described later) as a support embedding.
[0070] In this regard, MBTI is text embeddings between classes The objective function that maximizes the distance is defined as in Equation 7.
[0071]
[0072] In mathematical formula 7, is the text embedding of the i-th class, and represents the loss function that maximizes the distance to the support embedding. Here, the target embedding is It corresponds to.
[0073] SES (Support Embedding Selection) is among the optimized embeddings from the 1st to the i-1th class The process involves identifying K text embeddings with high cosine similarity with as support embeddings, and in each iteration, support embeddings can be selected based on the distance between embeddings.
[0074] That is, the electronic device (100) can identify a text embedding that is close to the target embedding as a support embedding by selecting a text embedding that has a high cosine similarity to each of the multiple text embeddings for the target embedding.
[0075] The electronic device (100) can optimize the target embedding for the target class by maximizing the distance between the target embedding and the support embedding based on a first image generated based on the target embedding and a second image generated based on the support embedding (S140).
[0076] In one embodiment, the electronic device (100) can acquire an image of a text embedding using a diffusion model.
[0077] Specifically, the electronic device (100) can input a target embedding into a diffusion model to obtain a first image of the target embedding from the diffusion model, and input a support embedding into a diffusion model to obtain a second image of the support embedding from the diffusion model.
[0078] For example, a diffusion model can convert an image into noise by performing forward diffusion, and generate an image by performing back diffusion on the noise.
[0079] Accordingly, the electronic device (100) can acquire noise added to the training image as the diffusion model performs forward diffusion on the training image associated with the target class, and can identify the training image converted into noise as reference noise as the forward diffusion process is completed.
[0080] Through this, the electronic device (100) can obtain a first noise that needs to be added to the reference noise at each time step during the process in which the diffusion model performs reverse diffusion on the reference noise to generate an image.
[0081] In one embodiment, the electronic device (100) can identify a first noise based on noise added to a training image as the diffusion model performs forward diffusion on the training image.
[0082] For example, the electronic device (100) can sort the noise added to the training image in the reverse order of the order in which the noise added to the training image is added to the training image as the diffusion model performs forward diffusion on the training image, and identify the first noise that needs to be added to the reference noise at each time step of the process of generating an image by the diffusion model performing reverse diffusion on the reference noise.
[0083] Additionally, the electronic device (100) can input a target embedding into a diffusion model, and in the process of generating a first image based on the target embedding by the diffusion model performing reverse diffusion on reference noise, it can obtain a second noise predicted at each time step.
[0084] Specifically, the electronic device (100) can obtain a second noise predicted at each time step through a denoising network, which is a function that predicts noise.
[0085] In addition, the electronic device (100) can input a support embedding into a diffusion model, and in the process of generating a first image based on the support embedding by having the diffusion model perform reverse diffusion on the reference noise, the third noise predicted at each time step can be obtained.
[0086] Specifically, the electronic device (100) can obtain a third noise predicted at each time step through a denoising network, which is a function that predicts noise.
[0087] Here, the electronic device (100) can obtain the loss function of the diffusion model (: Equation 5) for minimizing the difference in noise distribution between forward diffusion and reverse diffusion performed by the diffusion model.
[0088] Accordingly, the electronic device (100) can calculate a first loss by applying a first noise and a second noise at each time step to the loss function of the diffusion model.
[0089] In addition, the electronic device (100) can obtain the loss function of the Diffusion-Based Distance Metric (DMM) (: Equation 8), which is a loss function based on the distance between embeddings.
[0090] The DMM, a diffusion-based distance metric, operates by estimating additional noise by optimizing the text embedding w. The DMM uses the same noise as a starting point in the back diffusion to compare the denoising characteristics of the text embedding at each time step.
[0091] Through these characteristics, image structures can be generated progressively from a coarse form to a fine form (Li et al. 2023), and the coarse structure appears more similar than the fine structure. This phenomenon occurs regardless of differences in text embeddings.
[0092] The characteristics of this diffusion model are adjusted by applying a penalty through scaling. Distance metric of text embeddings is defined as in Equation 8, which corresponds to the cross-class loss described in Fig. 2.
[0093]
[0094] In mathematical formula 8, is an image sample of the i-th class, is the denoising network, t is a specific time step in the diffusion process, T is the maximum time step in the entire diffusion process, is the text embedding of the i-th class for training, vk is It refers to the k-th highest support embedding with cosine similarity.
[0095] vk and The distance between them is calculated as the average for k∈{1,2,...,K}, where K is a hyperparameter representing the number of supported embeddings and can be a preset value. Since it is defined with K fixed, for classes where i > K Apply.
[0096] Accordingly, the electronic device (100) can calculate a second loss by applying a second noise and a third noise per time step and a penalty set per time step to the DMM loss function. Here, the penalty set per time step is a time-adaptive penalty, and the value of the penalty may increase as time elapses.
[0097] Additionally, the electronic device (100) can obtain a time-adaptive normalization loss function (: Equation 9) to prevent overfitting of the diffusion model.
[0098] Depending on the time step t Regularization loss function for scaling and correction is defined as in mathematical formula 9.
[0099]
[0100] In mathematical equation 9, scaling is a value used to stabilize training and prevent image generation failure, and is the distance loss Since it varies significantly with the time step, as t approaches 0, the model can be forced to clearly generate finer class images.
[0101] Accordingly, the electronic device (100) can calculate a third loss by applying a second noise and a third noise per time step and a penalty set per time step to a time-adaptive normalized loss function.
[0102] At this time, the electronic device (100) can calculate the final loss based on the first to third losses.
[0103] Specifically, the electronic device (100) can apply a first hyperparameter having a positive value to a second loss and a second hyperparameter having a positive value to a third loss, and calculate a final loss by subtracting the second loss to which the first hyperparameter has been applied from the value calculated by summing the first loss and the third loss to which the second hyperparameter has been applied, and each of the first hyperparameter and the second hyperparameter may be a preset value.
[0104] MBTI uses the average distance between support embeddings defined in Equation 8 to learn fine-grained class text embeddings. It aims to maximize, and the overall learning goal of MBTI is defined as Equation 10.
[0105]
[0106] In mathematical formula 10, is the denoising loss of the potential diffusion model defined in mathematical equation (5), and is a positive hyperparameter for balancing the three loss items, and the first hyperparameter is , the second hyperparameter is It means.
[0107] Here, the electronic device (100) can adjust the target embedding in a direction where the final loss is reduced and the second loss is increased.
[0108] For example, the electronic device (100) can adjust the target embedding so that the cosine similarity of the supporting embedding to the target embedding is reduced.
[0109] The electronic device (100) can optimize the target embedding for the target class by iterating the above-described operation to adjust the target embedding so that it moves away from text embeddings that are close to the target embedding.
[0110] At this time, after the target embedding is optimized for the target class, the electronic device (100) can perform verification of the image classification model based on the target embedding.
[0111] To this end, the electronic device (100) may include at least one image classification model. The image classification model is a model that predicts the class of an image, and the image classification model outputs a probability distribution of several classes for an input image, while maintaining the sum of the total probabilities at a constant value (e.g., 1, 100), and selects the class with the highest probability as the final predicted class.
[0112] For example, the image classification model may be a machine learning-based model (e.g., SVM (Support Vector Machine), k-NN (k-Nearest Neighbors), Random Forest, etc.) that extracts feature information from an image based on algorithms such as SIFT (Scale-Invariant Feature Transform), HOG (Histogram of Oriented Gradients), and SURF (Speeded-Up Robust Features) and predicts the class of the image based on the extracted feature information, or it may be a CNN (Convolutional Neural Network)-based model or a Transformer-based model, but is not limited thereto. The image classification model may be a model stored in the memory (110) of the electronic device (100), or it may be a model stored on an external server accessible to the electronic device (100).
[0113] Here, the electronic device (100) may identify a machine learning-based image classification model as a first classification model, identify a CNN-based image classification model as a second classification model, and identify a Transformer-based image classification model as a third classification model. At this time, the first to third classification models may correspond to models that have been pre-trained for a plurality of classes including a target class.
[0114] In one embodiment, the electronic device (100) inputs a target image generated by a diffusion model based on a target embedding into a first to third classification model, and obtains a final prediction class and a probability for the final prediction class predicted by each of the first to third classification models.
[0115] At this time, the electronic device (100) can verify the classification performance of each of the first to third classification models based on the final predicted class and the probability for the final predicted class predicted by each of the first to third classification models.
[0116] For example, the electronic device (100) can add the probability of the final predicted class predicted by the classification model that predicted the final predicted class that matches the target class among the first to third classification models to the classification success history by matching it with the corresponding classification model, and add the probability of the final predicted class predicted by the classification model that predicted the final predicted class that does not match the target class to the classification failure history by matching it with the corresponding classification model.
[0117] Here, the electronic device (100) can identify the number of classification successes for each classification model based on the classification success history and calculate the average value of the probability matched with each of the first to third classification models for each classification model. Accordingly, the electronic device (100) can calculate a success coefficient for each classification model by applying the number of successes to the average value of the probability.
[0118] That is, the electronic device (100) can calculate the success coefficient of the first classification model based on the number of successful classifications of the first classification model and the average value of the probability matched with the first classification model, calculate the success coefficient of the second classification model based on the number of successful classifications of the second classification model and the average value of the probability matched with the second classification model, and calculate the success coefficient of the third classification model based on the number of successful classifications of the third classification model and the average value of the probability matched with the third classification model.
[0119] Additionally, the electronic device (100) can identify the number of classification failures for each classification model based on the classification failure history and calculate the average value of the probability matched with each of the first to third classification models for each classification model. Accordingly, the electronic device (100) can calculate a failure coefficient for each classification model by applying the number of failures to the average value of the probability.
[0120] That is, the electronic device (100) can calculate the failure coefficient of the first classification model based on the number of classification failures of the first classification model and the average value of the probability matched with the first classification model, calculate the failure coefficient of the second classification model based on the number of classification failures of the second classification model and the average value of the probability matched with the second classification model, and calculate the failure coefficient of the third classification model based on the number of classification failures of the third classification model and the average value of the probability matched with the third classification model.
[0121] At this time, the electronic device (100) calculates a value by subtracting the failure coefficient from the success coefficient for each classification model, sets the classification model with the lowest calculated value as the target model, and can retrain the target model based on retraining data including images generated by applying the MBTI technique.
[0122] Specifically, since there is a possibility that generalization performance may be degraded when training only with image data generated by applying the MBTI technique, the electronic device (100) can perform retraining of the target model based on the existing training data used to train the target model and the data generated by applying the MBTI technique.
[0123] To this end, the electronic device (100) can construct new training data including existing training data and data generated through the MBTI technique, and can perform sampling while maintaining the balance of the data. Accordingly, the electronic device (100) can perform retraining on the (pre-trained) target model based on the new training data.
[0124] After retraining is completed, the electronic device (100) can verify the performance of the retrained image classification model using test data corresponding to the existing training data and readjust the retrained image classification model as needed.
[0125] Through this, the electronic device (100) can continuously improve the classification performance of the image classification model for fine-grained classes and optimize it to maintain high performance even in few-shot training environments.
[0126] FIG. 2 is a diagram illustrating the operation of an electronic device optimizing text embeddings according to one embodiment of the present disclosure.
[0127] Referring to Figure 2, regarding the learning method for converting images into pseudo-words, the existing method according to (a) optimizes text embeddings by learning each class individually, whereas MBTI according to (b) can learn by considering the relationships between different classes in the text embedding space through cross-class loss.
[0128] The electronic device (100) can improve the distinction between classes by applying a cross-class loss (: third loss) to maximize the distance between text embeddings for each different class and optimizing each text embedding for each class.
[0129] FIG. 3 is a diagram illustrating the operation performed by an electronic device to train a diffusion model according to one embodiment of the present disclosure.
[0130] Referring to FIG. 3, the electronic device (100) can perform (a) Support Embedding Selection (SES) which repeatedly selects support embeddings close to the target embedding, and (b) optimize the target embedding through time-adaptive scaling or diffusion-based distance loss with applied penalty according to the Diffusion-Based Distance Metric (DMM).
[0131] FIG. 4 is a drawing for explaining an image generated based on MBTI and an original image performed by an electronic device according to one embodiment of the present disclosure.
[0132] The qualitative results of text-guided image generation were evaluated by comparing MBTI with Textual Inversion (Gal et al. 2022), and the quantitative results of detailed image classification were analyzed by comparing it with Real Guidance (He et al. 2023) and DA-Fusion (Trabucco et al. 2024) (Tables 1 to 4). Textual Inversion is a method that optimizes text embeddings by learning each class individually, and DA-Fusion is a variation of Textual Inversion that trains an image classification model using image augmentation techniques, employing a method of training by weighted summing the original image and the generated image.
[0133] Real Guidance and MBTI use the same image augmentation techniques as DA-Fusion, but they differ in how they combine image generation and classification. Additionally, existing image processing techniques (e.g., rotation, flipping, enlargement, color conversion, etc.) were utilized as standard methods.
[0134] Here, for detailed image classification, Stanford Cars (Krause et al. 2013), FGVC-Aircraft (Maji et al. 2013), NABirds (Van Horn et al. 2015), Flowers102 (Nilsback and Zisserman 2008), and custom datasets collected from the internet were utilized. Unless otherwise noted, 20 classes were randomly selected from each class, and experiments were conducted using 1, 2, 4, or 8 samples per class; for the Flowers102 dataset, the entire dataset was used for the experiments. Additionally, COCO (Lin et al. 2014) and PASCAL VOC (Everingham et al. 2010) were utilized as datasets for conventional image classification.
[0135] Referring to Figure 4, when comparing the original image (: Original) with the image generated based on MBTI and the image generated based on the existing learning method DA-Fusion regarding the feature parts of the original image identified by Grad-CAM, it can be seen that the image generated based on MBTI is similar to the original image. In other words, the image generated based on MBTI retains the features that the classification model considers important regions, whereas the existing method fails to accurately reproduce important regions.
[0136]
[0137] Table 1 is a few-shot classification performance evaluation table based on RESNET (Residual Network)-50 for a finely divided classification dataset. Referring to Table 1, MBTI showed higher performance than other methods across all datasets and number of examples. When using 8 examples in the Flowers102 dataset, MBTI recorded an accuracy of 0.865 (the degree to which the generated image matches the class), which is 10.5% higher than DA-Fusion.
[0138] Accordingly, MBTI can effectively learn fine differences between similar classes, and performance improves as the number of training samples increases; this trend is pronounced in the Flowers102 and NA-Bird datasets, indicating that MBTI can effectively utilize additional examples per class.
[0139]
[0140] Table 2 is a few-shot classification performance evaluation table based on RESNET-50 for two user-defined datasets (: Unseen Concept, BMW with New), each dataset using 10 classes, 5 samples per class as the test set, and 4 examples for training.
[0141] Referring to Table 2, it can be seen that MBTI recorded higher performance than existing methods across all datasets. Therefore, this implies that MBTI learns new concepts with less training data and has higher performance than existing methods in distinguishing differences between similar classes.
[0142]
[0143] Table 3 is a few-shot classification performance evaluation table for a general dataset. Referring to Table 3, an improvement in MBTI's classification performance was observed even on a general dataset.
[0144]
[0145] Table 4 is a few-shot classification performance table for the Flower102 dataset using a backbone based on DeiT (Data-Efficient Image Transformer) (Touvron et al. 2021), and is based on data from experiments conducted to evaluate the performance of MBTI across various neural network architectures. Referring to Table 4, it was confirmed that MBTI demonstrates excellent performance not only in Convolutional Neural Network (CNN)-based models but also in classification using Transformer models.
[0146] FIG. 5 is a graph for comparing performance when some of the elements constituting MBTI according to one embodiment of the present disclosure are removed.
[0147] Referring to Figure 5, the Y-axis of each of the graphs 5a to 5c represents the classification accuracy in the GVA-Aircraft dataset, graph 5a shows the classification accuracy according to the method used to select support embeddings, graph 5b shows the classification accuracy according to the number of support embeddings, and graph 5c shows the classification accuracy according to the type of loss used to calculate the final loss.
[0148] Graph 5a shows the classification accuracy for each case of selecting a support embedding using Random (selecting a support embedding randomly), MSE (selecting a text embedding with the lowest MSE score and identifying it as a support embedding), R-CS (selecting a text embedding with the lowest cosine similarity and identifying it as a support embedding), and Ours (selecting a text embedding with the highest cosine similarity and identifying it as a support embedding). Referring to Graph 5a, the accuracy for Random was 0.371, the accuracy for MSE was 0.356, the accuracy for R-CS was 0.366, and the accuracy for MBTI was approximately 0.38. This demonstrates that the strategy of selecting a text embedding with the highest cosine similarity to the target embedding of the target class as a support embedding is effective. Furthermore, since the classification accuracy varies depending on the method of selecting the support embedding, it can be confirmed that the method of selecting the support embedding affects image classification performance.
[0149] Graph 5b shows the classification accuracy for cases where the number of support embeddings (: K explained through Equation 8) is 1, 2, 5, and 10, respectively. Referring to Graph 5b, it can be seen that the classification accuracy increases as the number of support embeddings increases. However, there is a problem in that the computational cost also increases rapidly as the number of support embeddings increases, so the present disclosure uses fewer than 10 support embeddings to account for the computational cost.
[0150] Graph 5c shows the case where only the first loss is used to calculate the final loss with the number of support embeddings fixed at 1 (: w / o & , in the case of using the first loss and the second loss(: w / o ), in the case of using the first and third losses(: w / o The classification accuracy for each case using all 1 to 3 losses (: Ours) is shown, and when all 1 to 3 losses were used, an accuracy of 0.376 was recorded, showing a 1% improvement in performance compared to when only the 1 loss was used. This suggests that using each loss function together is more effective than using each loss function alone.
[0151] FIG. 6 is a diagram illustrating the operation of an electronic device training a diffusion model based on an image of a target class according to an additional embodiment of the present disclosure.
[0152] Referring to FIG. 6, the electronic device (100) can acquire at least one image of a target class as a training image and generate a target embedding for the target class through the training image.
[0153] At this time, the electronic device (100) can optimize the target embedding for the target class using MBTI. Accordingly, the electronic device (100) can input the target embedding into a diffusion model to obtain an image for the target class.
[0154] FIG. 7 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure.
[0155] Referring to FIG. 7, the electronic device (100) may include a memory (110), a processor (120), and a communication interface (130).
[0156] The memory (110) is configured to store at least one instruction or data related to an operating system (OS) for controlling the overall operation of the components of the electronic device (100) and the components of the electronic device (100).
[0157] The memory (110) stores various programs or data temporarily or non-temporarily and transmits the stored information to the processor upon the processor's call. Additionally, the memory (110) can store various information required for the processor's calculation, processing, or control operations in an electronic format.
[0158] For example, the memory (110) may store a diffusion model that generates an image by performing reverse diffusion based on embeddings.
[0159] The memory (110) may include, for example, at least one of a main memory and an auxiliary memory. The main memory may be implemented using a semiconductor storage medium such as ROM and / or RAM. The ROM may include, for example, a conventional ROM, EPROM, EEPROM and / or MASK-ROM. The RAM may include, for example, a DRAM and / or SRAM. The auxiliary memory may be implemented using at least one storage medium capable of storing data permanently or semi-permanently, such as a flash memory device, an SD (Secure Digital) card, a solid state drive (SSD), a hard disk drive (HDD), an optical recording medium such as a magnetic drum, a compact disc (CD), a DVD, or a laser disc, a magnetic tape, a magneto-optical disc and / or a floppy disk.
[0160] The processor (120) controls the overall operation of the electronic device (100). Specifically, the processor (120) is connected to the configuration of the electronic device (100) including memory as described above, and can control the overall operation of the electronic device (100) by executing at least one instruction stored in the memory (110) as described above.
[0161] In one embodiment, the processor (120) obtains a plurality of text embeddings optimized for each of different classes and a target embedding for a target class, identifies at least one of the plurality of text embeddings as a support embedding based on the cosine similarity of each of the plurality of text embeddings to the target embedding, and optimizes the target embedding for the target class by maximizing the distance between the target embedding and the support embedding based on a first image generated based on the target embedding and a second image generated based on the support embedding.
[0162] To this end, the processor (120) can control a text embedding module, an image embedding module, a similarity evaluation module, an image generation module, an embedding adjustment module, etc. Each of these modules may correspond to a module of a functional unit implemented in software and / or hardware.
[0163] The text embedding module is a module that vectorizes text. For example, the text embedding module can identify vector values in a latent space corresponding to text and obtain text embeddings corresponding to the identified vector values.
[0164] An image embedding module is a module that vectorizes images. For example, to learn meaningful representations in a latent space by converting images into vector forms, an image embedding module can utilize models such as a Convolutional Neural Network (CNN) or a Vision Transformer (Vision Transformer) to extract low-level features (e.g., color, shape, texture) and high-level features (e.g., objects, scenes, context), and can generate image embeddings by mapping the extracted features into a high-dimensional vector space.
[0165] The similarity evaluation module is a module for calculating cosine similarity between embeddings. For example, the similarity evaluation module calculates the cosine similarity of each of a plurality of text embeddings with respect to a target embedding, and can select at least one of the plurality of text embeddings as a support embedding based on the calculated cosine similarity.
[0166] The image generation module is a module that generates images based on text embeddings. For example, the image generation module can generate an image for a text embedding by performing a reverse diffusion process through a diffusion model.
[0167] The embedding adjustment module is a module that maximizes the distance between the target embedding and the support embedding based on the final loss. For example, the embedding adjustment module can adjust the target embedding so that the cosine similarity of the support embedding to the target embedding decreases, while adjusting the target embedding in a direction where the final loss decreases and the second loss increases.
[0168] In addition, the processor (120) can be implemented as a single processor, as well as as multiple processors.
[0169] The processor (120) may be implemented in various ways. For example, one or more processors (120) may include one or more of a CPU (Central Processing Unit), GPU (Graphics Processing Unit), APU (Accelerated Processing Unit), MIC (Many Integrated Core), DSP (Digital Signal Processor), NPU (Neural Processing Unit), hardware accelerator, or machine learning accelerator. One or more processors (120) may control one or any combination of other components of an electronic device and may perform operations or data processing related to communication. One or more processors (120) may execute one or more programs or instructions stored in memory. For example, one or more processors (120) may perform a method according to one embodiment of the present disclosure by executing one or more instructions stored in memory.
[0170] In embodiments of the present disclosure, the processor (120) may mean a system-on-chip (SoC) in which one or more processors (120) and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, GPU, APU, MIC, DSP, NPU, hardware accelerator, or machine learning accelerator, but the embodiments of the present disclosure are not limited thereto.
[0171] The communication interface (130) may include a wireless communication interface, a wired communication interface, or an input interface. The wireless communication interface may communicate with various external devices using wireless communication technology or mobile communication technology. Such wireless communication technologies may include, for example, Bluetooth, Bluetooth Low Energy, CAN communication, Wi-Fi, Wi-Fi Direct, ultrawide band (UWB), Zigbee, infrared data association (IrDA), or near field communication (NFC), and mobile communication technologies may include 3GPP, Wi-Max, LTE (Long Term Evolution), 5G, etc.
[0172] A wireless communication interface can be implemented using an antenna, a communication chip, a substrate, etc., capable of transmitting electromagnetic waves to the outside or receiving electromagnetic waves transmitted from the outside.
[0173] A wired communication interface can communicate with various devices based on a wired communication network. Here, the wired communication network can be implemented using physical cables, such as, for example, pair cables, coaxial cables, fiber optic cables, or Ethernet cables.
[0174] Depending on the embodiment, either the wireless communication interface or the wired communication interface may be omitted. Accordingly, the electronic device (100) may include only a wireless communication interface or only a wired communication interface. In addition, the electronic device (100) may be equipped with an integrated communication interface (130) that supports both wireless connection via the wireless communication interface and wired connection via the wired communication interface.
[0175] The electronic device (100) is not limited to having one communication interface that performs a communication connection in one manner, but may include multiple communication interfaces that perform communication connections in multiple manners.
[0176] FIG. 8 is a drawing illustrating the configuration of a system according to one embodiment of the present disclosure.
[0177] Referring to FIG. 8, the system (1000) may include an electronic device (100) and a user terminal (200).
[0178] The electronic device (100) can perform training of a diffusion model that generates images based on embeddings. It may be implemented as a desktop PC, laptop PC, smartphone, tablet PC, netbook computer, mobile device, wearable device, etc., or may be implemented as a server device or system including at least one computer.
[0179] A user terminal (200) is a device carried or stored by a user and may include at least one of a desktop PC, a laptop PC, a smartphone, a tablet PC, a netbook computer, a mobile device, and a wearable device, but is not limited thereto.
[0180] In one embodiment, the user terminal (200) can transmit text related to the target class to the electronic device (100).
[0181] Accordingly, the electronic device (100) identifies a vector value in a potential space corresponding to text received from a user terminal (200), obtains a target embedding corresponding to the identified vector value, obtains a plurality of text embeddings optimized for each of different classes, identifies at least one of the plurality of text embeddings as a support embedding based on the cosine similarity of each of the plurality of text embeddings to the target embedding, and optimizes the target embedding for the target class by maximizing the distance between the target embedding and the support embedding based on a first image generated based on the target embedding and a second image generated based on the support embedding.
[0182] Meanwhile, the various embodiments described above may be implemented by combining two or more embodiments, provided that they do not conflict or contradict each other.
[0183] Specifically, each of the multiple operations, steps, and configurations for implementing one embodiment may be embodied in another embodiment, or the operations and steps of another embodiment may be followed by the last operation or step of one embodiment, but are not limited thereto.
[0184] Meanwhile, computer instructions or computer programs for performing processing operations in the various embodiments of the present disclosure described above may be stored in a non-transitory computer-readable medium. When such computer instructions or computer programs stored in the non-transitory computer-readable medium are executed by a processor of a specific device, the specific device described above performs processing operations according to the various embodiments described above.
[0185] A non-transient computer-readable medium refers to a medium that stores data semi-permanently and can be read by a device, unlike media that store data for a short period of time such as registers, caches, and memory. Specific examples of non-transient computer-readable media include CDs, DVDs, hard disks, Blu-ray discs, USBs, memory cards, and ROMs.
[0186] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0187] Although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.
Claims
1. In a method of operating an electronic device, A step of obtaining multiple text embeddings optimized for each of different classes; A step of obtaining a target embedding for a target class that does not match the above different classes; A step of identifying at least one of the plurality of text embeddings as a support embedding based on the cosine similarity of each of the plurality of text embeddings to the target embedding; and A method of operation comprising: a step of optimizing the target embedding for the target class by maximizing the distance between the target embedding and the support embedding based on a first image generated based on the target embedding and a second image generated based on the support embedding.
2. In Paragraph 1, The step of obtaining a target embedding for the above target class is, Acquire text related to the above target class, identify vector values in the latent space corresponding to the acquired text, and A method of operation for obtaining the target embedding corresponding to the identified vector value.
3. In Paragraph 1, The step of identifying at least one of the above plurality of text embeddings as a support embedding is, A method of operation comprising: identifying the rank of each of the plurality of text embeddings based on the cosine similarity of each of the plurality of text embeddings to the target embedding, and setting the text embedding whose identified rank is greater than or equal to a threshold value as the support embedding.
4. In Paragraph 3, The method of operation of the above electronic device is, In the process of a diffusion model generating an image based on text embeddings, a step of acquiring a first noise that needs to be added to the reference noise at each time step; A step of inputting the target embedding into the above diffusion model, and in the process of generating the first image based on the target embedding by the above diffusion model performing reverse diffusion on reference noise, acquiring the second noise predicted for each time step; and A method of operation comprising: a step of inputting the support embeddings into the above-mentioned diffusion model, and in the process of generating the second image based on the support embeddings by the above-mentioned diffusion model performing reverse diffusion on reference noise, wherein the third noise predicted for each time step is obtained.
5. In Paragraph 4, The method of operation of the above electronic device is, A step of calculating a first loss according to the loss function of the diffusion model based on the first noise and second noise for each time step; A step of calculating a second loss according to the loss function of a Diffusion-Based Distance Metric (DMM) based on the second noise and third noise for each time step and a penalty set for each time step; and A method of operation comprising: a step of calculating a third loss according to a time-adaptive normalization loss function based on the first noise and second noise for each time step and a penalty set for each time step.
6. In Paragraph 5, The step of optimizing the above target embedding for the above target class is, A step of calculating a final loss based on the first to third losses above; and A method of operation comprising the step of adjusting the target embedding in a direction such that the final loss calculated above is reduced and the second loss above is increased.
7. In Paragraph 6, The step of calculating the final loss above is, A step of applying a first hyperparameter having a positive value to the second loss; A step of applying a second hyperparameter having a positive value to the third loss; and A method of operation comprising the step of calculating the final loss by subtracting the second loss to which the first hyperparameter is applied from the value calculated by summing the first loss and the third loss to which the second hyperparameter is applied.
8. In Paragraph 6, The step of adjusting the above target embedding is, A method of operation for adjusting the target embedding such that the cosine similarity of the support embedding to the target embedding is reduced.
9. In electronic devices, Memory in which at least one instruction is stored; and A processor connected to the above memory; including, The above processor is, Obtain multiple text embeddings optimized for each different class and target embeddings for the target class, and Based on the cosine similarity of each of the plurality of text embeddings with respect to the target embedding, at least one of the plurality of text embeddings is identified as a support embedding, and An electronic device that optimizes a target embedding for a target class by maximizing the distance between the target embedding and the support embedding based on a first image generated based on the target embedding and a second image generated based on the support embedding.
10. In the system, A user terminal that transmits text related to a target class to an electronic device; and An electronic device comprising: identifying a vector value in a latent space corresponding to text received from the user terminal; obtaining a target embedding corresponding to the identified vector value; obtaining a plurality of text embeddings optimized for each of different classes; identifying at least one of the plurality of text embeddings as a support embedding based on the cosine similarity of each of the plurality of text embeddings to the target embedding; and optimizing the target embedding for the target class by maximizing the distance between the target embedding and the support embedding based on a first image generated based on the target embedding and a second image generated based on the support embedding.