Learning method for fine-tuning a classification model in a few-shot environment and a computing device for performing the same
Patent Information
- Application Number
- KR1020250115330
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2026-09-02
- Estimated Expiration
- 2045-08-19
Smart Images

Figure 112025094684406-PAT00049_ABST
Abstract
Description
Technology Field
[0001] An embodiment of the present invention relates to a learning technique for fine-tuning a classification model in a few-shot environment. Background Technology
[0003] Recently, with the emergence of Large Language Models (LLMs) and Foundation Models, there has been a rapidly increasing demand for fine-tuning techniques to effectively apply pre-trained models to various downstream tasks. Large language models contain hundreds of millions of parameters, and traditional fine-tuning methods update all parameters of the model, which is not only inefficient in terms of computational resources but also leads to overfitting and storage space issues.
[0004] To overcome these limitations, Parameter Efficient Fine-Tuning (PEFT) techniques are being actively researched. PEFT secures both resource efficiency and generalization performance simultaneously by selectively training only a small number of parameters while keeping most of the neural network model's parameters fixed. Representative PEFT techniques include Adapter, Prompt Tuning, BitFit, and LoRA. Among these, Low-Rank Adaptation (LoRA) is widely used because it can be naturally integrated into Transformer structures and maintain performance with very low computational load.
[0005] However, LoRA is not always the best option in environments with extremely limited data (e.g., few-shot settings). That is, in situations where meaningful representations must be learned from a small number of samples, LoRA can be sensitive to noise or fail to sufficiently isolate key representations. Therefore, there is a need for training methods that can utilize Low-Rank Adaptation (LoRA) techniques in data-scarce environments. Prior art literature
[0007] Korean Published Patent Application No. 10-2025-0063040 (May 8, 2025) The problem to be solved
[0008] An embodiment of the present invention is intended to provide a new learning technique for fine-tuning a classification model in a few-shot environment. means of solving the problem
[0010] A learning method for fine-tuning a classification model according to an embodiment of the present invention is performed on a computing device having one or more processors and a memory storing one or more programs executed by said one or more processors, and comprises the steps of: inputting a target image into a classification model to extract a feature vector; passing the feature vector through a first layer composed of previously learned weights and a second layer composed of new weights, respectively, of said classification model; applying Multi-Head Attention to the output of the feature vector that has passed through the first layer and the second layer, respectively, to generate an attention representation; connecting the output of the feature vector that has passed through the first layer and the second layer, respectively, with the attention representation to generate a connection vector; inputting the connection vector into a gating network to calculate a gate vector and generating a gating adaptation output based on said gate vector; and inputting a final output, which is the sum of the output of the feature vector that has passed through the first layer and the gating adaptation output, into a classifier to classify the class of said target image.
[0011] A computing device according to one disclosed embodiment comprises one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include: a command for inputting a target image into a classification model to extract a feature vector; a command for passing the feature vector through a first layer composed of previously learned weights and a second layer composed of new weights, respectively, of the classification model; a command for generating an attention representation by applying Multi-Head Attention to the output of the feature vector that has passed through the first layer and the second layer, respectively; a command for generating a connection vector by connecting the output of the feature vector that has passed through the first layer and the second layer, respectively, with the attention representation; and a command for inputting the connection vector into a gating network to calculate a gate vector and generating a gating adaptation output based on the gate vector. and includes a command to input the final output, which is the sum of the output of the feature vector passed through the first layer and the gating adaptation output, into a classifier to classify the class of the target image. Effects of the invention
[0013] According to the disclosed embodiment, only important features can be dynamically selected and reflected in the training of a classification model through information selection gating in a few-shot environment, and key features in the target image can be well distinguished while being robust against noise through pre-set loss functions. In addition, the performance of the classification model can be improved by enhancing adaptability in a few-shot environment while maintaining a lightweight structure based on LoRA. Brief explanation of the drawing
[0015] FIG. 1 is a block diagram illustrating a computing environment including a computing device suitable for use in exemplary embodiments. FIG. 2 is a flowchart illustrating a classification model learning method for a few-shot environment according to an embodiment of the present invention. FIG. 3 is a schematic diagram showing the structure during fine tuning in a classification model according to one embodiment of the present invention. Specific details for implementing the invention
[0016] Hereinafter, specific embodiments of the present invention will be described with reference to the drawings. The following detailed description is provided to facilitate a comprehensive understanding of the methods, apparatus, and / or systems described herein. However, this is merely illustrative and the present invention is not limited thereto.
[0017] In describing the embodiments of the present invention, detailed descriptions of known technologies related to the present invention are omitted if it is determined that such detailed descriptions may unnecessarily obscure the essence of the present invention. Furthermore, the terms described below are defined in consideration of their functions within the present invention, and these may vary depending on the intentions or practices of the user or operator. Therefore, such definitions should be based on the content throughout this specification. Terms used in the detailed description are intended merely to describe the embodiments of the present invention and should not be limiting in any way. Unless explicitly stated otherwise, expressions in the singular form include the meaning of the plural form. In this description, expressions such as "include" or "comprise" are intended to refer to certain characteristics, numbers, steps, actions, elements, parts thereof, or combinations thereof, and should not be interpreted to exclude the existence or possibility of one or more other characteristics, numbers, steps, actions, elements, parts thereof, or combinations thereof other than those described.
[0018] Additionally, terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. These terms may be used for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component.
[0020] FIG. 1 is a block diagram illustrating a computing environment (10) including a computing device suitable for use in exemplary embodiments. In the illustrated embodiments, each component may have different functions and capabilities in addition to those described below, and may include additional components in addition to those described below.
[0021] The illustrated computing environment (10) includes a computing device (12). In one embodiment, the computing device (12) may be a device for few-shot learning of a classification model using LoRA (Low-Rank Adaptation). In one embodiment, the computing device (12) may be a device for learning a classification model in a few-shot environment using LoRA. In this case, the classification model may be a deep learning-based model.
[0022] The computing device (12) includes at least one processor (14), a computer-readable storage medium (16), and a communication bus (18). The processor (14) can cause the computing device (12) to operate according to the exemplary embodiment described above. For example, the processor (14) can execute one or more programs stored in the computer-readable storage medium (16). The one or more programs may include one or more computer-executable instructions, and the computer-executable instructions may be configured to cause the computing device (12) to perform operations according to the exemplary embodiment when executed by the processor (14).
[0023] A computer-readable storage medium (16) is configured to store computer-executable instructions or program code, program data and / or other suitable forms of information. A program (20) stored in the computer-readable storage medium (16) includes a set of instructions executable by a processor (14). In one embodiment, the computer-readable storage medium (16) may be memory (volatile memory such as random access memory, non-volatile memory, or a suitable combination thereof), one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other forms of storage media that are accessed by a computing device (12) and capable of storing desired information, or a suitable combination thereof.
[0024] The communication bus (18) interconnects various other components of the computing device (12), including the processor (14) and the computer-readable storage medium (16).
[0025] The computing device (12) may also include one or more input / output interfaces (22) and one or more network communication interfaces (26) that provide interfaces for one or more input / output devices (24). The input / output interfaces (22) and network communication interfaces (26) are connected to a communication bus (18). The input / output devices (24) may be connected to other components of the computing device (12) through the input / output interfaces (22). An exemplary input / output device (24) may include an input device such as a pointing device (such as a mouse or trackpad), a keyboard, a touch input device (such as a touchpad or touchscreen), a voice or sound input device, various types of sensor devices and / or imaging devices, and / or an output device such as a display device, a printer, a speaker and / or a network card. An exemplary input / output device (24) may be included inside the computing device (12) as a component constituting the computing device (12), or it may be connected to the computing device (12) as a separate device distinct from the computing device (12).
[0027] FIG. 2 is a flowchart illustrating a method for training a classification model in a few-shot environment according to an embodiment of the present invention, and FIG. 3 is a diagram schematically illustrating the structure during fine-tuning of a classification model according to an embodiment of the present invention. Although the method is described in the illustrated flowchart by dividing it into multiple steps, at least some of the steps may be performed in a different order, combined with other steps and performed together, omitted, divided into detailed steps, or performed with one or more steps not illustrated added.
[0028] Referring to FIG. 2, the computing device (12) can generate a patch embedding sequence by inputting an image to be classified (target image) into a classification model (S 101). In one embodiment, the classification model may be a Vision Transformer, but is not limited thereto. The classification model may be a model that has been trained to perform a classification task.
[0029] For example, a target image of size 224×224 can be divided into 196 patches of size 16×16 by a classification model. The classification model can generate a patch embedding sequence by performing embedding for each patch. The classification model can include position embeddings in the patch embedding sequence to maintain position information for each patch embedding.
[0030] Next, the computing device (12) can input the patch embedding sequence into the encoder of a classification model (e.g., the encoder of a vision transformer) to extract a feature vector (x) for the target image (S 103).
[0031] Next, referring to FIG. 3, the computing device (12) can pass a feature vector (x) through a first layer (51) composed of previously learned weights (W) of a classification model and a second layer (53) composed of new weights, respectively (S 105). The first layer (51) and the second layer (53) are neural network layers of a classification model located behind the encoder. The second layer (53) is composed of new weights different from the weights of the first layer (51) (i.e., learned weights (W)).
[0032] Here, the second layer (53) composed of new weights is a layer for fine-tuning using LoRA (Low-Rank Adaptation) and may be a layer composed of projection matrices (A, B).
[0033] Projection matrices (A, B) include a low-dimensional projection matrix (A) and a high-dimensional projection matrix (B). The low-dimensional projection matrix (A) and the high-dimensional projection matrix (B) are learnable matrices. The low-dimensional projection matrix (A) may be a matrix that down-projects a feature vector (x) to a lower dimension (rank r) than the original dimension. The high-dimensional projection matrix (B) may be a matrix that up-projects a feature vector from the lower dimension (rank r) to the original dimension. For example, if the dimension of the feature vector (x) is (n×k), the dimension of matrix A may be (n×r) and the dimension of matrix B may be (r×k). The feature vector (x) is restored to its original dimension by sequentially passing through matrix A and matrix B.
[0034] The output of the feature vector (x) that has passed through the first layer (51) and the second layer (53), respectively, can be represented as the output of a general LoRA by the following mathematical formula 1.
[0035] (Mathematical Formula 1)
[0036]
[0037] W: Weight of the first layer
[0038] A: Low-dimensional projection matrix
[0039] B: High-dimensional projection matrix
[0040] Next, the computing device (12) outputs the feature vector (x) that has passed through the first layer (51) and the second layer (53), respectively. An attention representation can be generated by applying Multi-Head Attention (MHA) (55) to the output of applying LoRA to the feature vector (i.e., the output of applying LoRA to the feature vector) (S 107).
[0041] For example, if the target image is a cat image, Head 1 can focus on the shape of the ears, Head 2 on the eyes, Head 3 on the fur pattern, and Head 4 on the overall contour, and an attention expression can be generated by combining the attention results of the four heads. In this case, the features of the cat can be analyzed from different perspectives using a Query, Key, and Value matrix for each head.
[0042] Next, the computing device (12) outputs the feature vector (x) that has passed through the first layer (51) and the second layer (53), respectively. ) and attention expression( A connection vector can be generated by concatenating ) (S 109).
[0043] Next, the computing device (12) can input the connection vector into the gating network (57) to generate a gating adaptive output (S 111).
[0044] Here, the gate network may be a sub-network that generates gate values by determining the importance of each element (or feature) of the input information based on the input information. In this case, the gate network can generate gate values between 0 and 1 for each element (or feature) of the input information. For example, gate values can be generated through a sigmoid function. The gate network can assign gate values closer to 1 as the importance of each element increases, and gate values closer to 0 as the importance of each element decreases.
[0045] The computing device (12) is a gating adaptive output ( ) can be calculated using the following mathematical formula 2. Through this gating adaptation output, important features in the connection vector are strengthened and unimportant features are suppressed.
[0046] (Mathematical Formula 2)
[0047]
[0048] : A gate vector composed of gate values
[0049] : Multiplication symbol for each dimension
[0050] Next, the computing device (12) outputs the feature vector that has passed through the first layer (51). = ) and gating adaptive output( The final output obtained by combining ) can be input into a classifier of a classification model to classify the class of the target image (S 113). Here, the final output can be represented by the following mathematical formula 3.
[0051] (Mathematical Formula 3)
[0052]
[0053] Meanwhile, the computing device (12) can train a classification model through a first loss function and a second loss function. Here, the first loss function is a basic loss function of the classification model, and is a loss function for minimizing the difference between the correct label and the label predicted by the classification model. In one embodiment, the first loss function (L ce ) can be expressed as the Cross-Entropy Loss by the following Equation 4.
[0054] (Mathematical Formula 4)
[0055]
[0056] N: Number of classes
[0057] : Correct answer label
[0058] : Labels predicted by the classification model
[0059] Also, the second loss function is the gate vector ( It is a loss function for maximizing the entropy of ). In this case, maximizing the entropy of the gate vector means ensuring that the gate vector is evenly distributed across the entire range rather than being concentrated in a certain range, thereby securing diversity of representation. In this case, the classification model becomes able to classify target images using various features without relying on specific features. In one embodiment, the second loss function (L entropy ) can be expressed by the following mathematical formula 5.
[0060] (Mathematical Formula 5)
[0061]
[0062] d model : Dimensions of a classification model
[0063] g[j] : gate value of the j-th dimension
[0064] Additionally, the computing device (12) can train a classification model by additionally using a third loss function in addition to the first loss function and the second loss function. The third loss function is a feature vector (x) and an attention representation ( It is a loss function designed to maintain a meaningful relationship between ). In other words, the third loss function is between the feature vector (x) and the attention representation ( The purpose is to maximize mutual information between them so that meaningful information is well transmitted between them and to ensure consistency and separability between representations. In one embodiment, the third loss function can be expressed as Equation 6 as an Information Noise-Contrastive Estimation (InfoNCE) based loss function.
[0065] (Mathematical Formula 6)
[0066]
[0067] B: Batch size
[0068] : i-th feature vector
[0069] : Attention expression for the i-th feature vector
[0070] Attention expressions other than the attention expression for the i-th feature vector
[0071] : Pre-set temperature parameter
[0072] In addition, the computing device (12) can train a classification model by additionally using a fourth loss function in addition to the first loss function and the second loss function. The fourth loss function is a loss function designed to cause samples belonging to the same class to become closer to each other in the vector space and samples belonging to different classes to become further apart from each other in the vector space when training the classification model. That is, the fourth loss function is a contrastive loss and can be represented by the following mathematical formula 7.
[0073] (Mathematical Formula 7)
[0074]
[0075] B: Batch size
[0076] h i : Feature vector of the i-th sample
[0077] h j : Feature vector of samples of the same class as the i-th sample
[0078] P(i): Set of samples of the same class as the i-th sample
[0079] h k : Feature vector of samples that are different in class from the i-th sample
[0080] Meanwhile, the computing device (12) can train a classification model using all of the first to fourth loss functions. At this time, the final loss function (L total ) can be expressed by the following mathematical formula 8.
[0081] (Mathematical Formula 8)
[0082]
[0083] , , : Pre-configured hyperparameters
[0084] According to the disclosed embodiment, only important features can be dynamically selected and reflected in the training of a classification model through information selection gating in a few-shot environment, and key features in the target image can be well distinguished while being robust against noise through pre-set loss functions. In addition, the performance of the classification model can be improved by enhancing adaptability in a few-shot environment while maintaining a lightweight structure based on LoRA.
[0086] Although representative embodiments of the present invention have been described in detail above, those skilled in the art will understand that various modifications can be made to the above-described embodiments without departing from the scope of the present invention. Therefore, the scope of the present invention should not be limited to the described embodiments, but should be defined by the claims set forth below as well as equivalents thereof. Explanation of the symbols
[0088] 10: Computing Environment 12: Computing device 14 : Processor 16: Computer-readable storage media 18: Communication bus 20 : Program 22 : Input / Output Interface 24 : Input / Output Devices 26: Network communication interface
Claims
Claim 1 A method performed in a computing device having one or more processors and a memory for storing one or more programs executed by said one or more processors, comprising: a step of inputting a target image into an encoder of a classification model to extract a feature vector; a step of passing said feature vector through a first layer composed of previously learned weights of said classification model and a second layer composed of new weights different from the weights of said first layer, respectively (said that the first layer and said second layer are neural network layers of said classification model located behind said encoder); a step of generating an attention representation by applying Multi-Head Attention to the output of the feature vector that has passed through said first layer and said second layer, respectively; a step of generating a connection vector by connecting the output of the feature vector that has passed through said first layer and said second layer, respectively, with said attention representation; a step of inputting said connection vector into a Gating Network to produce a gate vector, and multiplying said connection vector by said gate vector to generate a gating adaptation output. A learning method for fine-tuning a classification model, comprising: a step of inputting a final output, which is the sum of the output of a feature vector passed through the first layer and the gating adaptation output, into a classifier of the classification model to classify the class of the target image; and a step of calculating a loss function based on a set correct label and the class classification result of the classifier, and updating the neural network weights of at least one of the second layer and the gate network to minimize the loss function. Claim 2 A learning method for fine-tuning a classification model according to claim 1, wherein the second layer is a LoRA (Low-Rank Adaptation) based layer comprising a low-dimensional projection matrix and a high-dimensional projection matrix, wherein the low-dimensional projection matrix projects the feature vector to a lower dimension lower than the original dimension, and the high-dimensional projection matrix projects the feature vector projected to the lower dimension back to the original dimension. Claim 3 In claim 2, the gating adaptive output ( ) is a learning method for fine-tuning a classification model, generated by the following mathematical formula. (Mathematical formula) : A gate vector composed of gate values : Multiplication symbol for each dimension : Attention expression : Output of feature vectors passed through the 1st and 2nd layers, respectively Claim 4 A learning method for fine-tuning a classification model according to claim 1, wherein the computing device trains the classification model through a first loss function for minimizing the difference between a preset correct label and a label predicted by the classification model and a second loss function for maximizing the entropy of the gate vector. Claim 5 In claim 4, the second loss function (L entropy ) is a learning method for fine-tuning a classification model, expressed by the following mathematical formula. (Mathematical formula) d model : Dimension of the classification model g[j] : Gate value of the j-th dimension Claim 6 A learning method for fine-tuning a classification model according to claim 4, wherein the computing device learns the classification model by adding a third loss function to maintain a meaningful relationship between the feature vector and the attention representation in addition to the first loss function and the second loss function. Claim 7 In claim 6, the third loss function (L InfoNCE ) is a learning method for fine-tuning a classification model, expressed by the mathematical formula below. (Mathematical formula) B: Batch size : i-th feature vector : Attention expression for the i-th feature vector Attention expressions other than the attention expression for the i-th feature vector : Pre-set temperature parameter Claim 8 A learning method for fine-tuning a classification model according to claim 6, wherein the computing device trains the classification model by adding a fourth loss function in addition to the first to third loss functions to cause samples belonging to the same class to become closer to each other in a vector space and samples of different classes to become farther apart from each other in a vector space. Claim 9 One or more processors; memory; and includes one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include: a command for inputting a target image into an encoder of a classification model to extract a feature vector; a command for passing the feature vector to a first layer composed of previously learned weights of the classification model and a second layer composed of new weights different from the weights of the first layer, respectively (the first layer and the second layer are neural network layers of the classification model located behind the encoder); a command for generating an attention representation by applying Multi-Head Attention to the output of the feature vector that has passed through the first layer and the second layer, respectively; a command for generating a connection vector by connecting the output of the feature vector that has passed through the first layer and the second layer, respectively, with the attention representation; and a command for inputting the connection vector into a Gating Network to calculate a gate vector and multiplying the connection vector by the gate vector to generate a gating adaptation output. A command for inputting the final output, which is the sum of the output of the feature vector passed through the first layer and the gating adaptation output, into the classifier of the classification model to classify the class of the target image; and a command for calculating a loss function based on a pre-set ground truth label and the class classification result of the classifier, and for updating the neural network weights of at least one of the second layer and the gate network to minimize the loss function, Computing device. Claim 10 The computing device of claim 9, wherein the computing device trains the classification model through a first loss function for minimizing the difference between a preset correct label and a label predicted by the classification model and a second loss function for maximizing the entropy of the gate vector. Claim 11 In claim 10, the second loss function (L entropy ) is a computing device expressed by the following mathematical formula. (Mathematical formula) d model : Dimension of the classification model g[j] : Gate value of the j-th dimension Claim 12 The computing device of claim 10, wherein the computing device trains the classification model by adding a third loss function to maintain a meaningful relationship between the feature vector and the attention representation in addition to the first loss function and the second loss function. Claim 13 In claim 12, the third loss function (L InfoNCE ) is a computing device represented by the following mathematical formula. (Mathematical formula) B: Batch size : i-th feature vector : Attention expression for the i-th feature vector Attention expressions other than the attention expression for the i-th feature vector : Pre-set temperature parameter Claim 14 The computing device of claim 12, wherein, in addition to the first to third loss functions, the computing device trains the classification model by adding a fourth loss function to cause samples belonging to the same class to become closer to each other in a vector space and samples of different classes to become farther apart from each other in a vector space. Claim 15 A computer program stored on a non-transitory computer-readable storage medium, wherein the computer program comprises one or more instructions, and when the instructions are executed by a computing device having one or more processors, the computing device comprises the steps of: inputting a target image into an encoder of a classification model to extract a feature vector; passing the feature vector through a first layer composed of previously learned weights of the classification model and a second layer composed of new weights different from the weights of the first layer, respectively (the first layer and the second layer are neural network layers of the classification model located behind the encoder); applying Multi-Head Attention to the output of the feature vector that has passed through the first layer and the second layer, respectively, to generate an attention representation; connecting the output of the feature vector that has passed through the first layer and the second layer, respectively, with the attention representation to generate a connection vector; inputting the connection vector into a gating network to produce a gate vector, and the connection vector A computer program comprising: a step of generating a gating adaptation output by multiplying with the gate vector; a step of inputting the final output, which is the sum of the output of the feature vector that has passed through the first layer and the gating adaptation output, into a classifier of the classification model to classify the class of the target image; and a step of calculating a loss function based on a pre-set correct label and the class classification result of the classifier, and updating the neural network weights of at least one of the second layer and the gate network to minimize the loss function.
Citation Information
Patent Citations
Representation learning method and system using multi-attention module
KR102513285B1
Gated attention neural networks
US20220366218A1
Information retrieval systems and methods with granularity-aware adaptors for solving multiple different tasks
US20240127104A1