A Few-Shot Fine-Tuning Method Based on Vision Self-Attention Model
By using the visual self-attention model ViT and the learningable conversion module norm adapter in the small sample image classification, the problem of efficient fine-tuning in the small sample situation is solved, and the computing and storage resources are saved is achieved, and the model performance is improved.
Patent Information
- Application Number
- CN202310867841.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-16
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-07-16
AI Technical Summary
In small sample image classification, it is difficult for the prior art to effectively utilize pre-training knowledge for efficient fine-tuning under limited computing and storage resources, especially when the base class data of the target domain is difficult to obtain, the computing and storage overhead is too large, which limits the application scenarios of the model.
The visual self-attention model ViT is used as the backbone network, and a learningable transformation module norm adapter is built. After the normalization layer, the gain and bias are corrected by element-by-element multiplication and addition. Combined with the prototype network ProtoNet classification head, it is pre-trained on a large-scale data set and fine-tuned on small-sample tasks. Only the norm adapter parameters are updated, and other parameters are frozen.
It realizes efficient fine-tuning in small samples, has fewer requirements for computing and storage resources, and has better performance than traditional methods, and is suitable for practical application scenarios.
Smart Images

Figure CN117036901B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pattern recognition, and particularly relates to a small-sample fine-tuning method based on a visual self-attention model. Background Art
[0002] Pre-trained models have been widely used in the fields of natural language processing (NLP) and computer vision (CV), and have greatly improved the performance of downstream tasks. Therefore, the pre-training-fine-tuning paradigm has been widely accepted, especially after the rise of the vision attention model (ViT). Due to the large scale of pre-trained models, how to effectively transfer pre-trained knowledge to downstream tasks with limited computational and storage overheads is still under research. Some methods have been proposed to solve this problem, called parameter-efficient fine-tuning (PEFT) methods, such as: Adapter, bias-tuning, visual prompt tuning, etc.
[0003] However, there is little research on parameter-efficient fine-tuning methods in small-sample image classification. Small-sample image classification is a basic task of few-shot learning. Few-shot learning can expand the application scope of deep learning models by imitating human intelligence and generalizing to new concepts with a small number of samples. In the small-sample setting, the test data is divided into many tasks, each task consisting of two parts: a support set and a query set. The support set contains N*K labeled samples, that is, data of N classes, with K samples in each class. Such a small-sample task is called in the form of "N-way K-shot"; the query set contains N*Q samples for evaluating the model.
[0004] Recently, Shell et al. first introduced pre-trained models into the field of small-sample classification. They adopted the process of pre-training, meta-training, and finally fine-tuning. First, the model is pre-trained on a large-scale dataset (such as the ImageNet dataset), then meta-trained on the base-class data in the target domain, and finally, in the fine-tuning process, all parameters of the model are updated (full-tuning) using a small number of samples. The pre-training-meta-training-fine-tuning process has greatly improved the performance of the model. However, the base-class data in the target domain for meta-training is not easily obtained. In most cases, only a very small number of labeled samples can be obtained. Therefore, in this case, meta-training cannot be carried out, and updating all parameters of the model (full-tuning) using a small number of samples cannot fully utilize pre-trained knowledge. Moreover, the computational and storage overheads brought by updating all parameters are very large, severely limiting its application scenarios. Therefore, how to perform efficient fine-tuning in the small-sample case remains an open question. Summary of the Invention
[0005] To overcome the deficiencies of the prior art, the present invention provides a few-shot fine-tuning method based on a vision self-attention model, which adopts a process of pre-training on a large-scale dataset and fine-tuning on a few-shot task. The vision self-attention model is used as the backbone network, and a learnable transformation module, namely norm adapter, is constructed. The norm adapter consists of two vectors and is used to correct the gain and bias of the normalization layer of the original vision self-attention model. The norm adapter is located after all the normalization layers of the vision self-attention model ViT and is implemented through element-wise multiplication and addition. During pre-training, the backbone network trained in a fully supervised or self-supervised manner on a large-scale dataset is used. During the fine-tuning process, a prototype network ProtoNet classification head is adopted. The present invention is computationally simple and can be achieved through element-wise multiplication and addition, so it occupies less storage and computing resources, which is conducive to the pre-trained model being put into practical application scenarios.
[0006] The technical solution adopted by the present invention to solve its technical problems includes the following steps:
[0007] Step 1: Construct the backbone network;
[0008] Adopt an improved vision self-attention model ViT as the backbone network;
[0009] The original vision self-attention model consists of a patch embedding layer and N transformer layers. After passing through the patch embedding layer, the input image is encoded into a certain number of token vectors. After adding the position encoding, the input token vectors together with the CLS token are fed into N transformer layers. Finally, after passing through N transformer layers and a normalization layer LayerNorm, the CLS token is used for classification or other purposes. Each transformer layer contains two normalization layers LayerNorm, an MLP block, and a multi-head self-attention block MHSA;
[0010] Construct a learnable transformation module, which consists of two vectors and is used to correct the gain gain and bias bias of the normalization layer LayerNorm of the original vision self-attention model. The learnable transformation module is called norm adapter. The norm adapter is located after all the normalization layers of the vision self-attention model ViT and is implemented through element-wise multiplication and addition. As shown in formula (1), Scale and Shift are the two learnable vectors of the norm adapter, y is the output of the normalization layer, and ⊙ represents element-wise multiplication:
[0011] h = Scale ⊙ y + Shift (1)
[0012] The structures of the Scale and Shift parameters of the norm adapter are the same as the gain and bias parameters of the normalization layer, and are respectively initialized as vectors of all 1s and all 0s; during fine-tuning, only the Scale and Shift parameters are updated, and other parameters are frozen after pre-training and not optimized;
[0013] Step 2: During pre-training, use a backbone network trained in a fully supervised or self-supervised manner on a large-scale dataset;
[0014] Step 3: During fine-tuning, adopt the ProtoNet classification head; this classification head generates a probability distribution according to the distance between the query image and the prototype in the embedding space, as shown in formula (2):
[0015]
[0016] where f φ is the backbone network that encodes the input into the feature space; c k is the prototype of class k, which is the average value of the features belonging to class k; d is the metric function; specifically, the prototypes of each class are calculated by taking the mean of each class of samples in the support set, and the augmented support set is used as the pseudo-query set, and then the loss is calculated by the cosine distance between the prototype and the pseudo-query set, and the parameters are updated;
[0017] The cross-entropy loss is selected as the loss function.
[0018] Preferably, for the self-supervised method, the DINO and MOCO v3 algorithms are used to train the backbone network on the ImageNet-1K dataset; for the fully supervised method, the backbone network is trained on the ImageNet-21K dataset.
[0019] Preferably, the cosine distance is used as the metric function.
[0020] The beneficial effects of the present invention are as follows:
[0021] (1) As a few-shot fine-tuning method, the present invention updates a small number of parameters, which is only 0.045% of the number of parameters to be updated required for full-tuning. The calculation is simple and can be achieved by element-wise multiplication and addition. Therefore, the storage and computing resources occupied are relatively small, which is conducive to the pre-trained model being put into practical application scenarios.
[0022] (2) The test results of the present invention on the four datasets of real, clipart, sketch, and quickdraw are significantly better than methods such as full-tuning, bias-tuning, and visual prompt tuning. Description of the Drawings
[0023] Figure 1 It is a schematic diagram of the transformer layer of the visual self-attention model ViT.
[0024] Figure 2 It is a schematic diagram of the transformer layer after adding the norm adapter. Detailed Implementation Manner
[0025] The present invention will be further described below in conjunction with the drawings and embodiments.
[0026] The present invention adopts a process of pre-training on a large-scale dataset and fine-tuning on a small-sample task, without training on the base-class data in the target domain. The visual self-attention model (ViT) is used as the backbone network. A common visual self-attention model consists of a patch embedding layer and N transformer layers. After passing through the patch embedding layer, the input image is encoded into a certain number of token vectors. After adding the position encoding, the input token vectors together with the CLS token are fed into the N transformer layers. Finally, after passing through the N transformer layers and a normalization layer (LayerNorm), the CLS token is used for classification or other purposes. Each transformer layer contains two normalization layers (LayerNorm), an MLP block, and a multi-head self-attention block (MHSA). Figure 1 It is the transformer layer of the visual self-attention model (ViT), corresponding to the full-tuning method. The normalization layer (LayerNorm), MLP block, and multi-head self-attention block (MHSA) in the transformer layer are all learnable.
[0027] The present invention proposes to use a learnable transformation module, consisting of two vectors, to correct the gain and bias of the Layer Normalization (LayerNorm), called "norm adapter". The "norm adapter" is located after all the normalization layers of the Vision Transformer (ViT) and scales and shifts the activation values in the same way as the gain and bias. Specifically, it is achieved through element-wise multiplication and addition, as shown in Equation (1), where Scale and Shift are the two learnable vectors of the "norm adapter", y is the output of the normalization layer, and ⊙ represents element-wise multiplication.
[0028] h = Scale ⊙ y + Shift (1)
[0029] The shapes of the parameters s1 and s2 of the "norm adapter" are the same as the gain and bias of the normalization layer, and are initialized as all-one and all-zero vectors respectively. Therefore, compared with the original pre-trained model before fine-tuning, the model with the "norm adapter" has no change in the calculation results. During fine-tuning, only the parameters Scale and Shift of the "norm adapter" are updated, and other parameters are frozen after pre-training and not optimized. Figure 2 For the transformer layer after adding the "norm adapter", corresponding to the fine-tuning method proposed by the present invention, only the parameters Scale and Shift of the "norm adapter" in the transformer layer are learnable.
[0030] During pre-training, a backbone network trained in a fully-supervised or self-supervised manner on a large-scale dataset is used. For self-supervised algorithms, the DINO and MOCO v3 algorithms are used to train the backbone network on the ImageNet-1K dataset; for fully-supervised algorithms, the backbone network is trained on the ImageNet-21K dataset.
[0031] During fine-tuning, a prototype network (ProtoNet) classification head is adopted. This classification head generates a probability distribution based on the distance between the query image and the prototype in the embedding space, as shown in Equation (2):
[0032]
[0033] f φ is the backbone network that encodes the input into the feature space. c kThe prototype for class k is the average of the features belonging to class k. d is the metric function, and the cosine distance is used here. Specifically, the prototype is calculated from the support set, and the data-augmented support set is used as the pseudo query set. Then, the loss is calculated based on the cosine distance between the prototype and the pseudo query set to update the parameters. The cross-entropy loss is chosen as the loss function.
[0034] This invention uses the Vision Transformer (ViT) as the backbone network, including ViT-Base / 16 and ViT-Small / 16. For ViT-Base / 16, we train it on the ImageNet-21K dataset using the supervised learning method and on the ImageNet-1K dataset using the MOCO-v3 algorithm to obtain the pre-trained backbone network; for ViT-Small / 16, it is trained on the ImageNet-1K dataset using the DINO algorithm.
[0035] When fine-tuning and evaluating on downstream tasks, four datasets, namely real, clipart, sketch, and quickdraw, are used. They are sub-datasets of DomainNet and contain the same class names.
[0036] During the fine-tuning and evaluation process, the 30-way 5-shot format is adopted to construct the few-shot tasks. Each task contains data of 5 classes, with 5 labeled samples and 15 query samples for each class; all images are resized to a resolution of 224*224; the random data augmentation used to generate the pseudo query set includes color jitter, horizontal flipping, and translation; there are three key hyperparameters during the fine-tuning process: learning rate, number of iterations, and optimizer. Since the samples in each task are limited, the final performance is sensitive to the choice of hyperparameters. Therefore, for various situations, the hyperparameters are selected according to the average accuracy of 50 tasks on the validation set. The optimizer is selected from Adam or SGD, and the learning rate and number of iterations are selected from the empirical range, which are [1e-1, 1e-2, 1e-3, 1e-4, 1e-5, 1e-6] and [20, 50, 80, 100] respectively; finally, 600 tasks are randomly selected from the test set for evaluation, and the average precision is calculated as the final result. All experiments use a fixed random number seed.
Claims
1. A small-sample fine-tuning method based on a visual self-attention model, characterized in that, It includes the following steps: Step 1: Construct a backbone network; Use the improved Vision Transformer (ViT) as the backbone network; The original Vision Transformer consists of a patch embedding layer and N Transformer layers. After passing through the patch embedding layer, the input image is encoded into a certain number of token vectors. After adding the position encoding, the input token vectors together with the CLS token are fed into N Transformer layers. Finally, after passing through N Transformer layers and a Layer Normalization layer (LayerNorm), the CLS token is used for classification or other purposes. Each Transformer layer contains two Layer Normalization layers, an MLP block, and a multi-head self-attention block (MHSA); Construct a learnable transformation module, which consists of two vectors, to correct the gain and bias of the Layer Normalization layer in the original Vision Transformer. This learnable transformation module is called the norm adapter. The norm adapter is located after all the Layer Normalization layers of the Vision Transformer and is implemented through element-wise multiplication and addition. As shown in formula (1), Scale and Shift are the two learnable vectors of the norm adapter, y is the output of the Layer Normalization layer, and ⊙ represents element-wise multiplication: h = Scale ⊙ y + Shift (1) The structures of the parameters Scale and Shift of the norm adapter are the same as the gain and bias parameters of the Layer Normalization layer, and are initialized as vectors of all 1s and all 0s respectively. During fine-tuning, only the parameters Scale and Shift are updated, and other parameters are frozen after pre-training and not optimized; Step 2: During pre-training, use the backbone network trained in a fully-supervised or self-supervised manner on a large-scale dataset; Step 3: During the fine-tuning process, adopt the ProtoNet classification head. This classification head generates a probability distribution based on the distance between the query image and the prototype in the embedding space, as shown in formula (2): Among them, f φ is the backbone network that encodes the input into the feature space; c k is the prototype of class k, which is the average of the features belonging to class k; d is the metric function; specifically, the prototypes of each class are calculated by taking the mean of each class of samples in the support set, and the augmented support set is used as the pseudo query set, and then the loss is calculated by the cosine distance between the prototype and the pseudo query set to update the parameters; The loss function is selected as the cross-entropy loss.
2. The few-shot fine-tuning method based on the visual self-attention model according to claim 1, wherein In the self-supervised manner, use the DINO and MOCO v3 algorithms to train the backbone network on the ImageNet-1K dataset. In the fully-supervised manner, the backbone network is trained on the ImageNet-21K dataset.
3. A few-shot fine-tuning method based on a visual self-attention model according to claim 1, characterized in that, The metric function uses the cosine distance.
Citation Information
Patent Citations
Algorithm library framework for unified small sample learning training
CN114494796A
Malicious software identification method based on visual Transform
CN115879109A