Fine adjustment method of visual language model

Through the jump fine-tuning method, the hierarchical jump and category jump mechanisms are used to solve the problems of performance degradation and low resource efficiency of the visual language model in zero-sample generalization tasks, and the effect of efficient adaptation to downstream tasks and improving generalization performance is achieved.

CN120148037APending Publication Date: 2025-06-13UESTC (SHENZHEN) ADVANCED RES INST
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510216413.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The performance of existing visual language models in zero-sample generalization tasks is degraded, and the traditional full-parameter fine-tuning method has low resource efficiency, making it difficult to meet the practical application needs.

Method used

A jump fine-tuning method is proposed to cut the length and width of the gradient through the hierarchical jump and category jump mechanism, reduce the computational complexity in model training, and adapt to downstream tasks.

Benefits of technology

Improves fine-tuning efficiency and model performance, reduces memory and time overhead, and enhances the generalization performance of visual language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148037A_ABST
    Figure CN120148037A_ABST
Patent Text Reader

Abstract

The invention discloses a fine tuning method for a visual language model, which comprises the following steps of: respectively dividing an image coding module and a text coding module in the visual language model into a shallow layer and a deep layer, respectively inputting each image and category mark into the corresponding shallow layer module to obtain an intermediate feature, and deleting the shallow layer module to obtain a fine tuning visual language model; in each fine tuning iteration, respectively inputting the middle features of the images and the category marks into corresponding deep modules to obtain image features and category mark features, and screening through the similarity of the image features and the category mark features to obtain an alternative category mark set of each image; and then performing fine tuning training on the fine tuning visual language model according to the image and the corresponding alternative category mark, and merging the shallow layer module with the current fine tuning visual language model after the fine tuning is finished to obtain a visual language model after the fine tuning is finished. According to the method, the length and width of the feature-gradient propagation flow in the model training process can be effectively reduced, and efficient adaptation to downstream tasks is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and more specifically, relates to a method for fine-tuning a vision-language model. Background Art

[0002] In recent years, with the rapid development of artificial intelligence and deep learning technologies, large-scale pre-trained vision-language models (VLMs) have made remarkable progress in the field of the combination of computer vision and natural language processing. These models, by jointly learning image and text features, have demonstrated excellent performance in open-domain visual concept recognition and cross-modal tasks. For example, the CLIP model uses an image-text matching (ITM) loss to align images with their corresponding text descriptions into a common feature space, which has become an important breakthrough in the VLM field.

[0003] Although VLMs perform excellently in open visual concept understanding, their performance in zero-shot generalization tasks has significant limitations. When the categories, distributions, or domains of downstream tasks are different from those of upstream training data, the generalization performance of VLMs drops significantly. In addition, how to efficiently adapt large-scale pre-trained models to specific downstream tasks is a core issue in current research and applications. Traditional full fine-tuning (FT) methods require training all the parameters of the model, resulting in high memory occupancy, large computational overhead, and a time-consuming training process. Although this method has a relatively high upper limit in task performance, its resource efficiency is low and it is difficult to meet the requirements of practical applications.

[0004] To solve the above problems, prompt tuning (PT), as a lightweight transfer learning method, has received extensive attention. Prompt tuning learns task-specific prompts, that is, a small group of context vectors, and only needs to train these context vectors while fixing the remaining parameters of the pre-trained model, thus significantly reducing the number of parameters to be optimized. However, prompt tuning has limited effects on improving memory usage and time efficiency, and its performance is lower than that of full fine-tuning in some scenarios. In addition, since prompt tuning freezes most of the weights of the model, it cannot fully utilize pre-trained knowledge, resulting in limited knowledge transfer efficiency and task performance.

[0005] Currently, existing tuning methods have not achieved a good balance among performance, memory usage efficiency, and time overhead. In the vast majority of application scenarios, the memory efficiency and time efficiency of the model are often more important than the parameter efficiency. Compared with the cost of storing parameters, the computing and storage resources used for training the model are much more expensive. Therefore, there is an urgent need for a new type of tuning method that can achieve higher resource utilization efficiency while ensuring the knowledge transfer effect and task performance. Summary of the Invention

[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a fine-tuning method for a vision-language model. By proposing a new tuning paradigm, it effectively reduces the length and width of the Feature-Gradient Propagation Flows (FGPFs) during the model training process, achieving efficient adaptation to downstream tasks.

[0007] To achieve the above-mentioned invention purpose, the fine-tuning method for the vision-language model of the present invention includes the following steps:

[0008] S1: Obtain a data sample set for fine-tuning the vision-language model according to actual needs. Each data sample includes an image x i and a class label y i , where i = 1, 2,..., N, and N represents the number of data samples; the class label y i ∈{c 1 , c 2 ,..., c M}, where c j represents the class marker of the j-th class, and j = 1, 2,..., M, and M represents the number of classes in the data sample set;

[0009] S2: Divide the image encoding module V and the text encoding module T in the vision-language model into shallow and deep parts respectively according to actual needs, obtaining a shallow image encoding module V 1:ω , a deep image encoding module V ω+1:L and a shallow text encoding module T 1:ω , a deep text encoding module T ω+1:L , where ω represents the dividing boundary layer, and L represents the total number of layers of the image encoding module and the text encoding module;

[0010] S3: Input the image x i in each data sample in the fine-tuning data sample set into the shallow image encoding module V 1:ω in the vision-language model, and save the intermediate image features 1:ω output by the shallow image encoding module V Input each class marker c j into the shallow text encoding module T in the vision-language model1:ω to save the intermediate category token features output by T in the shallow text encoding module 1:ω

[0011] S4: Save the shallow image encoding module V 1:ω and the shallow text encoding module T 1:ω , and then delete these two shallow modules from the vision - language model to obtain a fine - tuned vision - language model;

[0012] S5: Input each intermediate category token feature into the deep text encoding module T in the current fine - tuned vision - language model ω+1:L to obtain the corresponding category token text feature C j = T ω+1:L (c j );

[0013] S6: Let the fine - tuning iteration number s = 1, and initialize the deep text encoding module deep image encoding module

[0014] S7: For each image x i , input its intermediate image feature into the deep image encoding module in the current fine - tuned vision - language model to obtain the corresponding image feature

[0015] S8: Calculate the similarity between the image feature and each category token text feature C j , and set the sampling probability for each category token c i of the image x j according to the similarity; the greater the similarity between the image feature and the category token text feature , the greater the sampling probability ; then sample m category tokens according to the sampling probability to form the set of candidate category tokens for the image x i in this round; the value of m is set according to actual needs;

[0016] S9: Input the intermediate category token features of the category tokens in the set of candidate category tokens i for each image x into the deep text encoding module of the current fine - tuned vision - language model to obtain the corresponding text feature ​​​Calculate the loss according to a preset loss function based on each image x i of the image features and the text features corresponding to each alternative class label and then update the parameters of the fine-tuned vision-language model. Denote the updated deep image encoding module as the deep text encoding module as

[0017] S10: Determine whether the fine-tuning end condition is satisfied. If so, go to step S11; otherwise, go to step S12;

[0018] S11: Let the iteration number s = s + 1, and return to step S7;

[0019] S12: Merge the shallow image encoding module V 1:ω and the shallow text encoding module T 1:ω with the current fine-tuned vision-language model to obtain the fine-tuned vision-language model.

[0020] For the fine-tuning method of the vision-language model of the present invention, the image encoding module and the text encoding module in the vision-language model are respectively divided into shallow and deep layers. Each image and class label are respectively input into the corresponding shallow module to obtain intermediate features. The shallow module is removed to obtain the fine-tuned vision-language model. In each fine-tuning iteration, the intermediate features of the image and class label are respectively input into the corresponding deep module to obtain image features and class label features. The set of alternative class labels for each image is obtained by screening based on the similarity of the image features and class label features. Then, the fine-tuned vision-language model is fine-tuned and trained according to the image and the corresponding alternative class labels. After the fine-tuning is completed, the shallow image encoding module and the shallow text encoding module are merged with the current fine-tuned vision-language model to obtain the fine-tuned vision-language model.

[0021] The present invention has the following beneficial effects:

[0022] 1) The present invention proposes a new Skip Tuning, which effectively reduces the computational complexity in model training by two mechanisms, Layer-wise Skipping (LSkip) and Class-wise Skipping (CSkip), by effectively trimming the length and width of the gradient, thus better adapting the vision-language model to downstream tasks, improving both the fine-tuning efficiency and ensuring the model performance;

[0023] 2) The present invention does not need to introduce additional context vectors or adaptation modules, and while maintaining parameter efficiency, it greatly improves memory and time efficiency, providing a new solution for the transfer of large-scale pre-trained vision-language models;

[0024] 3) Through experimental verification, the present invention has good generalization on different data sets in multi-tasks, effectively improving the generalization performance of the vision-language model. Description of the Drawings

[0025] Figure 1 is a flowchart of the specific implementation of the fine-tuning method of the vision-language model of the present invention;

[0026] Figure 2 is a schematic diagram of the model change of the fine-tuning method of the vision-language model in the present invention. Specific Embodiments

[0027] The following describes the specific embodiments of the present invention with reference to the drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.

[0028] Embodiment

[0029] Figure 1 is a flowchart of the specific implementation of the fine-tuning method of the vision-language model of the present invention. As Figure 1 shown, the specific steps of the fine-tuning method of the vision-language model of the present invention include:

[0030] S101: Obtain a fine-tuning data sample set:

[0031] According to actual needs, obtain a data sample set for fine-tuning the vision-language model. Each data sample includes an image x i and a class label y i , i = 1, 2,..., N, where N represents the number of data samples. The class label y i ∈{c 1 , c 2 ,..., c M}, and c j represents the class label of the jth class, j = 1, 2,..., M, where M represents the number of classes in the data sample set. In this embodiment, the format of the class label is "a photo of a[CLS]", and CLS represents the word corresponding to the class.

[0032] S102: Divide the vision-language model:

[0033] According to actual needs, divide the image encoding module V and the text encoding module T in the vision-language model into shallow and deep parts respectively, to obtain a shallow image encoding module V 1:ω , a deep image encoding module V ω+1:L and a shallow text encoding module T 1:ω , a deep text encoding module T ω+1:L, where ω represents the dividing line layer and L represents the total number of layers of the image encoding module and the text encoding module.

[0034] S103: Generate intermediate features:

[0035] For each data sample x in the fine-tuning data sample set i Input it into the shallow image encoding module V in the vision-language model 1:ω , and save the intermediate image features 1:ω output by the shallow image encoding module V Input each class label c j into the shallow text encoding module T in the vision-language model 1:ω , and save the intermediate class label features 1:ω output by the shallow text encoding module T The present invention saves the above intermediate features and uses them as inputs to update the parameters of subsequent deep layers in the fine-tuning iteration.

[0036] S104: Construct the fine-tuning vision-language model:

[0037] Save the shallow image encoding module V 1:ω and the shallow text encoding module T 1:ω , and then delete these two shallow modules from the vision-language model to obtain the fine-tuning vision-language model.

[0038] S105: Generate class label text features:

[0039] Input each intermediate class label feature into the deep text encoding module T in the current fine-tuning vision-language model ω+1:L to obtain the corresponding class label text feature C j = T ω+1:L (c j ).

[0040] S106: Initialize the fine-tuning parameters:

[0041] Let the fine-tuning iteration number s = 1, and initialize the deep text encoding module the deep image encoding module

[0042] S107: Generate the image features for this time:

[0043] For each image x i , input its intermediate image features into the deep image encoding module in the current fine-tuning vision-language model to obtain the corresponding image features

[0044] S108: Screen candidate class labels:

[0045] To avoid over-reliance on the same class subset in multiple rounds of training, the present invention will screen the class labels corresponding to each image in the data sample. The specific method is as follows:

[0046] Calculate the image features and the text features C of each class label j The similarity is calculated, and according to the similarity, the sampling probability for each class label c i of the image x j is set. The greater the similarity between the image features and the text features of the class label , the greater the sampling probability . Then, m class labels are sampled according to the sampling probability to form the set of candidate class labels for the image x i in this round. The value of m is set according to actual needs.

[0047] In this embodiment, the sampling probability is calculated according to the following formula:

[0048]

[0049] Among them, represents the serial number corresponding to the class label c j after sorting according to the similarity in this iteration, r represents a preset proportional parameter, and r×M<m is satisfied. represents a preset attenuation function. In this embodiment, the attenuation function e represents the natural constant.

[0050] When sampling with the sampling probability calculated by the above formula, the top r×M most relevant class labels are selected, and at the same time, a certain probability is used to sample from the remaining class labels to increase the diversity of the training data.

[0051] S109: Update the parameters of the fine-tuned vision-language model:

[0052] Input the intermediate class label features of the class labels in the set of candidate class labels i for each image x into the deep text encoding module of the current fine-tuned vision-language model to obtain the corresponding text features Adopt a preset loss function according to the image features i of each image x and text features corresponding to each alternative class label Calculate the loss, and then update the parameters of the fine-tuned vision-language model. Denote the updated deep image encoding module as the deep text encoding module as The calculation formula of the loss function can be set according to actual needs. In this embodiment, the image-text matching loss (Image-Text Matching Loss) L is adopted ITM 。

[0053] S110: Determine whether the fine-tuning end condition is satisfied. If so, enter step S111; otherwise, enter step S110. The fine-tuning end condition can be set according to actual needs. Generally, it can be set to whether the maximum fine-tuning iteration number is reached, or the model parameters converge

[0054] S111: Let the iteration number s = s + 1, and return to step S107

[0055] S112: Obtain the fine-tuned vision-language model

[0056] Merge the shallow image encoding module V 1:ω and the shallow text encoding module T 1:ω with the current fine-tuned vision-language model to obtain the fine-tuned vision-language model

[0057] According to the above description, it can be known that the fine-tuning method of the vision-language model of the present invention is a SkipTuning method, which includes two Skip methods: Layer-wise Skipping (LSkip) and Class-wise Skipping (CSkip) Figure 2 is a schematic diagram of the model change of the fine-tuning method of the vision-language model in the present invention. As Figure 2 shown, the two Skip methods in the present invention can be summarized as follows

[0058] Layer-wise Skipping is to divide the vision-language model into shallow and deep parts, delete the shallow image encoding module and the shallow text encoding module, and update the parameters of the remaining part. Its goal is to reduce the computational amount of Full Gradient Forward Propagation (FGFP) without significantly reducing the model performance, thereby improving the fine-tuning training efficiency

[0059] Class skipping dynamically filters class labels for each image in the data samples during each fine-tuning iteration, filtering out class labels that are irrelevant to loss calculation, avoiding the model's dependence on the same subset of classes in different training rounds, enhancing data diversity and generalization performance. And only K class labels are sampled for each image for image-text matching, thus significantly reducing the computational overhead of the text encoder.

[0060] The fine-tuning method obtained by combining the above two Skip methods realizes the efficient adaptation of the vision-language model to specific downstream tasks, thus better meeting the requirements of downstream tasks.

[0061] To better illustrate the technical effects of the present invention, specific examples are used to experimentally verify the present invention. In this embodiment, 9 algorithms that were relatively excellent in the field of vision-language model training at that time were selected as comparison methods, namely: CoOp, CoCoOp, ProGrad, KgCoOp, MaPLe, PromptSRC, TCP, DePT, CoPro. This embodiment was tested on 7 publicly available datasets, including datasets such as ImageNet, OxfordPets, and Caltech, covering tasks such as generalization from basic to new tasks, cross-dataset transfer, domain generalization, and few-shot learning. In addition, for the domain generalization task, ImageNet was used as the source dataset, and its 4 variants (ImageNetV2, ImageNet-Sketch, ImageNet-A, and ImageNet-R) were used as target datasets.

[0062] In this embodiment, the present invention uses ViT-B / 16 as the backbone network of the CLIP model; uses the SGD optimizer with a learning rate of 2e -5 , a batch size of 4. In the tasks of generalization from basic to new tasks, cross-dataset transfer, and domain generalization, the number of training rounds is set to 20, and in the few-shot learning task, the number of training rounds is set to 40; the hyperparameters ω, r, and λ are tuned in the experiment. All experiments in this embodiment are completed on an NVIDIA V100 GPU, and the experimental results are the average of three runs.

[0063] Table 1 is a comparison table of the generalization results from the benchmark to new data of the present invention and the comparison methods on 7 datasets in this embodiment.

[0064]

[0065]

[0066] Table 1

[0067] Table 2 is a comparison table of the cross-dataset generalization results of the present invention and the comparison methods on 11 datasets in this embodiment.

[0068]

[0069] Table 2

[0070] Table 3 is a comparison table of the domain generalization results of the present invention and the comparative method on ImageNet in this embodiment.

[0071]

[0072] Table 3

[0073] In Tables 1 to 3, Base represents the average accuracy on the base classes in the training set, New represents the average accuracy on the novel classes not included in the training set, and H represents the harmonic mean of the two, which is used to measure the comprehensive performance of the model on the base classes and novel classes. The calculation method is H = 2 / (1 / base + 1 / new). Avg ACC represents the average accuracy, and Cost represents the fine-tuning time. As shown in Tables 1 to 3, the fine-tuning method proposed by the present invention shows significant advantages in the generalization from the base task to the new task, cross-dataset generalization, and domain generalization, as follows:

[0074] 1. Generalization ability from the base task to the new task

[0075] As shown in Table 1, the method of the present invention has better performance than the prior art on 7 datasets. Compared with the state-of-the-art PromptSRC method, the generalization performance of the method of the present invention on the new task has achieved a 0.7% improvement in H-ACC, and at the same time, the time efficiency has been improved by 7.2 times, and the memory efficiency has been improved by 5.1 times. In addition, the present invention can achieve the best performance on both the base task and the new task, solving the problem of accuracy trade-off between the base task and the new task in the prior art, thereby effectively alleviating the overfitting problem during the migration of large-scale visual models.

[0076] 2. Cross-dataset generalization ability

[0077] As shown in Table 2, the method of the present invention also shows excellent performance in cross-dataset generalization. In the experiment with ImageNet as the source dataset and the other datasets as the target datasets, compared with the PromptSRC method, the accuracy of the method of the present invention on the source dataset and the target dataset has been improved by 1.5% and 1.2% respectively. At the same time, its time efficiency has been improved by 44 times, and the memory efficiency has been improved by 21 times. The experimental results show that the method of the present invention has significant robustness in dealing with the distribution drift problem and can maintain the performance stability of the source dataset while improving the performance of the target dataset.

[0078] 3. Domain generalization ability

[0079] As shown in Table 3, in the application scenario of domain generalization, the present invention uses ImageNet as the source domain and other ImageNet variants as the target domain for verification. Compared with the PromptSRC method, the accuracy of the method of the present invention in the source domain and the target domain is increased by 1.5% and 0.55% respectively, and the time efficiency and memory efficiency are increased by 44 times and 21 times respectively. In addition, the performance of the method of the present invention on all target domains is better than that of the prior art, and it will not damage the performance of the source domain, further proving the effectiveness and robustness of the present invention in dealing with the domain drift problem.

[0080] In summary, the present invention shows significant advantages in generalization performance, time efficiency and memory efficiency, can effectively solve the overfitting and distribution drift problems in the model migration process of the prior art, and has broad application prospects and practical values.

[0081] Although the above description of the illustrative specific embodiments of the present invention is provided for the convenience of those skilled in the art of the present technology to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.

Claims

1. A method for fine-tuning a visual language model, characterized in that: The following steps are involved: S1: Get a data sample set for fine-tuning the visual language model according to actual needs, each data sample includes an image x i and the class label y i , i = 1, 2, ..., N, N represents the number of data samples; category label y i ∈{c1,c2,…,c M },c j Represents the category label of the jth category, j = 1, 2, ..., M, M represents the number of categories in the data sample set; S2: According to actual needs, the image encoding module V and text encoding module T in the visual language model are divided into shallow and deep parts respectively, and the shallow image encoding module V is obtained. 1:ω , Deep Image Coding Module V ω+1:L and shallow text encoding module T 1:ω , Deep Text Encoding Module T ω+1:L , where ω represents the dividing boundary layer, and L represents the total number of layers of the image encoding module and the text encoding module; S3: The image x in each data sample in the fine-tuning data sample set i Input to the shallow image encoding module V in the visual language model 1:ω , save the shallow image encoding module V 1:ω The output intermediate image features Label each category c j Input to the shallow text encoding module T in the visual language model 1:ω , save the shallow text encoding module T 1:ω The output intermediate category label features S4: Save the shallow image encoding module V 1:ω and shallow text encoding module T 1:ω , and then delete these two shallow modules from the visual language model to obtain a fine-tuned visual language model; S5: Label each intermediate category with features Input the deep text encoding module T in the current fine-tuned visual language model ω+1:L , get the corresponding category label text feature C j =T ω+1:L (c j ); S6: Set the number of fine-tuning iterations s = 1 and initialize the deep text encoding module Deep Image Coding Module S7: For each image x i , its intermediate image features Input to the deep image encoding module in the current fine-tuned visual language model Get the corresponding image features S8: Calculate image features And each category label text feature C j The similarity of image x is set according to the similarity. i Each category is labeled c j The sampling probability Image features and category label text feature C j s The greater the similarity, the greater the sampling probability The larger the sampling probability, the Sample m category labels to form image x i The set of candidate category tags for this round The value of m is set according to actual needs; S9: Transform each image x i Alternative category tag set Intermediate category label features for medium category labels Input to the deep text encoding module of the current fine-tuned visual language model Get the corresponding text features Use the preset loss function to calculate the loss of each image x i Image features and the text features corresponding to each candidate category label Calculate the loss, then update the parameters of the fine-tuned visual language model, and record the updated deep image encoding module as The deep text encoding module is S10: Determine whether the fine-tuning end condition is met, if yes, go to step S11, otherwise go to step S12; S11: Set the number of iterations s=s+1, and return to step S7; S12: The shallow image encoding module V 1:ω and shallow text encoding module T 1:ω Merge with the current fine-tuned visual language model to obtain the fine-tuned visual language model.

2. The method for fine-tuning a visual language model according to claim 1, characterized in that: The format of the category tag in step S1 is "a photo of a[CLS]", where CLS represents the word corresponding to the category.

3. The method for fine-tuning a visual language model according to claim 1, characterized in that: The sampling probability in step S8 Calculate according to the following formula: in, Indicates the category label c after sorting by similarity in this iteration j The corresponding serial number, r represents the preset ratio parameter, and satisfies r×M<K, Represents the preset decay function.

4. The method for fine-tuning a visual language model according to claim 3, characterized in that: The decay function e represents a natural constant.

5. The method for fine-tuning a visual language model according to claim 1, characterized in that: The loss function in step S7 adopts image-text matching loss.

Citation Information

Cited By

  • Dynamic scene generation methods, electronic devices and computer-readable storage media

    CN122574174A