Task adaptive ViT model compression method based on microarchitecture search

By combining the microarchitecture search and multiple loss functions in the ViT model, the model structure is automatically adjusted, and the existing ViT model compression methods are solved in terms of task adaptability and efficiency, achieving more efficient model lightweighting and inference speed.

CN119992278APending Publication Date: 2025-05-13ANHUI UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411860201.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing ViT model compression method fails to fully combine the search and loss function of the microarchitecture, resulting in the need to improve the lightweight effect and inference speed, especially the problem of insufficient adaptability and efficiency under different tasks.

Method used

The task adaptive ViT model compression method based on microarchitecture search is adopted, combining target task loss, task-oriented knowledge distillation loss and model efficiency perception loss, and the best search space architecture is found through microarchitecture search, and the model structure is automatically adjusted to achieve the lightweight and efficient model.

Benefits of technology

It significantly improves the lightweight effect and inference speed of the model, while maintaining high task adaptability and prediction performance, and can automatically find suitable model structures in different downstream tasks to ensure high efficiency and performance performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992278A_ABST
    Figure CN119992278A_ABST
Patent Text Reader

Abstract

The invention discloses a task adaptive ViT model compression method based on microarchitecture search, and the method comprises the steps: pre-training a ViT model on a target data set Dt, and obtaining a fine-tuned ViTt model; and on the basis of the fine-tuned ViTt model, searching an optimal search space architecture by adopting microarchitecture search and taking minimization of a loss function as a target, and compressing the fine-tuned ViTt model into a self-adaptive ViT model of a self-adaptive target task on the basis of the optimal search space architecture. According to the adaptive ViT model provided by the invention, the lightweight of the ViT model can be adaptively carried out for different downstream tasks, and the adaptive ViT model can automatically and effectively adjust the model structure by combining microarchitecture search and a loss function (task-oriented knowledge distillation loss and model efficiency perception loss); and high task adaptability and prediction performance are kept while the model efficiency is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ViT (Vision Transformer) model compression optimization, and in particular to a task-adaptive ViT model compression method based on differentiable architecture search. Background Art

[0002] Self-attention based architectures, especially the Transformer, have become the models of choice in Natural Language Processing (NLP). Inspired by the successful application of Transformers in NLP, standard Transformers are directly applied to images with minimal modifications, achieving promising results. Although these models are very effective, they are built on large-scale datasets and often have billions of parameters. For example, the ViT-Base and ViT-Large models have 86 million and 307 million parameters, respectively, making it difficult to deploy such large-scale models in real-time applications with strict constraints on computational resources and inference time.

[0003] In recent years, with the widespread application of deep learning in the field of computer vision, large models have gradually been applied to real-time applications. In order to address this challenge, some existing solutions have begun to explore how to compress the ViT (Vision Transformer) model to reduce the amount of calculation and increase the inference speed.

[0004] Existing solutions mainly compress the ViT model into a task-independent model structure, that is, the same compressed ViT model structure is used for all different tasks. The ViT model learns various knowledge through a large-scale image library, and specific downstream tasks only need to acquire some of the knowledge. There are differences in the level of image knowledge learned by the hidden layers of different ViT models, and for different tasks, the importance of the attention head in the ViT model also changes with the change of tasks. This shows the need for a lightweight task-adaptive ViT model. Different image classification tasks utilize the ViT model in different ways, so it is necessary to compress large-scale models such as ViT separately for specific downstream tasks.

[0005] For example, the invention application with application number 202410340536.1 discloses a ViT model compression method and structure that is adaptive to different tasks. Through this application solution, the training time can be effectively shortened, storage space can be saved, and the lightweight ViT model can be achieved. However, this application solution also has the problem that it does not fully consider the combination of differentiable architecture search and loss function, and there is still room for improvement in lightweight effect and inference speed. Summary of the invention

[0006] In view of the above-mentioned problems, the purpose of the present invention is to provide a task-adaptive ViT model compression method based on differentiable architecture search, combining differentiable architecture search and loss function to improve the model lightweight effect and inference speed.

[0007] An embodiment of the present invention provides a task-adaptive ViT model compression method based on differentiable architecture search.

[0008] A first aspect: A task-adaptive ViT model compression method based on differentiable architecture search, comprising the steps of:

[0009] In the target dataset D t Pre-trained ViT model, get the fine-tuned ViT t Model; based on fine-tuned ViT t The model uses a differentiable architecture search to minimize the loss function to find the best search space architecture. Based on the best search space architecture, the fine-tuned ViT t The model is compressed into an adaptive ViT model for the adaptive target task.

[0010] As an optional implementation method, the loss function includes: target task loss, task-oriented knowledge distillation loss and model efficiency perception loss.

[0011] As an optional implementation method, the differentiable architecture search seeks the best search space architecture with the goal of minimizing the loss function, and the formula is expressed as:

[0012] L=(1-γ)L CE (α, ω α , D t )+γL KD (α, ω α , BERT t )+βL E (α)

[0013] Among them, α is the search space architecture of the differentiable architecture search, L is the loss function, and L CE , L KD and L E are the target task, task-oriented knowledge distillation, and model efficiency-aware losses, respectively. γ and β are the hyperparameters for balancing these loss terms. α are the trainable network weights of architecture α.

[0014] As an optional implementation manner, the search space architecture α is expressed as:

[0015] α={K,α c}, K≤K max

[0016] Among them, K is the number of unit stacking layers for differentiable architecture search, α c is the search space architecture parameter;

[0017] As an optional implementation mode, the search space architecture parameter α c , the formula is:

[0018] α c =[o 0,2 ,o 1,2 ,...,o i,j ,...,o N+1,N+2 ]

[0019] Among them, i,j It is the state transition operation from unit node i to node j.

[0020] As an optional implementation method, the task-oriented knowledge distillation loss is expressed as follows:

[0021]

[0022] Among them, Z t is the logarithm of the teacher model, Z s is the student model logarithm, τ is the distillation temperature, λ is the Kullback-Leibler divergence loss and the target task loss L CE The coefficient of , ψ represents the softmax function.

[0023] As an optional implementation method, the model efficiency perception loss is expressed as follows:

[0024]

[0025] Where SIZE(·) is the normalized parameter size and FLOPs(·) is the number of floating point operations per operation.

[0026] As an optional implementation method, based on the optimal search space architecture, the fine-tuned ViT t The model is compressed into an adaptive ViT model for the adaptive target task. The steps are:

[0027] S1. Model the search space architecture α as a one-hot discrete variable classification sample, expressed as:

[0028]

[0029] Among them, [1,K max ] is the one-hot variable range for searching the number of stacking layers of shared units; O is the candidate operation set of search space architecture parameters;

[0030] S2. Use the Gumbel-Softmax function to relax the discrete variable classification samples into continuous variable classification samples. The formula is:

[0031]

[0032] Among them, g i is the random noise drawn from the Gumbel (0, 1) distribution, and τ is the temperature that controls how closely Gumbel-Softmax approaches argmax.

[0033] S3. In the forward phase of the compression process, the argmax(·) function is used to process continuous variable classification samples, and in the back propagation phase, the continuous y K and o .

[0034] A second aspect: An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method provided in the first aspect when executing the program.

[0035] A third aspect: A non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method provided in the first aspect.

[0036] Beneficial effects of the present invention:

[0037] 1. The adaptive ViT model proposed in the present invention can adaptively lightweight the ViT model for different downstream tasks. By combining differentiable architecture search and loss function (task-oriented knowledge distillation loss and model efficiency perception loss), the adaptive ViT model can automatically and effectively adjust the model structure, significantly improving the model efficiency while maintaining high task adaptability and prediction performance.

[0038] 2. The present invention successfully draws on the technologies of differentiable architecture search and knowledge distillation in the field of ViT model compression, providing a unique approach that helps solve the balance problem between model lightweighting and task adaptation. It can not only significantly improve task efficiency, but also ensure the best model adaptation for different downstream tasks.

[0039] 3. In the evaluation of multiple image classification data sets, the adaptive ViT model of the present invention has excellent performance, successfully improved efficiency, and automatically found a suitable model structure in different tasks. The adaptive ViT can achieve performance similar to that of the original ViT model while ensuring high efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1It is a principle flow chart of the task adaptive ViT model compression method based on microarchitecture search of the present invention;

[0041] Figure 2 It is a principle flow chart of the micro-architecture search of the present invention;

[0042] Figure 3 It is a schematic structural diagram of the electronic device of the present invention. DETAILED DESCRIPTION

[0043] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar symbols throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0044] The adaptive compression of the existing ViT (Vision Transformer) model does not effectively combine differentiable architecture search and loss function, and there is room for improvement in task adaptability and prediction performance.

[0045] In view of the above problems, the present invention provides a task-adaptive ViT model compression method based on microarchitecture search. Figure 1 A principle flow chart of a task-adaptive ViT model compression method based on microarchitecture search provided by an embodiment of the present invention, the method comprising:

[0046] In the target dataset D t Pre-trained ViT model, get the fine-tuned ViT t Model; based on fine-tuned ViT t The model uses a differentiable architecture search to minimize the loss function to find the best search space architecture. Based on the best search space architecture, the fine-tuned ViT t The model is compressed into an adaptive ViT model for the adaptive target task.

[0047] ViT after fine-tuning t The model is a trained ViT model. As a teacher model, it usually has huge parameters. In real-time applications with strict constraints on computing resources and inference time, it is difficult to deploy such a large-scale model. Therefore, it is necessary to perform lightweight parameter compression.

[0048] The present invention aims to fine-tune ViT t The model is lightweight and compressed into a task-adaptive ViT model that is effective and efficient for a specific task, achieving lightweight parameters of the task-adaptive ViT model.

[0049] ViT after fine-tuning tThe model loss function can be: target task loss, task-oriented knowledge distillation loss and model efficiency perception loss. Differentiable architecture search is used to minimize the loss function and find the best search space architecture. The formula is expressed as:

[0050] L=(1-γ)L CE (α, ω α , D t )+γL KD (α, ω α , BERT t )+βL E (α)

[0051] Among them, α is the search space architecture of the differentiable architecture search, L is the loss function, and L CE , L KD and L E are the target task, task-oriented knowledge distillation, and model efficiency-aware losses, respectively. Specifically, L CE is the target task loss data D t The cross entropy loss in L KD is the task-oriented knowledge distillation loss, which provides hints for finding a suitable structure for the task, L E It provides model efficiency-aware loss, which helps search for lightweight and efficient model structures through constraints. γ and β are hyperparameters that balance these loss terms. α are the trainable network weights (e.g., weights of the feed-forward layers) of the differentiable architecture α.

[0052] like Figure 2 As shown in the figure, it is a schematic diagram of the principle structure of differentiable architecture search. Most differentiable architecture search methods focus on the search space based on searching shared units; that is, the search goal is to search for shared units, where the search shared unit structure parameter α c Shared for all layers.

[0053] In addition to considering the search for shared unit shared parameters α c In addition to stacking, the number of unit stacking layers K∈[1, 2, ..., K max ] to search; K is crucial to find a trade-off between model expressiveness and efficiency, as larger K leads to higher model capacity but slower inference speed.

[0054] The search space architecture α is expressed as:

[0055] α={K,α c}, K≤K max

[0056] Where K is the number of stacked layers of search sharing units for differentiable architecture search, αc is the search space architecture parameter.

[0057] like Figure 2 As shown, the search shared unit is represented as a directed acyclic graph, and each node in the unit represents a potential state h, o i,j The edge from node i to node j represents h i Convert to h j operation, O represents operation o i,j A collection of .

[0058] For the unit at the kth layer (k>1), the two input nodes c k-2 and c k-1 It is defined as a layer-by-layer residual connection, and the output node c is obtained by processing on all intermediate nodes k ; For the first layer units, nodes 0 and 1 are the task-related input embeddings.

[0059] Denote O as operation o i,j Assuming the topological order between N intermediate nodes, when i<j and j>1, there exists o i,j ∈O.

[0060] Search space architecture parameter α c , the formula is:

[0061] α c =[o 0,2 ,o 1,2 ,...,o i,j ,...,o N+1,N+2 ]

[0062] Among them, i,j It is the state transition operation from unit node i to node j.

[0063] In order to promote the adaptability of the learned structure to the target task, the present invention introduces task-oriented knowledge distillation to guide the process of structure search. Knowledge distillation is widely used in model compression methods. By allowing a small student model to learn knowledge from a large teacher model and learn knowledge from real labels to approach the performance of the teacher model, the model is compressed. During the training process, by applying knowledge distillation, the knowledge related to specific tasks in the original large ViT model can be learned and redundant knowledge can be removed.

[0064] The task-oriented knowledge distillation loss is formulated as:

[0065]

[0066] Among them, Z t For the teacher model (fine-tuned ViTt Model) logarithm, Z s is the logarithm of the student model (adaptive ViT model), τ is the distillation temperature, and λ is the balance between the Kullback-Leibler divergence loss (KL) and the cross entropy (L ce ), and ψ represents the softmax function.

[0067] The Kullback-Leibler divergence between the softmax of the teacher model and the softmax of the student model is minimized through knowledge distillation.

[0068] The purpose of the present invention is to adjust the ViT t The model is lightweight and compressed into an efficient adaptive ViT model. So, what kind of model is considered efficient? Therefore, a quantity is needed to constrain the number of parameters and inference speed of the lightweight model to ensure the efficiency of the lightweight model.

[0069] To achieve this goal, the present invention introduces efficiency-aware loss, which incorporates model efficiency from two aspects, parameter size and inference time, into the loss function to evaluate the parameter size and inference speed of the compression model.

[0070] Model efficiency perception loss function, the formula is expressed as:

[0071]

[0072] Where SIZE(·) is the normalized parameter size and FLOPs(·) is the number of floating point operations per operation.

[0073] The model efficiency perception loss mainly consists of two items. The first item is the SIZE(·) function, which is used to calculate the normalized parameter size of each compression operation; the second item is the FLOPs(·) function, which is used to calculate the number of floating-point operations. Here, the sum of FLOPs operations is used as an approximation of the actual inference time of the compression model.

[0074] like Figure 2 The space of differentiable architecture search utilized by the method of the present invention is shown. In order to make it easy to stack units and search network layers, the input and output nodes of each layer are kept in the same shape.

[0075] For candidate operations in a unit, lightweight CNN-based operations can be adopted. CNN is effective in computer vision tasks. CNN-based operations have an advantage in inference speed over RNN-based models and self-attention-based models, mainly because CNN is a parallel-friendly operation.

[0076] Specifically, the candidate operation set O includes convolution, pooling, skip connection, and zero operation setting search space; for convolution operations, it includes 1D standard convolution and dilated convolution with kernel size {3, 5, 7}, where dilated convolution can be used to enhance the ability to capture long dependency information, and each convolution is used as a Relu-Conv-BatchNorm structure; pooling operations include average pooling and maximum pooling with kernel size of 3; skip connections and zero operations are used to construct residual connections and discard operations, respectively, which helps to reduce network redundancy. In addition, for convolution and pooling operations, "SAME" padding is applied to make the output length the same as the input.

[0077] Directly optimizing the total loss function L by brute-force enumeration of all candidate operations is difficult due to the huge search space of combined operations and time-consuming training.

[0078] The present invention transforms the search space architecture α={K,α c}, K≤K max Modeling as a one-hot discrete variable classification sample, subject to discrete probability distribution, to solve this problem, the formula is expressed as:

[0079]

[0080] Among them, [1,K max ] is the search range of shared unit stacking layers; O is the search space architecture parameter candidate operation set.

[0081] K and o i,j The classification samples are modeled as one-hot discrete variables and are respectively from the layer range [1,K max ] and sample from the candidate operation set O; because the loss function L is not differentiable, the discrete sampling process prevents the gradient from backpropagating back to the learnable parameters P K and P O .

[0082] Therefore, the present invention uses the Gumbel-Softmax function to relax the classification samples into a continuous sample vector and o ∈R |O| , the formula is:

[0083]

[0084] Among them, g i is the random noise extracted from the Gumbel (0, 1) distribution, and τ is the temperature that controls the degree to which Gumbel-Softmax approaches argmax, that is, when τ approaches 0, the sample becomes one-hot. In this way, y K and ois a differentiable proxy variable for discrete samples, and can directly use gradient information to effectively optimize L.

[0085] Specifically, in the forward phase of the compression process, the one-hot sample vector argmax(y K ) and argmax(y O ), while continuous y is used in the back-propagation phase K and o , and make the forward process in training consistent with testing.

[0086] The present invention also provides a task-adaptive ViT model compression device based on microarchitecture search, such as Figure 1 As shown, the device comprises:

[0087] Fine-tuning module, using the target dataset D t Pre-train ViT model and obtain fine-tuned ViT t Model;

[0088] Differentiable search module, based on fine-tuned ViT t Model loss function, using differentiable architecture search to find the best search space architecture;

[0089] The adaptive module,utilizes the optimal search space architecture to form an adaptive,ViT model for the adaptive target task.

[0090] The present invention also provides an electronic device, Figure 3 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, such as Figure 3 As shown, the electronic device may include: a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory communicate with each other via the communications bus. The processor may call the logic instructions in the memory, for example, to execute the following method:

[0091] In the target dataset D t Pre-trained ViT model, get the fine-tuned ViT t Model; based on fine-tuned ViT t The model uses a differentiable architecture search to minimize the loss function to find the best search space architecture. Based on the best search space architecture, the fine-tuned ViT t The model is compressed into an adaptive ViT model for the adaptive target task.

[0092] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0093] An embodiment of the present invention further provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in each of the above embodiments is implemented, for example, including:

[0094] In the target dataset D t Pre-trained ViT model, get the fine-tuned ViT t Model; based on fine-tuned ViT t The model uses a differentiable architecture search to minimize the loss function to find the best search space architecture. Based on the best search space architecture, the fine-tuned ViT t The model is compressed into an adaptive ViT model for the adaptive target task.

[0095] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0096] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A task-adaptive ViT model compression method based on differentiable architecture search, characterized in that: include: In the target dataset D t Pre-trained ViT model, get the fine-tuned ViT t Model; based on fine-tuned ViT t The model uses a differentiable architecture search to minimize the loss function to find the best search space architecture. Based on the best search space architecture, the fine-tuned ViT t The model is compressed into an adaptive ViT model for the adaptive target task.

2. The compression method according to claim 1, characterized in that: The loss function includes: target task loss, task-oriented knowledge distillation loss and model efficiency perception loss.

3. The compression method according to claim 2, characterized in that: The differentiable architecture search aims to find the best search space architecture by minimizing the loss function, and the formula is expressed as: L=(1-γ)L CE (oh, oh) α ,D t )+γL KD (oh, oh) α , BERT t )+βL E (a) Among them, α is the search space architecture of the differentiable architecture search, L is the loss function, and L CE , L KD and L E are the target task, task-oriented knowledge distillation, and model efficiency-aware losses, respectively. γ and β are the hyperparameters for balancing these loss terms. α are the trainable network weights of the search space architecture α.

4. The compression method according to claim 3, characterized in that: The search space architecture α is formulated as: α={K,α c },K≤K max Among them, K is the number of unit stacking layers for differentiable architecture search, α c is the search space architecture parameter.

5. The compression method according to claim 4, characterized in that: The search space architecture parameter α c , the formula is: α c =[the 0,2 ,the 1,2 ,...,the i,j ,...,the N+1,N+2 ] Among them, i,j It is the state transition operation from unit node i to node j.

6. The compression method according to claim 2, characterized in that: The task-oriented knowledge distillation loss is formulated as: Among them, Z t is the logarithm of the teacher model, Z s is the student model logarithm, τ is the distillation temperature, λ is the Kullback-Leibler divergence loss and the target task loss L CE The coefficient of , ψ represents the softmax function.

7. The compression method according to claim 2, characterized in that: The model efficiency perception loss is expressed as: Where SIZE(·) is the normalized parameter size and FLOPs(·) is the number of floating point operations per operation.

8. The compression method according to claim 4, characterized in that: Based on the optimal search space architecture, the fine-tuned ViT t The model is compressed into an adaptive ViT model for the adaptive target task. The steps are: S1. Model the search space architecture α as a one-hot discrete variable classification sample, expressed as: Among them, [1,K max ] is the one-hot variable range for searching the number of stacking layers of shared units; O is the candidate operation set of search space architecture parameters; S2. Use the Gumbel-Softmax function to relax the discrete variable classification samples into continuous variable classification samples. The formula is: Among them, g i is the random noise extracted from the Gumbel (0, 1) distribution, τ is the temperature that controls the degree to which Gumbel-Softmax approaches argmax; S3. In the forward phase of the compression process, the argmax(·) function is used to process continuous variable classification samples, and in the back propagation phase, the continuous y K and o .

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • ViT model compression method and structure adaptive to different tasks

    CN118172640A