Continuous visual instruction fine tuning method based on separable mixed low-rank adaptation
By adopting separable hybrid low-rank adaptation technology and adaptive fusion technology in the continuous visual instruction fine-tuning method, the dual catastrophic forgetting problem in the continuous visual instruction fine-tuning scenario is solved, and the efficient and stable performance of the model in the multi-task scenario is achieved.
Patent Information
- Application Number
- CN202510288601.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The prior art has dual catastrophic forgetting problems in the continuous visual instruction fine-tuning scenario, including the forgetting of visual understanding and the forgetting of instruction following, resulting in the model being unable to correctly identify the input image content and the inability to correctly give the answers to text instructions.
Using a continuous visual instruction fine-tuning method based on separable hybrid low-rank adaptation, the low-rank adaptation matrix is dynamically selected using separable routing technology, and the visual understanding and instruction follow output is processed through adaptive fusion technology.
The dual catastrophic forgetting problem is effectively solved, ensuring that the model maintains efficient and stable performance between multiple tasks, reducing waste of computing resources, and improving the flexibility and adaptability of the model.
Smart Images

Figure CN120219926A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of model deep learning and continual learning, and in particular, to a continual visual instruction fine-tuning method based on separable hybrid low-rank adaptation. Background Art
[0002] The prior art has problems such as serious resource consumption, flexible usage scenarios, and weak forgetting ability, mainly including the following points: The regularization method has a weak forgetting ability, and for models with a larger number of parameters, it requires huge training resources and is difficult to effectively achieve reliable continual instruction fine-tuning; The memory replay method needs to save part of the data of each past task, greatly increasing the memory resource occupation, which is not feasible in actual application scenarios; Hybrid low-rank adaptation usually uses the calculation of a routing network to allocate different low-rank adaptation matrices for inputs. In the scenario of continual visual instruction fine-tuning, the continuous dynamic changes of tasks make it impossible for the routing network to effectively allocate the most reasonable low-rank adaptation matrix for inputs, resulting in the problem of catastrophic forgetting.
[0003] In view of the above drawbacks, two catastrophic forgetting phenomena in the scenario of continual visual instruction fine-tuning are first noted: the forgetting of visual understanding, that is, the model cannot correctly recognize the content in the input picture; the forgetting of instruction following, that is, for the input text instruction, the model cannot correctly give the corresponding correct answer format. Therefore, it is urgent to effectively alleviate the double forgetting problem in continual visual instruction fine-tuning. Summary of the Invention
[0004] The purpose of the present invention is to provide a continual visual instruction fine-tuning method based on separable hybrid low-rank adaptation, which uses hybrid low-rank adaptation and separable routing technologies to efficiently solve the double catastrophic forgetting problem without sacrificing the model's previous learning ability.
[0005] To achieve the above purpose, the present invention provides a continual visual instruction fine-tuning method based on separable hybrid low-rank adaptation, including the following steps:
[0006] S1. Create a pre-trained model M and a low-rank adaptation module containing several adaptation matrices W i ;
[0007] S2. Input the continual visual instruction fine-tuning dataset sequence T = {t1, t2,..., t n-1 , t n} into the pre-trained model M, where n ∈ N;
[0008] S3. Select a suitable low-rank adaptation matrix according to the continual visual instruction fine-tuning dataset t i and perform fine-tuning, where i ∈ n;
[0009] S4. Obtain the visual understanding routing selection G of S3 using the separable routing technology vu (Z uv ) and the instruction following routing selection G if (Z if );
[0010] S5. Through G in S4 vu (Z uv ) and G if (Z if ) and its corresponding adaptability matrix W i Perform weighted operations to obtain two corresponding outputs
[0011] S6. Use the adaptive fusion technology to obtain the final output of the low-rank adaptation module
[0012]
[0013] Preferably, the adaptability matrix W of the low-rank adaptation module described in S1 i Is updated by the following formula:
[0014]
[0015] Where G(z) k Represents the weight assigned by the routing network to the k-th adaptability matrix based on the input z.
[0016] Preferably, the continuous visual instruction fine-tuning dataset sequence in S2 includes an image feature dataset, a video feature dataset, and a text feature dataset.
[0017] Preferably, the visual understanding routing selection G in S4 vu (Z uv ) is calculated as follows:
[0018] G vu (Z uv ) = softmax(top k (R uv ·Avg(x l-1 ))); (2)
[0019] Where R uv Represents the trainable routing matrix saved in the visual understanding module, x l-1 Is the input of this module, softmax is the normalization function, l is the current layer number of the model, Z uv Is the input of the visual understanding routing selection, top k Is to only take the top k largest values and set other values to 0, Avg is the average value.
[0020] Preferably, in S4, the instruction following routing selects G if (Z if ) is calculated as follows:
[0021] G if (Z if ) = softmax(top k (R if ·f(X ins ))); (3)
[0022] where X ins is the instruction for the current visual instruction fine-tuning task, Z if is the input of the instruction following routing, R if is the weight matrix, and f is the embedding operation.
[0023] Preferably, the calculation process of the output in S5 is as follows: The calculation process of the output in S5 is as follows:
[0024]
[0025] Preferably, the calculation process of the output in S5 is as follows: The calculation process of the output in S5 is as follows:
[0026]
[0027] Preferably, the adaptive fusion technology fusion algorithm in S6 is as follows:
[0028]
[0029] where α and β are weight coefficients learned during the training process.
[0030] Therefore, the present invention adopts a continuous visual instruction fine-tuning method based on separable hybrid low-rank adaptation as described above. Compared with the prior art, it has the following beneficial effects:
[0031] 1. By dividing the task into two independent modules of visual understanding and instruction following, interference between tasks is avoided, and each module can be ensured to focus on the optimization of specific tasks, effectively solving the two aspects of catastrophic forgetting problems;
[0032] 2. By dynamically selecting and allocating low-rank adaptation modules, the adaptation ability of the model is adjusted according to the needs of different tasks, with high flexibility and efficiency. It not only reduces the waste of computing resources but also can efficiently adapt to multiple tasks, ensuring that the model maintains an efficient and stable performance among multiple tasks;
[0033] 3. Through specific model architectures and routing optimization strategies, the double catastrophic forgetting problem is effectively alleviated, ensuring that the model can continuously learn new tasks while retaining the knowledge and performance of old tasks.
[0034] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0035] Figure 1 is the overall framework diagram of a continuous visual instruction fine-tuning method based on separable hybrid low-rank adaptation of the present invention;
[0036] Figure 2 is the separable hybrid low-rank adaptation module diagram of a continuous visual instruction fine-tuning method based on separable hybrid low-rank adaptation of the present invention. Detailed Embodiments
[0037] Embodiment
[0038] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of the present invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0039] As Figure 1 - Figure 2 shown, a continuous visual instruction fine-tuning method based on separable hybrid low-rank adaptation of the present invention includes the following steps:
[0040] S1. Create a pre-trained model M and a low-rank adaptation module containing several adaptation matrices W i ; The adaptation matrix W i is updated by the following formula:
[0041]
[0042] where G(z) k represents the weight assigned to the kth adaptation matrix by the routing network based on the input z; The adaptation matrix W i varies with the pre-trained model M, and the dynamic nature of the task requires a more adaptive routing strategy to avoid loss of knowledge.
[0043] S2. Input the continuous visual instruction fine-tuning data set sequence T = {t1, t2,..., t n-1 , t n} into the pre-trained model M, n ∈ N, N is the set of positive integers, and the continuous visual instruction fine-tuning data set sequence includes an image feature data set, a video feature data set, and a text feature data set; It is responsible for processing visual inputs and extracting information related to visual tasks;
[0044] S3. Fine-tune the dataset t according to continuous visual instructions i Select a suitable low-rank adaptation matrix and perform fine-tuning, i ∈ n;
[0045] S4. Use the separable routing technique to obtain the visual understanding routing selection G of S3 vu (Z uv ) and the instruction following routing selection G if (Z if ); Separate the mixed low-rank adaptation to handle the differences between tasks. In the visual understanding routing selection, use the average feature of the input to dynamically select a suitable low-rank adaptation module; in the instruction following routing, use the embedding of the current task instruction as the input to select an appropriate low-rank adaptation module for the task that requires instruction interpretation;
[0046] Visual understanding routing selection G vu (Z uv ) is calculated as follows:
[0047] G vu (Z uv ) = softmax(top k (R uv ·Avg(x l-1 ))); (2)
[0048] Where R uv represents the trainable routing matrix saved in the visual understanding module, x l-1 is the input of this module, softmax is the normalization function, l is the current layer number of the model, Z uv is the input of the visual understanding routing selection, top k is to only take the top k largest values and set other values to 0, Avg is the average value;
[0049] Instruction following routing selection G if (Z if ) is calculated as follows:
[0050] G if (Z if ) = softmax(top k (R if ·f(X ins ))); (3)
[0051] Where X ins is the instruction of the current visual instruction fine-tuning task, Z if is the input of the instruction following routing selection, R if is the weight matrix, f is the embedding operation;
[0052] S5. Through G in S4 vu(Z uv ) and G if (Z if ) and its corresponding adaptation matrix W i Perform a weighted operation to obtain two corresponding outputs
[0053] Among them The calculation process is as follows:
[0054]
[0055] Among them The calculation process is as follows:
[0056]
[0057] S6. Use the adaptive fusion technology to obtain the final output of the low-rank adaptation module The adaptive fusion technology fusion algorithm is as follows:
[0058]
[0059] Among them, α and β are weight coefficients learned during the training process. After the routing is completed, the output results of the visual understanding module and the instruction following module Will be adaptively fused. This fusion method assigns a weight to the output of each module to ensure that the most relevant output is used first;
[0060] In the specific implementation process, assume that there is a pre-trained model M and multiple adaptation matrices W in the low-rank adaptation module i . The continuous visual instruction fine-tuning dataset sequence T = {t1, t2,..., t n-1 , t n}; When the i-th instruction fine-tuning dataset t i arrives, it is necessary to select an appropriate low-rank adaptation matrix and perform fine-tuning to meet the requirements of the current task; Use the separable routing technology to obtain the visual understanding routing selection G vu (Z uv ) and the instruction following routing selection G if (Z if );
[0061] Through G vu (Z uv ) and G if (Z if ) and the corresponding adaptation matrix W i Perform a weighted operation to obtain the outputs of the two modules Finally, use the adaptive fusion technology to obtain the final output of the low-rank adaptation module Optimize the low-rank adaptation module through backpropagation to efficiently solve the double catastrophic forgetting problem without sacrificing the model's previous learning ability.
[0062] Therefore, the present invention adopts a continuous visual instruction fine-tuning method based on separable hybrid low-rank adaptation of the above content, using hybrid low-rank adaptation and separable routing technology to efficiently solve the double catastrophic forgetting problem without sacrificing the model's previous learning ability.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for continuous visual command fine-tuning based on separable mixed low-rank adaptation, characterized by: S1. Create a pre-trained model M and several adaptability matrices W i Low-rank adaptation module; S2, fine-tune the continuous visual instruction dataset sequence T = {t1, t2, ..., t n-1 ,t n } Input to the pre-trained model M, n∈N; S3. Fine-tuning dataset t based on continuous visual instructions i Select a suitable low-rank adaptation matrix and fine-tune it, i∈n; S4, using separable routing technology to obtain the visual understanding of S3 routing selection G vu (Z uv ) and instruction follow routing selection G if (Z if ); S5, through S4 in G uv (Z uv ) and G if (Z if ) and its corresponding adaptability matrix W i Perform weighted operations to obtain two corresponding outputs S6. Use adaptive fusion technology to obtain the final output of the low-rank adaptation module 2. The method for continuous visual command fine-tuning based on separable mixed low-rank adaptation according to claim 1, characterized in that: The adaptability matrix W of the low-rank adaptation module described in S1 i Updated by the following formula: Where G(z) k represents the weight assigned by the routing network to the kth adaptability matrix based on the input z.
3. The method for continuous visual command fine-tuning based on separable mixed low-rank adaptation according to claim 2, characterized in that: The continuous visual instruction fine-tuning dataset sequence in S1 includes image feature dataset, video feature dataset, and text feature dataset.
4. The method for continuous visual command fine-tuning based on separable mixed low-rank adaptation according to claim 3, characterized in that: Visual understanding routing in S4 vu (Z uv ) is calculated as follows: G vu (Z uv )=softmax(top k (R uv ·Avg(x l-1 ))); (2) where R uv represents the trainable routing matrix stored in the visual understanding module, x l-1 is the input of the module, softmax is the normalization function, l is the current layer number of the model, Z uv Input for visual understanding routing, top k To take only the first k largest values, set the other values to 0, and Avg is the average value.
5. The method for continuous visual command fine-tuning based on separable mixed low-rank adaptation according to claim 4, characterized in that: Instruction follower routing selection G in S4 if (Z if ) is calculated as follows: G if (Z if )=softmax(top k (R if ·f(X ins ))); (3) Where X ins is the instruction for the current visual instruction fine-tuning task, Z if The input for the command following the routing selection, R if is the weight matrix and f is the embedding operation.
6. The method for continuous visual command fine-tuning based on separable mixed low-rank adaptation according to claim 5, characterized in that: Output as described in S5 The calculation process is as follows:
7. The method for continuous visual command fine-tuning based on separable mixed low-rank adaptation according to claim 6, characterized in that: Output as described in S5 The calculation process is as follows:
8. The method for continuous visual command fine-tuning based on separable mixed low-rank adaptation according to claim 7, characterized in that: The adaptive fusion technology fusion algorithm in S6 is as follows: Where α and β are the weight coefficients learned during the training process.
Citation Information
Patent Citations
Image classification pre-training model continuous learning method based on low-rank adaptive combination
CN117611913A
Monocular self-supervision depth estimation method and system for laparoscope video image
CN117876453A
Multi-task processing method and system for large language model fused with LoRA
CN119416143A
Multi-domain question and answer task fine tuning method based on low-rank adaptive matrix and mixed experts
CN119537554A
Low-Rank Adaptation of Neural Network Models
US20220383126A1