A continuous evolutionary learning method for cue fine-tuning in the null space

By updating the prompt word symbols in zero space, the catastrophic forgetting problem caused by training interventions in continuous evolutionary learning is solved, achieving higher accuracy and lower forgetting rates.

CN118657989BActive Publication Date: 2025-06-06NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410740338.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-08
Publication Date
2025-06-06
Estimated Expiration
2044-06-08

AI Technical Summary

Technical Problem

Existing continuous evolutionary learning methods based on tip-tuning are prone to changes in the characteristics of the old task when training new tasks, resulting in catastrophic forgetting.

Method used

Update the prompt word character to be trained in zero space to associate it with the image character of the old task, thereby eliminating the changes in the image characteristics of the old task by updating the prompt word character.

Benefits of technology

Effectively eliminate training intervention problems, significantly alleviate catastrophic forgetting, improve accuracy and reduce forgetting rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118657989B_ABST
    Figure CN118657989B_ABST
Patent Text Reader

Abstract

The present invention provides a continuous evolutionary learning method for prompt fine-tuning in null space, so that the prompt word token to be trained is updated in the null space related to the image word token of the old task, thereby eliminating the change of the old task image features caused by the update of the prompt word token, achieving the purpose of eliminating the training intervention problem and theoretically avoiding catastrophic forgetting. The present invention maps the gradient of the prompt word token to the null space of the old task image word token, thereby avoiding the change of the old task image features caused by the update of the prompt word token, solving the technical difficulties of training intervention in continuous evolutionary learning based on prompt fine-tuning, and significantly reducing catastrophic forgetting; on the existing public continuous evolutionary learning test benchmark, the present invention improves the accuracy by 4% to 10% and reduces the forgetting rate by 3% to 17% under the setting of class incremental continuous learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pattern recognition, and in particular to a continuous evolutionary learning method for prompt fine-tuning in null space. Background Art

[0002] Faced with an open and dynamic real-world environment, the static learning paradigm based on the closed world assumption encounters severe challenges. Under the static learning paradigm, the model can only generalize to known categories. When faced with test samples from new categories or other distributions with large differences, the generalization ability of the model will be seriously degraded. If the model is updated with data from new categories, it will lead to catastrophic forgetting of the ability to discriminate old categories. These problems seriously restrict the application potential of deep learning models in practice. Continuous evolutionary learning aims to enable the model to be continuously updated to learn new categories or tasks while maintaining good performance on the learned categories or tasks. With the widespread use of large models in recent years, prompt fine-tuning has also become a popular method to transfer the generalization ability of large models to downstream tasks. Continuous evolutionary learning methods based on prompt fine-tuning are also gradually receiving more research.

[0003] Existing continuous evolutionary learning methods based on prompt fine-tuning first query and select a set of learnable prompts for sample instances, then combine the prompts with the input image tokens, and send them to the pre-trained visual transformer (ViT) to extract fine features for classification. For example, a method called L2P first creates a prompt pool with several prompt words, then selects a set of appropriate prompts from the prompt pool based on the distance between the rough features extracted by the pre-trained ViT and the query key corresponding to the prompt, and then sends them to the pre-trained ViT to extract fine image features, and finally sends them to the classifier for classification. CODA-Prompt transforms the process of selecting several prompts from the prompt pool based on an attention process into a process of weighted fusion of all prompts, making the selection of prompts a process that can be optimized end-to-end. Consistent-Prompting selects prompts from the prompt pool as evenly as possible during the training process, so that the difference between the prompt selection in the test phase and the training phase is reduced, thereby improving the robustness of prompt selection. These methods all have the problem of training intervention, that is, when training a new task, the parameter update of the prompt also means that the optimized prompt on the old task will change, causing the features corresponding to the old task samples to change, thus leading to catastrophic forgetting.

[0004] There are also some methods that do not use the prompt pool, but are based on other strategies to prevent forgetting in continuous evolutionary learning. For example, using a smaller accuracy update as the backbone network ViT, or using a larger momentum to slowly update an offline expert model, thereby reducing the rate of change of features; using extended network branches, using a separate learnable branch for each new task, and performing certain semantic compensation on the branches of the old tasks to reduce the degree of feature change in the old task branches; studying the use of more discriminative classifiers, such as using energy model-based alignment, discriminative space based on a large number of random mappings, without paying attention to the changes in features caused by training intervention problems.

[0005] According to the above analysis, the existing continuous evolutionary learning technology based on cue fine-tuning has not yet solved the problem of how to eliminate the training intervention of new tasks on old tasks during the training process, which will lead to serious catastrophic forgetting. Summary of the invention

[0006] In order to overcome the shortcomings of the prior art, the present invention provides a continuous evolutionary learning method for cue fine-tuning in null space. The present invention updates the cue token to be trained in the null space related to the image token of the old task, thereby eliminating the change of the old task image features caused by the update of the cue token, achieving the purpose of eliminating the training intervention problem and theoretically avoiding catastrophic forgetting.

[0007] The technical solution adopted by the present invention to solve the technical problem specifically includes the following steps:

[0008] Step 1: Randomly initialize the prompt token P t,s , t represents the task number, s represents the number of iterations in the training process, the total number of tasks is represented by T′, and the maximum number of iterations preset in each task is represented by N; initialize the first non-centralized covariance matrix C 1,t = 0, where 0 represents the zero matrix, initialize the second non-centered covariance matrix C 2,t =0; Initialize the first null space mapping matrix B 1,t =I, where I represents the identity matrix, initialize the second null space mapping matrix B 2,t =I;

[0009] Step 2: Get the tth task A training sample set, each training sample in the training sample set contains an image and its corresponding category label; the number of categories in the training sample set is C t ;

[0010] Step 3: Calculate the prompt token P t,N ;

[0011] Step 4: Utilize Tasks The training sample set and prompt word Pt,N Update the first non-centered covariance matrix C 1,t and the second non-centered covariance matrix C 2,t ;

[0012] Step 5: Use singular value decomposition to calculate C separately 1,t and C 2,t The basis of the approximate null space of and

[0013] Step 6: Use a basis that approximates the null space and Update the null space mapping matrix B 1,t+1 and B 2,t+1 :

[0014]

[0015]

[0016] where η 1 With η 2 are two hyperparameters between 0.9 and 1, ||·|| F Represents the Frobenius norm of the matrix;

[0017] Updated B 1,t+1 and B 2,t+1 For subsequent tasks Medium gradient G t,s Perform null space mapping;

[0018] Step 7: Determine whether the learning of the task is completed, that is, whether t = T' is satisfied. If t = T', end the training process and go to step 8; if t is not equal to T', go to step 2 and start the t+1th task Learning continues until t = T′ is satisfied;

[0019] Step 8: Output the optimized prompt word P T,N , P T,N That is, the optimized prompt token obtained through the continuous evolutionary learning method is used to infer the category label of the test image in practice.

[0020] Step 3 calculates the prompt word P t,N The steps are:

[0021] Judgment Task Whether it is the first learning task, that is, whether t=1 is satisfied. If it is the first learning task, go to step S3.1, otherwise go to step 3.4;

[0022] Step 3.1: Change the prompt word P t,sThe training sample set obtained in step 2 is sent to the given pre-trained ViT model for forward propagation, thereby calculating the classification loss L 1 , and then obtain P through back propagation t,s The gradient G t,s ;

[0023] Step 3.2: Use the prompt word P t,s The gradient G t,s Update the prompt word to get the prompt word P of the s+1th iteration t,s+1 , assuming the learning rate is γ, the update process is expressed as:

[0024] P t,s+1 =P t,s -γG t,s (1)

[0025] Step 3.3: Determine whether the number of iterations s+1 reaches the target number of iterations N. If it reaches the target number of iterations N, keep the prompt word P obtained in the last iteration. t,N , then go to step 4. If the number of iterations s+1 does not reach the target number of iterations N, go to step 3.1;

[0026] Step 3.4: When t>1, that is For the 2nd to Tth tasks, the prompt word P t,s The training sample set obtained in step 2 is sent to the pre-trained ViT model for forward propagation, thereby calculating the classification loss L 1 , and calculate the distribution invariant loss L 2 , and then use the sum of classification loss and distribution invariance loss L 1 +L 2 Perform back propagation to obtain P t,s The gradient G t,s ; In the task In the example, the prompt token of the Sth iteration obtained after the training is represented as P t-1,N ,use Indicates P t-1,N The mean of each row vector in , Indicates P t-1,N The standard deviation of each row vector in is calculated using Indicates P t,s The mean of each row vector in , Indicates P t,s The standard deviation of each row vector in, distribution invariant loss L 2 Calculated by the following formula:

[0027]

[0028] in and As fixed target distributions, and is the distribution parameter optimized by gradient descent;

[0029] L 1 With L 2 Add together to get the total loss L 3 for:

[0030] L 3 =L 1 +ξL 2

[0031] Where ξ is a loss coefficient greater than 0, which is manually set as a hyperparameter to obtain the total loss L 3 After that, through back propagation calculation, we get the prompt word P t,s The gradient G t,s ;

[0032] Step 3.5: Convert the gradient G t,s With the first null space mapping matrix B 1,t and the second null space mapping matrix B 2,t Multiply them together to get the mapped gradient ΔP t,s for:

[0033] ΔP t,s =B 2,t G t,s B 1,t

[0034] Step 3.6: Use the mapped gradient ΔP t,s Update P with learning rate γ t,s , and get P t,s+1 :

[0035] P t,s+1 =P t,s -γΔP t,s

[0036] Step 3.7: Determine whether the number of iterations s+1 reaches the target number of iterations N. If the number of iterations s+1 reaches the target number of iterations N, retain the prompt word P obtained in the last iteration. t,N , then go to step 4. If the number of iterations s+1 does not reach the target number of iterations N, go to step 3.4.

[0037] Step 3.1 Obtain P t,s The gradient G t,s The specific steps are:

[0038] In the sth training iteration, firstly, the image M in the training sample set is tDivide into multiple blocks and obtain the corresponding image token X t ; Then X t With the prompt word P t,s They are sent together to the pre-trained ViT model for forward propagation of the network to obtain the representative C t A vector of class prediction scores Then get the category label y in the training sample set t , category label y t is a one-hot encoded vector, and the true category of the image is in y t The value of the corresponding category number in is 1, and the value of other category numbers is 0; then calculate y t and The cross entropy loss L between 1 As classification loss:

[0039]

[0040] where y t,i Represents y t The value of the i-th category in , express The value of the i-th category in , after calculating the cross entropy loss, the prompt word P is calculated through back propagation t,s The gradient G t,s .

[0041] Step 4: Utilize Tasks The training sample set and prompt word P t,N Update the first non-centered covariance matrix C 1,t and the second non-centered covariance matrix C 2,t ; The specific steps are:

[0042] First, the image M in the training sample set t Divide into multiple blocks and obtain the corresponding image token X t ; Then X t With the prompt word P t,N are sent to the pre-trained ViT model for network forward propagation. In the process of layer-by-layer forward propagation, the input image token X is first recorded. t Image query tokens after normalization and linear transformation

[0043]

[0044] in and Respectively represent X t The mean and standard deviation of all row vectors, α and β are the pre-trained vector parameters in layer normalization (LayerNorm), Wq , b q It is the pre-trained weight matrix and pre-trained offset vector in the query token conversion layer. ⊙ represents element-wise multiplication. When vectors and matrices are operated, they are broadcasted as matrices of the same dimension.

[0045] Then record the softmax normalized activation attention map To obtain First calculate the prompt key word

[0046]

[0047] in and Respectively represent P t,N The mean and standard deviation of all row vectors, W k , b k is the pre-trained weight matrix and pre-trained offset vector in the key-to-word conversion layer; then calculate the softmax activated attention map

[0048]

[0049] Where D represents the dimension of a token and T is the transposition;

[0050] The task The number of samples in the training sample set is represented as n, and the image query tokens of all images are calculated. Then the query image tokens are respectively compared with W k Multiply them together and concatenate all the result matrices into one matrix J 1,t :

[0051]

[0052] Calculate the attention map of all images And concatenate into a matrix J 2,t :

[0053]

[0054] Then use J 1,t and J 2,t Update the first non-centered covariance matrix C separately 1,t and the second non-centered covariance matrix C 2,t :

[0055]

[0056] Among them, C 1,t-1 With C 2,t-1 Respectively in the task The first non-centered covariance matrix and the second non-centered covariance matrix obtained in .

[0057] In step 5, for C 1,t , first perform singular value decomposition:

[0058]

[0059] where Λ 1,t represents the singular values ​​[λ 1,t ,λ 2,t , …, λ D,t ] is a diagonal matrix, V 1,t Represents the matrix composed of left singular vectors, U 1,t Represents the matrix composed of right singular vectors, and then find Λ 1,t R 1,t The smallest singular value, R 1,t Determined by the following formula:

[0060]

[0061] Use [u 1,t ,…,u D,t ] indicates U 1,t Column vector in, take out the right singular matrix U 1,t R 1,t right singular vectors Composition C 1,t The basis of the approximate null space of

[0062] For C 2,t , use with C 1,t The same process is used to obtain C 2,t The basis of the approximate null space of

[0063] An electronic device includes one or more processors and a memory; one or more programs are stored in the memory and are configured to be executed by the processor to implement the above method.

[0064] A computer-readable storage medium stores program code, wherein the above method is executed when the program code is executed by a processor.

[0065] The beneficial effect of the present invention lies in that by adopting the null space mapping technical means, the gradient of the prompt word is mapped to the null space of the old task image word, thereby avoiding the change of the old task image characteristics caused by the update of the prompt word, solving the technical difficulties of training intervention in continuous evolutionary learning based on prompt fine-tuning, and significantly reducing catastrophic forgetting; on the existing public continuous evolutionary learning test benchmark, the present invention improves the accuracy by 4% to 10% and reduces the forgetting rate by 3% to 17% under the setting of incremental continuous learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a program flow chart of the present invention. DETAILED DESCRIPTION

[0067] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0068] The technical solution adopted by the present invention to solve the technical problem specifically includes the following steps:

[0069] Step S101: Randomly initialize the prompt word P t,s , t represents the task number, s represents the number of iterations in the training process, the total number of tasks is represented by T′, and the maximum number of iterations preset in each task is represented by N; initialize the first non-centralized covariance matrix C 1,t = 0, where 0 represents the zero matrix, initialize the second non-centered covariance matrix C 2,t =0; Initialize the first null space mapping matrix B 1,t =I, where I represents the identity matrix, initialize the second null space mapping matrix B 2,t =I;

[0070] Step S102: Get the tth task A training sample set, each training sample in the training sample set contains an image and its corresponding category label; the number of categories in the training sample set is C t ;

[0071] Step S103: Determine the task Whether it is the first learning task, that is, whether t=1 is satisfied. If it is the first learning task, proceed to step S104, otherwise proceed to step S107;

[0072] Step S104: The prompt word P t,s The training sample set obtained in step S102 is sent to the given pre-trained ViT model for forward propagation, thereby calculating the classification loss L 1 , and then obtain P through back propagation t,s The gradient G t,sSpecifically, in the sth training iteration, firstly, the image M in the training sample set is t Divide into multiple blocks and obtain the corresponding image token X t ; Then X t With the prompt word P t,s They are sent together to the pre-trained ViT model for forward propagation of the network to obtain the representative C t A vector of class prediction scores Then get the category label y in the training sample set t , category label y t is a one-hot encoded vector, and the true category of the image is in y t The value of the corresponding category number in is 1, and the value of other category numbers is 0; then calculate y t and The cross entropy loss L between 1 As classification loss:

[0073]

[0074] In formula (1), y t,i Represents y t The value of the i-th category in , express The value of the i-th category in is calculated according to formula (1) to obtain the cross entropy loss, and then the prompt word P is calculated through back propagation. t,s The gradient G t,s ;

[0075] Step S105: Use the prompt word P t,s The gradient G t,s Update the prompt word to get the prompt word P of the s+1th iteration t,s+1 , assuming the learning rate is γ, the update process is expressed as:

[0076] P t,s+1 =P t,s -γG t,s (2)

[0077] Step S106: Determine whether the number of iterations s+1 reaches the target number of iterations N. If it reaches the target number of iterations N, retain the prompt word P obtained in the last iteration. t,N , then go to step S111, if the number of iterations s+1 does not reach the target number of iterations N, go to step S104;

[0078] Step S107: When t>1, that is For the 2nd to Tth tasks, the prompt word P t,sThe training sample set obtained in step S102 is sent to the pre-trained ViT model for forward propagation, thereby calculating the classification loss L 1 , and calculate the distribution invariant loss L 2 , and then use the sum of classification loss and distribution invariance loss L 1 +L 2 Perform back propagation to obtain P t,s The gradient G t,s ; In the task In the example, the prompt token of the Sth iteration obtained after the training is represented as P t-1,N ,use Indicates P t-1,N The mean of each row vector in , Indicates P t-1,N The standard deviation of each row vector in is calculated using Indicates P t,s The mean of each row vector in , Indicates P t,s The standard deviation of each row vector in, distribution invariant loss L 2 Calculated by formula (3):

[0079]

[0080] In formula (3) and As fixed target distributions, and are the distribution parameters to be optimized via gradient descent;

[0081] L 1 With L 2 Add together to get the total loss L 3 for:

[0082] L 3 =L 1 +ξL 2 (4)

[0083] In formula (4), ξ is a loss coefficient greater than 0, which is manually set as a hyperparameter to obtain the total loss L 3 After that, through back propagation calculation, we get the prompt word P t,s The gradient G t,s ;

[0084] Step S108: The gradient G t,s With the first null space mapping matrix B 1,t and the second null space mapping matrix B 2,t Multiply them together to get the mapped gradient ΔP t,s for:

[0085] ΔPt,s =B 2,t G t,s B 1,t (5)

[0086] Step S109: Using the mapped gradient ΔP t,s Update P with learning rate γ t,s , and get P t,s+1 :

[0087] P t,s+1 =P t,s -γΔP t,s

[0088] Step S110: Determine whether the number of iterations s+1 reaches the target number of iterations N. If the number of iterations s+1 reaches the target number of iterations N, retain the prompt word P obtained in the last iteration. t,N , then go to step S111, if the number of iterations s+1 does not reach the target number of iterations N, go to step S107;

[0089] Step S111: Utilize Task The training sample set and prompt word P t,N Update the first non-centered covariance matrix C 1,t and the second non-centered covariance matrix C 2,t ; The specific steps are:

[0090] First, the image M in the training sample set t Divide into multiple blocks and obtain the corresponding image token X t ; Then X t The prompt word P obtained in step S106 or S110 t,N are sent to the pre-trained ViT model for network forward propagation. In the process of layer-by-layer forward propagation, the input image token X is first recorded. t Image query tokens after normalization and linear transformation

[0091]

[0092] In formula (6) and Respectively represent X t The mean and standard deviation of all row vectors, α and β are the pre-trained vector parameters in layer normalization (LayerNorm), W q 、b q are the pre-trained weight matrix and pre-trained offset vector in the query token conversion layer, ⊙ represents element-wise multiplication; in formula (6), the vector and matrix are broadcasted as matrices of the same dimension when performing operations;

[0093] Then record the softmax normalized activation attention map To obtain First calculate the prompt key word

[0094]

[0095] In formula (7) and Respectively represent P t,N The mean and standard deviation of all row vectors, W k 、b k is the pre-trained weight matrix and pre-trained offset vector in the key-to-word conversion layer; then calculate the softmax activated attention map

[0096]

[0097] In formula (8), D represents the dimension of a token, and T is the transposition;

[0098] The task The number of samples in the training sample set is denoted as n, and the image query tokens of all images are calculated using formula (6): Then the query image tokens are respectively compared with W k Multiply them together and concatenate all the result matrices into one matrix J 1,t :

[0099]

[0100] Then use equations (7) and (8) to calculate the attention map of all images And concatenate into a matrix J 2,t :

[0101]

[0102] Then use J 1,t and J 2,t Update the first non-centered covariance matrix C separately 1,t and the second non-centered covariance matrix C 2,t :

[0103]

[0104] In formulas (11) and (12), C 1,t-1 With C 2,t-1 Respectively in the task The first non-centered covariance matrix and the second non-centered covariance matrix obtained in ;

[0105] Step S112: Use singular value decomposition to calculate C 1,t and C 2,t The basis of the approximate null space of and For C 1,t , first perform singular value decomposition:

[0106]

[0107] In formula (13), Λ 1,t represents the singular values ​​[λ 1,t ,λ 2,t , …, λ D,t ] is a diagonal matrix, V 1,t Represents the matrix composed of left singular vectors, U 1,t Represents the matrix composed of right singular vectors, and then find Λ 1,t R 1,t The smallest singular value, R 1,t Determine by formula (14):

[0108]

[0109] Use [u 1,t ,…,u D,t ] indicates U 1,t Column vector in, take out the right singular matrix U 1,t R 1,t right singular vectors Composition C 1,t The basis of the approximate null space of

[0110] For C 2,t , using the same process as equations (13) and (14), we get C 2,t The basis of the approximate null space of

[0111] Step S113: Using the basis of the approximate null space and Update the null space mapping matrix B 1,t+1 and B 2,t+1 :

[0112]

[0113] In formulas (15) and (16), η 1 With η 2 are two hyperparameters between 0.9 and 1, ||·|| F Represents the Frobenius norm of the matrix;

[0114] Updated B1,t+1 and B 2,t+1 For subsequent tasks Use formula (5) to calculate the gradient G t,s Perform null space mapping;

[0115] Step S114: Determine whether the task learning is completed, that is, whether t=T′ is satisfied. If t=T′, end the training process and enter step S115; if t is not equal to T′, enter step S102 and start the t+1th task. Learning continues until t = T′ is satisfied;

[0116] Step S115: Outputting the optimized prompt word P T,N , P T,N That is, the optimized prompt token obtained through the continuous evolutionary learning method is used to infer the category label of the test image in practice.

[0117] The implementation example of the present invention is implemented on the public dataset 10-split ImageNet-R, using the ViT model pre-trained on ImageNet-21k, and using the VPT-Deep method to add a prompt of length 4 for fine-tuning training. The flowchart of the technical solution adopted is as follows: Figure 1 As shown, Figure 1 Description: First, initialize the trainable cue parameters, two non-centered covariance matrices, and two null space mapping matrices. If the current task is the first learning task, that is, t = 1, the update amount of the cue parameters is calculated through classification loss optimization and directly used to update the cue parameters until the training converges. Then, two intermediate feature matrices are calculated for all samples in the task and used to update the two non-centered covariance matrices respectively. Then, the right singular vectors corresponding to the singular values ​​close to 0 are obtained through singular value decomposition, which are used as the basis of the null space corresponding to the two non-centered covariance matrices. Then, the two null space mapping matrices are calculated using the basis of the null space. Starting from the second task, that is, when t>1, in addition to the classification loss, the loss function during training also adds a cue parameter distribution invariant loss, which is used to calculate the total loss value and back propagation. After the optimizer gives the candidate cue parameter update, it is first mapped through the two null space mapping matrices, and then the parameters are updated. At the end of each task training, the data of the task is also used to calculate two intermediate feature matrices, and the two non-centered covariance matrices are updated. The two null space mapping matrices are further calculated through singular value decomposition and used to map the learnable prompt parameters in the next learning task. This cycle continues until all learning tasks are completed.

[0118] The specific steps include:

[0119] Step 1: Randomly initialize the trainable hint parameters P t, initialize two non-centered covariance matrices C 1 =0 and C 2 = 0 is the zero matrix, initialize two zero space mapping matrices B 1 =I and B 2 =I is the unit matrix.

[0120] Step 2: Input the data of the tth learning task, consisting of images and corresponding labels.

[0121] Step 3: Use image-label pairs to train the hint parameter P t , until convergence. Specifically including the following steps:

[0122] Step 3-1: Place the image Divide into multiple tiles and obtain the corresponding image token X t , and the prompt word P t Send them to the pre-trained ViT together, perform the forward propagation process of the network, and get the predicted category Calculate the loss of the current sample sent to the network, including the predicted category The cross entropy classification loss with the true label y is:

[0123]

[0124] And the loss of the limit of the change of the parameter distribution is indicated:

[0125]

[0126] The distribution change loss function, i.e., formula (18), is only used when task t>1. and As fixed target distributions, and Optimization is performed by gradient descent. In the first task, the loss function only uses equation (17), and from the second task onwards, both equations (17) and (18) are used:

[0127] L total =L ce +L ln (19)

[0128] Step 3-2: Get candidate hint parameters from the optimizer through back propagation to update P G , after two null space mappings B 1 and B 2 Get the true parameter update ΔP:

[0129] ΔP=B 2 P G B 1 (20)

[0130] Step 3-3: Update the prompt parameter P t :

[0131] P t ←P t -ΔP (21)

[0132] Repeat steps 3-1 to 3-3 until convergence to complete the training of the current task.

[0133] Step 4: Update the two non-centralized covariance matrices C using the training data of task t 1 and C 2 , and update the null space mapping matrix B used in the t+1 task 1 and B 2 The specific steps include:

[0134] Step 4-1: Obtain two intermediate variables of the current task data to update two non-centralized covariance matrices C 1 and C 2 First, for all images in the current task, the corresponding image tokens are obtained, and then they are compared with the P t are sent to the pre-trained ViT for forward propagation. During the forward propagation, the input token X is recorded. t Query tokens after normalization and linear transformation and the pre-trained keyword weights W k The product matrix of and the softmax normalized partial activation feature map

[0135] Calculate query terms

[0136]

[0137] Where α and β are the pre-trained vector parameters in LayerNorm, W q 、b q are the weight matrix and offset vector in the query token conversion layer. ⊙ represents element-wise multiplication. The vector and matrix are automatically broadcast to matrices of the same dimension when performing operations.

[0138] Calculate the softmax normalized partial activation feature map

[0139]

[0140] In formula (23), W k 、b k is the pre-trained weight matrix and offset vector in the key-to-word conversion layer, bk Automatically broadcast to matrices of corresponding dimensions for addition; in formula (24), the softmax operation is performed row by row, and D represents the dimension of a token.

[0141] For all samples in task t, the tokens obtained Use equations (22)-(24) to calculate the corresponding matrices one by one. Then, the concatenation matrix J of all samples is obtained for the two intermediate matrices: 1 and J 2 :

[0142]

[0143] Use J again 1 and J 2 Update the two non-centered covariance matrices C 1 and C 2 :

[0144]

[0145] Step 4-2: Use singular value decomposition to calculate C 1 and C 2 Calculate the basis U of the approximate null space respectively 1,0 and U 2,0 For C 1 , first perform singular value decomposition:

[0146]

[0147] where Λ 1 Represents the singular values ​​in descending order [λ 1 ,λ 2 , …, λ D ] is a diagonal matrix, V 1 Represents the matrix composed of left singular vectors, U 1 A matrix representing the right singular vectors.

[0148] Then find Λ 1 R 1 The smallest singular value is considered to be a singular value close to 0. 1 Determined based on the maximum second-order derivative of the curve composed of singular values:

[0149]

[0150] According to formula (14), R 1 Then select U 1 The corresponding R 1 right singular vectors, forming C 1 The basis U of the null space 1,0 .

[0151] For C 2 , using the same process as equations (13) and (14), calculate the corresponding null space basis U 2,0 .

[0152] Step 4-3: Use U 1,0 and U 2,0 Update the null space mapping matrix B 1 and B 2 :

[0153]

[0154] In this embodiment, in formula (31) and (32), n 1 With η 2 Both are set to 0.94.

[0155] Updated B 1 and B 2 It is used to perform null space mapping on candidate parameter updates using equation (4) in the subsequent t+1 task.

[0156] Step 5: Determine whether the learning of all tasks is completed. If completed, end the training; otherwise, enter the t+1th task and repeat steps 2 to 5.

[0157] The comparison of the present invention with the baseline model on four incremental continuous learning tasks 10-split CIFAR-100, 20-split CIFAR-100, 10-split ImageNet-R, and 10-split DomainNet-200 is shown in Table 1, where n-split represents that the corresponding dataset is evenly divided into n continuous learning tasks. In the table, VPT-Seq and CLIP-Seq represent sequential baseline methods using VPT and CLIP models, which do not use the loss function L in formula (3). ln , and B in formula (5) 1 and B 2 The identity matrix is ​​maintained consistently, that is, no null space mapping is used; VPT-NSP 2 and CLIP-NSP2 represent models that have the method proposed in the present invention added to them respectively; the upper bound represents the performance of joint training of all tasks, which can be considered as the upper bound of continuous learning performance; the higher the accuracy, the better, and the lower the forgetting rate, the better. It can be seen from the table that the present invention can significantly improve the accuracy (increase by 4% to 10%) and reduce the forgetting rate (decrease by 3% to 17%).

[0158] Table 1 Comparison with baseline methods

[0159]

[0160] The present invention (VPT-NSP 2 ) is compared with other methods on four types of incremental continuous learning tasks as shown in Table 2. The value after ± represents the standard deviation during multiple experiments. It can be seen from the table that this method has achieved the best performance so far.

[0161] Table 2 Comparison with other methods

[0162]

[0163]

[0164] The performance of the implementation example of the present invention on 10-split ImageNet-R is shown in Table 3:

[0165] Table 3 Comparison of implementation examples with other methods and baseline method (VPT-Seq) on 10-split ImageNet-R

[0166]

[0167] Among them, VPT-Seq is the baseline method for sequential training, VPT-NSP2 is the implementation example of this method, and the others are existing advanced methods. It can be seen that this implementation example achieves an accuracy improvement of 6.42% and a forgetting rate reduction of 14.35% compared to the baseline, and exceeds other existing methods.

Claims

1. A continuous evolutionary learning method for prompt fine-tuning in null space, characterized by The steps include: Step 1: Randomly initialize the prompt token P t,s , t represents the task number, s represents the number of iterations in the training process, the total number of tasks is represented by T′, and the maximum number of iterations preset in each task is represented by N; Initialize the first non-centered covariance matrix C 1,t = 0, where 0 represents the zero matrix, initialize the second non-centered covariance matrix C 2,t =0; Initialize the first null space mapping matrix B 1,t =I, where I represents the identity matrix, initialize the second null space mapping matrix B 2,t =I; Step 2: Get the tth task A training sample set, each training sample in the training sample set contains an image and its corresponding category label; the number of categories in the training sample set is C t ; Step 3: Calculate the prompt token P t,N ; Step 4: Utilize Tasks The training sample set and prompt word P t,N Update the first non-centered covariance matrix C 1,t and the second non-centered covariance matrix C 2,t ; Step 5: Use singular value decomposition to calculate C separately 1,t and C 2,t The basis of the approximate null space of and In step 5, for C 1,t , first perform singular value decomposition: where Λ 1,t represents the singular values ​​[λ 1,t ,λ 2,t , …, λ D,t ] is a diagonal matrix, V 1,t Represents the matrix composed of left singular vectors, U 1,t Represents the matrix composed of right singular vectors, and then find Λ 1,t R 1,t The smallest singular value, R 1,t Determined by the following formula: Use [u 1,t ,…u D,t ] indicates U 1,t Column vector in, take out the right singular matrix U 1,t R 1,t Right singular vectors Composition C 1,t The basis of the approximate null space of For C 2,t , use with C 1,t The same process is used to obtain C 2,t The basis of the approximate null space of Step 6: Use a basis that approximates the null space and Update the null space mapping matrix B 1,t+1 and B 2,t+1 : Where η1 and η2 are two hyperparameters between 0.9 and 1, respectively, ||·|| F Represents the Frobenius norm of the matrix; Updated B 1,t+1 and B 2,t+1 For subsequent tasks Medium gradient G t,s Perform null space mapping; Step 7: Determine whether the learning of the task is completed, that is, whether t = T' is satisfied. If t = T', end the training process and go to step 8; if t is not equal to T', go to step 2 and start the t+1th task Learning continues until t = T′ is satisfied; Step 8: Output the optimized prompt word P T,N , P T,N That is, the optimized prompt token obtained through the continuous evolutionary learning method is used to infer the category label of the test image in practice.

2. The continuous evolutionary learning method for prompt fine-tuning in null space according to claim 1, characterized in that: Step 3 calculates the prompt word P t,N The steps are: Judgment Task Whether it is the first learning task, that is, whether t=1 is satisfied. If it is the first learning task, go to step S3.1, otherwise go to step 3.4; Step 3.1: Change the prompt word P t,s The training sample set obtained in step 2 is sent to the given pre-trained ViT model for forward propagation, thereby calculating the classification loss L1, and then obtaining P through back propagation t,s The gradient G t,s ; Step 3.2: Use the prompt word P t,s The gradient G t,s Update the prompt word to get the prompt word P of the s+1th iteration t,s+1 , assuming the learning rate is γ, the update process is expressed as: P t,s+1 =P t,s -γG t,s (1) Step 3.3: Determine whether the number of iterations s+1 reaches the target number of iterations N. If it reaches the target number of iterations N, keep the prompt word P obtained in the last iteration. t,N , then go to step 4. If the number of iterations s+1 does not reach the target number of iterations N, go to step 3.1; Step 3.4: When t>1, that is For the 2nd to Tth tasks, the prompt word P t,s The training sample set obtained in step 2 is sent to the pre-trained ViT model for forward propagation, thereby calculating the classification loss L1 and the distribution invariant loss L2, and then the sum of the classification loss and the distribution invariant loss L1+L2 is used for back propagation to obtain P t,s The gradient G t,s ; In the task In the example, the prompt token of the Sth iteration obtained after the training is represented as P t-1,N ,use Indicates P t-1,N The mean of each row vector in , Indicates P t-1,N The standard deviation of each row vector in is calculated using Indicates P t,s The mean of each row vector in , Indicates P t,s The standard deviation of each row vector in , the distribution invariant loss L2 is calculated by the following formula: in and As fixed target distributions, and are the distribution parameters to be optimized via gradient descent; Adding L1 and L2, the total loss L3 is: L3=L1+ξL2 Where ξ is a loss coefficient greater than 0, which is manually set as a hyperparameter. After the total loss L3 is obtained, the prompt word P is obtained through back propagation calculation. t,s The gradient G t,s ; Step 3.5: Convert the gradient G t,s With the first null space mapping matrix B 1,t and the second null space mapping matrix B 2,t Multiply them together to get the mapped gradient ΔP t,s for: ΔP t,s =B 2,t G t,s B 1,t Step 3.6: Use the mapped gradient ΔP t,s Update P with learning rate γ t,s , and get P t,s+1 : P t,s+1 =P t,s -γΔP t,s Step 3.7: Determine whether the number of iterations s+1 reaches the target number of iterations N. If the number of iterations s+1 reaches the target number of iterations N, retain the prompt word P obtained in the last iteration. t,N , then go to step 4. If the number of iterations s+1 does not reach the target number of iterations N, go to step 3.

4.

3. The continuous evolutionary learning method for prompt fine-tuning in null space according to claim 2, characterized in that: Step 3.1 Obtain P t,s The gradient G t,s The specific steps are: In the sth training iteration, firstly, the image M in the training sample set is t Divide into multiple blocks and obtain the corresponding image token X t ; Then X t With the prompt word P t,s They are sent together to the pre-trained ViT model for forward propagation of the network to obtain the representative C t A vector of class prediction scores Then get the category label y in the training sample set t , category label y t is a one-hot encoded vector, and the true category of the image is in y t The value of the corresponding category number in is 1, and the value of other category numbers is 0; then calculate y t and The cross entropy loss L1 between them is used as the classification loss: where y t,i Represents y t The value of the i-th category in , express The value of the i-th category in , after calculating the cross entropy loss, the prompt word P is calculated through back propagation t,s The gradient G t,s .

4. The continuous evolutionary learning method for prompt fine-tuning in null space according to claim 2, characterized in that: Step 4: Utilize Tasks The training sample set and prompt word P t,N Update the first non-centered covariance matrix C 1,t and the second non-centered covariance matrix C 2,t ; The specific steps are: First, the image M in the training sample set t Divide into multiple blocks and obtain the corresponding image token X t ; Then X t With the prompt word P t,N are sent to the pre-trained ViT model for network forward propagation. In the process of layer-by-layer forward propagation, the input image token X is first recorded. t Image query tokens after normalization and linear transformation in and Respectively represent X t The mean and standard deviation of all row vectors, α and β are the pre-trained vector parameters in layer normalization (LayerNorm), w q 、b q It is the pre-trained weight matrix and pre-trained offset vector in the query token conversion layer. ⊙ represents element-wise multiplication. When vectors and matrices are operated, they are broadcasted as matrices of the same dimension. Then record the softmax normalized activation attention map To obtain First calculate the prompt key word in and Respectively represent P t,N The mean and standard deviation of all row vectors, W k 、b k is the pre-trained weight matrix and pre-trained offset vector in the key-to-word conversion layer; then calculate the softmax activated attention map Where D represents the dimension of a token and T is the transposition; The task The number of samples in the training sample set is represented as n, and the image query tokens of all images are calculated. Then the query image tokens are respectively compared with W k Multiply them together and concatenate all the result matrices into one matrix J 1,t : Calculate the attention map of all images And concatenate into a matrix J 2,t : Then use J 1,t and J 2,t Update the first non-centered covariance matrix C separately 1,t and the second non-centered covariance matrix C 2,t : Among them, C 1,t-1 With C 2,t-1 Respectively in the task The first non-centered covariance matrix and the second non-centered covariance matrix obtained in .

5. An electronic device, characterized in that: include: comprising one or more processors; There is also storage; One or more programs are stored in the memory and are configured to cause the processor to execute the method according to any one of claims 1 to 4.

6. A computer-readable storage medium, characterized in that: A computer-readable storage medium having program code stored therein, wherein when the program code is executed by a processor, the method according to any one of claims 1 to 4 is executed.

Citation Information

Patent Citations

  • Image classification pre-training model continuous learning method based on low-rank adaptive combination

    CN117611913A

  • Cross-device incremental bearing fault diagnosis method based on continuous learning

    WO2024021246A1