Prompt generation and prompt transformation-based class incremental learning method and system

By generating and transforming model hints, combining the key-value pairs of historical tasks and newly trained optical image features, the optical image classification model is optimized, which solves the problem of increased training volume and reduced accuracy in incremental learning, and achieves more efficient model training and more accurate classification results.

CN120339676APending Publication Date: 2025-07-18SHENZHEN POLYTECHNIC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510241614.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When the existing incremental learning methods process new data, there are problems such as the increase in model training volume and decrease in accuracy. In particular, the network fine-tuning method forgets more old tasks, the regularization method relies on the correlation between new and old tasks and has high computing costs, and the playback method requires additional storage space and computing resources.

Method used

By generating model prompts, combining the key-value pairs of historical tasks and the features of newly trained optical images, the prompt generation module and the prompt transformation module generate preparatory model prompts, optimize the optical image classification model, reduce the training amount and improve data relevance.

Benefits of technology

It effectively reduces the amount of optimization training of the optical image classification model, improves the data correlation of the model prompts, and thus improves the classification accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339676A_ABST
    Figure CN120339676A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and discloses a class incremental learning method and system based on prompt generation and prompt transformation, and the method comprises the steps: obtaining a first optical image classification model and key value pairs corresponding to a plurality of historical tasks; inputting a new training optical image of the current incremental task into the first optical image classification model, and extracting a plurality of layers of features and high-level semantic features; constructing a prompt generation module, and generating a preparatory model prompt according to the plurality of layers of features and the high-level semantic features; and a prompt conversion module is constructed, and information fusion and model prompt generation are carried out on the preparatory model prompt according to the key value pair corresponding to each historical task, so that the first optical image classification model is optimized, a second optical image classification model is generated, the input to-be-classified optical images are classified, and a classification result is obtained. The optimization training amount of the optical image classification model can be reduced, the data relevance prompted by the model is improved, and the accuracy of the optical image classification model is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a class incremental learning method and system based on prompt generation and prompt transformation. Background Art

[0002] The rapid development of information technology has promoted the incremental generation of data. Current data processing algorithms are often trained based on closed data sets and cannot effectively handle the processing tasks of newly added data. Incremental learning aims to enable the model to have the ability to process continuously added data. The current state-of-the-art incremental learning method is generally a prompt-based method, which trains a set of corresponding prompts for the newly added data separately, enabling the pre-trained model to quickly adapt to new tasks and new data. However, this method causes the prompt storage amount of data to increase exponentially as the newly added data increases; moreover, during the application test phase, it is very difficult to find the corresponding prompts for the data, which affects the application performance of the final algorithm.

[0003] Current incremental learning methods are mainly divided into three categories: network fine-tuning methods, regularization methods, and replay methods. Most of the network fine-tuning-based methods only fine-tune the fully connected layers related to the new task. This type of method forgets a lot about the old tasks and increases the network volume at the same time; the regularization-based methods ensure that the old knowledge is not covered by the new knowledge by imposing constraints on the loss function of the new task. This type of method highly depends on the correlation between the old and new tasks, and the training time increases linearly with the number of learning tasks; the replay method retains a part of the representative old data for the model to review the old knowledge when training the new task. This type of method is sensitive to the selected old data, and the computational cost increases exponentially as the number of tasks increases, and additional computational resources and storage space are required. At the same time, there are also potential risks in data privacy. Therefore, the current incremental learning methods have relatively high requirements for the correlation between the old and new training data of the model, and there is a defect that the model training amount increases as the newly added data increases, resulting in a decrease in the accuracy of the model. Summary of the Invention

[0004] The present invention provides a class incremental learning method and system based on prompt generation and prompt transformation, which effectively reduces the optimization training amount of the optical image classification model by generating model prompts, and improves the data correlation of the model prompts by combining historical classified optical images, thereby improving the accuracy of the optimized optical image classification model.

[0005] To solve the above technical problems, the present invention provides a class incremental learning method based on prompt generation and prompt transformation, including:

[0006] Obtain a first optical image classification model and key-value pairs corresponding to a plurality of historical tasks of the first optical image classification model;

[0007] Input the new training optical image of the current incremental task into the first optical image classification model, and extract several layers of features and high-level semantic features of the new training optical image;

[0008] Construct a prompt generation module in the first optical image classification model, and use the prompt generation module to generate a preliminary model prompt according to several layers of features and high-level semantic features of the new training optical image;

[0009] Construct a prompt transformation module in the first optical image classification model, and use the prompt transformation module to perform information fusion on the preliminary model prompt according to the key-value pairs corresponding to several historical tasks to generate a model prompt;

[0010] Optimize the first optical image classification model based on the model prompt to generate a second optical image classification model;

[0011] Obtain the optical image to be classified, input the optical image to be classified into the second optical image classification model, so that the second optical image classification model performs size change and sliding block cutting processing on the optical image to be classified, generates an optical image vector to be classified, and classifies the optical image vector to be classified to obtain the classification result of the optical image to be classified.

[0012] The present invention obtains a first optical image classification model and key-value pairs corresponding to multiple historical tasks executed by the model; when the first optical image classification model receives a new training optical image of the current incremental task, it extracts each layer of features and high-level semantic features of the new training optical image; constructs a prompt generation module in the first optical image classification model, and uses the prompt generation module to process each layer of features and high-level semantic features of the new training optical image to generate a preliminary model prompt; constructs a prompt transformation module in the first optical image classification model, and performs information fusion on the preliminary model prompt generated by the prompt generation module based on the key-value pairs corresponding to multiple historical tasks executed by the first optical image classification model, so as to effectively obtain the relevance between the new training optical image and the historical classified optical image; optimize the first optical image classification model based on the model prompt to obtain a second optical image classification model, so that the second optical image classification model classifies the optical image to be classified to obtain the classification result of the optical image to be classified. When the present invention receives a new training optical image, it optimizes the first optical image classification model by extracting the model prompt, which can reduce the optimization training amount of the model; when generating the model prompt, it combines the historical classified optical images, improves the data relevance of the generated model prompt, and thus improves the classification accuracy of the optimized second optical image classification model.

[0013] Further, the obtaining of the first optical image classification model and the key-value pairs corresponding to several historical tasks of the first optical image classification model are specifically as follows:

[0014] Obtain several historical tasks of the first optical image classification model;

[0015] Obtain the high-level semantic features and classification results of each historical classified optical image in each of the historical tasks;

[0016] Generate key-value pairs corresponding to each historical task according to the high-level semantic features and classification results of each historical classified optical image.

[0017] Further, the inputting of the new training optical image of the current incremental task into the first optical image classification model and the extraction of several layers of features and high-level semantic features of the new training optical image are specifically as follows:

[0018] Input the new training optical image into the first optical image classification model;

[0019] Use several hierarchical structures in the first optical image classification model to respectively extract features from the new training optical image to obtain several layers of features of the new training optical image;

[0020] Use the last hierarchical structure in the first optical image classification model to extract features from the new training optical image to obtain the high-level semantic features of the new training optical image.

[0021] Further, the construction of a prompt generation module in the first optical image classification model and the generation of a preliminary model prompt by using the prompt generation module according to several layers of features and high-level semantic features of the new training optical image are specifically as follows:

[0022] Construct a prompt generation module in the first optical image classification model; the prompt generation module includes a normalization layer, several layers of perceptrons, a cross-attention layer, and a fully connected layer;

[0023] Input the high-level semantic features of the new training optical image into the normalization layer and several layers of perceptrons to generate mapped high-level semantic features;

[0024] Input several layers of features and mapped high-level semantic features of the new training optical image into the cross-attention layer to generate an attention result;

[0025] Process the attention result by using residual connection to form enhanced high-level semantic features of the new training optical image;

[0026] Input the enhanced high-level semantic features of the new training optical image into the fully connected layer to generate a preliminary model prompt.

[0027] Further, inputting the several layers of features of the newly trained optical image and the mapped high-level semantic features into the cross-attention layer to generate an attention result specifically includes:

[0028] Input the several layers of features of the newly trained optical image and the mapped high-level semantic features into the cross-attention layer, set the several layers of features of the newly trained optical image as the query vector, set the mapped high-level semantic features of the newly trained optical image as the key vector and value vector, and calculate the attention result.

[0029] Further, constructing a prompt transformation module in the first optical image classification model, and using the prompt transformation module to perform information fusion on the preliminary model prompt according to the key-value pairs corresponding to several historical tasks to generate a model prompt, specifically including:

[0030] Construct a prompt transformation module in the first optical image classification model;

[0031] Input the high-level semantic features of the newly trained optical image into the prompt transformation module to generate a downsampled feature;

[0032] Generate a scaling parameter and a translation coefficient based on the key-value pairs corresponding to several historical tasks and the downsampled feature;

[0033] Based on the scaling parameter and the translation coefficient, perform a linear transformation on the preliminary model prompt to generate a model prompt, specifically:

[0034] P fl = s c ·P l + s f

[0035] In the formula, P fl is the model prompt; P l is the preliminary model prompt; s c is the scaling parameter; s f is the translation coefficient.

[0036] Further, inputting the high-level semantic features of the newly trained optical image into the prompt transformation module to generate a downsampled feature specifically includes:

[0037] Input the high-level semantic features of the newly trained optical image into the prompt transformation module, extract the downsampled feature from the high-level semantic features of the newly trained optical image to generate an initial CLS feature;

[0038] Use several layers of perceptrons in the prompt transformation module to downsample the initial CLS feature to generate a downsampled feature.

[0039] Further, generating a scaling parameter and a translation coefficient based on the key-value pairs corresponding to several of the historical tasks and the downsampled features specifically includes:

[0040] Obtaining several key vectors from the key-value pairs corresponding to several of the historical tasks;

[0041] Calculating several similarities between the downsampled features and the several key vectors respectively;

[0042] Based on the several similarities, performing a weighted sum of the several value vectors corresponding to the several historical tasks to obtain a summation result;

[0043] Performing dimensionality splitting on the summation result to generate a scaling parameter and a translation coefficient.

[0044] Further, optimizing the first optical image classification model based on the model hint to generate a second optical image classification model specifically includes:

[0045] Setting the model hint in several hierarchical structures of the first optical image classification model, setting the output dimension of the last-layer classifier of the first optical image classification model to the total number of categories of the current incremental task, and controlling the first optical image classification model to output the classification prediction result of the newly trained optical image;

[0046] Calculating the cross-entropy loss between the classification prediction result of the newly trained optical image and the true class value, and optimizing the first optical image classification model according to the cross-entropy loss to generate a second optical image classification model.

[0047] Correspondingly, the present invention provides a class-incremental learning system based on hint generation and hint transformation, including: a data acquisition module, a feature extraction module, a first generation module, a second generation module, a model optimization module, and a model application module;

[0048] The data acquisition module is used to acquire the first optical image classification model and the key-value pairs corresponding to several historical tasks of the first optical image classification model;

[0049] The feature extraction module is used to input the newly trained optical image of the current incremental task into the first optical image classification model to extract several layers of features and high-level semantic features of the newly trained optical image;

[0050] The first generation module is used to construct a hint generation module in the first optical image classification model, and use the hint generation module to generate a preliminary model hint according to the several layers of features and high-level semantic features of the newly trained optical image;

[0051] The second generation module is used to construct a prompt transformation module in the first optical image classification model, and use the prompt transformation module to perform information fusion on the preliminary model prompt according to the key-value pairs corresponding to a plurality of the historical tasks to generate a model prompt;

[0052] The model optimization module is used to optimize the first optical image classification model based on the model prompt to generate a second optical image classification model;

[0053] The model application module is used to obtain an optical image to be classified, input the optical image to be classified into the second optical image classification model, so that the second optical image classification model performs size change and sliding block cutting processing on the optical image to be classified to generate a vector of the optical image to be classified, and classify the vector of the optical image to be classified to obtain a classification result of the optical image to be classified. Description of the Drawings

[0054] Figure 1 It is a schematic flowchart of an embodiment of the class incremental learning method based on prompt generation and prompt transformation provided by the present invention;

[0055] Figure 2 It is a schematic flowchart of another embodiment of the class incremental learning method based on prompt generation and prompt transformation provided by the present invention;

[0056] Figure 3 It is a schematic structural diagram of an embodiment of the class incremental learning system based on prompt generation and prompt transformation provided by the present invention. Detailed Embodiment

[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0058] The flowchart shown in the drawings is only an example illustration, and does not necessarily include all the content and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.

[0059] Next, some embodiments of the present invention will be described in detail in conjunction with the drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0060] Embodiment 1

[0061] AsFigure 1 As shown, it is a schematic flowchart of an embodiment of a class incremental learning method based on prompt generation and prompt transformation provided by the present invention. The method includes steps 101 to 106, and the specific steps are as follows:

[0062] Step 101: Obtain a first optical image classification model and key-value pairs corresponding to several historical tasks of the first optical image classification model.

[0063] Step 102: Input the new training optical images of the current incremental task into the first optical image classification model, and extract several layers of features and high-level semantic features of the new training optical images.

[0064] Step 103: Build a prompt generation module in the first optical image classification model, and use the prompt generation module to generate a preliminary model prompt according to several layers of features and high-level semantic features of the new training optical images.

[0065] Step 104: Build a prompt transformation module in the first optical image classification model, and use the prompt transformation module to perform information fusion on the preliminary model prompt according to the key-value pairs corresponding to several historical tasks to generate a model prompt.

[0066] Step 105: Optimize the first optical image classification model based on the model prompt to generate a second optical image classification model.

[0067] Step 106: Obtain the optical image to be classified, input the optical image to be classified into the second optical image classification model, so that the second optical image classification model performs size change and sliding block cutting processing on the optical image to be classified, generates a vector of the optical image to be classified, and classifies the vector of the optical image to be classified to obtain the classification result of the optical image to be classified.

[0068] In the embodiment of the present invention, when there is a new incremental task for the first optical image classification model, obtain the first optical image classification model and key-value pairs corresponding to multiple historical tasks executed by the first optical image classification model. By obtaining the key-value pairs corresponding to historical tasks, the key vectors and value vectors corresponding to historical classified optical images can be known. Applying these historical data in the subsequent generation of model prompts for new training optical images can effectively improve the relevance of model prompts.

[0069] In the embodiments of the present invention, the newly trained optical images in the new incremental tasks are input into the first optical image classification model, and the features of each layer and the high-level semantic features of the newly trained optical images can be extracted. The features of each layer of the newly trained optical images are obtained by extracting features using each hierarchical structure in the first optical image classification model. The high-level semantic features of the newly trained optical images are obtained by extracting features using the last hierarchical structure in the first optical image classification model.

[0070] In the embodiments of the present invention, a prompt generation module is constructed in the first optical image classification model, and the features of each layer and the high-level semantic features of the newly trained optical images are used as the input of the prompt generation module. A preliminary model prompt is generated through attention calculation, residual operation, and full connection layer processing.

[0071] In the embodiments of the present invention, a prompt transformation module is constructed in the first optical image classification model. Based on the key-value pairs corresponding to multiple historical tasks executed by the first optical image classification model, the similarity between the CLS feature in the high-level semantic features of the newly trained optical images and the key vectors corresponding to each task history is calculated, and the value vectors of each historical task are weighted and summed based on this similarity to generate the scaling coefficient and translation coefficient of the current incremental task. Based on the scaling coefficient and translation coefficient, a linear transformation is performed on the preliminary model prompt generated by the prompt generation module to complete information fusion and generate a model prompt. By combining historical data to generate a model prompt, the present invention can effectively obtain the correlation between the newly trained optical images and historical data, thereby improving the data correlation of the model prompt.

[0072] In the embodiments of the present invention, after the model prompt is generated, the first optical image classification model is optimized based on the model prompt to generate a second optical image classification model. The second optical image classification model performs size change and sliding block cutting processing on the input optical image to be classified to obtain the optical image vector to be classified, and the collaborative effect of each module in the second optical image classification model is used to classify the optical image to be classified to obtain the classification result. By optimizing the first optical image classification model that has completed the current incremental task into the second optical image classification model, the present invention can further improve the classification accuracy of the optical image classification model.

[0073] In summary, the embodiment of the present invention provides a class incremental learning method based on prompt generation and prompt transformation, which obtains a first optical image classification model and key-value pairs corresponding to several historical tasks of the first optical image classification model; inputs the new training optical images of the current incremental task into the first optical image classification model to extract several layers of features and high-level semantic features of the new training optical images; constructs a prompt generation module in the first optical image classification model, and uses the prompt generation module to generate a preliminary model prompt according to the several layers of features and high-level semantic features of the new training optical images; constructs a prompt transformation module in the first optical image classification model, and uses the prompt transformation module to perform information fusion on the preliminary model prompt according to the key-value pairs corresponding to several historical tasks to generate a model prompt; optimizes the first optical image classification model based on the model prompt to generate a second optical image classification model; obtains the optical image to be classified, and inputs the optical image to be classified into the second optical image classification model to obtain the classification result of the optical image to be classified. The present invention effectively reduces the optimization training amount of the optical image classification model by generating a model prompt, improves the data relevance of the model prompt by combining historical classified optical images to generate the model prompt, and further improves the accuracy of the optimized optical image classification model.

[0074] Embodiment 2

[0075] As Figure 1 shown, it is a schematic flowchart of an embodiment of the class incremental learning method based on prompt generation and prompt transformation provided by the present invention. The method includes steps 101 to 106, and the specific steps are as follows:

[0076] Step 101: Obtain a first optical image classification model and key-value pairs corresponding to several historical tasks of the first optical image classification model.

[0077] Further, in the embodiment of the present invention, obtaining a first optical image classification model and key-value pairs corresponding to several historical tasks of the first optical image classification model is specifically:

[0078] Obtain several historical tasks of the first optical image classification model;

[0079] Obtain the high-level semantic features and classification results of each historical classified optical image in each of the historical tasks;

[0080] Generate key-value pairs corresponding to each historical task according to the high-level semantic features and classification results of each historical classified optical image.

[0081] In an embodiment of the present invention, the first optical image classification model is a classification model with ViT as the backbone network and trained. After the first optical image classification model executes historical tasks, it records the high-level semantic features and classification results of each historical task, and then generates key vectors and value vectors for each historical task to form learnable key-value pairs corresponding to each historical task. Therefore, when the first optical image classification model has a new incremental task, the first optical image classification model and the key-value pairs corresponding to multiple historical tasks executed by the first optical image classification model are obtained. By obtaining the key-value pairs corresponding to historical tasks, the key vectors and value vectors corresponding to historical classified optical images can be known. Applying these historical data to the generation of model prompts for newly trained optical images in the future can effectively improve the relevance of model prompts.

[0082] Step 102: Input the newly trained optical image of the current incremental task into the first optical image classification model, and extract several layers of features and high-level semantic features of the newly trained optical image.

[0083] Furthermore, in an embodiment of the present invention, inputting the newly trained optical image of the current incremental task into the first optical image classification model and extracting several layers of features and high-level semantic features of the newly trained optical image specifically includes:

[0084] Input the newly trained optical image into the first optical image classification model;

[0085] Use several hierarchical structures in the first optical image classification model to extract features from the newly trained optical image respectively to obtain several layers of features of the newly trained optical image;

[0086] Use the last hierarchical structure in the first optical image classification model to extract features from the newly trained optical image to obtain the high-level semantic features of the newly trained optical image.

[0087] In an embodiment of the present invention, when the first optical image classification model receives the newly trained optical image of the current incremental task, input the newly trained optical image into the first optical image classification model, and use the hierarchical structures in the first optical image classification model to extract features from the newly trained optical image respectively to obtain each layer of features and high-level semantic features of the newly trained optical image. Among them, each layer of features of the newly trained optical image is obtained by using each hierarchical structure in the first optical image classification model for feature extraction. And the high-level semantic features of the newly trained optical image are obtained by using the last hierarchical structure in the first optical image classification model for feature extraction.

[0088] Step 103: Construct a prompt generation module in the first optical image classification model, and use the prompt generation module to generate a preliminary model prompt according to several layers of features and high-level semantic features of the newly trained optical image.

[0089] Further, in the embodiment of the present invention, constructing a prompt generation module in the first optical image classification model, and using the prompt generation module to generate a preliminary model prompt according to several layers of features and high-level semantic features of the newly trained optical image, specifically:

[0090] Construct a prompt generation module in the first optical image classification model; the prompt generation module includes a normalization layer, several layers of perceptrons, a cross-attention layer, and a fully connected layer;

[0091] Input the high-level semantic features of the newly trained optical image into the normalization layer and several layers of perceptrons to generate mapped high-level semantic features;

[0092] Input several layers of features and the mapped high-level semantic features of the newly trained optical image into the cross-attention layer to generate an attention result;

[0093] Use residual connection to process the attention result to form enhanced high-level semantic features of the newly trained optical image;

[0094] Input the enhanced high-level semantic features of the newly trained optical image into the fully connected layer to generate a preliminary model prompt.

[0095] In the embodiment of the present invention, after obtaining the features of each layer and high-level semantic features of the newly trained optical image, construct a prompt generation module in the first optical image classification model, use the features of each layer and high-level semantic features of the newly trained optical image as the input of the prompt generation module, and generate a preliminary model prompt through attention calculation, residual operation, and fully connected layer processing.

[0096] Further, in the embodiment of the present invention, inputting several layers of features and the mapped high-level semantic features of the newly trained optical image into the cross-attention layer to generate an attention result, specifically:

[0097] Input several layers of features and the mapped high-level semantic features of the newly trained optical image into the cross-attention layer, set several layers of features of the newly trained optical image as query vectors, set the mapped high-level semantic features of the newly trained optical image as key vectors and value vectors, and calculate the attention result.

[0098] Preferably, the prompt generation module includes a normalization layer, several layers of perceptrons, a cross-attention layer, and a fully connected layer. First, the high-level semantic features of the newly trained optical image are used as the input of the normalization layer and the multi-layer perceptron, so that the normalization layer and the multi-layer perceptron output the mapped high-level semantic features. The features of each layer of the newly trained optical image and the mapped high-level semantic features are used as the input of the cross-attention layer. The features of each layer are set as the Q (query vector) of the attention mechanism, and the mapped high-level semantic features are set as the K (key vector) and V (value vector), and the corresponding attention results are calculated through a lightweight cross-attention layer. Then, the residual operation (Res) between the attention result and the features of each layer is further set, and the two are added to obtain the enhanced high-level semantic features. Finally, the enhanced high-level semantic features are input into the fully connected layer (FC) to output the final preliminary model prompt of the current layer.

[0099] As an example of the embodiment of the present invention, the application process of the prompt generation module can be represented by the following formula:

[0100] P l =FC(Res(LA(F l ,K l ,V l ),F l )) T

[0101] In the formula, FC(·) is the fully connected layer operation; Res(·) is the residual operation; LA(·) is the attention calculation operation; P l is the preliminary model prompt; F l is the feature of the l-th layer; K l is the key vector of the l-th layer; V l is the value vector of the l-th layer.

[0102] Among them, K l =V l =MLP l (Norm(F)), F is the high-level semantic feature, F, F r and F l have dimensions of and contain d-dimensional features with a length of L t . The corresponding obtained prompt P l has dimensions of and contains d-dimensional prompts with a length of L p .

[0103] Step 104: Construct a prompt transformation module in the first optical image classification model, and use the prompt transformation module to perform information fusion on the preliminary model prompt according to the key-value pairs corresponding to a plurality of the historical tasks to generate a model prompt.

[0104] Further, in the embodiment of the present invention, a prompt transformation module is constructed in the first optical image classification model, and the prompt transformation module is used to perform information fusion on the preliminary model prompt according to the key-value pairs corresponding to a plurality of the historical tasks to generate a model prompt. Specifically:

[0105] Construct a prompt transformation module in the first optical image classification model;

[0106] Input the high-level semantic features of the newly trained optical image into the prompt transformation module to generate downsampled features;

[0107] Generate scaling parameters and translation coefficients based on the key-value pairs corresponding to a plurality of the historical tasks and the downsampled features;

[0108] Perform a linear transformation on the preliminary model prompt based on the scaling parameters and the translation coefficients to generate a model prompt. Specifically:

[0109] P fl = s c ·P l + s f

[0110] In the formula, P fl is the model prompt; P l is the preliminary model prompt; s c is the scaling parameter; s f is the translation coefficient.

[0111] In the embodiment of the present invention, a prompt transformation module is constructed in the first optical image classification model. Based on this prompt transformation module, the CLS features of the high-level semantic features of the newly trained optical image can be generated. Based on the key-value pairs corresponding to multiple historical tasks executed by the first optical image classification model and the CLS features, the scaling coefficient and the translation coefficient of the current incremental task can be generated. Furthermore, based on the scaling coefficient and the translation coefficient, a linear transformation is performed on the preliminary model prompt generated by the prompt generation module to complete information fusion and generate a model prompt that combines the fusion instance and task information. By combining historical data to generate a model prompt, the present invention can effectively obtain the correlation between the newly trained optical image and the historical data, thereby improving the data correlation of the model prompt.

[0112] Further, in the embodiment of the present invention, inputting the high-level semantic features of the newly trained optical image into the prompt transformation module to generate downsampled features, specifically:

[0113] Input the high-level semantic features of the newly trained optical image into the prompt transformation module, extract the downsampled features in the high-level semantic features of the newly trained optical image, and generate initial CLS features;

[0114] Using several layers of perceptrons in the prompt transformation module, downsample the initial CLS feature to generate a downsampled feature.

[0115] In the embodiment of the present invention, the CLS feature in the high-level semantic feature of the newly trained optical image can be extracted by using the prompt transformation module to obtain the initial CLS feature. Further, to reduce the parameter calculation amount, the initial CLS feature can be downsampled by setting a multi-layer perceptron (MLP) to obtain the downsampled feature.

[0116] Further, in the embodiment of the present invention, based on the key-value pairs corresponding to several historical tasks and the downsampled feature, a scaling parameter and a translation coefficient are generated, specifically:

[0117] Obtain several key vectors from the key-value pairs corresponding to several historical tasks;

[0118] Calculate several similarities between the downsampled feature and several key vectors;

[0119] Based on several similarities, perform weighted summation on several value vectors corresponding to several historical tasks to obtain a summation result;

[0120] Perform dimensionality splitting on the summation result to generate a scaling parameter and a translation coefficient.

[0121] In the embodiment of the present invention, after obtaining the CLS feature of the high-level semantic feature of the newly trained optical image, calculate the similarity between the downsampled feature and the key vector corresponding to each task history, and perform weighted summation on the value vectors of each historical task based on this similarity, so as to obtain a summation result. Continuing to perform dimensionality splitting on this summation result can generate a scaling parameter and a translation coefficient. For example, when the matrix dimension of the summation result is 1536, the summation result can be evenly divided in dimension to obtain two matrices with a matrix dimension of 768, which are the scaling parameter and the translation coefficient respectively.

[0122] The scaling coefficient and translation coefficient of the current incremental task are used for linear transformation of the preliminary model prompt.

[0123] Step 105: Optimize the first optical image classification model based on the model prompt to generate a second optical image classification model.

[0124] Further, in the embodiment of the present invention, optimizing the first optical image classification model based on the model prompt to generate a second optical image classification model is specifically:

[0125] Set the model prompt at several hierarchical structures of the first optical image classification model, set the output dimension of the last layer classifier of the first optical image classification model to the total number of categories of the current incremental task, and control the first optical image classification model to output the classification prediction result of the newly trained optical image;

[0126] Calculate the cross-entropy loss between the classification prediction result of the newly trained optical image and the class true value, and optimize the first optical image classification model according to the cross-entropy loss to generate a second optical image classification model.

[0127] In the embodiment of the present invention, after generating the model prompt, the first optical image classification model can be optimized based on the model prompt to generate a second optical image classification model, so that the second optical image classification model classifies the input optical image to be classified and obtains the classification result of the optical image to be classified. The first optical image classification model is optimized to the second optical image classification model after completing the current incremental task, which can further improve the classification accuracy of the optical image classification model.

[0128] Preferably, the optimization process of the first optical image classification model can be expressed as: select the frozen pre-trained ViT in the first optical image classification model as the backbone network, set the above-obtained model prompt at each hierarchical structure of the backbone network, modify the output dimension of the last layer classifier to the total number of categories of the current incremental task, and output the classifier as the classification prediction result of the current newly trained optical image. Among them, in the optimization process of the first optical image classification model, to avoid interference between tasks, it can be set that the fully connected layer of the prompt generation module and the multi-layer perceptron layer of the prompt transformation module are frozen after training, and at the same time, the learnable key vector and value vector corresponding to the historical task are frozen after learning. In the process of optimizing the first optical image classification model, the cross-entropy loss between the classification prediction result and the class true value can be used as the classification loss, and the prompt generation module, prompt transformation module and classifier can be optimized simultaneously through the classification loss in an end-to-end manner.

[0129] As an example of the embodiment of the present invention, see Figure 2 , which is a schematic flowchart of another embodiment of the class incremental learning method based on prompt generation and prompt transformation provided by the present invention. When the first optical image classification model receives the newly trained optical image of the current incremental task, the features F of each layer are extracted by using the first optical image classification model l and the high-level semantic feature F. Let F lSet as the query vector; through the normalization layer (Norm) and the multi-layer perceptron (MLP), set the high-level semantic feature F as the key vector and the value vector. After setting the query vector, key vector, and value vector, use the cross-attention layer to perform attention calculation. Perform a residual operation on the obtained attention result, and output the instance-specific prompt P through the fully connected layer (FC) for the result of the residual calculation l On the other hand, after extracting the high-level semantic features of the newly trained optical image, extract the initial CLS feature F of the high-level semantic features cls , and downsample the initial CLS feature through the multi-layer perceptron (MLP) to obtain the downsampled feature F d . Calculate the similarity between the downsampled feature F d and the key vectors corresponding to each task history, and perform a weighted sum on the value vectors of each historical task based on each similarity, so as to generate the scaling coefficient and translation coefficient of the current incremental task. After obtaining the scaling coefficient and translation coefficient, perform a linear transformation on the instance-specific prompt P l , and finally obtain the prompt P that fuses instance and task information fl .

[0130] Step 106: Obtain the optical image to be classified, input the optical image to be classified into the second optical image classification model, so that the second optical image classification model performs size change and sliding block cutting processing on the optical image to be classified, generate the optical image vector to be classified, and classify the optical image vector to be classified to obtain the classification result of the optical image to be classified

[0131] In the embodiment of the present invention, after the first optical image classification model is optimized to generate the second optical image classification model, the second optical image classification model is used to classify the optical image to be classified, so as to obtain a more accurate model classification result. Specifically, when the optical image to be classified is input into the second optical image classification model, first perform a size change on the optical image to be classified, and the specific methods include direct scaling or random cropping. Then perform sliding block cutting processing on the image. For example, when the size of the optical image to be classified is 224x224, non-overlapping sliding blocks can be made according to the size of 16x16, and a total of 196 image blocks can be obtained after division. Each image block has three channels and a dimension of (16, 16, 3), and then it is flattened into a vector with a length of 768. Each vector is regarded as a separate input, and there are 196 vectors in total. Then use the collaborative effect of each module in the second optical image classification model to classify the optical image to be classified to obtain the classification result

[0132] In summary, the embodiment of the present invention provides a class incremental learning method based on prompt generation and prompt transformation, which obtains a first optical image classification model and key-value pairs corresponding to several historical tasks of the first optical image classification model; inputs new training optical images of the current incremental task into the first optical image classification model to extract several layers of features and high-level semantic features of the new training optical images; constructs a prompt generation module in the first optical image classification model, and uses the prompt generation module to generate a preliminary model prompt according to the several layers of features and high-level semantic features of the new training optical images; constructs a prompt transformation module in the first optical image classification model, and uses the prompt transformation module to perform information fusion on the preliminary model prompt according to the key-value pairs corresponding to several historical tasks to generate a model prompt; optimizes the first optical image classification model based on the model prompt to generate a second optical image classification model; obtains an optical image to be classified, inputs the optical image to be classified into the second optical image classification model, and obtains the classification result of the optical image to be classified. The present invention effectively reduces the optimization training amount of the optical image classification model by generating model prompts, improves the data relevance of the model prompts by combining historical classified optical images to generate model prompts, and further improves the accuracy of the optimized optical image classification model.

[0133] Embodiment 3

[0134] See Figure 3 FIG. is a schematic structural diagram of an embodiment of a class incremental learning system based on prompt generation and prompt transformation provided by the present invention. The system includes a data acquisition module 201, a feature extraction module 202, a first generation module 203, a second generation module 204, a model optimization module 205, and a model application module 206;

[0135] The data acquisition module 201 is used to obtain a first optical image classification model and key-value pairs corresponding to several historical tasks of the first optical image classification model;

[0136] The feature extraction module 202 is used to input new training optical images of the current incremental task into the first optical image classification model to extract several layers of features and high-level semantic features of the new training optical images;

[0137] The first generation module 203 is used to construct a prompt generation module in the first optical image classification model, and use the prompt generation module to generate a preliminary model prompt according to the several layers of features and high-level semantic features of the new training optical images;

[0138] The second generation module 204 is used to construct a prompt transformation module in the first optical image classification model, and use the prompt transformation module to perform information fusion on the preliminary model prompt according to the key-value pairs corresponding to several historical tasks to generate a model prompt;

[0139] The model optimization module 205 is used to optimize the first optical image classification model based on the model prompt, and generate a second optical image classification model;

[0140] The model application module 206 is used to obtain the optical image to be classified, input the optical image to be classified into the second optical image classification model, so that the second optical image classification model performs size change and sliding block cutting processing on the optical image to be classified, generate a vector of the optical image to be classified, and classify the vector of the optical image to be classified to obtain the classification result of the optical image to be classified.

[0141] Further, in the embodiment of the present invention, obtaining the first optical image classification model and the key-value pairs corresponding to a plurality of historical tasks of the first optical image classification model specifically includes:

[0142] Obtaining a plurality of historical tasks of the first optical image classification model;

[0143] Obtaining the high-level semantic features and classification results of each historical classified optical image in each of the historical tasks;

[0144] Generating key-value pairs corresponding to each historical task according to the high-level semantic features and classification results of each historical classified optical image.

[0145] Further, in the embodiment of the present invention, inputting the newly trained optical image of the current incremental task into the first optical image classification model, and extracting several layers of features and high-level semantic features of the newly trained optical image specifically includes:

[0146] Inputting the newly trained optical image into the first optical image classification model;

[0147] Using several hierarchical structures in the first optical image classification model to respectively extract features from the newly trained optical image to obtain several layers of features of the newly trained optical image;

[0148] Using the last hierarchical structure in the first optical image classification model to extract features from the newly trained optical image to obtain the high-level semantic features of the newly trained optical image.

[0149] Further, in the embodiment of the present invention, a prompt generation module is constructed in the first optical image classification model, and the prompt generation module is used to generate a preliminary model prompt according to several layers of features and high-level semantic features of the newly trained optical image, specifically including:

[0150] Constructing a prompt generation module in the first optical image classification model; the prompt generation module includes a normalization layer, several layers of perceptrons, a cross-attention layer and a fully connected layer;

[0151] Input the high-level semantic features of the newly trained optical image into a normalization layer and several layers of perceptrons to generate the mapped high-level semantic features;

[0152] Input the several layers of features of the newly trained optical image and the mapped high-level semantic features into a cross-attention layer to generate an attention result;

[0153] Use residual connection to process the attention result to form the enhanced high-level semantic features of the newly trained optical image;

[0154] Input the enhanced high-level semantic features of the newly trained optical image into a fully connected layer to generate a preliminary model prompt.

[0155] Furthermore, in the embodiment of the present invention, inputting the several layers of features of the newly trained optical image and the mapped high-level semantic features into a cross-attention layer to generate an attention result, specifically:

[0156] Input the several layers of features of the newly trained optical image and the mapped high-level semantic features into a cross-attention layer, set the several layers of features of the newly trained optical image as query vectors, set the mapped high-level semantic features of the newly trained optical image as key vectors and value vectors, and calculate the attention result.

[0157] Furthermore, in the embodiment of the present invention, construct a prompt transformation module in the first optical image classification model, and use the prompt transformation module to perform information fusion on the preliminary model prompt according to the key-value pairs corresponding to several historical tasks to generate a model prompt, specifically:

[0158] Construct a prompt transformation module in the first optical image classification model;

[0159] Input the high-level semantic features of the newly trained optical image into the prompt transformation module to generate downsampled features;

[0160] Generate scaling parameters and translation coefficients based on the key-value pairs corresponding to several historical tasks and the downsampled features;

[0161] Perform a linear transformation on the preliminary model prompt based on the scaling parameters and the translation coefficients to generate a model prompt, specifically:

[0162] P fl =s c ·P l +s f

[0163] In the formula, P fl is the model prompt; P l is the preliminary model prompt; s cis the scaling parameter; s f is the translation coefficient.

[0164] Furthermore, in the embodiment of the present invention, the high-level semantic features of the new training optical image are input into the prompt transformation module to generate downsampled features, specifically:

[0165] The high-level semantic features of the new training optical image are input into the prompt transformation module, the downsampled features in the high-level semantic features of the new training optical image are extracted, and initial CLS features are generated;

[0166] Using several layers of perceptrons in the prompt transformation module, the initial CLS features are downsampled to generate downsampled features.

[0167] Furthermore, in the embodiment of the present invention, based on the key-value pairs corresponding to several historical tasks and the downsampled features, a scaling parameter and a translation coefficient are generated, specifically:

[0168] Obtain several key vectors from the key-value pairs corresponding to several historical tasks;

[0169] Calculate several similarities between the downsampled features and several key vectors respectively;

[0170] Based on several similarities, perform weighted summation on several value vectors corresponding to several historical tasks to obtain a summation result;

[0171] Perform dimension splitting on the summation result to generate a scaling parameter and a translation coefficient.

[0172] Furthermore, in the embodiment of the present invention, based on the model prompt, the first optical image classification model is optimized to generate a second optical image classification model, specifically:

[0173] Set the model prompt in several hierarchical structures of the first optical image classification model, set the output dimension of the last layer classifier of the first optical image classification model to the total number of categories of the current incremental task, and control the first optical image classification model to output the classification prediction result of the new training optical image;

[0174] Calculate the cross-entropy loss between the classification prediction result of the new training optical image and the class true value, and optimize the first optical image classification model according to the cross-entropy loss to generate a second optical image classification model.

[0175] In summary, the embodiment of the present invention provides a class incremental learning system based on prompt generation and prompt transformation. Based on the organic combination of modules, a first optical image classification model and key-value pairs corresponding to several historical tasks of the first optical image classification model are obtained; the new training optical images of the current incremental task are input into the first optical image classification model to extract several layers of features and high-level semantic features of the new training optical images; a prompt generation module is constructed in the first optical image classification model, and the prompt generation module is used to generate a preliminary model prompt according to several layers of features and high-level semantic features of the new training optical images; a prompt transformation module is constructed in the first optical image classification model, and the prompt transformation module is used to perform information fusion on the preliminary model prompt according to the key-value pairs corresponding to several historical tasks to generate a model prompt; the first optical image classification model is optimized based on the model prompt to generate a second optical image classification model; the optical image to be classified is obtained, and the optical image to be classified is input into the second optical image classification model to obtain the classification result of the optical image to be classified. The present invention effectively reduces the optimization training amount of the optical image classification model by generating a model prompt, improves the data relevance of the model prompt by combining historical classified optical images to generate the model prompt, and further improves the accuracy of the optimized optical image classification model.

[0176] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A class incremental learning method based on prompt generation and prompt transformation, characterized in that, Including: Obtain a first optical image classification model and key-value pairs corresponding to several historical tasks of the first optical image classification model; Input a newly trained optical image of the current incremental task into the first optical image classification model, and extract several layers of features and high-level semantic features of the newly trained optical image; Construct a prompt generation module in the first optical image classification model, and use the prompt generation module to generate a preliminary model prompt according to several layers of features and high-level semantic features of the newly trained optical image; Construct a prompt transformation module in the first optical image classification model, and use the prompt transformation module to perform information fusion on the preliminary model prompt according to key-value pairs corresponding to several historical tasks to generate a model prompt; Optimize the first optical image classification model based on the model prompt to generate a second optical image classification model; Obtain an optical image to be classified, input the optical image to be classified into the second optical image classification model, so that the second optical image classification model performs size change and sliding block cutting processing on the optical image to be classified, generates a vector of the optical image to be classified, and classifies the vector of the optical image to be classified to obtain a classification result of the optical image to be classified.

2. The class incremental learning method based on prompt generation and prompt transformation according to claim 1, characterized in that, The obtaining of the first optical image classification model and key-value pairs corresponding to several historical tasks of the first optical image classification model is specifically: Obtain several historical tasks of the first optical image classification model; Obtain high-level semantic features and classification results of each historical classified optical image in each historical task; Generate key-value pairs corresponding to each historical task according to the high-level semantic features and classification results of each historical classified optical image.

3. The class incremental learning method based on prompt generation and prompt transformation according to claim 1, wherein The inputting of the newly trained optical image of the current incremental task into the first optical image classification model and extracting several layers of features and high-level semantic features of the newly trained optical image is specifically: Input the newly trained optical image into the first optical image classification model; Use several hierarchical structures in the first optical image classification model to extract features from the newly trained optical image respectively to obtain several layers of features of the newly trained optical image; Use the last hierarchical structure in the first optical image classification model to extract features from the newly trained optical image to obtain high-level semantic features of the newly trained optical image.

4. The class incremental learning method based on prompt generation and prompt transformation according to claim 1, characterized in that, The constructing of a prompt generation module in the first optical image classification model and using the prompt generation module to generate a preliminary model prompt according to several layers of features and high-level semantic features of the newly trained optical image is specifically: Construct a prompt generation module in the first optical image classification model; the prompt generation module includes a normalization layer, several layers of perceptrons, a cross-attention layer and a fully connected layer; Input the high-level semantic features of the newly trained optical image into the normalization layer and several layers of perceptrons to generate mapped high-level semantic features; Input several layers of features and mapped high-level semantic features of the newly trained optical image into the cross-attention layer to generate an attention result; Process the attention result using a residual connection to form enhanced high-level semantic features of the new training optical image; Input the enhanced high-level semantic features of the new training optical image into a fully connected layer to generate a preliminary model prompt.

5. The class incremental learning method based on prompt generation and prompt transformation according to claim 4, characterized in that, The step of inputting several layers of features and the mapped high-level semantic features of the new training optical image into a cross-attention layer to generate an attention result is specifically: Input several layers of features and the mapped high-level semantic features of the new training optical image into a cross-attention layer, set several layers of features of the new training optical image as query vectors, set the mapped high-level semantic features of the new training optical image as key vectors and value vectors, and calculate the attention result.

6. The class incremental learning method based on prompt generation and prompt transformation according to claim 1, characterized in that The step of constructing a prompt transformation module in the first optical image classification model and using the prompt transformation module to perform information fusion on the preliminary model prompt according to key-value pairs corresponding to several historical tasks to generate a model prompt is specifically: Construct a prompt transformation module in the first optical image classification model; Input the high-level semantic features of the new training optical image into the prompt transformation module to generate downsampled features; Generate a scaling parameter and a translation coefficient based on key-value pairs corresponding to several historical tasks and the downsampled features; Perform a linear transformation on the preliminary model prompt based on the scaling parameter and the translation coefficient to generate a model prompt, specifically: P fl = s c · P l + s f Wherein, P fl is the model prompt; P l is the preliminary model prompt; s c is the scaling parameter; s f is the translation coefficient.

7. The class incremental learning method based on prompt generation and prompt transformation according to claim 6, characterized in that The step of inputting the high-level semantic features of the new training optical image into the prompt transformation module to generate downsampled features is specifically: Input the high-level semantic features of the new training optical image into the prompt transformation module, extract the downsampled features from the high-level semantic features of the new training optical image to generate initial CLS features; Use several layers of perceptrons in the prompt transformation module to downsample the initial CLS features to generate downsampled features.

8. The class incremental learning method based on prompt generation and prompt transformation according to claim 6, characterized in that The step of generating a scaling parameter and a translation coefficient based on key-value pairs corresponding to several historical tasks and the downsampled features is specifically: Obtain several key vectors from key-value pairs corresponding to several historical tasks; Calculate several similarities between the downsampled features and several key vectors respectively; Based on several similarities, perform a weighted sum of several value vectors corresponding to several historical tasks to obtain a summation result; Perform dimensional splitting on the summation result to generate a scaling parameter and a translation coefficient.

9. The class incremental learning method based on prompt generation and prompt transformation according to claim 1, wherein The step of optimizing the first optical image classification model based on the model prompt to generate a second optical image classification model is specifically: Set the model prompt in several hierarchical structures of the first optical image classification model, set the output dimension of the last layer classifier of the first optical image classification model to the total number of categories of the current incremental task, and control the first optical image classification model to output the classification prediction result of the new training optical image; Calculate the cross-entropy loss between the classification prediction result of the new training optical image and the class true value, and optimize the first optical image classification model according to the cross-entropy loss to generate a second optical image classification model.

10. A class incremental learning system based on prompt generation and prompt transformation, characterized in that, including: A data acquisition module, a feature extraction module, a first generation module, a second generation module, a model optimization module, and a model application module; The data acquisition module is used to acquire a first optical image classification model and key-value pairs corresponding to a number of historical tasks of the first optical image classification model; The feature extraction module is used to input a newly trained optical image of the current incremental task into the first optical image classification model, and extract several layers of features and high-level semantic features of the newly trained optical image; The first generation module is used to construct a prompt generation module in the first optical image classification model, and use the prompt generation module to generate a preliminary model prompt according to several layers of features and high-level semantic features of the newly trained optical image; The second generation module is used to construct a prompt transformation module in the first optical image classification model, and use the prompt transformation module to perform information fusion on the preliminary model prompt according to key-value pairs corresponding to a number of the historical tasks to generate a model prompt; The model optimization module is used to optimize the first optical image classification model based on the model prompt to generate a second optical image classification model; The model application module is used to acquire an optical image to be classified, input the optical image to be classified into the second optical image classification model, so that the second optical image classification model performs size change and sliding block cutting processing on the optical image to be classified, generate an optical image vector to be classified, and classify the optical image vector to be classified to obtain a classification result of the optical image to be classified.