Method and system for prompt adaptation during testing of visual language model

By using the method of exchanging prompts when testing visual language models, using the interaction and learning of online prompts and target prompts, the problem of failure to fully utilize the pre-trained model representation ability in the prior art is solved, and efficient adaptive performance during testing is achieved.

CN120032155APending Publication Date: 2025-05-23THE HONG KONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411690931.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-23
Filing Date
2024-11-25
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art fails to fully utilize the representation ability of pre-trained models in the testing of visual language models, and relies on computationally intensive backbone network fine-tuning, resulting in lower performance than supervised prompt adaptive methods.

Method used

A method of exchanging cues is proposed, adopting a dual-cues paradigm, including online cues and target cues. Through the exchange prediction mechanism and exponential moving average update strategy, the online cues and target cues can interact and learn from each other, thereby adapting to the test domain.

Benefits of technology

By exchanging prompts, the representation ability of the pre-trained model can be effectively utilized to improve the adaptive performance during testing, and achieve the effect comparable to the supervised prompt adaptive method, while reducing the computational density.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032155A_ABST
    Figure CN120032155A_ABST
Patent Text Reader

Abstract

The invention provides a method for prompt self-adaption during testing of a visual language model, and the method comprises the steps that an exchange prompt is used during testing of the visual language model, and the exchange prompt is a double-prompt normal form and comprises an online prompt and a target prompt. The online cues are trained through comparative learning using an exchange prediction mechanism. The target cue is updated based on the online cue, and the online cue can be optimally trained according to the target cue such that the online cue and the target cue can interact and learn each other. By using the scheme of the invention, a series of prompts of which the quality is improved over time can be created, so that the prompts can learn more knowledge from the strong representation capability of the pre-training model and can better adapt to downstream image classification tasks, thereby realizing high accuracy based on exchange prompts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to methods and systems for test-time prompt adaptation of visual-language models. Background Art

[0002] Test-time adaptation (TTA) is a special and practical setting in unsupervised domain adaptation, which allows a pre-trained model in a source domain to adapt to unlabeled test data in another target domain. To avoid the computationally intensive backbone network fine-tuning process, the zero-shot generalization potential of emerging pre-trained visual-language models (e.g., CLIP, CoOp) is exploited to only adapt to the runtime cues of the unseen test domain. However, existing schemes have not fully exploited the representation power of pre-trained models, as they only focus on entropy-based optimization and perform much worse than supervised cue adaptation methods, such as CoOp.

[0003] The generalization performance of deep neural networks can suffer when there is a difference between the distribution of training and test data [1, 2, 3]. The focus of domain adaptation is to build models that can adapt to changes in data distribution by transferring knowledge from the source domain to a new and related target domain, which usually requires both source and target domain data during the training phase [4, 5]. However, in real-world scenarios, one usually only has a model trained in the source domain without access to the source data or authorization to change the original training process [6, 7, 8]. To address this issue, test-time adaptation (TTA) [9, 10] has been proposed and has shown the potential to adapt models to the target / unseen domain by leveraging only unlabeled test data streams. Existing work has developed techniques such as entropy minimization [7, 11], category prototypes [12, 13], image generation [14, 15], and self-supervised training

[10] , which have shown remarkable performance.

[0004] Although traditional model-based TTA methods have been shown to be effective, they often rely on computationally intensive tuning of the model backbone parameters. The situation becomes worse with the advent of vision-language pre-trained models (e.g., CLIP

[16] , CoOp

[17] , and CoCoOp

[18] ), which have a large number of parameters and are difficult to optimize. Therefore, it is desirable to explore effective techniques to fine-tune only a small set of parameters for adapting the model to the new domain during testing while keeping the backbone fixed.

[0005] Test-time adaptation refers to reducing the performance gap when the source model is deployed on a different target domain from the test data. The challenge of this problem is that only the source model and unlabeled test data are available, and the training process and the source data should not be accessible. Many schemes have been proposed to address this problem, namely, minimizing the entropy of model predictions [7, 11], maintaining a set of dynamic prototypes and measuring the similarity between the test sample and each prototype [12, 13], generating new data similar to the target domain to assist model adaptation [14, 15], and leveraging the idea of ​​self-supervised training to improve generalization ability

[10] . However, test-time adaptation in vision-language models has not been fully explored.

[0006] Pre-trained visual language models are trained on a large number of image-text pairs, which introduces a powerful paradigm and provides new insights to solve the above problems. A simple approach is to leverage the strong zero-shot learning ability of pre-trained visual language models to distinguish the domains of test data by fine-tuning on labeled data of downstream tasks. However, this approach may not be feasible in TTA scenarios where labeled downstream data is unavailable. Shu et al.

[19] proposed test-time prompt tuning (TPT) to address the label scarcity problem at test time. However, it is limited in practice to tune instance-specific prompts from directly minimizing entropy, which may lead to the risk of over-trusting the model (i.e., generating high confidence in wrong results).

[0007] Hint learning in visual language models first appeared in the field of natural language processing (NLP), aiming to enhance the performance of pre-trained models by leveraging different hints. With the emergence of visual language models that integrate visual and textual modalities, inspired by hint learning in NLP, CoOp

[16] was proposed for hint learning, which transforms handcrafted hints into learnable continuous hints and adapts them to downstream tasks. CoCoOp

[17] (an improvement on CoOp) adopts a meta-network and image features to generate individual hints for each image. In addition, there are some other methods (such as CLIP adapter

[22] and Tip adapter

[23] ) that do not modify the hints but add additional classification layers after the backbone model. What they have in common is that they all rely heavily on a set of labeled data, which makes them unsuitable for test-time settings. Another work focuses on implementing hint learning in an unsupervised manner during the training process, namely unsupervised hint learning (UPL)

[24] . However, it only extends the pseudo-labeling method to visual language models without fully utilizing the powerful representation power of pre-trained models.

[0008] Therefore, a solution for providing adaptive prompts during testing is needed to solve at least one of the above problems. Summary of the invention

[0009] The present disclosure aims to propose a solution that can solve or alleviate at least one of the above problems.

[0010] According to a first aspect of the present invention, a method for prompt adaptation during testing of a visual language model is provided, the method comprising:

[0011] When testing the visual language model, an exchange prompt is used, wherein the exchange prompt is a double prompt demonstration format, including: an online prompt and a target prompt;

[0012] wherein the online prompt is trained by contrastive learning using a swap prediction mechanism,

[0013] The target prompt is updated based on the online prompt, and the online prompt can be optimized and trained according to the target prompt, so that the online prompt and the target prompt can interact and learn from each other.

[0014] In one embodiment, the exchange prediction mechanism comprises predicting the category assignment of the target cue to two or more enhanced views of the same test image based on the trained online cue and / or the updated target cue.

[0015] In one embodiment, the target prompt is arranged to be averaged from the online prompts, such that the exchange prompt updates the target prompt using an exponential moving average of the online prompts, wherein the target prompt is represented as:

[0016] t t =∈t t +(1-∈)t O ,

[0017] Among them, t o Indicates online prompt, t t represents the target cue, ∈ represents the decay rate of the target cue, ∈∈[0,1].

[0018] In one embodiment, the exchange prediction mechanism further includes using text features generated by the target cue as a prototype, using the prototype to generate two or more enhanced views of the test image to obtain category assignments or predictions of the target cue for the two or more enhanced views.

[0019] In one embodiment, the contrastive learning adopts a loss function, wherein the loss function includes a first loss function and a second loss function, wherein the sum of the first loss function and the second loss function is used as the loss function for the test image x established under the exchange prediction mechanism. i The loss function of .

[0020] In one embodiment, the first loss function is a prompt exchange prediction loss function L swap , and the second loss function is the pseudo-label loss function L pseudo ,in,

[0021]

[0022] Wherein, m and n represent any two different enhancement methods for obtaining the first enhanced view and the second enhanced view based on the test image, Represents the test image x i Prediction of online prompts for the first enhanced view, Represents the test image x i The prediction of the online prompt of the second enhanced view, Represents the test image x i The prediction of the target cue of the first augmented view, Represents the test image x i The prediction of the target cue of the second enhanced view, Represents the test image x i The pseudo-marker obtained, l 1 is a function representing the difference between the predictions generated with respect to the online cue and the predictions generated with respect to the target cue, l 2 Function representing the difference between predictions and pseudo-labelings generated with respect to online prompts.

[0023] In one embodiment, the method further comprises: in establishing a i The first loss function and the second loss function and before obtaining the pseudo label are based on at least the test image x i The confidence of the relevant test data is selected, and the top K test data with the highest confidence are used to form the confidence matrix for the test image x. i Adaptive set.

[0024] In one embodiment, the method further comprises: further establishing a method for the test image x based on the adaptive set under the exchange prediction mechanism. i The loss function L for the adaptive set adapt :

[0025] L adapt (x i )=αL swap (x i )+βL pseudo (x i ),

[0026] Among them, α and β are given trade-off hyperparameters.

[0027] In one embodiment, the method further includes: when the test data of the test image x i is online test data that arrives in the form of batch test data less than a threshold number, based on at least each batch of test data that arrives, perform a corresponding new confidence ranking, and select the top k (k < K) test data with the highest confidence to obtain a corresponding new adaptive set.

[0028] According to a second aspect of the present invention, there is provided a system for test-time prompt adaptation for a vision language model, characterized in that the system includes:

[0029] An online prompt module configured to be trained for online prompts of test images by contrastive learning using an exchange prediction mechanism, and

[0030] A target prompt module configured to update the target prompt based on the online prompt,

[0031] wherein the online prompt module is further configured to optimize the training of the online prompt according to the target prompt, such that the online prompt and the target prompt can interact and learn from each other, and

[0032] wherein the system is configured to use an exchange prompt including the online prompt and the target prompt during the test of the vision language model.

[0033] According to a third aspect of the present invention, there is provided a computer device including a memory and a processor, and computer instructions are stored on the memory, and when the computer instructions are executed by the processor, the method described above is caused to be executed.

[0034] According to a fourth aspect of the present invention, there is provided a non-transitory computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the method described above is caused to be executed.

[0035] According to a fifth aspect of the present invention, there is provided a computer program product including computer instructions, and when the computer instructions are executed by a processor, the method described above is caused to be executed

[0036] The solution of the present invention adopts an exchange prompt and an exchange prediction mechanism including online prompts and target prompts when testing a visual language model, wherein the online prompts and the target prompts can interact and learn from each other, and can create a series of prompts whose quality improves over time, which enables the prompts to learn more knowledge from the powerful representation ability of the pre-trained model, and through contrast learning strategies to enable the prompts to better adapt to downstream image classification tasks, thereby achieving high accuracy based on exchange prompts. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Non-limiting and non-exhaustive embodiments of the present invention are described, by way of example, with reference to the following drawings, in which:

[0038] Figure 1a is an example model architecture in the prior art that uses CoOp to adjust prompts based on labeled data;

[0039] Figure 1b is an example VL-type architecture in the prior art that uses TPT to optimize prompts by minimizing marginal entropy;

[0040] Figure 1c is an example model architecture for exchanging cues using self-supervised contrastive learning to facilitate cue adaptation at test time according to an embodiment of the present invention;

[0041] Figure 2 A flowchart of a method for test-time prompt adaptation of a visual language model according to an embodiment of the present invention is illustrated;

[0042] Figure 3 A schematic block diagram of a system for test-time prompt adaptation of a visual language model according to an embodiment of the present invention is illustrated;

[0043] Figure 4 A schematic diagram of an exemplary exchange prompt framework used in a method for prompt adaptation during testing of a visual language model according to an embodiment of the present invention is illustrated;

[0044] Figure 5a A diagram illustrating the average accuracy over all 14 datasets and compared with CoOp as the upper bound; and

[0045] Figure 5b The average accuracy of the exchange prompt on the five datasets is shown as the different adaptation rounds change. The accuracy of UPL and TPT is the average accuracy of the last round. DETAILED DESCRIPTION

[0046] Figure 1a is an example model architecture in the prior art that uses CoOp to adjust prompts based on labeled data; Figure 1bis an example model VL-type architecture in the prior art that uses TPT to optimize hints by minimizing marginal entropy. Figure 1c is an example model architecture for exchanging cues using self-supervised contrastive learning to facilitate cue adaptation at test time according to an embodiment of the present invention.

[0047] Specifically, this paper proposes swapping prompts, a new test-time prompt adaptation method, such as Figure 1c As shown in the example. Compared with the previous existing methods (ref. Figure 1a and Figure 1b ), the proposed exchange hints utilize a self-supervised contrastive learning strategy in the test domain, which, in one embodiment, may include at least one of the following two key features: exponential moving average (EMA) hints and a hint exchange prediction mechanism. The EMA mechanism adopts a dual hint paradigm: target hint and online hint. In one embodiment, the online hint is optimized while the target hint is gradually updated through a slow moving average process, which incorporates past information to improve stability and effectiveness. The hint exchange prediction mechanism benefits from the unsupervised learning method SwAV

[20] . Based on the enhanced view of an image and the online hints, the exchange hints predict the category assignment of the enhanced view of the same image. This enables the online hints to learn more representation knowledge. The basic principle behind the exchange prediction strategy is that two different enhanced views of the same image should have similar category assignments. Contrastive representation learning methods are used to generate better decision boundaries

[21] .

[0048] In addition to the loss function for self-supervised contrastive learning, in one embodiment, the traditional cross-entropy loss as in CLIP and CoOp is also adopted, which uses the high-confidence pseudo-labels generated by zero-shot learning CLIP to adjust the prompts. In one embodiment, the method can also be used in online test-time scenarios, where the test data arrives in the form of a small batch stream. In one embodiment, the operation of all test data is decomposed into multiple small batches, which will be discussed in detail in the experimental section. In one embodiment, the method of the present disclosure is evaluated on various test-time adaptation benchmarks, including ImageNet and four natural distribution offset datasets based on it, as well as nine fine-grained classification datasets. Experimental results show that the method of the present disclosure achieves state-of-the-art test-time adaptation performance. In this article, the main contributions are as follows:

[0049] We propose Swapping Cues, a new test-time cue adaptation method that uses a self-supervised contrastive learning strategy to make the cue more adaptable to the downstream image classification task.

[0050] Firstly, unsupervised representation learning is applied to prompt adaptation of pre-trained visual language models. EMA prompts and prompt exchange prediction strategies are introduced, which enable prompts to learn more knowledge from the powerful representation ability of pre-trained models.

[0051] In one embodiment, extensive experiments are conducted on ImageNet and its four variants as well as nine other image classification datasets. Empirical evaluations show that the method of the present disclosure significantly outperforms the current TPT method and can even compete with supervised prompt adaptation methods on most datasets.

[0052] In one embodiment, the present disclosure proposes SwapPrompt, a new framework that can effectively utilize self-supervised contrastive learning to promote test-time prompt adaptation. SwapPrompt adopts a dual prompt paradigm, namely an online prompt and a target prompt that averages (average) from the online prompts to retain historical information. In addition, SwapPrompt applies a swap prediction mechanism that leverages the representation power of a pre-trained model to enhance the online prompts via contrastive learning. Specifically, the online prompts and the enhanced views of the input image are used to predict the class assignments generated by the target prompt and the alternative enhanced views of the same image. The proposed SwapPrompt can be easily deployed on a visual language model without additional requirements, and experimental results show that it achieves state-of-the-art test-time adaptation performance on ImageNet and nine other datasets. The study also shows that SwapPrompt can even achieve performance comparable to supervised prompt adaptation methods.

[0053] Figure 2 The flowchart of a method 100 for test-time prompt adaptation of a visual language model according to an embodiment of the present invention is illustrated. The method 100 comprises steps S102, S104 and S106.

[0054] In step S102, a swap prompt is used when testing the visual language model. The swap prompt is a double prompt model, including: an online prompt and a target prompt.

[0055] In step S104, the online prompt is trained by contrastive learning using a swap prediction mechanism.

[0056] In step S106, the target prompt is updated based on the online prompt, and the online prompt can be optimized and trained according to the target prompt, so that the online prompt and the target prompt can interact and learn from each other.

[0057] It should be understood that although Figure 2The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0058] Figure 3 A schematic block diagram of a system 50 for prompt adaptation during testing of a visual language model according to an embodiment of the present invention is illustrated. The system 50 may include an online prompt module 502 and a target prompt module 504. The online prompt module 502 may be configured so that the online prompt is trained by contrastive learning using an exchange prediction mechanism. The target prompt module 504 may be configured so that the target prompt is updated based on the online prompt. The online prompt can be optimized and trained according to the target prompt so that the online prompt and the target prompt can interact and learn from each other. The system 50 may be configured to use an exchange prompt including the online prompt and the target prompt during testing of the visual language model.

[0059] In one embodiment, the exchange prediction mechanism comprises predicting the category assignment of the target cue to two or more enhanced views of the same test image based on the trained online cue and / or the updated target cue.

[0060] In one embodiment, the target prompt module is configured to obtain the target prompt by averaging the online prompts, so that the exchange prompt updates the target prompt using an exponential moving average of the online prompts.

[0061] In one embodiment, the exchange prediction mechanism further includes using text features generated by the target cue as a prototype, using the prototype to generate two or more enhanced views of the test image to obtain category assignments or predictions of the target cue for the two or more enhanced views.

[0062] In one embodiment, the contrastive learning adopts a loss function, which includes a first loss function and a second loss function, wherein the sum of the first loss function and the second loss function is used as the loss function for the test image established under the exchange prediction mechanism.

[0063] As those skilled in the art will appreciate, the system for test-time prompt adaptation of a visual language model described herein may be similarly extended using any implementation of the method for test-time prompt adaptation of a visual language model described herein, which will not be described in detail herein.

[0064] In addition, in this section, we introduce a preliminary definition of test-time cue adaptation and elaborate on the proposed swap cue framework, which leverages self-supervised contrastive learning to facilitate cue adaptation, as described below. Figure 4 The corresponding workflow is also described later.

[0065] Figure 4 An exemplary schematic diagram of an exchange prompt framework used in a method for test-time prompt adaptation of a visual language model according to an embodiment of the present invention is illustrated. In the exchange prompt framework, the present disclosure uses text features generated by a target prompt as prototypes, and assigns image features of an enhanced view of an image to these prototypes to obtain soft category assignments. Online prompts are trained to predict this category assignment using different enhanced views of the same image. The EMA of the online prompts is used to update the target prompt.

[0066] In one embodiment, we focus on test-time prompt adaptation of a pre-trained visual language model (e.g., CLIP), where the visual language model is trained on the source domain, but the test data belongs to the target domain. In this case, zero-shot learning CLIP with general prompts (e.g., "a photo of [CLS]") has some zero-shot generalization capabilities, but these hand-crafted prompts cannot fully extract the rich knowledge learned by CLIP from the large-scale and diverse pre-training datasets. Optimized prompts can further improve the ability of CLIP to retrieve knowledge about the target domain. There have been some related works on supervising target domain data, including the well-known and effective CoOp method, and part of the method of the present disclosure also incorporates the ideas of CoOp. Then, we will briefly review CoOp and define the symbols used in this article.

[0067] Context Optimization (CoOp). CoOp is a cue adaptation method based on CLIP. Similar to CLIP, CoOp includes an image encoder and a text encoder, which are denoted as f(·) and g(·), respectively. Assume is a dataset on the target domain, where x i is the i-th input data sample, y i is the corresponding label, and for the C-class classification problem y i∈{1, 2, ..., C}, N is the size of the dataset. Let t denote a learnable continuous prompt and {t; c} be the input for the cth category of the text encoder. Then, z i =f(x i ) and w c = g({t;c}) are defined as the output feature embeddings of the image encoder and the text encoder respectively. i The predicted probability of the cth class is calculated as:

[0068]

[0069] Where τ represents the temperature parameter and sim(·) represents the cosine similarity. For all training data, CoOp calculates the probabilities of all categories by formula (1) and minimizes the cross entropy loss to adjust the prompt.

[0070] Test-time prompt adaptation. In the test-time scenario, labeled data from the target domain is not available, so prompts cannot be optimized as in CoOp. Consider the test dataset There is no labeled information. The goal of prompt adaptation during testing can be expressed as:

[0071]

[0072] in is the cross entropy loss function. The purpose of this disclosure is to design an unsupervised prompt adaptation method to promote the prompt t and The compatibility of the target domain enables CLIP to obtain more target domain knowledge.

[0073] Swap Tips Overview

[0074] In this section, the test-time adaptive method exchange prompts proposed in the present disclosure is introduced, which contains at least one of two key features: exponential moving average (EMA) prompts and prompt exchange predictions.

[0075] As in Figure 4 As shown in , the exchange hint adopts a dual hint paradigm: online hint and target hint. These two types of hints will interact and learn from each other by applying the EMA update strategy to adapt to the target domain. In addition, unlike most previous frameworks that only match one hint to one image, the exchange hint adopts a hint exchange prediction mechanism to establish self-supervised representation learning in hint adaptation, where the goal is to assign similar categories to two different augmented views of the same image.

[0076] Exponential Moving Average Tips

[0077] The goal of exchanging hints is to learn an online hint t that can be used on a test dataset.o As mentioned before, the exchange cue has a target cue t t Take the online tips to guide o The main motivation behind this design is: from a given target hint t t In the above example, we can predict the t The generated representation knowledge is used to train new and possibly improved prompts. By repeatedly using subsequent online prompts as new target prompts for further training, a series of prompts that improve in quality over time can be created. In practice, the target prompt is set to be averaged from the online prompts, so that the exchange prompt updates the target prompt using a slow exponential moving average of the online prompts, and the following update is performed after each training step, i.e., the target prompt is updated to:

[0078] t t =∈t t +(1-∈)t o (3)

[0079] where ∈∈[0,1] is the decay rate of the target cue.

[0080] Tip Exchange Forecast

[0081] In the field of self-supervised learning, cross-view prediction has been widely used in many existing works. These methods usually cast the prediction problem into a representation space, and then learn representations by predicting different enhanced views generated from the same image. It is assumed that different enhanced views of an image should be relatively close in the representation space.

[0082] Among these methods, SwAV

[20] proposed a different approach. Instead of directly enforcing a consistent mapping between image features in the representation space, SwAV clusters a set of image features, uses cluster centroids as prototypes, and matches different augmented views of the image to these prototypes to compute their cluster assignments. By comparing their cluster assignments rather than their features, SwAV performs contrastive learning between multiple image views.

[0083] In contrast, in one embodiment, a prototype is used to assign an augmented view of an image to obtain its soft class assignment, and another augmented view of the same image is used to predict the class assignment. This approach is well suited for CLIP because it has a natural prototype: the text features output by a text encoder. Specifically, for the test dataset All categories in Target prompt t will form C inputs for the text encoder This will generate C text features of different categories Due to the supervised contrastive pre-training on large-scale image-text pairs, the text features generated by CLIP have high similarity with image features of the same category and low similarity with text features of different categories. Therefore, these C text features can be used as a set of high-quality prototypes. In other words, the text features here is generated by the target prompt. The target prompt in the present disclosure is updated more stably (because it adopts exponential moving average update), so the text feature generated by it can be used as a prototype. The text feature w in formula (1) c It is generated by the hints in the CoOp method, making its updates unstable and volatile.

[0084] In one embodiment, a data augmentation method [A m ...A n ...], all enhancement methods need to be paired in pairs, where m and n represent any two different enhancement methods for obtaining the first enhanced view and the second enhanced view based on the test image (for example, m=1, n=2 represent the first enhancement method and the second enhancement method, respectively). i When there are two or more enhanced views, the teaching or concept of the present disclosure can be used to operate between any two or more enhanced views to obtain the image x. i The loss function of .

[0085] In one embodiment, the loss function includes a first loss function and a second loss function, wherein the sum of the first loss function and the second loss function is used as the loss function for the test image x established under the exchange prediction mechanism. i Specifically, for any two different enhancement methods among the two or more different data enhancement methods, when the image x i When there are more than two enhanced views, operations are performed between any two enhanced views according to the teachings or concepts of the present disclosure to obtain the image x. i The loss function L includes a first loss function L swap and the second loss function L pseudo :

[0086]

[0087] Among them, L swap Specifically, it represents the prompt exchange prediction loss function, and L pseudo Specifically represents a pseudo-label loss function, and wherein m and n represent any two different enhancement methods for obtaining a first enhanced view and a second enhanced view based on a test image, Represents the test image x i Prediction of online prompts for the first enhanced view, Represents the test image x i The prediction of the online prompt of the second enhanced view, Represents the test image x i The prediction of the target cue of the first augmented view, Represents the test image x i The prediction of the target cue of the second enhanced view, Represents the test image x i The pseudo-marker obtained, l 1 is a function representing the difference between the predictions generated with respect to the online cue and the predictions generated with respect to the target cue, l 2 Function representing the difference between predictions and pseudo-labelings generated with respect to online prompts.

[0088] In one embodiment, for example, for an image x i , two different data augmentation methods can be used to obtain two enhanced views A 1 (x i ) and A 2 (x i ), for example, m=1, n=2 respectively represent the first enhancement method and the second enhancement method. The image features generated by the image encoder are and By prototyping Applying formula (1) above, we can obtain the corresponding category assignments (or predictions generated by target cues) of the two enhanced image views:

[0089]

[0090] Where Pt(c|A) represents the probability of the prediction result being the cth class generated by the target prompt, and p t (c|A 1 (x i ) and p t (c|A 2 (x i ) can be expressed as:

[0091]

[0092] Similarly, the predictions generated by online prompts are defined as:

[0093]

[0094] Where Po(c|A) represents the probability of the prediction result being the cth class generated by the online prompt,

[0095] In one embodiment, the image x i The prompt exchange prediction loss function is established as:

[0096]

[0097] Where l is a function that measures the difference between the predictions and class assignments generated by the online prompt. In one embodiment, a cross entropy loss function is used, i.e.,

[0098]

[0099] Online Tips o It will be optimized by formula (8), and the target prompt t t Will not be updated by this loss function.

[0100] Optimization through pseudo-labeling hints

[0101] In addition to optimizing prompts using self-supervised representation learning methods, in one embodiment, a set of labeled data similar to CoOp is also used to optimize online prompts. However, unlike CoOp where labels are available for the target domain, the test data is unlabeled in the test-time scenario. Therefore, in one embodiment, the test data is first reasoned on using hand-crafted prompts (e.g., "a photo of [CLS]") to obtain their pseudo-labels. And then use formula (7) and cross entropy loss as:

[0102]

[0103] It should be noted that the pseudo labels obtained by this method may contain noise. Therefore, Equation (10) cannot be applied to all test data. The process of data selection is discussed in Section 3.3.

[0104] Algorithm workflow

[0105] In this subsection, we will explain the overall cue adaptation process for swapping cues and the training method when facing online test samples.

[0106] When using formula (4) (e.g., formula (8) and formula (10)) and pseudo-labeling Before training, data selection is required to filter out potential noisy pseudo-labels. Specifically, in one embodiment, zero-shot learning CLIP and handcrafted hints are first used to obtain pseudo-labels and classification confidences for test data. Then, for each category, only the top K test data with the highest confidence are selected. These selected test data form the adaptive set It is a subset. For all Given the trade-off hyperparameters α and β, the following loss function will be used for prompt adaptation:

[0107] L adapt (x i ) = αL swap (x i ) + βL pseudo (x i ) (11)

[0108] By adopting formula (11), adding the trade-off hyperparameters can further improve the prediction accuracy of the present disclosure.

[0109] When the target images arrive in small batches (i.e., less than a threshold number, e.g., less than 10% of the entire dataset, such as 4% of the entire dataset), that is, the test data is online, it is not possible to rank the confidence of the entire test dataset. However, it is still possible to perform confidence-based ranking on the small batches to select the top k (k < K) test data with the highest confidence while keeping the rest of the training process unchanged. When new small batches of test data arrive, the available test data is sorted by confidence again to obtain new

[0110] experiments

[0111] Experimental settings

[0112] Datasets. In one embodiment, the proposed swap prompt was evaluated on fourteen datasets, including ImageNet

[25] and its four variants (ImageNet-V2

[26] , ImageNet-A

[27] , ImageNet-R

[28] , and ImageNet-Sketch

[29] ), as well as nine other publicly available image classification datasets used in CLIP (Caltech101

[30] , DTD

[31] , Flowers 102

[32] , Oxford-Pets

[33] , UCF101

[34] , StanfordCars

[35] , Food101

[36] , EuroSAT

[37] , and SUN397

[38] ). These datasets cover a wide variety of visual classification tasks, including general objects, fine-grained categories, and even texture classification, forming a comprehensive benchmark. In one embodiment, only the test data is used for adaptation and they are used to evaluate the model.

[0113] Baselines. In an experiment, the performance of the exchanged prompt is compared with the state-of-the-art methods. In one instance, in addition to CLIP

[16] for zero-shot learning, it also includes: TPT

[19] , a test-time prompt adaptation method that minimizes the marginal entropy of the test data; UPL

[24] , an unsupervised prompt learning method with some modifications to adapt to the test-time setting; CoOp

[17] , a supervised few-shot prompt adaptation method. Some labeled data from the same domain as the test data is used during the training of this baseline in order to use it as the upper bound of the test-time prompt adaptation performance.

[0114] Implementation details. In all experimental examples listed above, the publicly available CLIP model and the ResNet-50

[39] visual encoder are used as the backbone model. Unless otherwise stated, the prompts are randomly initialized using 4 learnable tokens in the exchange prompt, UPL, and CoOp. For TPT, the prompt is initialized to the default "a photo of a". When comparing the performance with the baseline, in one embodiment, the top 16 test data with the highest confidence are selected to train the exchange prompt and UPL. For the exchange prompt, the decay rate of the target prompt is 0.99, and α and β are both 1. The same image augmentation method as SimCLR

[40] can be used to generate two different enhanced views for the image. In one embodiment, the prompt is optimized for 50 epochs using the SGD optimizer and the cosine decay learning rate scheduler, with an initial learning rate of 0.002. The batch size of images on all datasets is 32.

[0115] In one embodiment, all of the above experiments were performed on a workstation with an RTX 3090 GPU, a 3.5GHZ Intel Core i9-11900K CPU, and 64GB RAM.

[0116] Performance Comparison

[0117] First, the proposed exchange hint is compared with baseline methods on fourteen benchmark datasets in one embodiment. Figure 5a A diagram illustrating the average accuracy on all 14 datasets and compared with CoOp as the upper boundary; Figure 5b The average accuracy of the exchange prompt on the five datasets is shown as the different adaptation rounds change. The accuracy of UPL and TPT is the average accuracy of the last round.

[0118] The classification accuracies are listed in Table 1, where Δ represents the gain of the swap hint relative to the better one of UPL and TPT, and “+Online” represents the swap hint with online test data.

[0119] Table 1: Comparison of test-time adaptation methods on 14 datasets

[0120]

[0121] It should be noted that CoOp uses labeled target domain data for training, and in one embodiment uses 4-shot data by category. From the results, it can be seen that the exchange prompt proposed in the present disclosure provides test-time adaptive performance that is better than the baseline on most datasets. In one embodiment, it not only outperforms better baselines in UPL and TPT, but also is very close to or even outperforms CoOp on many datasets (e.g., Caltech101, Oxford-Pets, Food101, ImageNet, ImageNet-R, and ImageNet-Sketch). Figure 5a The average accuracy over 14 datasets for all baselines is shown. The swap prompt is 2.31% and 2.17% more accurate than TPT and UPL, respectively.

[0122] Strong performance in the adaptive setting when testing online. In one embodiment, the results of the exchange hint are also added to the online test data. In this setting, mini-batches with only a small fraction of the data are received. For example, in DTD, the mini-batch is only 64 test data samples, which is less than 4% of the entire dataset. On most datasets, the accuracy drops slightly because it is not possible to hold enough test data at the beginning of training to learn appropriate hints to classify the first test data that appears. However, the online exchange hint still outperforms UPL and TPT on most datasets, as well as outperforming Figure 5a The average accuracy on all datasets in . It should be noted that on Food101, the accuracy is better for the online exchange hint because the hints in the middle stage are better than those in the last stage. This may be due to the presence of high-confidence noise in the pseudo-labeling during the last stage, as discussed in Section 4.3.

[0123] Ablation studies

[0124] In this section, a detailed analysis is presented to help understand the superiority of the exchange hint method of the present disclosure, including the trade-off between accuracy and efficiency, and the relationship between the objective function L swap and L pseudo analysis, the decay rate ∈ of the target cue, the impact of the value of K in data selection, and the impact of the context length and initialization of the cue.

[0125] Tradeoff between accuracy and efficiency. The main factor affecting the efficiency of exchanging cues is the number of cue adaptation rounds. Figure 5b The relationship between the number of rounds and the average accuracy of the exchange prompt on 5 datasets (Caltech101, DTD, Flowers102, Oxford-Pets, and UCF101) is shown. It can be seen that the accuracy of the exchange prompt increases rapidly in the first 3 rounds, then reaches its highest value around 20 rounds and stabilizes there. Therefore, when the adaptation time is limited, the exchange prompt can trade off between accuracy and adaptation rounds, for example, only training 3 rounds for fast inference. It is worth noting that the exchange prompt outperforms the accuracy of the last round of UPL only in round 2, and outperforms the accuracy of TPT in round 1.

[0126] Analysis of objective functions. In one embodiment, two objective functions L for exchanging prompts were evaluated on 5 datasets. swap and L pseudo The results in Table 2 present a clear ablation study to demonstrate the effectiveness of the objective function proposed in this disclosure. First, UPL is used as a basic baseline with confident test data selection and objective function L pseudo Then, UPL+AUG means just adding image enhancement to the baseline so that L pseudo Applied to 2 augmented image views. It can be observed that the accuracy has improved on all datasets, which proves the benefit of data augmentation. Compared with swapping prompts, UPL+AUG has no L swap function, and the performance is worse than the swap hint. Finally, the case using all two objective functions (i.e., the complete swap hint) has the best performance, which indicates that the loss function L for the swap prediction mechanism swap The prompt can be further improved.

[0127] Table 2: Analysis of objective function

[0128]

[0129] Analysis of the decay rate of the target prompt. The exchange prompt updates the target prompt by a slow moving average of the online prompt, so the target prompt represents a delayed and more stable version of the online prompt. A higher decay rate means that more historical information is retained. As shown in formula (3), when the decay rate ∈ is 0, the target prompt is immediately updated to the online prompt at each step. It should be noted that in L swap Loss (i.e. ) does not work because the gradients will collapse, an alternative is to maintain a target hint that remains the same as the online hint (i.e., ∈=0). When the decay rate ∈ is 1, the target hint is never updated and remains at a constant value corresponding to its initialization. In this case, in one embodiment the target hint is initialized to “a photo of…” so it still has basic zero-shot learning generalization capabilities. There is a trade-off between updating the target too frequently and updating the target too slowly.

[0130] Table 3 shows the results of different decay rates on 6 datasets. When ∈=0 and ∈=1, the performance is worse than the other three compromise values. All values ​​of decay rate between 0.9 and 0.999 have their best applicable datasets, and the decay rate of 0.99 has the highest average accuracy. In addition, compared with maintaining a fixed target cue (∈=1), the EMA strategy can improve the average accuracy by 1.35%.

[0131] Table 3: Results for different decay rates ∈.

[0132]

[0133] Analysis of Top-K confidence data selection. Before prompting adaptation, in one embodiment, data selection is performed to filter potential noise pseudo-labels. Only the top K confidence test data will be used in test-time adaptation. Too much data will not only reduce model performance, but also slightly increase training time. Table 4 provides the performance of exchange prompts using different K values ​​on 5 data sets, and also includes the results without data selection. In general, larger K values ​​show better performance, and the highest accuracy on all 5 data sets is achieved when K is set to 8 or 16, where K=16 has the best average accuracy. On the other hand, the performance using the entire test data will not exceed the performance when data selection is adopted. Data selection can increase the average accuracy by 2.5%. This shows that the negative impact of noise pseudo-labels exceeds the positive impact.

[0134] Table 4: Results for different K in data selection. “None” means no data selection.

[0135]

[0136] Analysis on context length and initialization of prompts. To explore whether the exchange prompts of the present disclosure are equally effective on prompts of different context lengths and initializations, in one embodiment, the experiments were repeated on 5 datasets by varying the context length from 4 to 8 to 16 and using three different types for initialization.

[0137] Results for different context lengths are shown in Table 5(a), which shows that having more context tags sometimes leads to a slight drop in accuracy. This may be due to shorter learning prompts, less overfitting, because the exchange prompt only uses a portion of the test data for adaptation, and too many parameters in the prompt may lead to overfitting on these data. Nevertheless, the exchange prompt still maintains high-level performance at different context lengths. In Table 5(b), three different initializations are: Hand-craft (“a photo of [CLS]”), Pre-ImageNet (prompt trained on 16 sample ImageNet using CoOp), and Random. It can be found that different initializations only slightly affect the final accuracy, and the exchange prompt shows robust performance due to the effective adaptation of the prompt, which shows that the method of the present disclosure does not rely on any previous source domain prompts.

[0138] Table 5: (a) Results for different context lengths.

[0139]

[0140] Table 5: (b) Results for different initializations.

[0141]

[0142] The present disclosure also provides a computer device, which includes a memory and a processor, wherein the memory stores computer instructions executable by the processor, and when the computer instructions are executed by the processor, the processor is instructed to execute the steps of the method for prompting self-adaptation during testing of a visual language model of the present invention. The executable computer instructions can be embodied and implemented in the form of an application program. The computer device can be broadly a server, a terminal, or any other electronic device with necessary computing and / or processing capabilities. In one embodiment, the computer device may include a processor, a memory, a network interface, a communication interface, etc. connected via a system bus. The processor of the computer device can be used to provide necessary computing, processing and / or control capabilities. The memory of the computer device may include a non-volatile storage medium and an internal memory. An operating system, a computer program, etc. may be stored in or on the non-volatile storage medium. The internal memory can provide an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface and the communication interface of the computer device can be used to connect and communicate with external devices through a network. When the computer program is executed by the processor, the steps of the method for prompting self-adaptation during testing of a visual language model of the present invention are executed.

[0143] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, causes the steps of the method for prompting self-adaptation during testing of a visual language model of the present invention to be executed. In one embodiment, the computer program is distributed on a plurality of computer devices or processors coupled to a network so that the computer program is stored, accessed, and executed in a distributed manner by one or more computer devices or processors. A single method step / operation, or two or more method steps / operations, may be performed by a single computer device or processor or by two or more computer devices or processors. One or more method steps / operations may be performed by one or more computer devices or processors, and one or more other method steps / operations may be performed by one or more other computer devices or processors. One or more computer devices or processors may perform a single method step / operation, or perform two or more method steps / operations.

[0144] It will be appreciated by those skilled in the art that all or part of the steps of the method for testing-time prompt adaptation of a visual language model of the present invention may be performed by instructing related hardware such as a computer device or a processor through a computer program, and the computer program may be stored in a non-transitory computer-readable storage medium, which causes the steps of the method of the present invention to be performed when the computer program is executed. Depending on the circumstances, any reference to memory, storage, database, or other media herein may include non-volatile and / or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state disk, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.

[0145] In summary, this disclosure proposes a new test-time prompt adaptation method (Exchange Prompt) to learn prompts adapted to the test domain for pre-trained visual language models. Specifically, this disclosure maintains online prompts and target prompts updated by EMA, which interact and learn from each other. A swap prediction mechanism is designed to train online prompts so that it can predict the category assignment of target prompts to the same image under different augmented views. Exchange Prompt can be easily deployed at test time of visual language models without any other requirements. Extensive empirical experiments are conducted on various datasets to verify the effectiveness and superior performance of Exchange Prompt.

[0146] In addition, the pre-trained visual language model has a wide range of application scenarios, for example, it is used in scenarios such as visual question answering system (Visual Q&A system), content audit (content Audit) or autonomous driving. Therefore, the method for prompting self-adaptation during testing of the visual language model of the present disclosure also has a correspondingly broad application prospect.

[0147] Various documents are mentioned and cited herein (as listed below), the contents of each of which are hereby incorporated by reference in their entirety.

[0148] The various technical features described above can be combined arbitrarily. Although all possible combinations of these technical features are not described, any combination of these technical features should be considered to be covered by this specification as long as there is no contradiction in such combination.

[0149] Although the present invention has been described in conjunction with the embodiments, it should be understood by those skilled in the art that the above description and the accompanying drawings are only exemplary and non-restrictive, and the present invention is not limited to the disclosed embodiments. Various modifications and variations are possible without departing from the spirit of the present invention.

[0150] Reference list:

[0151] [1] J Quinonero Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil DLawrence. Dataset shift in machine learning. The MITPress, 1:5, 2009.

[0152] [2] Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. In International Conference on Learning Representations, 2019.

[0153] [3]Pang Wei Koh,Shiori Sagawa,Henrik Marklund,Sang Michael Xie,MarvinZhang,Akshay Balsubramani,Weihua Hu,Michihiro Yasunaga,Richard LanasPhillips,Irena Gao,et al.Wilds:A benchmark of in-the-wild distributionshifts.In International Conference on Machine Learning,pages 5637-5664.PMLR,2021.

[0154] [4]Mathilde Caron,Piotr Bojanowski,Armand Joulin,and MatthijsDouze.Deep clustering for unsupervised learning of visual features.InProceedings of the European conference on computer vision(ECCV),pages 132-149,2018.

[0155] [5]Judy Hoffman,Eric Tzeng,Taesung Park,Jun-Yan Zhu,Phillip Isola,Kate Saenko,Alexei Efros,and Trevor Darrell.Cycada:Cycle-consistentadversarial domain adaptation.In International conference on machinelearning,pages 1989-1998.Pmlr,2018.

[0156] [6]Yu Sun,Xiaolong Wang,Zhuang Liu,John Miller,Alexei Efros,andMoritz Hardt.Test-time training with self-supervision for generalizationunder distribution shifts.In International conference on machine learning,pages 9229-9248.PMLR,2020.

[0157] [7]Dequan Wang,Evan Shelhamer,Shaoteng Liu,Bruno Olshausen,and TrevorDarrell.Tent:Fully test-time adaptation by entropy minimization.InInternational Conference on Learning Representations,2021.

[0158] [8]Dian Chen,Dequan Wang,Trevor Darrell,and SaynaEbrahimi.Contrastive test-time adaptation.In Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition,pages 295-305,2022.

[0159] [9]Malik Boudiaf,Romain Mueller,Ismail Ben Ayed,and LucaBertinetto.Parameter-free online test-time adaptation.In Proceedings of theIEEE / CVF Conference on Computer Vision and Pattern Recognition,pages 8344-8353,2022.

[0160]

[10] Fatemeh Azimi,Sebastian Palacio,Federico Raue, Hees,LucaBertinetto,and Andreas Dengel.Self-supervised test-time adaptation on videodata.In Proceedings of the IEEE / CVF Winter Conference on Applications ofComputer Vision,pages 3439-3448,2022.

[0161]

[11] Chaithanya Kumar Mummadi,Robin Hutmacher,Kilian Rambach,EvgenyLevinkov,Thomas Brox,and Jan Hendrik Metzen.Test-time adaptation todistribution shift by confidence maxi-mization and input transformation,arXivpreprint arXiv:2106.14999,2021.

[0162]

[12] Jian Liang,Dapeng Hu,and Jiashi Feng.Do we really need to accessthe source data?source hypothesis transfer for unsupervised domainadaptation.In International Conference on Machine Learning,pages 6028-6039.PMLR,2020.

[0163]

[13] Yusuke Iwasawa and Yutaka Matsuo.Test-time classifier adjustmentmodule for model-agnostic domain generalization.Advances in NeuralInformation Processing Systems,34:2427-2440,2021.

[0164]

[14] Jogendra Nath Kundu,Naveen Venkat,R Venkatesh Babu,etal.Universal source-free domain adaptation.In Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition,pages 4544-4553,2020.

[0165]

[15] Rui Li,Qianfen Jiao,Wenming Cao,Hau-San Wong,and Si Wu.Modeladaptation:Unsuper-vised domain adaptation without source data.In Proceedingsof the IEEE / CVF conference on computer vision and pattern recognition,pages9641-9650,2020.

[0166]

[16] Alec Radford,Jong Wook Kim,Chris Hallacy,Aditya Ramesh,GabrielGoh,Sandhini Agarwal,Girish Sastry,Amanda Askell,Pamela Mishkin,Jack Clark,etal.Learning transferable visual models from natural language supervision.InInternational conference on machine learning,pages 8748-8763.PMLR,2021.

[0167]

[17] Kaiyang Zhou,Jingkang Yang,Chen Change Loy,and Ziwei Liu.Learningto prompt for vision-language models.International Journal of ComputerVision,130(9):2337-2348,2022.

[0168]

[18] Kaiyang Zhou,Jingkang Yang,Chen Change Loy,and ZiweiLiu.Conditional prompt learning for vision-language models.In Proceedings ofthe IEEE / CVF Conference on Computer Vision and Pattern Recognition,pages16816-16825,2022.

[0169]

[19] Manli Shu,Weili Nie,De-An Huang,Zhiding Yu,Tom Goldstein,AnimaAnandkumar,and Chaowei Xiao.Test-time prompt tuning for zero-shotgeneralization in vision-language models.In Alice H.Oh,Alekh Agarwal,DanielleBelgrave,and Kyunghyun Cho,editors,Advances in Neural Information ProcessingSystems,2022.

[0170]

[20] Mathilde Caron,Ishan Misra,Julien Mairal,Priya Goyal,PiotrBojanowski,and Armand Joulin.Unsupervised learning of visual features bycontrasting cluster assignments.Advances in neural information processingsystems,33:9912-9924,2020.

[0171]

[21] Connor Shorten,Taghi M Khoshgoftaar,and Borko Furht.Text dataaugmentation for deep learning.Journal of big Data,8:1-34,2021.

[0172]

[22] Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao. Clip-adapter: Better vision-language models with feature adapters, arXiv preprint arXiv:2110.04544, 2021.

[0173]

[23] Renrui Zhang, Rongyao Fang, Wei Zhang, Peng Gao, Kunchang Li, Jifeng Dai, Yu Qiao, and Hongsheng Li. Tip-adapter: Training-free clip-adapter for better vision-language modeling. arXiv preprint arXiv:2111.03930, 2021.

[0174]

[24] Tony Huang, Jack Chu, and Fangyun Wei. Unsupervised prompt learning for vision-language models, arXiv preprint arXiv:2204.03649, 2022.

[0175]

[25] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei.Imagenet: A large-scale hierarchical image database.In 2009 IEEEconference on computer vision and pattern recognition, pages 248-255.Ieee, 2009.

[0176]

[26] Benjamin Recht,Rebecca Roelofs,Ludwig Schmidt,and VaishaalShankar.Do imagenet classifiers generalize to imagenet?In Internationalconference on machine learning,pages5389-5400.PMLR,2019.

[0177]

[27] Dan Hendrycks,Kevin Zhao,Steven Basart,Jacob Steinhardt,and DawnSong.Natural adversarial examples.In Proceedings of the IEEE / CVF Conferenceon Computer Vision and Pattern Recognition,pages 15262-15271,2021.

[0178]

[28] Dan Hendrycks,Steven Basart,Norman Mu,Saurav Kadavath,Frank Wang,Evan Dorundo,Rahul Desai,Tyler Zhu,Samyak Parajuli,Mike Guo,et al.The manyfaces of robustness:A critical analysis of out-of-distributiongeneralization.In Proceedings of the IEEE / CVF International Conference onComputer Vision,pages 8340-8349,2021.

[0179]

[29] Haohan Wang,Songwei Ge,Zachary Lipton,and Eric P Xing.Learningrobust global repre-sentations by penalizing local predictive power.Advancesin Neural Information Processing Systems,32,2019.

[0180]

[30] Li Fei-Fei,Rob Fergus,and Pietro Perona.Learning generativevisual models from few training examples:An incremental bayesian approachtested on 101 object categories.In 2004conference on computer vision andpattern recognition workshop,pages 178-178.IEEE,2004.

[0181]

[31] Mircea Cimpoi,Subhransu Maji,Iasonas Kokkinos,Sammy Mohamed,andAndrea Vedaldi.Describing textures in the wild.In Proceedings of the IEEEconference on computer vision and pattern recognition,pages 3606-3613,2014.

[0182]

[32] Maria-Elena Nilsback and Andrew Zisserman.Automated flowerclassification over a large number of classes.In 2008 Sixth Indian Conferenceon Computer Vision,Graphics&Image Processing,pages 722-729.IEEE,2008.

[0183]

[33] Omkar M Parkhi,Andrea Vedaldi,Andrew Zisserman,and CVJawahar.Cats and dogs.In 2012IEEE conference on computer vision and patternrecognition,pages 3498-3505.IEEE,2012.

[0184]

[34] Khurram Soomro,Amir Roshan Zamir,and Mubarak Shah.Ucf101:Adataset of 101 human actions classes from videos in the wild.arXivpreprintarXiv:1212.0402,2012.

[0185]

[35] Jonathan Krause,Michael Stark,Jia Deng,and Li Fei-Fei.3d objectrepresentations for fine-grained categorization.In Proceedings of the IEEEinternational conference on computer vision workshops,pages 554-561,2013.

[0186]

[36] Lukas Bossard,Matthieu Guillaumin,and Luc Van Gool.Food-101-mining discriminative components with random forests.In Computer Vision-ECCV2014:13th European Conference,Zurich,Switzerland,September 6-12,2014,Proceedings,Part VI 13,pages 446-461.Springer,2014.

[0187]

[37] Patrick Helber,Benjamin Bischke,Andreas Dengel,and DamianBorth.Eurosat:A novel dataset and deep learning benchmark for land use andland cover classiiication.IEEE Journal of Selected Topics in Applied EarthObservations and Remote Sensing,12(7):2217-2226,2019.

[38] Jianxiong Xiao,JamesHays,Krista A Ehinger,Aude Oliva,and Antonio Torralba.Sun database:Large-scale scene recognition from abbey to zoo.In 2010 IEEE computer societyconference on computer vision and pattern recognition,pages 3485-3492.IEEE,2010.

[0188]

[39] Kaiming He,Xiangyu Zhang,Shaoqing Ren,and Jian Sun.Deep residuallearning for image recognition.In Proceedings of the IEEE conference oncomputer vision and pattern recognition,pages 770-778,2016.

[0189]

[40] Ting Chen,Simon Kornblith,Mohammad Norouzi,and Geoffrey Hinton.Asimple framework for contrastive learning of visual representations.InInternational conference on machine learning,pages 1597-1607.PMLR,2020。

Claims

1. A method for prompting self-adaptation during testing of a visual language model, characterized in that: The method comprises: When testing the visual language model, an exchange prompt is used, wherein the exchange prompt is a double prompt demonstration format, including: an online prompt and a target prompt; wherein the online prompt is trained by contrastive learning using a swap prediction mechanism, The target prompt is updated based on the online prompt, and the online prompt can be optimized and trained according to the target prompt, so that the online prompt and the target prompt can interact and learn from each other.

2. The method according to claim 1, characterized in that The exchange prediction mechanism includes predicting class assignments of the target cue to two or more enhanced views of the same test image based on the trained online cue and / or the updated target cue.

3. The method according to claim 1 or 2, characterized in that: The target hint is set to be averaged from the online hints, such that the exchange hint updates the target hint using an exponential moving average of the online hints, where the target hint is represented as: t t =and t +(1-e)t o , Among them, t o Indicates online prompt, t t represents the target cue, ∈ represents the decay rate of the target cue, ∈∈[0,1].

4. The method according to claim 3, characterized in that The exchange prediction mechanism further includes using text features generated by the target cue as prototypes, and using the prototypes to generate two or more enhanced views of the test image to obtain category assignments or predictions of the target cue for the two or more enhanced views.

5. The method according to claim 4, characterized in that The contrastive learning adopts a loss function, wherein the loss function includes a first loss function and a second loss function, wherein the sum of the first loss function and the second loss function is used as the loss function for the test image x established under the exchange prediction mechanism. i The loss function of .

6. The method according to claim 4, characterized in that The first loss function is the prompt exchange prediction loss function L swap , and the second loss function is the pseudo-label loss function L pseudo ,in, Wherein, m and n represent any two different enhancement methods for obtaining the first enhanced view and the second enhanced view based on the test image, Represents the test image x i Prediction of online prompts for the first enhanced view, Represents the test image x i The prediction of the online prompt of the second enhanced view, Represents the test image x i The prediction of the target cue of the first augmented view, Represents the test image x i The prediction of the target cue of the second enhanced view, Represents the test image x i The pseudo-labels obtained, l1 represents the function of the difference between the predictions generated by the online prompt and the predictions generated by the target prompt, and l2 represents the function of the difference between the predictions generated by the online prompt and the pseudo-labels.

7. The method according to claim 6, characterized in that The method further comprises: In creating the test image x i The first loss function and the second loss function and before obtaining the pseudo label are based on at least the test image x i The confidence of the relevant test data is selected, and the top K test data with the highest confidence are used to form the confidence matrix for the test image x. i Adaptive set.

8. The method according to claim 7, characterized in that The method further comprises: Under the exchange prediction mechanism, based on the adaptive set, a method for the test image x is further established. i The loss function L for the adaptive set adapt : L adapt (x i )=αL swap (x i )+βL pseudo (x i ), Among them, α and β are given trade-off hyperparameters.

9. The method according to claim 7 or 8, characterized in that: The method further comprises: When the test data of the test image x i is online test data that arrives in the form of batch test data with a quantity less than a threshold number, a corresponding new confidence ranking is performed based on at least each batch of test data that arrives, and the top k (k < K) test data with the highest confidence are selected to obtain a corresponding new adaptive set.

10. A system for prompting self-adaptation during testing of a visual language model, characterized in that: The system comprises: an online hint module configured to train online hints for a test image by contrastive learning using an exchange prediction mechanism, and a target prompt module, the target prompt module being configured to update the target prompt based on the online prompt, wherein the online prompt module is further configured to optimize and train the online prompt according to the target prompt, so that the online prompt and the target prompt can interact and learn from each other, and The system is configured to use an exchange prompt including the online prompt and the target prompt when testing the visual language model.

11. The system according to claim 10, characterized in that The exchange prediction mechanism includes predicting class assignments of the target cue to two or more enhanced views of the same test image based on the trained online cue and / or the updated target cue.

12. The system according to claim 10 or 11, characterized in that The target hint module is configured to obtain the target hint by averaging the online hints, so that the exchange hint updates the target hint using an exponential moving average of the online hints.

13. The system according to claim 12, characterized in that The exchange prediction mechanism further includes using text features generated by the target cue as prototypes, and using the prototypes to generate two or more enhanced views of the test image to obtain category assignments or predictions of the target cue for the two or more enhanced views.

14. The system according to claim 13, characterized in that The contrastive learning adopts a loss function, which includes a first loss function and a second loss function, wherein the sum of the first loss function and the second loss function is used as the loss function for the test image established under the exchange prediction mechanism.

15. A computer device comprising a memory and a processor, wherein the memory stores computer instructions, and when the computer instructions are executed by the processor, the method according to any one of claims 1 to 9 is performed.

16. A non-transitory computer-readable storage medium having computer instructions stored thereon, which when executed by a processor cause the method according to any one of claims 1 to 9 to be performed.

17. A computer program product comprising computer instructions which, when executed by a processor, cause the method according to any one of claims 1 to 9 to be performed.