Open set identification method based on CLIP target perception double prompt learning

The dual prompt learning method for CLIP-based open-set recognition addresses issues of overexpanded boundaries and distribution discrepancies by using visual prototypes and a target-aware enhancement module, enhancing the model's adaptability and accuracy.

CN120318804APending Publication Date: 2025-07-15NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510396650.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing open set recognition method based on CLIP has the problem of semantics that excessive attention to image semantics and neglecting physical properties, resulting in the expansion of decision boundaries, image background interfering with learning, and the difference in the distribution of training data and task data, affecting the accuracy of the model's open set recognition.

Method used

Using the CLIP-based target perception dual prompt learning method, learnable semantic and visual prototype prompts are constructed, combined with the target perception enhancement module and data adaptive mechanism, open-set recognition is performed through visual semantic joint inference scores.

Benefits of technology

It effectively reduces the expansion of known class decision boundaries, reduces the interference of image background, reduces the distribution offset between training data and task data, and improves the adaptability and accuracy of model identification in open sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318804A_ABST
    Figure CN120318804A_ABST
Patent Text Reader

Abstract

The invention discloses an open set recognition method based on CLIP target perception double prompt learning. The open set recognition method comprises the following steps: 1, constructing a learnable semantic prototype prompt; 2, constructing a learnable visual prototype prompt; 3, designing a target perception enhancement module, obtaining a target perception known class sample set and a difficult pseudo-unknown class sample set by using the training set, and meanwhile, configuring a data self-adaptive mechanism; 4, designing a loss function to train the model; and 5, performing open set identification by utilizing visual semantic joint inference scores. According to the method, a visual prototype of a known class is fused into an open set recognition method based on CLIP prompt learning, and the limitation that only semantic prompt is depended on is made up. Besides, the target perception enhancement module provided by the invention eliminates the interference of the image background, and can relieve the distribution offset between the CLIP training data and the open set recognition target task data, so that the model is a more compact decision boundary for known class learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of open-set recognition, and mainly relates to an open-set recognition method based on CLIP object-aware dual prompt learning. Background Art

[0002] In traditional classification tasks, the class labels of the training set and the test set are the same, and such problems are called "closed-set classification". Existing classification methods perform excellently in closed-set tasks. However, in practical applications, the model usually faces a severe challenge: the number of classes during training is limited. After the model is deployed, it may encounter classes that never appeared during training, resulting in misclassification and affecting the reliability of the model. To address this challenge, the model needs to be able to correctly classify known classes and also determine previously unseen classes as unknown classes, and the latter is the core issue. Such tasks are defined as "open-set recognition". Compared with closed-set recognition, open-set recognition is closer to real-world situations and is very important in fields with variable data and security sensitivity, such as autonomous driving, medical image analysis, etc.

[0003] Traditional open-set recognition methods can be divided into discriminative and generative methods. For discriminative methods, various techniques are often used to establish a more compact decision boundary for known classes to obtain an excellent classifier. For generative methods, generative models are often used to generate fake unknown class samples to simulate real unknown classes. The core of both is to make the model learn a compact decision boundary for known classes so as to be able to effectively reject unknown classes.

[0004] With the demonstration of powerful capabilities of large models such as CLIP in various downstream tasks, the use of CLIP for open-set recognition has received extensive attention. Existing methods often use techniques such as semantic prompt learning to enhance the adaptability of CLIP to open-set recognition tasks, thereby improving the recognition ability of the model. Compared with traditional open-set recognition methods, although CLIP-based methods have achieved significant performance improvements, most methods only learn good semantic prototypes for known classes to match their visual features. Semantic prompts usually focus more on understanding images rather than their direct physical attributes (such as color, texture, etc.), which inevitably over-expands the decision boundary of known classes, thus increasing the OSR risk. Secondly, the image background easily interferes with the learning of semantic prompts, which is usually ignored in most existing methods. In addition, there is a distribution difference between the training data of CLIP and the target task data, which may further exacerbate the above deficiencies. Summary of the Invention

[0005] Objective of the Invention: Most of the existing open-set recognition methods based on CLIP only focus on semantic cues, and there are the following three deficiencies: i) Semantic cues usually focus more on understanding images rather than their direct physical attributes (such as color), which inevitably over-expands the decision boundaries of known classes, thus increasing the OSR risk; ii) The image background easily interferes with the learning of semantic cues, which is usually ignored in most existing methods; iii) There is a distribution difference between the training data of CLIP and the target task data, which may further exacerbate the above deficiencies. In view of the above three main problems, the present invention proposes an open-set recognition method based on CLIP target-aware dual-cue learning.

[0006] Technical Solution: To achieve the above objective, the present invention adopts the following technical solutions:

[0007] Step S1, construct a learnable semantic prototype cue;

[0008] Step S2, construct a learnable visual prototype cue;

[0009] Step S3, design a target-aware enhancement module, use the training set to obtain a target-aware known-class sample set and a difficult pseudo-unknown-class sample set, and at the same time equip it with a data adaptation mechanism;

[0010] Step S4, design a loss function to train the model;

[0011] Step S5, perform open-set recognition using the visual-semantic joint inference score.

[0012] The present invention has the following advantages:

[0013] As described above, the present invention proposes an open-set recognition method based on CLIP target-aware dual-cue learning. This method adopts visual prototype cues, integrates the visual prototypes of known classes into cue learning, and effectively makes up for the deficiencies of semantic cues. In addition, by using a target-aware enhancement module, it solves the interference caused by the image background during the learning process, reduces the distribution shift between the CLIP training data and the OSR task data, and enhances the adaptability of CLIP to the OSR task. Description of the Drawings

[0014] Figure 1 is a flowchart of the open-set recognition method based on CLIP target-aware dual-cue learning in an embodiment of the present invention.

[0015] Figure 2 is an architecture diagram of the model in an embodiment of the present invention.

[0016] Figure 3 is a visualization diagram of the samples obtained by the target-aware enhancement module in an embodiment of the present invention. Detailed Embodiment

[0017] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments:

[0018] As Figure 1 and Figure 2 shown, the method includes the following steps:

[0019] Step S1: Construct learnable positive semantic prompts and negative semantic prompts for each known class respectively.

[0020] The positive semantic prompt includes two parts: a fixed class name and a learnable prefix, which can be expressed by the formula:

[0021]

[0022] where, V i (i ∈ {1, 2,..., L s}) is a series of learnable vectors with the same dimension as the text encoding, L s is the number of prefix tokens, and [CLASS] is the known class name encoded by CLIP.

[0023] The negative semantic prompt also includes two parts: a class name and a prefix, and its construction form is basically the same as that of the positive semantic prompt. The difference is that the class name part of the negative semantic prompt is also learnable.

[0024] Step S2: Construct a learnable visual prototype prompt to make up for the deficiency of relying solely on semantic prototype prompts. Specifically, following the design method of semantic prompts, construct a learnable visual prototype prompt, which includes two parts: a fixed visual prototype and a learnable prefix.

[0025] For the fixed visual prototype part, replace the [class name] part of the semantic prompt with [visual prototype], that is, use the image encoder Enc img (·) of CLIP to encode the images of each class in and average the obtained features. The specific formula is:

[0026]

[0027] where, is the sample set of the i-th known class, proto i ∈R 1×d is the visual prototype of the i-th known class, and d is the dimension of the visual prototype.

[0028] For the learnable prefix part, use the same method as the semantic prompt prefix, randomly initialize it, and obtain the final prefix representation through the training of the model. Finally, the visual prototype prompt can be expressed as:

[0029]

[0030] Among them, v i (i ∈ {1, 2,..., L v}) is a series of learnable vectors, d and L v represent the length of the CLIP visual feature encoding and the length of the visual prompt prefix respectively. To further improve the expressive ability of the visual prototype prompt and facilitate calculation, the above visual prototype prompt also needs to be projected through a linear layer, and the formula is:

[0031]

[0032] Among them, represents the linear layer, is the visual prototype prompt vp i after projection.

[0033] Step S3: Design a target perception enhancement module to obtain a target perception known class sample set and a difficult pseudo-unknown class sample set using the training set, and at the same time equip it with a data adaptation mechanism. The specific implementation steps of this module include:

[0034] Target perception sample acquisition

[0035] Step S3.1: Determine the training set Among them, is the known class label, and N is the number of samples.

[0036] Step S3.2: Apply c random cropping operations to each image x in to obtain a sequence of cropped images i

[0037] Step S3.3: For each known class y i construct a semantic prompt t i = a photo of a [y i . Calculate the cosine similarity between and the text prompt corresponding to its true class.

[0038] Step S3.4: Sort the cosine similarities in Step S3.3 in descending order, and take the m cropped images with the highest and lowest similarities as the target perception known class samples and the difficult pseudo-unknown class samples

[0039] ​Step S3.5: To further expand the diversity of the difficult pseudo-unknown class samples X btm in Step S3.4, apply the Mixup operation m times to obtain m mixed images. The specific expression is:

[0040] x neg = λ · x1 + (1 - λ) · x2

[0041] where λ is the mixing coefficient, and x1 and x2 are any two images in X btm .

[0042] After performing the above operations on all the images in the training set, the final target-aware sample set and the difficult pseudo-unknown class sample set are generated. The generated target-aware samples and difficult pseudo-unknown class samples are as Figure 3 shown.

[0043] Data Adaptive Mechanism

[0044] The image sample x is first input into the Patch Embedding layer to obtain:

[0045] x′ = [x img , x cls

[0046] where x img is the token after image encoding, and x cls is the class token of ViT pre-training. Subsequently, a series of learnable prompts will be concatenated to obtain the final input x′, which is formalized as:

[0047] x′ = [x img , P adpt , x cls

[0048] where d t represents the length of the image token, K is the number of known classes, and [·, ·] represents the concatenation operation.

[0049] Step S4: Design a loss function to train the model.

[0050] For the visual prototype prompt, use the target-aware sample set to train, and the loss function is defined as:

[0051]

[0052] where sim(·, ·) represents the cosine similarity, f​​img = Enc img (x) represents the features of the image, is the true visual prototype hint z corresponding to the image * features.

[0053] For the semantic prototype hint, the loss function is defined as:

[0054]

[0055] where,

[0056]

[0057] where, and respectively represent the features of the true positive semantic hint and the negative semantic hint. and respectively represent the features of the i-th positive and negative semantic hints.

[0058] The total loss function is defined as:

[0059]

[0060] Step S5, use the visual-semantic joint inference scoring function for open-set recognition, and the decision score is defined as:

[0061] S(x) = maxα·P sem +(1-α)·P vis

[0062] where, α represents the weight coefficient, P sem is the similarity score between the input image and the positive semantic prototype hint, P vis is the similarity score between the input image and the visual prototype hint, and are respectively defined as:

[0063]

[0064] In practical applications, it is necessary to define an appropriate threshold ∈ according to the task scenario. If S(x) ≥ ∈, the sample to be tested is a known class, otherwise it is an unknown class.

[0065] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to listing the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.

Claims

1. An open-set recognition method based on CLIP target-aware dual prompt learning, characterized in that The method includes the following steps: Step S1, construct a learnable semantic prototype prompt; Step S2, construct a learnable visual prototype prompt; Step S3, design a target-aware enhancement module, obtain a target-aware known class sample set and a difficult pseudo-unknown class sample set by using a training set, and at the same time equip with a data adaptation mechanism; Step S4, design a loss function to train the model; Step S5, perform open-set recognition by using the visual-semantic joint inference score.

2. The open-set recognition method based on CLIP target-aware dual prompt learning according to claim 1, characterized in that, In the said Step S1, a learnable positive semantic prompt and a learnable negative semantic prompt are constructed for each known class respectively; The positive semantic prompt includes two parts: a fixed class name and a learnable prefix, which is expressed by the formula: Among them, V i (i ∈ {1, 2,..., L s} is a series of learnable vectors with the same dimension as the text encoding, and L s is the number of prefix tokens, and [CLASS] is the known class name after CLIP encoding; The negative semantic prompt also includes two parts: a class name and a prefix, and its construction form is basically the same as that of the positive semantic prompt. The difference is that the class name part of the negative semantic prompt is also learnable.

3. The open-set recognition method based on CLIP target-aware dual prompt learning according to claim 1, characterized in that In the said Step S2, following the design method of the semantic prompt, a learnable visual prompt is constructed, which includes two parts: a fixed visual prototype and a learnable prefix; For the fixed visual prototype part, replace the [category name] part of the semantic prompt with [visual prototype], that is, use the image encoder Enc of CLIP img (·) pair Encode the images of each category in and average the obtained features. The specific formula is: Among them, is the sample set of the i-th known class, and proto i ∈R 1×d is the visual prototype of the i-th known class, and d is the dimension of the visual prototype; For the learnable prefix part, the same method as the semantic prompt prefix is adopted, and it is randomly initialized and the final prefix representation is obtained through the training of the model; finally, the visual prototype prompt can be expressed as: Among them, v i (i ∈ {1, 2,..., L v}) is a series of learnable vectors, d and L v respectively represent the length of the CLIP visual feature encoding and the number of visual prompt prefix tokens; to further improve the expressive ability of the visual prototype prompt and facilitate calculation, the above visual prototype prompt also needs to be projected through a linear layer, and the formula is: Among them, represents a linear layer, is the visual prototype prompt z i the projected feature.

4. The open-set recognition method based on CLIP object-aware dual prompt learning according to claim 1, characterized in that, In the said Step S3, the target-aware enhancement module includes two parts: target-aware sample acquisition and a data adaptation mechanism. The specific implementation method is as follows: 4.1 Target-aware sample acquisition Step S3.1, determine the training set wherein is the known class label and N is the number of samples; Step S3.

2. For each image x in apply c times of random cropping operations to obtain a sequence of cropped images i ​ Step S3.

3. For each known category y i Construct a semantic hint t i = a photo of a[y i ; Calculate the cosine similarity between and the text hint corresponding to its true category; Step S3.4: Sort the cosine similarities in step 3.3 in descending order, and take the m cropped images with the highest and lowest similarities as the target perception known class samples and difficult pseudo-unknown class samples Step S3.5: To further expand the diversity of the difficult pseudo-unknown class samples X in Step 3.4 btm apply m Mixup operations to them to obtain m mixed images, and the specific expression is as follows: x neg = λ·x1 + (1 - λ)·x2 where λ is the mixing coefficient, and x1 and x2 are any two images in X btm ; After performing the above operations on all the images in the training set, the final target perception sample set is obtained and the difficult pseudo-unknown class sample set 4.2 Data adaptation mechanism The image sample x is first input into the Patch Embedding layer to obtain: x′ = [x img , x cls ​ where x img is the token after image encoding, and x cls is the class token for ViT pre-training; subsequently, a series of learnable prompts will be concatenated to obtain the final input, which is formulated as: x′ = [x img , P adpt , x cls ​ where d t represents the length of the image token, K is the number of known categories, and [·, ·] represents the concatenation operation.

5. The open-set recognition method based on CLIP target-aware dual prompt learning according to claim 1, wherein In the said Steps S2 and S3, it is necessary to design a loss function to train the model; For the semantic prototype prompt, the loss function is defined as: where, Among them, and respectively represent the features of the true positive semantic hint and the negative semantic hint; and respectively represent the features of the i-th positive and negative semantic hints; For visual prototype cues, use the target perception sample set for training, and the loss function is defined as: Among them, sim(·, ·) represents cosine similarity, and f img = Enc img (x) represents the feature of the image, is the feature of the true visual prototype prompt z * corresponding to the image; The total loss function is defined as:

6. The open-set recognition method based on CLIP target-aware dual prompt learning according to claim 1, characterized in that In the said Step S5, the visual prototype prompt and the semantic prototype prompt are jointly used in the inference stage, and the decision score is defined as: S(x) = max α·P sem + (1 - α)·P vis where α represents the weight coefficient, and P sem is the similarity score between the input image and the positive semantic prototype prompt, and P vis is the similarity score between the input image and the visual prototype prompt, which are respectively defined as: