Zero sample image description method based on filtering hybrid representation and hallucination suppression
By using filtered mixed representation and hallucination suppression modules in the zero-sample image description method, the problems of modal gap and object hallucination in the prior art are solved, and more accurate and more generalized image description is achieved.
Patent Information
- Application Number
- CN202510289608.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The existing zero-sample image description methods have shortcomings in mitigating modal gaps and alleviating object hallucinations, resulting in inaccurate descriptions and insufficient generalization capabilities.
Filtered mixed representation (FMR) and hallucination suppression (HD) modules are used to obtain key entities in the image, and the image encoder is used to obtain the sub-region features and category tokens of the entity, project and mix, and combine the hallucination suppression method to obtain the original generation probability of word elements to form an image description.
It effectively reduces the modal gap, reduces the probability of object hallucination generation, and improves the accuracy and generalization ability of image description.
Smart Images

Figure CN120219879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image description method, in particular to a zero-shot image description method based on filtered hybrid representation and hallucination suppression. Background Art
[0002] Supervised Image Captioning (SIC) relies on image-text pair data for training and is usually difficult to generalize to images in the wild. Therefore, Zero-Shot Image Captioning (ZIC) has emerged as a key area of research, aiming to generate captions using pre-trained models without further training on a carefully curated dataset. According to the time of visual information injection, ZIC can be divided into two categories, namely the late-guidance strategy and the early-guidance strategy. The late-guidance strategy usually utilizes the language generation ability of large language models (LLMs) to integrate visual information into the model after the initial word prediction stage, but there is a problem of object hallucination. In contrast, the early-guidance-based method provides explicit guidance for word generation of LLMs by adding visual cues to the text token prefix, which can reduce modality bias and improve the alignment between the input image and the generated caption. Due to overfitting of visual cues caused by limited corpora, ViECap and IFCap extract entities as hard cues to make up for the deficiency of visual cues (i.e., soft cues). Although hard cues and soft cues are integrated to improve domain generalization ability, there is still a modality gap, which leads to the generation of hallucinated words. Summary of the Invention
[0003] The present invention provides a method that uses an FMR (Filtered Hybrid Representation) and an HD (Hallucination Degradation) module to reduce the modality gap and alleviate object hallucination.
[0004] A zero-shot image description method based on filtered hybrid representation and hallucination suppression, comprising:
[0005] Step S100, obtaining key entities in the image;
[0006] Step S200, obtaining sub-region features and class tokens of the entities through an image encoder, projecting the sub-region features and class tokens to obtain projected sub-region features and global image representations, filtering the projected sub-region features and mixing them with the global image representations;
[0007] Step S300, using a hallucination suppression method to obtain the original generation probability of the word tokens to form an image description.
[0008] Further, the specific process of step S100 is as follows:
[0009] Step S101, using a text encoder to obtain text embedding vectors of each sentence in the text corpus;
[0010] Step S102, use an image encoder to embed the image to obtain an image embedding vector;
[0011] Step S103, obtain the cosine similarity between the image embedding vector and each text embedding vector;
[0012] Step S104, sort the cosine similarities in descending order and select the top k text embedding vectors;
[0013] Step S105, identify all nouns in the selected text embedding vectors and select the top k' most frequent nouns as the detected key entities
[0014] Further, the specific process of step S200 is as follows:
[0015] Step S201, obtain the sub-region features and class tokens of the key entities;
[0016] Step S202, use the image encoder to project the sub-region features and class tokens of the key entities into the image-text shared space to obtain the projected sub-region features and global image representation;
[0017] Step S203, use the text encoder to encode the prompt constructed by the key entities to obtain a prompt embedding;
[0018] Step S204, obtain the cosine similarity between the projected sub-region features and the prompt embedding, and adaptively filter the projected sub-region features;
[0019] Step S205, mix the filtered representation with the global image representation.
[0020] Further, in step S204, if the cosine similarity between the projected sub-region features and the prompt embedding is greater than the threshold, filter out the projected sub-region features that are not similar to the prompt embedding.
[0021] Further, the mixing of the filtered representation and the global image representation in step S205 is achieved through the following formula
[0022]
[0023] where, I fmr is the final filtered and mixed representation, I fr is the filtered representation, N fr is the number of vectors to be merged.
[0024] Further, the specific process of step S300 is as follows:
[0025] Step S301, use the hallucination capture module to extract the hallucination words in the mixed representation and form a hallucination word list;
[0026] Step S302: Embed all the words and the list of hallucinated words through word embedding, calculate the cosine similarity between the embedding of each hallucinated word and each word in all the words. If the cosine similarity exceeds the threshold, add the word and its corresponding cosine similarity score to the negative candidate words.
[0027] Step S303: If the next token matches a negative candidate word, the original generation probability of the next token is suppressed using the following formula
[0028]
[0029] where x <t represents the previously generated token, p = p(x t |x <t ) represents the original generation probability of the next token, ω is a hyperparameter used to control the degree of suppression, t nc is the negative candidate word, and sim i is the cosine similarity between the embedding of the i-th hallucinated word and each word in all the words.
[0030] Compared with the prior art, the present invention has the following advantages: (1) The filtering hybrid representation combines the global representation with the entity-aware sub-region attention weights to reduce the modality gap; (2) The hallucination degradation module can effectively utilize the detected words to filter out the hallucinated words and reduce the generation probability of object hallucinations.
[0031] The present invention will be further described below with reference to the accompanying drawings of the specification. Description of the Drawings
[0032] Figure 1 is a schematic flowchart of the method of step S200 of the present invention.
[0033] Figure 2 is a schematic flowchart of the method of step S300 of the present invention.
[0034] Figure 3 is a heat map of the similarity between FMR, MR, and ORI and the text features respectively.
[0035] Figure 4 is a schematic diagram of the statistical analysis of hallucinated words before and after applying HD. Detailed Embodiments
[0036] Combined with Figure 1 and Figure 2 , a zero-shot image captioning method based on filtering hybrid representation and hallucination suppression includes:
[0037] Step S100: Obtain the key entity E in the image;
[0038] Step S200, extract sub-region features from entity E obtained by the image encoder, and mix the sub-region features with the global image representation;
[0039] Step S300, use the hallucination suppression method to obtain the original generation probability of the token and form an image description.
[0040] Entity extraction methods mainly include retrieval-based methods, classification-based methods, and detection-based methods. The accuracy of entities is crucial for filtering sub-region attention weights and identifying hallucinations in generation-based sentences. In this specific embodiment, a retrieval-based method is used to obtain entities, and this method shows higher accuracy than other methods. The specific process of step S100 is as follows:
[0041] Step S101, use the text encoder CLIP to obtain the text embedding vector T of each sentence S in the text corpus C of the text, where the text corpus C is expressed as i where N represents the total number of sentences and i is the index of the sentence; i
[0042] Step S102, use the image encoder CLIP to embed the image I to obtain an image embedding vector;
[0043] Step S103, obtain the cosine similarity between the image embedding vector and each text embedding vector T i ;
[0044] Step S104, sort the cosine similarities in descending order and select the top k text embedding vectors;
[0045] Step S105, use a syntax parsing tool to identify all nouns in the text embedding vectors selected in step S104, and select the top k' highest-frequency-ranked nouns as the detected key entity E.
[0046] In step S200, the filtered sub-region features are mixed with the global image representation. This process helps to improve the model's understanding and expression ability of key elements in the image. By combining local and global information, the image description becomes more accurate and rich. The specific process is as follows:
[0047] Step S201, use the vision transformers encoder (ViT) to filter the sub-region (patch) features and class tokens of entity E; among them, the output of the last layer of ViT The output of the last layer contains patch features and a class token I ct where N pf is the number of patch features, and D v is the output dimension of ViT;
[0048] Step S202, use the projection layer proj of the Image Encoder (CLIP Image Encoder) to project the sub-region features and class tokens into the image-text shared space respectively; the patch features and class tokens I ct are projected into the image-text shared space, and the projected class tokens are the global representation where D is the dimension of the image-text shared space, and the projected patch features are The projection process can be expressed as:
[0049]
[0050] I ppf is the projected patch feature, and I g is the global image representation;
[0051] Step S203, use the Text Encoder (CLIP Text Encoder) to encode the prompt constructed by the entity E to obtain the prompt embedding where the prompt is the prompt statement composed of the entity E. For example, the term "There are cake, fork in the image" is usually expressed by hard prompt;
[0052] Step S204, obtain the cosine similarity between I ppf and T pe , and adaptively filter the projected patch features. This process can be expressed as
[0053]
[0054] where the filtering is to filter out a part of I pe that is not similar to T ppf through the threshold α. α is a balance parameter used to control the mixing ratio of local features and global features, and N ppf represents the number of I ppf embedding vectors;
[0055] Step S205, mix the filtered representation I fr with the global image representation I g . This process can be expressed as
[0056]
[0057] where I fr is the filtered representation, I fmr is the final filtered and mixed representation, and N fr is the number of vectors that need to be concat (merged).
[0058] Considering the problems of synonyms and inflections in the vocabulary of LLMs, it is not enough to simply reduce the probability of a single token. Therefore, this specific implementation needs to identify all words similar to the hallucinated words in the vocabulary as negative candidates. The specific process of step S300 is as follows:
[0059] Step S301, use the Hallucination Capture Module (HCM) to extract the hallucinated word list L h
[0060] L h = HMC(C ini ) (4)
[0061] where C ini is the description generated using I fmr features;
[0062] Step S302, all words V d in the LLMs vocabulary and the hallucinated word list L h are embedded through the word embedding wte, and the cosine similarity between the embedding of each hallucinated word and each word in the entire vocabulary V d is calculated. If the cosine similarity exceeds the threshold β, the i-th word in the vocabulary and its corresponding cosine similarity score sim i are added to the negative candidate list t nc , i ∈ [0, N all ), N all is the total number of words in the vocabulary V d . This process can be expressed as
[0063]
[0064] Step S303, during the process of recursively generating description statements, if the next token x t matches a negative candidate, the original generation probability of the next token is suppressed using formula (6)
[0065]
[0066] where the match means that the cosine similarity between the embedding vector of the generated token and the word embedding vector in the negative candidate list is greater than β, x <t represents the previously generated token, p = p(x t |x <t ) represents the original generation probability of the next token, and ω is a hyperparameter used to control the degree of suppression;
[0067] Step S304, decode the generated tokens to obtain words.
[0068] Example
[0069] This example is conducted on three widely used image captioning benchmark datasets, namely NoCaps, COCO, and Flickr30k. For COCO and Flickr30k, the common Karpathy split is followed. For NoCaps, this example trains the model on the COCO training set and reports the results on the validation set as per the OSCAR's suggestion. The results are reported using common caption evaluation metrics, including BLEU@n (B@n), METEOR (M), CIDEr (C), and SPICE (S).
[0070] In hallucination detection (HD), the threshold β for word similarity is set to 0.75. The prompt for the entity list is “There are... in the image”. For caption generation, this example uses beam search with a beam size of 5.
[0071] To verify the generality of the proposed modules of the present invention, this example integrates these modules into five established pre - guiding methods (CapDec, Knight, Decap, ViECap, and IFCap), and evaluates their effects on zero - shot image captioning (ZIC) in different domain settings. It should be noted that due to the fixed number of image representations in the Knight method, it is not suitable for applying the FMR (Filtered Sub - region Features mixed with Global Image Representation) module. In addition, the HD module can also be applied to the post - guiding strategy.
[0072] Table 1
[0073]
[0074] Table 1 shows the in - domain caption generation results on the COCO test set and Flickr30K test set. The values on the left side of the slash represent the original results, and the values on the right side represent the results after applying our method. The symbol explanations are as follows: represents the results of our re - implementation; * represents the use of a memory bank; represents the use of a text - to - image generation model during training; represents the use of an object detector during training and inference.
[0075] Table 2
[0076]
[0077] Among them, X→Y refers to the source domain to the target domain.
[0078] Cross - domain evaluation in Table 3
[0079]
[0080] Table 3 shows the cross - domain description generation results on the NoCaps validation set. Note that SmallCap reports the results on the NoCaps test set, while other methods report the results on the NoCaps validation set.
[0081] 1. Intra - domain description generation
[0082] This experiment evaluation was conducted on the intra - domain description task, that is, the training and test data come from the same dataset. The generated results were compared with the late - guided method and the early - guided method trained on synthetic data. For fair comparison, the LLMs in all methods were only trained with plain text using the corresponding datasets.
[0083] Table 1 shows that applying the FMR and HD modules during the inference process led to performance improvements in most metrics for all methods on different datasets. Specifically, the CIDEr score of the current state - of - the - art method IFCap increased by 0.5 on the COCO dataset and by 4 on the Flickr30K dataset. This improvement can be attributed to the FMR module further reducing the modality gap and the HD module reducing the occurrence of hallucinated words. Other methods, such as CapDec and DeCap, did not hinder the effectiveness of FMR and HD even with additional operations on the input image features. This indicates that the method of the present invention has a certain robustness. This embodiment shows the generated descriptions in Figure 2 and provides more results in the appendix.
[0084] 2. Cross - domain description generation
[0085] The evaluation of cross - domain descriptions is divided into two parts: experiments between COCO and Flickr30K, and experiments from COCO to NoCaps.
[0086] Table 2 shows the cross - domain evaluation results between COCO and Flickr30K. It can be seen that IFCap significantly improved the CIDEr score in different cross - domain settings (from 59.2 to 63.7, from 76.3 to 82.4). Other methods also had a certain degree of improvement, such as CapDec from COCO to Flickr30K (from 32.0 to 33.1), and ViECap from Flickr30K to COCO (from 53.1 to 56.5). These results indicate that the cross - domain generalization ability of most methods has been enhanced after integrating the FMR and HD modules. Both Table 2 and Table 3 demonstrate the significant applicability and improvement effect of our method under cross - domain conditions.
[0087] 3. Ablation experiment
[0088] Taking IFCap as an example, this paper explores the effects of FMR, HD, and their related hyperparameters on the results. The main hyperparameters include the threshold α for filtering attention weights and the threshold ω for controlling the hallucination probability.
[0089] (1) Effects of FMR with different alphas
[0090] Table 4 shows the effects of constructing the FMR module using different thresholds α. The results indicate that the performance of FMR on the COCO dataset gradually improves as α increases. When α is 0.4, the CIDEr score reaches 108.2. However, using MR leads to a decrease in the CIDEr score. Although MR has a positive impact on the Flickr30k dataset, FMR significantly outperforms the baseline method in all cases. Table 5 shows the cross-domain results of FMR under different thresholds. Compared with MR, FMR achieves better results in both cross-modal settings, demonstrating its stronger cross-modal ability and effectiveness in reducing the modality gap, thus improving the zero-shot generation performance.
[0091] Table 4
[0092]
[0093] Table 5
[0094]
[0095] (2) Effects of HD with different omegas
[0096] To further explore the effects of different weights on the generated captions, this embodiment uses four different weight values and evaluates the generated captions in different domain settings. The results are shown in Tables 6 and 7. It can be seen that relatively smaller weights achieve better results in both settings. This is because smaller weights can better balance visual guidance and the degradation of hallucinated words. In the cross-domain setting, it significantly outperforms the baseline method under all weight values, indicating that the hallucination phenomenon is more prevalent under cross-domain conditions, thus making the HD module play a greater role.
[0097] Table 6
[0098]
[0099] Table 7
[0100]
[0101] (3) Analysis of each module
[0102] To further evaluate the discrimination ability of FMR, eight similar images were selected in this embodiment, and the similarity heatmaps of FMR (fine-grained mixed representation), MR (mixed representation), and ORI (original image representation) with text features were obtained respectively. Figure 3 The first row is from the COCO dataset, and the second row is from the Flickr30K dataset. Through this comparison, the advantages of FMR in capturing subtle differences between images and texts can be more intuitively understood. Figure 3 The similarity heatmaps of FMR, MR, and ORI on eight similar images are shown. It can be seen from the figure that the difference between the diagonal similarity and the non-diagonal similarity of FMR is more obvious than that of MR and ORI. This indicates that FMR is more powerful in multi-modal discrimination ability. Figure 4 The changes in object hallucination before and after applying the HD module in the cross-domain setting from Flickr30K to COCO are shown. The results show that the HD module can effectively reduce the appearance of hallucinated words and improve the quality of the generated sentences.
Claims
1. A zero-shot image description method based on filtered hybrid representation and hallucination suppression, characterized in that: include: Step S100, obtaining key entities in the image; Step S200, obtaining sub-region features and category tokens of the entity through an image encoder, projecting the sub-region features and category tokens to obtain projected sub-region features and a global image representation, filtering the projected sub-region features and mixing them with the global image representation; Step S300: using a hallucination suppression method to obtain the original generation probability of a word unit to form an image description.
2. The method according to claim 1, characterized in that The specific process of step S100 is: Step S101, using a text encoder to obtain a text embedding vector for each sentence in a text corpus; Step S102, using an image encoder to embed the image to obtain an image embedding vector; Step S103, obtaining the cosine similarity between the image embedding vector and each text embedding vector; Step S104, sorting the cosine similarities from high to low, and selecting the first k text embedding vectors; Step S105, identifying all nouns in the selected text embedding vector, and selecting the top k' highest frequency nouns as detected key entities.
3. The method according to claim 1, characterized in that The specific process of step S200 is: Step S201, obtaining sub-region features and category tokens of key entities; Step S202, using an image encoder to project the sub-region features and category tokens of the key entity into an image-text shared space to obtain projected sub-region features and a global image representation; Step S203, using a text encoder to encode the prompt constructed by the key entity to obtain a prompt embedding; Step S204, obtaining the cosine similarity between the projected sub-region features and the hint embedding, and adaptively filtering the projected sub-region features; Step S205, mixing the filtered representation with the global image representation.
4. The method according to claim 3, characterized in that: In step S204, if the cosine similarity between the projected sub-region feature and the hint embedding is greater than a threshold, the projected sub-region feature that is not similar to the hint embedding is filtered out.
5. The method according to claim 3, characterized in that: The filtered representation in step S205 is mixed with the global image representation by the following formula: Among them, I fmr is the final filtered mixed representation, I fr is the filtered representation, N fr is the number of vectors that need to be merged.
6. The method according to claim 1, characterized in that The specific process of step S300 is: Step S301, using a hallucination capture module to extract hallucination words in the mixed representation and form a hallucination word list; Step S302, embed all words and the hallucination word list through word embedding, calculate the cosine similarity between the embedding of each hallucination word and each word in all words, and if the cosine similarity exceeds a threshold, add the word and its corresponding cosine similarity score to the negative candidate words; Step S303, in the process of recursively generating a description sentence, if the next word element matches a negative candidate word, the original generation probability of the next word element is suppressed using the following formula: Among them, x <t represents the previously generated word, p=p(x t |x <t ) represents the original generation probability of the next word, ω is a hyperparameter used to control the degree of inhibition, t nc is a negative candidate word, sim i is the cosine similarity between the embedding of the i-th hallucinated word and each word in all the vocabularies; Step S304, decoding the generated word-grams to obtain words and generate image descriptions.
Citation Information
Cited By
Image description method and device based on prompt vector and CLIP reward and punishment mechanism
CN121883658A
Image description method and device based on prompt vector and CLIP reward and punishment mechanism
CN121883658B