An evaluation method for the alignment constraints of vision-language models
By building alignment limit benchmark datasets and evaluating the performance of vision-language models, AlignVLM solves the challenges of vision-language models in image-text alignment, significantly improving the performance and accuracy of downstream tasks.
Patent Information
- Application Number
- CN202411747937.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-02
AI Technical Summary
Existing visual-language models have significant challenges in the alignment between images and text, especially when encountering similar visual-language pairing data, with significant performance declines.
A method of evaluation of visual-language model alignment limitation (AlignVLM) is proposed to evaluate the performance of visual-language model by constructing text-text-to-image alignment limitation benchmark dataset (TT2I) and image-image-to-text alignment limitation benchmark dataset (II2T), and using recall R@K as an evaluation indicator.
By improving the alignment of visual and linguistic data, AlignVLM has significantly improved the performance and accuracy of various downstream tasks, such as in the applications of medical image analysis, autonomous driving, e-commerce platforms and online education, which has significantly improved the accuracy of image-text retrieval and visual question-and-answer.
Smart Images

Figure CN119249115B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and natural language processing, and relates to the research of vision-language models (VLMs), and in particular to an evaluation method of vision-language model alignment restrictions for alignment analysis of image and text data. Background Art
[0002] Vision-Language Models (VLMs) have achieved remarkable success in recent years and have attracted extensive research attention. These models mainly achieve cross-modal learning and coverage of open vocabulary knowledge by aligning large-scale image-text pairing data and building a shared embedding space. As a result, VLMs have shown excellent capabilities in various vision-language downstream tasks such as image-text retrieval and visual question answering (VQA). Despite the exciting performance improvements of vision-language models, there is still a significant research gap in exploring the fundamental principles of their functionality.
[0003] Recently, some studies have begun to explore the fundamentals of VLMs, focusing on visual deficiencies and language limitations. Visual deficiencies refer to the fact that VLMs encode different visual images in a similar embedding space, which may lead to ambiguity in the encoding of at least one image. Language limitations refer to failure modes in pre-trained text encoders that may lead to ambiguity in text-guided generative models. While these studies address the problems of VLMs in terms of visual deficiencies and language limitations, they ignore the alignment problem between images and text, which is the biggest challenge facing VLMs. Most vision-language models (VLMs) perform significantly worse when they encounter similar visual-language paired data. This observation is called the "alignment limitation". Summary of the invention
[0004] The purpose of the present invention is to solve the above-mentioned problems existing in the prior art and to provide an evaluation and analysis method for visual-language model alignment constraints (AlignVLM), which is applicable to a variety of downstream tasks, such as image-text retrieval and visual question answering, and can be widely used in medical image analysis, autonomous driving, e-commerce platforms, online education and other fields.
[0005] The specific technical solution for achieving the purpose of the present invention is:
[0006] A method for evaluating a visual-language model alignment constraint, the method comprising the following steps:
[0007] Step 1: Construct a text-to-image alignment restriction benchmark dataset, namely the TT2I strategy, which includes:
[0008] Step 1.1: Text embedding extraction, use the pre-trained CLIP text encoder to embed the input text data x and obtain the text representation vector: ;
[0009] Step 1.2: Calculate the similarity using the formula: in Represents text With text Similarity express According to the set similarity threshold, text pairs with similarity greater than the threshold are selected to form a data set of text transition benchmarks;
[0010] Step 1.3: Corresponding to the text pairs of the text transition benchmark dataset, select the images associated with the text pairs , and use the formula To calculate the similarity between images, Representing images With image Similarity express The standardized features of the dataset are obtained by filtering images by a set similarity threshold, retaining image pairs with similarity less than the threshold and their associated text pairs to form a text-text-to-image alignment restriction benchmark dataset; the texts in the resulting dataset are similar, but the image representations are not similar; among them, Represents text The transpose of the eigenvector of Representing images The eigenvector transpose of ;
[0011] Step 2: Construct an image-to-image to text alignment restriction benchmark dataset, namely the II2T strategy, which includes:
[0012] Step 2.1: Image embedding extraction: Use the pre-trained CLIP image encoder to embed the input image data and obtain the image representation vector ;
[0013] Step 2.2: Calculate image similarity and construct a transition benchmark dataset, using the similarity calculation formula: To calculate the similarity between images, Representing images With image Similarity express Based on the standardized features of the image, according to the set similarity threshold, the image pairs with similarity greater than the threshold are selected to form a dataset of image transition benchmarks;
[0014] Step 2.3: Corresponding to the image pairs in the Image Transition Benchmark dataset, select the text associated with these image pairs , , and use the formula To calculate the similarity between texts, Represents text With text Similarity express Standardized features; Filter texts by setting a similarity threshold, retain text pairs with similarity less than the threshold and their associated image pairs, and form an image-image-to-text alignment restriction benchmark dataset; The images in the resulting dataset are similar, but the texts are not similar; Representing images The eigenvector transpose of Represents text The eigenvector transpose of ;
[0015] Step 3: Evaluate the performance of visual-language models (VLMs) on the two constructed alignment-restricted benchmark datasets, using recall R@K as the evaluation metric, where R@K is defined as the proportion of correctly retrieved images or texts in the top K results; the higher R@K, the better the model performance.
[0016] The AlignVLM method of the present invention has shown significant effects in practical applications in multiple fields. By improving the alignment ability of visual and language data, the performance and accuracy of various downstream tasks can be significantly improved. For example: Case analysis: Through AlignVLM, CT images can be more accurately aligned with medical record descriptions, thereby helping doctors identify abnormal lesions more quickly. Diagnostic assistance: AlignVLM can provide more accurate multimodal data matching in auxiliary diagnosis systems and reduce misdiagnosis rates. Autonomous driving systems rely on accurate alignment of visual and language data to achieve accurate understanding of the vehicle's surroundings. AlignVLM can help autonomous driving systems improve the alignment accuracy of visual and language information when processing data of similar scenes, thereby improving the safety and reliability of driving decisions. For example: Scene recognition: When processing similar traffic signs or road signs, AlignVLM can help the system understand the content of the signs more accurately and avoid misidentification. Voice navigation: By improving the alignment effect of visual and voice commands, AlignVLM can improve the accuracy of voice navigation systems and ensure that vehicles travel along the expected route. In e-commerce platforms, the alignment accuracy of images and text descriptions directly affects users' shopping experience and sales conversion rates. AlignVLM can help improve the alignment of product images and text descriptions, improve the accuracy of the recommendation system and user satisfaction. For example: Product recommendation: AlignVLM can more accurately match product images and text descriptions, and improve the recommendation quality of the recommendation system. Search optimization: AlignVLM can help optimize search results so that users can get more accurate matching results when searching for similar products, improving user experience. On social media platforms, user-generated content contains a large amount of image and text data. AlignVLM can help improve the alignment of image and text content, and improve the accuracy of content recommendation and review. For example: Content recommendation: AlignVLM can more accurately match images and text content posted by users, and improve the quality of personalized recommendations. Content review: AlignVLM can help the review system more accurately identify inappropriate content in images and text, and improve the efficiency and accuracy of platform content review. In the field of education, the alignment of visual and language data is crucial for the understanding and dissemination of multimedia teaching content. AlignVLM can help improve the alignment of teaching videos and text handouts, improve teaching effectiveness and students' learning experience. For example: Multimedia courseware: AlignVLM can more accurately align the image content in teaching videos with text handouts, and improve the quality of courseware. Online Q&A: AlignVLM can help online education platforms improve the matching effect of images and text content, and improve the accuracy and efficiency of student Q&A. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1Constructing a restricted benchmark data set for the AlignVLM method of the present invention;
[0018] Figure 2 Schematic diagram of performance comparison of different vision-language models constructed by the AlignVLM method of the present invention. DETAILED DESCRIPTION
[0019] In order to make the above-mentioned purpose, features and advantages of the present invention more obvious and easy to understand, the specific implementation mode of the present invention is described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. The technical features in each embodiment of the present invention can be combined accordingly without conflicting with each other.
[0020] The method of the present invention comprises the steps of:
[0021] Step 1: Construct a text-to-image alignment restricted benchmark dataset, namely the TT2I strategy, such as Figure 1 As shown, specifically including:
[0022] Step 1.1: Text embedding extraction, use the pre-trained CLIP text encoder to embed the input text data x and obtain the text representation vector: ;
[0023] Step 1.2: Calculate the similarity using the formula: in Represents text With text Similarity express According to the set similarity threshold, text pairs with similarity greater than the threshold are selected to form a data set of text transition benchmarks;
[0024] Step 1.3: Corresponding to the text pairs of the text transition benchmark dataset, select the images associated with the text pairs , and use the formula To calculate the similarity between images, Representing images With image Similarity express Standardized features; Screen images by setting a similarity threshold, retain image pairs with similarity less than the threshold and their associated text pairs, and form a text-text-to-image alignment restriction benchmark dataset; The texts in the resulting dataset are similar, but the image representations are not similar;
[0025] Step 2: Construct an image-to-text alignment restricted benchmark dataset, namely the II2T strategy, such as Figure 1 As shown, specifically including:
[0026] Step 2.1: Image embedding extraction: Use the pre-trained CLIP image encoder to embed the input image data and obtain the image representation vector ;
[0027] Step 2.2: Calculate image similarity and construct a transition benchmark dataset, using the similarity calculation formula: To calculate the similarity between images, Representing images With image Similarity express Based on the standardized features of the image, according to the set similarity threshold, the image pairs with similarity greater than the threshold are selected to form a dataset of image transition benchmarks;
[0028] Step 2.3: Corresponding to the image pairs in the Image Transition Benchmark dataset, select the text associated with these image pairs , , and use the formula To calculate the similarity between texts, Represents text With text Similarity express Standardized features; Filter texts by setting a similarity threshold, retain text pairs with similarity less than the threshold and their associated image pairs, and form an image-image-to-text alignment restriction benchmark dataset; The images in the resulting dataset are similar, but the texts are not similar;
[0029] Step 3: If Figure 2 The figure shows an evaluation of the performance of different visual-language models (VLMs) on two constructed alignment-restricted benchmark datasets, with recall R@K as the evaluation metric, where R@K is defined as the proportion of correctly retrieved images or texts in the top K results; the higher the R@K, the better the model performance; in Figure (a), the candidate images are quite different, and the model can easily identify the target image with a purple outline, and its corresponding text-image retrieval result is shown on the right; however, when faced with a similar candidate image in Figure (b), the model failed to correctly identify the target image, but instead selected a non-target image with a red outline, and the performance dropped significantly, as shown on the right.
[0030] Example
[0031] The corpus datasets mainly used in this embodiment are Flickr30k and MSCOCO. The overall process is as follows:
[0032] 1) Dataset selection
[0033] To evaluate the performance, two different vision and language downstream tasks are considered: image-text retrieval and visual question answering (VQA). The downstream datasets used for each task are described below:
[0034] S1. Image-text retrieval: Experiments are conducted on two of the most widely used benchmark datasets, namely Flickr30K (1K test set) and MSCOCO (5K test set). Each benchmark dataset contains multiple image-text pairs, where each image is described by five corresponding sentences. Specifically, the Flickr30K dataset contains 31,789 images, and according to the split, there are 29,000 images in the training set, 1,000 images in the validation set, and the remaining 1,000 images for the test set. For the MSCOCO dataset, the training, validation, and test sets contain 113,287, 5,000, and 5,000 images, respectively.
[0035] S2. Visual Question Answering (VQA): VQAv1 is the first version of the Visual Question Answering (VQA) dataset, which includes a diverse set of images, each with a natural language question. The dataset contains about 200,000 images, three questions per image, and about 600,000 question-answer pairs in total. VQAv2 is the second version of the Visual Question Answering (VQA) dataset, which aims to address some limitations in the original dataset, such as bias in questions and answers. It contains a larger and more balanced set of images, about 265,000 images, three questions per image, and about 800,000 question-answer pairs in total.
[0036] 2) Data processing
[0037] S1. In TT2I, the first step is to extract text (or image) embeddings using CLIP’s pre-trained text (or image) encoder. The second step is to calculate the similarity between any two texts. Select all text information whose similarity is greater than a given text-to-text threshold, which is a hyperparameter, to form a new dataset called transition benchmark. The third step is to select image information corresponding to the text such that it is less than a given image-to-image threshold based on the transition benchmark. This aims to obtain an alignment-restricted benchmark where text representations are similar but image embeddings are dissimilar.
[0038] S2. In the II2T setting, the first step is the same as TT2I for extracting feature representations. However, in the second step, the similarity between any two images is calculated. Then, all image representations are selected to form a transition benchmark that is larger than a given image-to-image threshold. In the third step, the text representation corresponding to the image based on the transition benchmark is selected to be smaller than a given text-to-text threshold. This aims to obtain an alignment-constrained benchmark where the image representations are similar but the text embeddings are dissimilar.
[0039] 1. Benchmark Vision-Language Methods: The proposed AlignVLM is evaluated on a large number of algorithms covering different learning strategies, including: 1. Pre-training learning: CLIP, CoCa, ALBEF, FLAVA, and BLIP. 2. Vision-Language Model Fine-tuning Learning: CLIP Full Parameter Fine-tuning, ALBEF Full Parameter Fine-tuning, BLIP Full Parameter Fine-tuning, CLIP-Adapter, and TaskRes.
[0040] 2. Experimental details: In the image-text retrieval task, experimental results using standard metrics are reported, including R@K (recall at K=1, 5, 10). R@K is defined as the proportion of correctly retrieved images or texts in the top K results. In the text-text-image (TT2I) setting, 0.9 and 0.4 are used as text-text and image-image similarity thresholds to select the TT2I subset. In the image-image-text (II2T) setting, 0.9 and 0.4 are used as image-image and text-text similarity thresholds to obtain the II2T subset. AlignVLM, as a model-agnostic method, can adopt any image encoder, and ViT-B / 16 is used as an example in the experiments. In Table 1, the image size comparison of the constructed alignment-restricted benchmark dataset using different similarity thresholds on four downstream datasets is also shown. For the visual question answering (VQA) task, samples from all constructed benchmark datasets are used as the test set, because the test results of the VQA task are independent of the number of samples. However, for the image-text retrieval task, 1,000 samples as Flickr30K or 5,000 samples as the test set of MSCOCO were randomly selected from the constructed datasets because the performance of the retrieval task is related to the number of samples.
[0041] 3. Evaluating the Zero-Shot Learning Performance of VLMs: We evaluate the zero-shot performance of visual language models (VLMs) on two downstream tasks, involving four benchmarks. The experimental results are shown in Table 1. Based on these tables, the following findings are drawn: (1) Compared with the original test set, the performance of all VLMs on the constructed alignment-restricted benchmarks drops significantly on retrieval-type downstream tasks and datasets, which indicates its motivation, i.e., AlignVLM, constructs a more challenging benchmark and that existing VLMs have serious alignment restrictions. (2) In the TT2I setting, its goal is to find data with different images but more similar texts, which theoretically affects the performance of image-text tasks rather than text-image tasks. Similarly, in the II2T setting, it performs better on image-text tasks because images are more similar but texts are not. The results in Flickr30k and MSCOCO confirm this phenomenon in all VLM baselines. For example, in Flickr30K, for CLIP, the performance drop for image-text (i.e., 25.9%) is more significant than that for text-image (i.e., 16.3%) in the TT2I setting using the R@1 metric. (3) In the visual question answering (VQA) task, the proposed dataset outperforms the original test set in some settings. We believe that the possible reason is that the VQA task is relatively difficult, so the selected image-text pairing helps to alleviate this problem. However, when averaged over all question sets, the performance of VLMs drops significantly.
[0042] Table 1 Experimental results
[0043]
[0044] 4. Evaluating the fine-tuning effect of VLMs: We evaluate the fine-tuning performance of visual language models (VLMs) on Flickr30K. Note that the VLMs are fine-tuned using the MSCOCO training data and tested on the Flickr30K test data. The reason is that if AlignVLM is only retrieved and constructed on the test set, it will result in a small sample size. Secondly, MSCOCO and Flickr30K are similar, and the images in them are all from the image sharing website Flickr. (1) Compared with the zero-shot experimental results, fine-tuning of VLMs not only improves the performance on the original data, but also improves the performance of the proposed alignment-restricted benchmark. For example, the performance of CLIP is improved from 62.2% to 67.3%. Surprisingly, the performance of CLIP is significantly improved from 36.3% to 41.6% in the proposed setting. This shows that fine-tuning the model can alleviate the alignment restriction. (2) Compared with different fine-tuning strategies, full-parameter fine-tuning has the best performance among the three visual language models, which provides researchers with a research direction to solve the alignment problem between vision and language.
[0045] The AlignVLM method proposed in this paper has demonstrated remarkable results in practical applications in multiple fields. By improving the alignment ability of visual and language data, it can significantly improve the performance and accuracy of various downstream tasks.
[0046] The above-described embodiment is only a preferred solution of the present invention, but it is not intended to limit the present invention. A person skilled in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.
Claims
1. A method for evaluating visual-language model alignment constraints, characterized in that: The method comprises the following steps: Step 1: Construct a text-to-image alignment restriction benchmark dataset, namely the TT2I strategy, which includes: 1.1: Text embedding extraction: Use the pre-trained CLIP text encoder to embed the input text data x and obtain the text representation vector: f(x); 1.2: Through the similarity calculation formula: Where d(x v ,x t ) represents the text x v With text x t similarity; Represents x t According to the set similarity threshold, text pairs with similarity greater than the threshold are selected to form a data set of text transition benchmarks; 1.3: Corresponding to the text pairs of the text transition benchmark dataset, select the image y associated with the text pair v ,y t , and use the formula To calculate the similarity between images, d(y v ,y t ) represents the image y v With image y t similarity; Represents y t Standardized features; Screen images by setting a similarity threshold, retain image pairs with similarity less than the threshold and their associated text pairs, and form a text-text-to-image alignment restriction benchmark dataset; The texts in the resulting dataset are similar, but the image representations are not similar; Step 2: Construct an image-to-image to text alignment restriction benchmark dataset, namely the II2T strategy, which includes: 2.1: Image embedding extraction: Use the pre-trained CLIP image encoder to embed the input image data and obtain the image representation vector f(y); 2.2: Use the similarity calculation formula: To calculate the similarity between images, select image pairs with similarity greater than a threshold to form a dataset of image transition benchmarks; 2.3: Corresponding to the image pairs in the image transition benchmark dataset, select the text x associated with the image pair v ,x t , and use the formula To calculate the similarity between texts, retain text pairs with similarity less than a threshold and their associated image pairs, forming an image-to-image alignment restricted benchmark dataset; the images in the resulting dataset are similar, but the texts are not similar; Step 3: Evaluate the performance of the visual-language model on the two constructed alignment-restricted benchmark datasets, using recall R@K as the evaluation metric, where R@K is defined as the proportion of correctly retrieved images or texts in the top K results; the higher R@K, the better the model performance.
Citation Information
Patent Citations
Prompt learning method for modal interaction enhancement of visual language model
CN116503683A
Visual language alignment method and system based on cross-modal structure consistency and pre-training technology
CN117557803A