An unsupervised image quality assessment method based on multi-modal prompt learning

By introducing multimodal prompt learning and two-stage training paradigm in reference-free image quality evaluation, the shortcomings of existing methods in complex image distortion and local distortion evaluation are solved, and more accurate and stable image quality evaluation is achieved.

CN117078656BActive Publication Date: 2025-06-24XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311131117.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-04
Publication Date
2025-06-24
Estimated Expiration
2043-09-04

AI Technical Summary

Technical Problem

Existing reference-free image quality evaluation methods are difficult to cope with complex and variable image distortion and quality changes, and the selection of input text prompts is limited, resulting in unstable performance and it is difficult to accurately evaluate local distorted images.

Method used

Using a multimodal prompt learning method, fine-grained text prompts and learnable image prompts are introduced, and the model's understanding and evaluation ability of image quality is enhanced through multi-task learning and two-stage training paradigm.

Benefits of technology

It improves the accuracy and stability of image quality evaluation, can better capture subtle differences in image quality, adapt to local distortion image evaluation, and achieve a comprehensive understanding and accurate prediction of image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117078656B_ABST
    Figure CN117078656B_ABST
Patent Text Reader

Abstract

An unsupervised image quality assessment method based on multi-modal prompt learning, belonging to the field of computer vision technology. No-reference image quality assessment aims to simulate human evaluation of image quality without a reference image (original image). The present invention gives full play to the potential of the pre-trained CLIP model in challenging image perception assessment tasks. First, multi-modal prompt learning is introduced, enabling flexible adjustment of the CLIP model's representation space in BIQA, thereby stimulating its potential in challenging image perception assessment tasks. Second, the previous method of text prompt learning is improved, replacing the antonym text prompt learning used in the previous method with a fine-grained text prompt learning, so as to be able to capture the fine-grained features of the image and obtain a more accurate quality assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a no-reference image quality assessment method based on multi-modal prompt learning. Background Art

[0002] Image quality assessment (IQA) is an important research direction in the field of computer vision, and its main goal is to quantitatively predict and evaluate the visual quality of images based on human visual perception characteristics. Image quality assessment has important practical significance in many applications such as image processing, image transmission, and video coding.

[0003] Over the years, a variety of IQA methods have been developed and evaluated. These methods are generally divided into three categories according to the amount of information provided by the original reference image: full reference (FullReference-IQA, FR-IQA), reduced reference (ReducedReference-IQA, RR-IQA), and no reference (NoReference-IQA, NR-IQA), and no reference is also called blind reference (BlindIQA, BIQA). Full reference image quality assessment can refer to the original undistorted image and obtain the image quality score of the distorted image based on the difference between the distorted image and the original image. Reduced reference image quality assessment uses partial information of the original image as a reference to predict the image quality. However, in practical applications, it is difficult to obtain the reference image, making the above two methods inapplicable. Therefore, recent work has gradually focused on the BIQA field.

[0004] Traditional no-reference image quality assessment methods usually rely on manually designed features and rules, and it is difficult to cope with complex and variable image distortion and quality change situations. With the rapid development of deep learning technology, image quality assessment methods based on deep neural networks have made significant progress in terms of accuracy and generalization ability. These methods utilize the powerful expressive ability of deep neural networks to automatically learn and extract complex perceptual features in images, thereby achieving accurate assessment of image quality. Through deep learning methods, the limitation of manually designing features in traditional methods can be overcome, and end-to-end training can be carried out on large-scale datasets, thus improving the performance and generalization ability of image quality assessment.

[0005] OpenAI proposed the CLIP model (Radford, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning (PMLR), pages 8748-8763, 2021.), which is a self-supervised pre-training framework that encodes images and text into a shared vector space without labeled data. CLIP demonstrates strong zero-shot transfer learning capabilities on various tasks. Wang et al. (Wang, et al. Exploring clip for assessing the look and feel of images. arXiv preprint arXiv:2207.12396 (2022)) first explored the potential of CLIP in challenging image perception assessment tasks. They proposed a prompt pairing strategy using antonym prompts (e.g., "good photo" and "bad photo") to reduce the ambiguity of the prompts. On this basis, Zhang et al. (Zhang, et al. Blind image quality assessment via vision-language correspondence: A multitask learning perspective. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14071-14081, 2023.) proposed multitask learning to fine-tune CLIP by combining image quality assessment with distortion type classification and scene classification tasks. This method automatically determines the sharing of model parameters and loss weights, leveraging auxiliary knowledge from other tasks. They used a Likert scale with five quality levels, and the model output gives the quality level as a score of a logical weighted sum.

[0006] However, the effectiveness of these methods is often limited by the selection of input text prompts. For image quality assessment, choosing appropriate text prompts is crucial, as different prompts may lead to performance instability and fluctuations. Additionally, some models in existing methods still predict high quality scores even when the main part of the image is clear while other parts are distorted, which does not match the actual situation. On the other hand, existing methods rarely consider the distortion conditions of different regions of the image and only focus on the perceptual evaluation of the overall image. This makes these methods have certain limitations when dealing with images with local distortions. Summary of the Invention

[0007] The object of the present invention is to provide a reference-free image quality assessment method based on multi-modal prompt learning, which fully exploits the potential of the CLIP model in challenging image perceptual evaluation tasks. This method introduces fine-grained text prompts, and by classifying the image quality assessment problem more precisely, enables the model to more accurately capture various subtle differences in image quality. At the same time, learnable image prompts are added to each layer of the visual branch, so that the model can better adjust its representation space in the image quality assessment task and more comprehensively understand the quality-related information in the image. In addition, a two-stage evaluation paradigm is proposed, enabling the model to gradually adapt from rough perception to detailed evaluation, thereby achieving a comprehensive understanding and accurate prediction of image quality.

[0008] The present invention includes the following steps:

[0009] 1) Grade each picture in the dataset according to its score and assign a level (or category) label. For example, an image with a score in the range of [49 - 50) is classified as level

[50] , corresponding to category

[50] .

[0010] 2) In the text branch of the model, introduce learnable text prompts to address the sensitivity of the CLIP model to prompts when migrating to downstream tasks; at the same time, in order to enable the model to better adjust its representation space in the image quality assessment task and more comprehensively understand the quality-related information in the image, learnable image prompts are also introduced in the image branch.

[0011] 3) The training of the model is divided into two stages: in the first model training stage, only the text prompts and image prompts are trained, and other parameters are frozen. With extremely few training parameters, the CLIP model migrated by the present invention is enabled to have the ability to perceive and evaluate image quality. In the second model training stage, the prompts of both branches are frozen, and only the image encoder is trained, enabling the model to gradually adapt from rough perception to detailed evaluation, thereby achieving a comprehensive understanding and accurate prediction of image quality.

[0012] Features and effects of the present invention:

[0013] The no-reference image quality assessment method based on multi-modal prompt learning proposed by the present invention solves the defect of performance bottleneck in previous work and achieves a great improvement in performance. The present invention introduces the concept of multi-modal prompt learning in the image quality assessment task. It not only introduces learnable text prompts in the text branch, but also adds learnable image prompts to each layer of the image branch. This enables the model to make full use of the information interaction between text and image when evaluating image quality, better understand the quality-related features in the image, and thus improve the performance of the model in the image quality assessment task. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a comparison chart between the present invention and previous methods.

[0015] Figure 2 It is a framework diagram of the present invention.

[0016] Figure 3 It is a visualization heat map of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0017] The following embodiments will further illustrate the present invention in conjunction with the accompanying drawings.

[0018] The flow of the method of the present invention is as Figure 2 shown. It includes two stages. In the first model training stage, only the text prompts and image prompts are trained, and other parameters are frozen. With extremely few training parameters, the transferred CLIP model is enabled to have the ability to perceive and evaluate image quality. In the second model training stage, the prompts of the two branches are frozen, and only the image encoder is trained, enabling the model to gradually adapt from rough perception to detailed evaluation, so as to achieve a comprehensive understanding and accurate prediction of image quality. Figure 2 The upper right corner is to introduce learnable prompts in the image branch, and add learnable image prompts to each transformer layer of the image encoder.

[0019] 1. Training Instructions

[0020] The embodiments of the present invention include the following steps:

[0021] 1) Grade each picture in the dataset according to its score and assign a level (or category) label. For example, an image with a score in the range of [49 - 50) is classified as level

[50] , corresponding to category

[50] . The model will calculate the similarity based on the output features of the text encoder and the input features of the image encoder, and obtain the final quality score q(x) in a weighted summation manner.

[0022]

[0023] Among them, P(c|x) represents the class probability after applying softmax, and C represents the total number of classes.

[0024] 2) In the text branch of the model, a set of learnable text prompts are introduced to make full use of the representation ability of the text encoder. The form of the text input is designed as follows: "[X]1[X]2[X]3...[X] M [class]." "[X]1[X]2[X]3...[X] M " is used to represent the prefix of the text input, where M represents the number of learnable tokens. Additionally, in the image branch, "depth visual prompts" are introduced, which involves adding prompts to each layer of the image encoder. The main purpose is to enhance the alignment between the perceptual features of the image and the quality-level text features. By introducing prompts at multiple levels, the learning process of the image features is more finely controlled, thus better adapting to the image quality assessment task.

[0025] A set of learnable tokens P are introduced, located between the classtoken of the image and the patch embeddings of the image. At each layer of the image encoder, the input is represented as [Cls i-1 ,P i-1 ,E i-1 , where Cls represents the classtoken and E represents the patch embeddings. After passing through the i-th Transformer layer L i , new learnable tokens P are continuously introduced and concatenated with the output Cls and E.

[0026] [Cls i ,○,E i = L i ([Cls i-1 ,P i-1 ,E i-1 )#(2)

[0027] Among them, ○ means not to be used as the input for the next Transformer layer. This design allows the present invention to introduce learnable tokens at each layer of the image encoder, enriching the model's ability to capture and represent important image features and quality characteristics, and ultimately contributing to more effective and accurate image quality assessment.

[0028] 3) The training of the model is divided into two stages. In the first stage of model training, only the text prompts and image prompts are trained, and other parameters are frozen. With very few training parameters, the transferred CLIP model is made to have the ability to perceive and evaluate image quality. The loss function used is:

[0029]

[0030] Among them, i represents the i-th picture, T and V respectively represent the output features of the text encoder and the image encoder, sim(,) represents the cosine similarity of two features, B represents a minibatch, and A represents the set of all pictures in the minibatch that belong to the same category as picture i.

[0031] In the second model training stage, the prompts of the two branches are frozen, and only the image encoder is trained, enabling the model to gradually adapt from rough perception to detailed evaluation, thereby achieving a comprehensive understanding and accurate prediction of image quality. The loss function used in this stage is:

[0032]

[0033] Introduce the fidelity loss to consider pairwise learning for ranking the model estimates. Additionally, use Smooth L1Loss and cross-entropy loss with label smoothing for optimization. Among them, α and β are coefficients for balancing and .

[0034]

[0035] p and p' respectively represent the true probability distribution and the predicted probability distribution.

[0036]

[0037] I' k =(1 - ε)I k +ε / C represents the value in the quality level target distribution, and P k represents the predicted logits for class k.

[0038] 2. Implementation details

[0039] 1) Model details

[0040] The present invention is implemented using the Pytorch framework. Both the image encoder and the text encoder of the present invention are derived from the CLIP framework. Specifically, the ViT-B / 16 architecture and a single-layer transformer decoder are adopted as the image encoder. This configuration contains 12 transformer layers, and the hidden size of each layer is 768 dimensions. In order to match the output of the text encoder, a linear layer is used to reduce the dimension of the image feature vector from 768 to 512.

[0041] 2) Training details

[0042] The training process is divided into two stages, each stage containing 60 epochs. In the first stage, only the learnable text prompts and image prompts are concerned, and other parameters are frozen. The Adam optimizer is used with an initial learning rate of 3×10 -5 , and then cosine learning rate scheduling is adopted for decay. To augment the training data, each original image is randomly cropped into 8 sub-images, each with a size of 3×224×224. In the second stage, the image encoder is optimized. The Adam optimizer is still used, but this stage includes a warm-up phase of up to 10 epochs. The learning rate linearly increases from 9.5×10 -7 to 5×10 -6 . The learning rate is decreased by multiplying by 0.1 at the 30th and 50th epochs. Data augmentation in this stage includes random horizontal flipping and random cropping. The coefficients α is set to 0.001 and β is set to 0.1.

[0043] According to the score label of each image, a new class label is assigned to each image. The batch size for the LIVE and CSIQ datasets is set to 32, while the batch size for other datasets is set to 64. The data is divided into 80% for training and 20% for testing. To reduce performance bias, each experiment is repeated 10 times, and the average PLCC and SRCC are calculated.

[0044] 3. Application Fields

[0045] The present invention can be applied to the field of no-reference image quality assessment to realize the judgment of the quality of distorted pictures without the original image.

[0046] Table 1 shows the performance comparison between the model of the present invention and previous models on 6 common image quality assessment datasets. These datasets include four real datasets and two synthetic datasets. The real datasets used include LIVE, CSIQ, TID2013, and KADID. The synthetic datasets include LIVEC and KonIQ. To evaluate the performance of the model, the Pearson linear correlation coefficient (PLCC) and the Spearman rank correlation coefficient (SRCC) are used as evaluation metrics. PLCC measures the accuracy of the model prediction, while SRCC evaluates the monotonicity of the BIQA algorithm prediction. The value ranges of both metrics are from 0 to 1, and the higher the value, the better the performance in terms of prediction accuracy and monotonicity.

[0047] As can be seen from Table 1, on all datasets, the model of the present invention exhibits competitive performance compared to the state-of-the-art methods. It is worth noting that on the CSIQ, TID2013, and KADID datasets, state-of-the-art performance is achieved, surpassing existing methods by increasing the PLCC metric by 1.0%, 1.7%, and 5.0%, respectively, and the SRCC metric by 1.9%, 2.2%, and 4.2%. These results highlight the effectiveness and leading performance of the model of the present invention in image quality assessment.

[0048] Table 1

[0049]

[0050] Table 1: Performance comparison is measurd by averages of SRCC andPLCC. Best results are highlighted in bold, second-best are underlined .

[0051] Table 2 shows the comparison of the generalization performance of the model of the present invention with previous models on cross-datasets. Specifically, a BIQA model is trained on one dataset and directly evaluated on another dataset without fine-tuning or parameter adjustment. Four datasets are used and median experimental results are reported. On the KonIQ and CSIQ datasets, the model of the present invention exhibits superior performance and remains competitive on other datasets. These experimental results highlight the robust generalization performance of the model.

[0052] Table 2

[0053]

[0054] Figure 1 A comparison chart of the present invention and previous methods is given. It can be seen that CLIP-IQA + introduces learnable text prompts on the basis of CLIP-IQA, but does not change the way of antonym classification. Compared with the former two, a fine-grained classification method according to quality score intervals is introduced. In addition, the present invention also introduces learnable prompts in the image branch.

[0055] Figure 3The visualization heat map of the present invention is given, and GradCAM is used to visualize the feature-attention map of the input image in DEIQT (Qin G, Hu R, Liu Y, et al. Data-Efficient Image Quality Assessment with Attention-Panel Decoder[J]. arXiv preprint arXiv:2304.04952, 2023.) and the model of the present invention. Focusing on low-scoring images, it aims to reveal the reason why the model of the present invention is superior to DEIQT on LIVEC. As Figure 3 shown, it is observed that DEIQT overly focuses on the distortions within the main content of the image. This results in DEIQT predicting a higher quality score even when the main content of the image remains clear, while there are serious distortions in other regions, which is clearly inconsistent with reality. In contrast, the model of the present invention considers the distortions in different regions of the image, thereby providing a more accurate assessment of the image quality. This advantage stems from the multi-modal prompt and two-stage training mode design of the present invention. First, the multi-modal prompt design allows the model to incorporate both text and visual information simultaneously, enabling a more comprehensive understanding of the inherent features in the image. Especially in low-quality images, the model of the present invention can effectively capture the subtle distortions that may be overlooked. Second, the two-stage training strategy of the present invention enables the model to gradually adapt to the BIQA task, from rough perception to detailed evaluation, achieving a comprehensive understanding of the image quality.

[0056] The above embodiments are only preferred embodiments of the present invention and should not be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made in accordance with the scope of the present invention application shall still fall within the scope covered by the patent of the present invention.

Claims

1. An unsupervised image quality assessment method based on multi-modal prompt learning, characterized in that It includes the following steps: 1) Classify each image in the dataset according to its score and assign a level or category label; 2) In the text branch of the model, introduce learnable text prompts to address the sensitivity of the CLIP model to prompts when migrating to downstream tasks. Specifically: Introduce a set of learnable text prompts to make full use of the representation ability of the text encoder. Additionally, in the image branch, introduce deep visual prompts, which involves adding prompts to each layer of the image encoder to enhance the alignment between the perceptual features of the image and the text features of the quality level. By introducing prompts at multiple levels, the learning process of image features can be more finely controlled, thus better adapting to the image quality assessment task; Introduce a set of learnable tokens P, located between the class token of the image and the patch embeddings of the image; at each layer of the image encoder, the input is represented as [Cls i-1 ,P i-1 ,E i-1 , where Cls represents the class token and E represents the patch embeddings; after passing through the i-th Transformer layer L i , continue to introduce new learnable tokens P and concatenate them with the output Cls and E; [Cls i , ○, E i = L i ([Cls i-1 , P i-1 , E i-1 )#(2) Among them, ○ indicates not being used as the input to the next Transformer layer; this design allows learnable tokens to be introduced at each layer of the image encoder, enriching the model's ability to capture and represent important image features and quality characteristics, and ultimately contributing to more effective and accurate image quality assessment; Meanwhile, to enable the model to better adjust its representation space in the image quality assessment task and more comprehensively understand the quality-related information in the image, learnable image prompts are also introduced in the image branch; 3) The training of the model is divided into two stages: In the first model training stage, only the text prompts and image prompts are trained, and other parameters are frozen. With extremely few training parameters, the migrated CLIP model is made to have the ability to perceive and evaluate image quality; The loss function used is: Among them, i represents the i-th image, T and V respectively represent the output features of the text encoder and the image encoder, sim(,) represents calculating the cosine similarity of two features, B represents a minibatch, and A represents the set of all images in the minibatch that belong to the same category as image i; In the second model training stage, the prompts of the two branches are frozen, and only the image encoder is trained to enable the model to gradually adapt from rough perception to detailed evaluation, thus achieving a comprehensive understanding and accurate prediction of image quality; The loss function used in this stage is: Introduce fidelity loss to consider pairwise learning for ranking model estimates; in addition, use smooth L1 loss Lsmo and cross-entropy loss with label smoothing for optimization; where α and β are coefficients for balancing and ; Among them, p and p' respectively represent the true probability distribution and the predicted probability distribution; where, I' k = (1 - ε)I k + ε / C represents a value in the mass level target distribution, and P k represents the predicted logits for class k.

2. The unsupervised image quality assessment method based on multi-modal prompt learning according to claim 1, wherein In step 1), when classifying each image in the dataset according to its score and assigning a level or category label, if an image with a score in the range of [49 - 50) is classified as level [50], it corresponds to category [50]; The model calculates the similarity based on the output features of the text encoder and the input features of the image encoder, and obtains the final quality score q(x) in the way of weighted summation: Among them, P(c|x) represents the class probability after applying softmax, and C represents the total number of classes.