A no-reference image quality assessment method based on text-image pairs
By adopting a two-stage training model based on text-image pairs in reference-free image quality evaluation, and introducing learnable text marking and quality perception modules, the existing methods have solved the performance and labeling cost problems, and efficient image quality evaluation has been achieved.
Patent Information
- Application Number
- CN202311131204.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-04
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-09-04
AI Technical Summary
The existing reference-free image quality evaluation methods have bottlenecks in performance, and require high manual labeling costs, making it difficult to capture the perceived features of complex images.
A reference-free image quality evaluation method based on text-image pairs is proposed, using a two-stage training model, introducing learnable text marking and quality perception modules, and achieving fine-grained quality-level hierarchy through multi-view feature extraction and fusion.
The performance and generalization capabilities of image quality evaluation are improved, the performance bottlenecks and manual labeling costs are solved, and the accurate evaluation of image quality is achieved.
Smart Images

Figure CN117115123B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a no-reference image quality assessment method based on text-image pairs. Background Art
[0002] Digital images are increasingly widely used in fields such as medical imaging, multimedia communication, and computer vision. Therefore, accurately assessing image quality has become crucial. Image Quality Assessment (IQA) aims to quantify the quality of an image objectively by simulating human visual perception. In many applications, especially those involving image transmission, image compression, and image processing, accurate assessment of image quality is essential for ensuring the accuracy and fidelity of visual perception.
[0003] Over the years, a variety of IQA methods have been developed and evaluated. These methods are generally divided into three categories according to the amount of information provided by the original reference image: Full Reference-IQA (FR-IQA), Reduced Reference-IQA (RR-IQA), and No Reference-IQA (NR-IQA), and no reference is also called Blind IQA (BIQA). Full-reference image quality assessment can refer to the original undistorted image and obtain the image quality score of the distorted image based on the difference between the distorted image and the original image. Reduced-reference image quality assessment uses partial information of the original image as a reference to predict the image quality. However, in practical applications, it is difficult to obtain a reference image, making the above two methods inapplicable. Therefore, recent work has gradually focused on the BIQA field.
[0004] Traditional BIQA methods usually rely on manually designed features, such as Natural Scene Statistics (NSS) and shallow feature learning based on codebooks. However, these methods often have difficulty capturing the perceptual features of complex images and require manual design and adjustment of feature extractors, which limits their generalization ability and applicability.
[0005] In recent years, with the development of deep learning technology, methods based on deep neural networks have made significant progress in the field of image quality assessment. These methods utilize the powerful expressive ability of deep neural networks to automatically learn and extract complex perceptual features in images, thereby achieving accurate assessment of image quality. Through deep learning methods, the limitations of traditional methods that require manually designed features can be overcome, and end-to-end training can be carried out on large-scale datasets, thus improving the performance and generalization ability of image quality assessment.
[0006] OpenAI proposed the CLIP model (Radford, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning (PMLR), pages 8748-8763, 2021.), which is a self-supervised pre-training framework that encodes images and text into a shared vector space without labeled data. CLIP demonstrates strong zero-shot transfer learning capabilities on various tasks. Wang et al. (Wang, et al. Exploring clip for assessing the look and feel of images. arXiv preprint arXiv:2207.12396 (2022)) first explored the potential of CLIP in challenging image perception assessment tasks. They proposed a prompt pairing strategy using antonym prompts (e.g., "good photo" and "bad photo") to reduce prompt ambiguity. On this basis, Zhang et al. (Zhang, et al. Blind image quality assessment via vision-language correspondence: A multitask learning perspective. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14071-14081, 2023.) proposed multitask learning to fine-tune CLIP by combining image quality assessment with distortion type classification and scene classification tasks. This method automatically determines the sharing of model parameters and loss weights, leveraging auxiliary knowledge from other tasks. They used a Likert scale with five quality levels, and the model output scores the quality level as a logical weighted sum. However, this multitask training method increases annotation costs, and there are performance bottlenecks in the coarse division of quality levels. Summary of the Invention
[0007] The objective of the present invention is to provide a reference-free image quality assessment method based on text-image pairs, giving full play to the potential of the CLIP model in challenging image perception assessment tasks. The present invention proposes a fine-grained quality level stratification strategy, making the learned features more closely related to image quality. A two-stage training model is proposed. In the model, a set of learnable text tokens are introduced to make full use of the representation ability of the text encoder. Meanwhile, a quality perception module is proposed to evaluate image quality from multiple perspectives and extract deep features closely related to the quality level.
[0008] The present invention includes the following steps:
[0009] 1) Grade each picture in the dataset according to its score and assign a level (or category) label. For example, an image with a score in the range of [49 - 50) is classified as level
[50] , corresponding to category
[50] .
[0010] 2) In the first model training stage, a set of learnable text tokens are introduced to represent fine-grained image quality levels, making full use of the representation ability of the text encoder. In this stage, the model is trained for the pairing ability of the quality score text and the image.
[0011] 3) In the second model training stage, a quality-perceivable module is introduced to fuse image features extracted from multiple perspectives. The model will calculate the similarity based on the output features of the text encoder and the input features of the image encoder, and obtain the final quality score in a weighted summation manner. In this stage, the model is trained for the ability to accurately predict the picture quality score.
[0012] Features and effects of the present invention:
[0013] The reference-free image quality assessment method based on text-image pairs proposed by the present invention solves the defects such as performance bottlenecks or the need for other manual annotation costs in previous work, achieving a significant improvement in performance. To solve the performance bottleneck of the Likert scale with five quality levels in previous work, a fine-grained quality level grading method is proposed. And, a set of learnable text tokens are introduced to represent the quality level labels of pictures. To extract image features closer to image quality, a quality perception module is introduced after the image encoder to extract multiple features from multiple perspectives and perform feature fusion. Brief Description of the Drawings
[0014] Figure 1 is the framework diagram of the present invention.
[0015] Figure 2 is the quality perception module diagram of the present invention. Detailed Description of the Invention
[0016] The present invention proposes a no-reference image quality assessment method based on text-image pairs. The present invention will be described in detail below with reference to the accompanying drawings.
[0017] Figure 1 The framework diagram of the present invention is given; the model includes two stages. In the first model training stage, a set of learnable text tokens is introduced to represent fine-grained image quality levels, so as to make full use of the representation ability of the text encoder. In this stage, the model's pairing ability for quality score text and images is trained. In the second model training stage, a quality-aware module is introduced to fuse image features extracted from multiple perspectives. The model will calculate the similarity based on the output features of the text encoder and the input features of the image encoder, and obtain the final quality score in a weighted summation manner. In this stage, the model's ability to accurately predict the picture quality score is trained.
[0018] 1. Training instructions
[0019] The embodiments of the present invention include the following steps:
[0020] 1) Each picture in the dataset is graded according to its score and a level (or category) label is assigned. For example, an image with a score in the range of [49 - 50) is classified into level
[50] , corresponding to category
[50] .
[0021] 2) In the first model training stage, a set of learnable text tokens is introduced to represent fine-grained image quality levels, so as to make full use of the representation ability of the text encoder. Specifically, the form of the text input is designed as follows: "A photo with a quality score of [X]1[X]2[X]3...[X] M ." "[X]1[X]2[X]3...[X] M " is used to represent the quality level, where M represents the number of learnable tokens. In this stage, the model's pairing ability for quality score text and images is trained. The loss function used is:
[0022]
[0023]
[0024] where i represents the i-th picture, T and V respectively represent the output features of the text encoder and the image encoder, sim(,) represents the cosine similarity of two features, B represents a minibatch, and A represents the set of all pictures in the minibatch that belong to the same category as picture i.
[0025] 3) In the second model training stage, a quality-aware module is introduced (such as Figure 2As shown in the figure, it fuses the image features extracted from multiple perspectives. After using the transformer decoder and a set of learnable parameters to reinterpret the features of the image encoder and obtaining the multi-perspective features, a feature fusion block is used for fusion.
[0026] (1) The image encoder generates a feature representation for image i, denoted as where P is the number of patches. Subsequently, Z i is sliced to obtain
[0027] (2) Let L be the number of views, and create an attention panel embedding In addition, the CLS token is extended L times to obtain
[0028] (3) Take (J + T) and Z' as inputs and input them into the transformer decoder to obtain multiple feature embedding outputs Then, input S into the feature fusion module to obtain the final multi-view image feature V i . Specifically, the L features are averaged, denoted as:
[0029]
[0030] The model will calculate the similarity based on the output features of the text encoder and the input features of the image encoder, and obtain the final quality score q(x) in the way of weighted summation.
[0031]
[0032] Among them, P(c|x) represents the class probability after applying softmax, and C represents the total number of classes. At this stage, the ability of the model to accurately predict the picture quality score is trained. The loss function used in this stage is:
[0033]
[0034] Among them, α and β are coefficients for balancing and .
[0035]
[0036] p and p' respectively represent the true probability distribution and the predicted probability distribution.
[0037]
[0038] I' k =(1 - ε)I k+ε / C represents the value in the quality level target distribution, P k represents the predicted logits for class k.
[0039] 2. Implementation details
[0040] 1) Model details
[0041] The present invention is implemented using the Pytorch framework. The basis of the image and text feature extractors are the image encoder V and text encoder T from CLIP respectively. Specifically, the ViT-B / 16 architecture is adopted as the image encoder, which consists of 12 Transformer layers with a hidden size of 768 dimensions. To be consistent with the output of T, the dimension of the image feature vector is reduced from 768 to 512 through a linear layer. In the quality perception module of the present invention, a single-layer Transformer decoder is adopted.
[0042] 2) Training details
[0043] In the first training stage, the Adam optimizer is used with an initial learning rate of 3×10 -5 , and cosine scheduling is used for decay. The training process includes a warm-up stage of 5 epochs, during which random cropping is applied as data augmentation. Each original image is randomly cropped into 8 images of size 224×224. In this stage, only the learnable token "[X]1[X]2[X]3...[X] M " is optimized. In the second training stage, the Adam optimizer is also used, but the warm-up stage is 10 epochs. The learning rate linearly increases from 9.5×10 -7 to 5×10 -6 . At the 30th and 50th epochs, the learning rate is reduced by 0.1 times. In this stage, random horizontal flipping and random cropping are applied as data augmentation. The coefficients α is set to 0.001 and β is set to 0.1 for Both training stages include 60 epochs.
[0044] 3. Application fields
[0045] The present invention can be applied to the field of no-reference image quality assessment to realize the judgment of the quality of distorted pictures without the original image.
[0046] Table 1 shows the performance comparison between the model of the present invention and previous models on six common image quality assessment datasets. These datasets include four real datasets and two synthetic datasets. The real datasets used include LIVE, CSIQ, TID2013, and KADID. The synthetic datasets include LIVEC and KonIQ. To evaluate the performance of the model of the present invention, the Pearson linear correlation coefficient (PLCC) and the Spearman rank correlation coefficient (SRCC) are used as evaluation metrics. PLCC measures the accuracy of model prediction, while SRCC evaluates the monotonicity of the BIQA algorithm prediction. The value ranges of both metrics are from 0 to 1, and a higher value indicates better performance in terms of prediction accuracy and monotonicity.
[0047] As can be seen from Table 1, on all datasets, the model of the present invention demonstrates competitive performance compared to the state-of-the-art methods. It is worth noting that state-of-the-art performance is achieved on the CSIQ, TID2013, and KADID datasets, surpassing the existing methods by increasing the PLCC metric by 0.7%, 3.6%, and 5.0% respectively, and the SRCC metric by 1.2%, 4.7%, and 4.8% respectively. These results highlight the effectiveness and leading performance of the model of the present invention in image quality assessment.
[0048] Table 1
[0049]
[0050] Table 2
[0051]
[0052] Table 2 shows the generalization performance comparison between the model of the present invention and previous models on cross-datasets. Specifically, a BIQA model is trained on one dataset and directly evaluated on another dataset without fine-tuning or parameter adjustment. Four datasets are utilized and the median experimental results are reported. On the KonIQ and CSIQ datasets, the model of the present invention demonstrates superior performance and remains competitive on other datasets. These experimental results highlight the robust generalization performance of the model.
[0053] The above embodiments are only preferred embodiments of the present invention and should not be considered as limiting the scope of implementation of the present invention. Any equivalent changes and improvements made within the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.
Claims
1. A no-reference image quality assessment method based on text-image pairs, characterized in that Including the following steps: 1) Classify each image in the dataset according to its score and assign a level or category label; 2) In the first model training stage, introduce a set of learnable text tokens to represent fine-grained image quality levels, so as to make full use of the representation ability of the text encoder; in this stage, train the model's pairing ability for quality score texts and images; The loss function used is: where i represents the i-th image, T and V respectively represent the output features of the text encoder and the image encoder, sim(,) represents the cosine similarity of two features, B represents a minibatch, and A represents the set of all images in the minibatch that belong to the same category as image i; 3) In the second model training stage, introduce a quality-aware module to fuse image features extracted from multiple perspectives; the model will calculate the similarity based on the output features of the text encoder and the input features of the image encoder, and obtain the final quality score in a weighted summation manner; in this stage, train the model's ability to accurately predict the image quality score; (1) The image encoder generates a feature representation for image i, denoted as where P is the number of patches; subsequently, Z i is sliced to obtain (2) Let L be the number of views, and create an attention panel embedding In addition, expand the CLS token L times to obtain (3) Use (J+T) and Z' as inputs and feed them into the transformer decoder to obtain multiple feature embedding outputs Then, input S into the feature fusion module to obtain the final multi-view image feature V i ; Average the L features, expressed as: The model will calculate the similarity based on the output features of the text encoder and the input features of the image encoder, and obtain the final quality score q(x) in a weighted summation manner; where P(c|x) represents the class probability after applying softmax, and C represents the total number of classes; in this stage, train the model's ability to accurately predict the image quality score; the loss function used in this stage is: where α and β are coefficients for balance and ; where p and p' respectively represent the true probability distribution and the predicted probability distribution; where, I' k = (1 - ε)I k + ε / C represents a value in the target distribution of the quality level, and P k represents the predicted logits for class k.
2. The method for no-reference image quality assessment based on text-image pairs according to claim 1, wherein, It is characterized in that In step 1), when classifying each image in the dataset according to its score and assigning a level or category label, if the score of an image is in the range of [49 - 50), the image is classified as level [50], corresponding to category [50].