Multi-token semi-supervised no-reference image quality assessment method based on cross-dataset distillation

Through cross-data set knowledge distillation and multi-token semi-supervised methods, the lack of reference-free image quality evaluation in the case of data scarcity is solved, the analysis and generalization capabilities of the model are enhanced, and more accurate and reliable image quality evaluation is achieved.

CN117173518BActive Publication Date: 2025-05-16XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311131030.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-04
Publication Date
2025-05-16
Estimated Expiration
2043-09-04

AI Technical Summary

Technical Problem

Existing reference-free image quality assessment methods are difficult to obtain satisfactory results when data is scarce, and insufficient knowledge fusion between students and teachers' models hinders the acquisition of key quality feature representations.

Method used

Using a multi-token semi-supervised approach to knowledge distillation across datasets, the analytical capabilities of the Vision Transformer architecture are enhanced by introducing distillation tokens and multi-class tokens, and an attention scoring mechanism is designed to alleviate the uncertainty of scoring.

Benefits of technology

Through knowledge distillation across datasets, the model is able to capture key information related to image quality perception, enhance generalization capabilities, reduce prediction uncertainty, and demonstrate excellent performance in actual production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117173518B_ABST
    Figure CN117173518B_ABST
Patent Text Reader

Abstract

Semi-supervised no-reference image quality assessment method based on multi-token distillation across datasets, involving computer vision technology. A NR-IQA method based on attention distillation is proposed. Effectively integrate knowledge from different datasets to enhance the representation of image quality and improve the accuracy of prediction. A distillation token is introduced in the Transformer encoder to enable the student model to learn from the teacher on different datasets. By leveraging knowledge from different source domains, the model is able to capture the basic features related to image distortion and enhance the generalization ability of the model. In order to refine the perceptual information from different perspectives, multiple class tokens simulating multiple reviewers are introduced. Improve the interpretability of the model and reduce the uncertainty of the prediction. A mechanism called attention scoring is introduced, which combines the attention score matrix from the encoder with the MLP head after the decoder to refine the final quality score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a multi-token semi-supervised reference-free image quality assessment method based on cross-dataset distillation. Background Art

[0002] In the modern digital age, images have become an integral part of our lives. We capture, share, and browse images through a variety of devices and platforms, whether sharing photos on social media, displaying product images in e-commerce, or using images in medical diagnosis and scientific research. Image quality is an important indicator for evaluating the visual performance and authenticity of images. High-quality images can provide a clear, accurate, and natural visual experience, while low-quality images may cause problems such as blur, noise, and distortion, reducing the usability and understandability of the image. The applications of image quality assessment are wide and diverse. In the field of image processing, quality assessment can be used to evaluate and optimize the performance of algorithms such as image enhancement, denoising, and restoration. In the field of image transmission and storage, quality assessment can help select appropriate compression algorithms and parameters to maximize the retention of image quality. In fields such as medical images and satellite images, quality assessment is essential to ensure accurate diagnosis and scientific research. In order to ensure the quality of images, researchers and engineers have been committed to developing methods and technologies for image quality evaluation. These evaluation methods are designed to measure the performance of image clarity, contrast, color accuracy, detail retention, etc., and provide reliable indicators to judge the quality of images.

[0003] No-reference image quality assessment (NR-IQA) is a branch of image quality assessment that does not rely on reference images or any prior knowledge, but instead performs assessment based on the characteristics of the image itself. With the development of deep learning and computer vision, no-reference image quality assessment has also ushered in new challenges and opportunities. Many methods based on machine learning and neural networks have been proposed to automatically and intelligently assess image quality. These methods are able to learn and understand the characteristics of images and compare them with human subjective evaluations, thereby achieving more accurate and reliable image quality assessments.

[0004] The most advanced methods in this field currently utilize pre-trained upstream backbones to extract semantic features and are subsequently fine-tuned on the NR-IQA dataset. However, the scarcity of IQA datasets poses a challenge as simple fine-tuning often produces unsatisfactory results. Therefore, many researchers are committed to making full use of the limited information available in IQA datasets. For example, DR-IQA (Zheng H, Yang H, Fu J, et al. Learning conditional knowledge distillation for degraded-reference image quality assessment [C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021: 10242-10251.) extracts reference information from degraded images by extracting knowledge from original quality images, thereby being able to capture deep image priors useful for quality assessment. CVRKD-IQA (Yin G, Wang W, Yuan Z, et al. Content-variant reference image quality assessment via knowledge distillation [C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2022, 36 (3): 3134-3142.) combines non-IQA datasets as reference images to expand the dataset, and uses knowledge distillation to transfer the distribution differences between various distorted images.

[0005] We believe that extracting knowledge only from the original dataset or other task datasets is not enough to obtain feature representations of critical quality, which also hinders the knowledge fusion between the student and the teacher. To overcome this challenge, we propose a new method that exploits knowledge distillation between datasets to obtain more essential quality representations. Specifically, we enhance the Vision Transformer architecture with a distillation token, which serves a similar purpose as the class token, but focuses on learning pseudo labels simulated by the teacher. It is worth noting that the student and teacher models are associated with different datasets. Through distillation, the student can obtain cross-dataset knowledge from the teacher and further enhance the analysis ability of the model through cross-dataset knowledge fusion. In addition, the present invention introduces multi-class tokens and a self-attention-based attention scoring mechanism to alleviate the uncertainty of scoring. Summary of the invention

[0006] The object of the present invention is to provide a method for semi-supervised no-reference image quality assessment based on multi-token distillation across datasets, which enhances the representation ability of image quality through knowledge distillation across datasets. The present invention introduces an additional distillation token to promote students to learn from teachers and realize knowledge distillation across datasets. The present invention regards class tokens as abstractions of quality-aware features and simulates multiple reviews by increasing the number of class tokens. This method helps to reduce the uncertainty of prediction. The present invention designs an attention scoring mechanism to refine the output of each class token. The method of the present invention demonstrates the potential to solve the image quality assessment problem in actual production.

[0007] The present invention comprises the following steps:

[0008] 1) Cut the input image into N patches;

[0009] 2) N patches are transformed into N patch tokens X = {x0, x1, ..., x n}∈R N×D ;

[0010] 3) Introduce M additional learnable Class Tokens C 0 ={c0,c1,...,c n}∈R M×D and 1 Distillation Token D 0 ={d0}∈R 1×D ;

[0011] 4) Concatenate all Tokens T = {C 0 ,D 0 ,X} and sent to Transformer Encoder and Decoder;

[0012] 5) The output of Class Tokens and the Attention Matrix obtained by the self-attention mechanism of Transformer Encoder are sent to Attention Scoring to calculate the score of Class Tokens;

[0013] 6) Feed the output of DistillationToken into MLP to calculate the score of DistillationTokens;

[0014] 7) Input the image into the teacher model across datasets to obtain the image quality score Y T ;

[0015] 8) Calculate L1 loss and distillation loss using the scores obtained in steps 5), 6) and 7) respectively;

[0016] 9) Given any image, input it into the model, and weighted sum the scores of Class Tokens and DistillationTokens to get its predicted final quality score Y final .

[0017] Features and effects of the present invention:

[0018] The present invention proposes a multi-token semi-supervised no-reference image quality assessment method based on cross-dataset distillation. The present invention introduces distillation tokens in the NR-IQA task for cross-dataset knowledge distillation. By integrating knowledge from different datasets, the model of the present invention can capture key information related to image quality perception. This distillation further enhances the generalization ability of the model. The present invention provides a new attention scoring mechanism, which combines the attention matrix obtained by the self-attention of the decoder with the MLP head. The present invention introduces multiple class tokens to simulate different image quality judgments, and introduces a new attention scoring mechanism to generate quality predictions. Multiple class labels reduce the randomness of the prediction, and the attention score deeply explores the relationship between the class label and the image, as well as the potential distortion information. The method of the present invention shows advanced performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a framework diagram of the present invention.

[0020] Figure 2 Comparison of the visualization and prediction results of the present invention and the benchmark model on some images. DETAILED DESCRIPTION

[0021] The present invention proposes a multi-token semi-supervised no-reference image quality assessment method based on cross-dataset distillation, which is described in detail below with reference to the accompanying drawings:

[0022] The process of the method of the present invention is as follows Figure 1 As shown, the embodiment of the present invention first cuts the image into N Patches, and projects the Patches into N tokens, and then introduces M Class Tokens and 1 Distillation Token, splices them and sends them to Transformer. The output of Class Tokens is subjected to an Attention Scoring to obtain the prediction score of ClassTokens, and the L1 loss is calculated with GroundTruth; the output of Distillation Tokens is subjected to MLP to obtain the prediction score of Distillation Tokens, and the distillation loss is calculated with the prediction result of the Teacher model across data sets.

[0023] The embodiment of the present invention specifically includes the following steps:

[0024] 1) Cut the input image into N patches;

[0025] 2) N patches are transformed into N patch tokens X = {x0, x1, ..., x n}∈R N×D ;

[0026] 3) Introduce M additional learnable Class Tokens C 0 ={c0,c1,...,c n}∈R M×D and 1 Distillation Token D 0 ={d0}∈R 1×D ;

[0027] 4) Concatenate all Tokens T = {C 0 ,D 0 ,X} and sent to TransformerEncoder and Decoder;

[0028] 41) All tokens are projected through three different layers to obtain Q, K, V∈R (M+N+1)×D , and then sent to Transformer Encoder to obtain the output feature F after Multi-Head Self-Attention (MHSA) and MLP o :

[0029] F M =MHSA(Q,K,V)+T (1)

[0030] F O =MLP(Norm(F M ))+F M (2)

[0031] where F o ={F o [0],F o [1],...,F o [M+N]}∈R (M+N+1)×D .

[0032] 42) The output of the Encoder is sent to the Decoder, after Multi-Head Self-Attention (MHSA) and Multi-Head Cross-Attention (MHCA):

[0033] Q d =MHSA(Norm(C 1 ,D 1 )+(C 1 +D 1 )) (3)

[0034] C 2 ,D 2 =MHCA(Norm(Q d ),K d ,V d )+Q d (4)

[0035] Among them, C 1 ={F o [0],F o [1],...,F o [M-1]}∈R M×D , D 1 ={F o [M]}∈R 1×D , respectively C 0 and D 0 The output, K d =V d ={F o [M+1],F o [M+2],...,F o [M+N]}∈R N×D .

[0036] 5) Send the output of ClassTokens and the AttentionMatrix obtained by the self-attention mechanism of Transformer Encoder to Attention Scoring to calculate the score of ClassTokens;

[0037] 51) From the global attention matrix Extract a token-to-token attention matrix A c2p =A t2t [0:M-1,M+1:M+N]∈R M×N ;

[0038] 52) Extract the attention map from the last K layers of the Transformer encoder and calculate the sum A c2i :

[0039]

[0040]

[0041] 53) Using A c2iand C 2 Calculate the final score of Class Tokens:

[0042]

[0043] 6) Feed the output of Distillation Token into MLP to calculate the score of Distillation Tokens;

[0044] 61) Using MLP and D 2 Calculate the final score of Distillation Token:

[0045] Y distillation_token =MLP(D 2 ) (8)

[0046] 7) Input the image into the teacher model across datasets to obtain the image quality score Y T ;

[0047] 8) Calculate L1 loss and distillation loss using the scores obtained in 5), 6), and 7) respectively;

[0048] Loss global =λ||Y class_token -Y||+(1-λ)||Y distillation_token -Y T || (9)

[0049] Among them, Y is the Ground Truth of the image, Y T is the output of the cross-dataset Teacher model for this image, and λ is the balancing factor.

[0050] 9) Given any image, input it into the model, and weighted sum the scores of Class Tokens and DistillationTokens to get its predicted final quality score Y final .

[0051] Y final =βY class_token +(1-β)Y distillation_token (10)

[0052] Among them, β is the balancing factor.

[0053] The present invention proposes a new NR-IQA method based on attention distillation. Effectively integrate knowledge from different datasets to enhance the representation of image quality and improve the accuracy of prediction. A distillation token is introduced in the Transformer encoder to enable the student model to learn from the teacher on different datasets. By leveraging knowledge from different source domains, the model of the present invention is able to capture the essential features related to image distortion and enhance the generalization ability of the model. In addition, in order to refine the perceptual information from different perspectives, multiple class tokens simulating multiple reviewers are introduced. This not only improves the interpretability of the model, but also reduces the uncertainty of the prediction. In addition, the present invention also introduces a mechanism called attention scoring, which combines the attention scoring matrix from the encoder with the MLP head behind the decoder to refine the final quality score.

[0054] Figure 2 The visualization and prediction results of the present invention and the benchmark model on some images are compared. Figure 2 It can be seen that the model of the present invention pays more attention to the features related to image distortion, and the predicted image quality score is closer to the true value. The numbers under each row of images represent the predicted values ​​of the model, and the numbers in brackets represent the distance from the true value. In general, the method of the present invention demonstrates the potential to solve the problem of image quality assessment in actual production.

Claims

1. A multi-token semi-supervised no-reference image quality assessment method based on cross-dataset distillation, characterized by The following steps are involved: 1) Cut the input image into N patches; 2) N patches are transformed into N patch tokens X = {x0, x1, ..., x n }∈R N×D ; 3) Introduce M additional learnable Class Tokens C 0 ={c0,c1,...,c n }∈R M×D and 1 DistillationToken D 0 ={d0}∈R 1×D ; 4) Concatenate all Tokens T = {C 0 ,D 0 ,X} and sent to Transformer Encoder and Decoder. The specific steps are: 41) All tokens are projected through three different layers to obtain Q, K, V∈R (M+N+1)×D , and then sent to TransformerEncoder to obtain the output feature F after multi-head self-attention and MLP o : F M =MHSA(Q,K,V)+T (1) F O =MLP(Norm(F M ))+F M (2) Among them, F o ={F o [0],F o [1],...,F o [M+N]}∈R (M+N+1)×D ; 42) The output of the Encoder is sent to the Decoder, after Multi-Head Self-Attention and Multi-Head Cross-Attention: Q d =MHSA(Norm(C 1 ,D 1 )+(C 1 +D 1 )) (3) C 2 ,D 2 =MHCA(Norm(Q d ),K d ,V d )+Q d (4) Among them, C 1 ={F o [0],F o [1],...,F o [M-1]}∈R M×D , D 1 ={F o [M]}∈R 1×D , respectively C 0 and D 0 The output, K d =V d ={F o [M+1],F o [M+2],...,F o [M+N]}∈R N×D ; 5) Send the output of Class Tokens and the AttentionMatrix obtained by the self-attention mechanism of Transformer Encoder to Attention Scoring to calculate the score of Class Tokens; 6) Feed the output of DistillationToken into MLP to calculate the score of DistillationTokens; 7) Input the image into the teacher model across datasets to obtain the image quality score Y T ; 8) Calculate L1 loss and distillation loss using the scores obtained in steps 5), 6) and 7) respectively; 9) Given any image, input it into the model, and weighted sum the scores of Class Tokens and DistillationTokens to get its predicted final quality score Y final .

2. The method for semi-supervised no-reference image quality assessment based on cross-dataset distillation multi-token as claimed in claim 1, characterized in that In step 5), the output of Class Tokens and the Attention Matrix obtained by the self-attention mechanism of Transformer Encoder are sent to Attention Scoring to calculate the score of Class Tokens. The specific steps are: 51) From the global attention matrix Extract a token-to-token attention matrix A c2p =A t2t [0:M-1,M+1:M+N]∈R M×N ; 52) Extract the attention map from the last K layers of the Transformer encoder and calculate the sum A c2i : 53) Using A c2i and C 2 Calculate the final score of Class Tokens:

3. The method for semi-supervised no-reference image quality assessment based on cross-dataset distillation multi-token as claimed in claim 1, characterized in that In step 6), the output of the Distillation Token is fed into the MLP to calculate the score of the DistillationTokens. The specific steps are: 61) Using MLP and D 2 Calculate the final score of Distillation Token: Y distillation_token =MLP(D 2 ) (8)。 4. The method for semi-supervised no-reference image quality assessment based on cross-dataset distillation multi-token as claimed in claim 1, characterized in that In step 8), the L1 loss and distillation loss are calculated to obtain the total loss: Loss global =λ||And class_token -Y||+(1-λ)||Y distillation_token -AND T || (9) Among them, Y is the Ground Truth of the image, Y T is the output of the cross-dataset Teacher model for this image, and λ is the balancing factor.

5. The method for semi-supervised no-reference image quality assessment based on cross-dataset distillation multi-token as claimed in claim 1, characterized in that In step 9), given any image, it is input into the model, and the weighted sum of the scores of Class Tokens and Distillation Tokens is used to obtain the predicted final quality score Y final ,as follows: AND final =βY class_token +(1-β)Y distillation_token (10) Among them, β is the balancing factor.