Semi-supervised no-reference image quality assessment method based on distillation learning and incremental learning

By adopting a semi-supervised approach of distillation learning and incremental learning in reference-free image quality assessment, the problem of insufficient data is solved, model performance is improved, and catastrophic forgetting is prevented, and more accurate image quality assessment is achieved.

CN117115121BActive Publication Date: 2025-05-13XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311121777.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-01
Publication Date
2025-05-13
Estimated Expiration
2043-09-01

AI Technical Summary

Technical Problem

The existing reference-free image quality evaluation method has poor performance in the case of insufficient data and is difficult to effectively solve the problem of insufficient data.

Method used

A semi-supervised method based on distillation learning and incremental learning is adopted to assign pseudo-labels to unlabeled data through a knowledge distillation algorithm, and an incremental learning algorithm is used to prevent catastrophic forgetting, ensuring the performance of the model in the case of insufficient data.

Benefits of technology

It effectively solves the problem of insufficient data in reference-free image quality evaluation, improves the model's performance on unlabeled data, and prevents catastrophic forgetting, ensuring the accuracy of image quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117115121B_ABST
    Figure CN117115121B_ABST
Patent Text Reader

Abstract

A semi-supervised no-reference image quality assessment method based on distillation learning and incremental learning belongs to the field of computer vision technology. No-reference image quality assessment aims to simulate human evaluation of image distortion and has a great demand for labeled data, but labeled data is far from enough in practice. The present invention proposes a unified semi-supervised and incremental learning framework to solve the above problems. When the training data is insufficient, semi-supervised learning is required to infer a large amount of unlabeled data. At the same time, multiple semi-supervised learning can easily lead to catastrophic forgetting problems, requiring incremental learning. Knowledge distillation is used to provide pseudo labels for unlabeled data to maintain analytical capabilities, thereby achieving semi-supervised learning. At the same time, by selecting representative examples in multiple semi-supervised learning processes, incremental learning is used to correct previous data, thereby ensuring that the model of the present invention will not degenerate. The present invention demonstrates its potential to solve image quality assessment problems in actual production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a semi-supervised reference-free image quality assessment method based on distillation learning and incremental learning. Background Art

[0002] At present, the images and videos involved in multimedia services are all digital products. In the process of image acquisition, compression, transmission and decompression, various types of distortion will inevitably be introduced due to the limitations of physical conditions such as storage and transmission, as well as the influence of processing algorithms, resulting in a decrease in image quality. However, in practical applications, people usually hope that the image quality is as high as possible. For example, in the medical field, higher quality images will mean higher disease diagnosis rates and treatment effects; while in the aerospace field, higher image quality will mean more valuable exploration and discovery. Therefore, establishing an effective method for evaluating image quality has become a research hotspot in the field of digital image processing. According to whether the reference image is complete and available, objective image quality assessment can be divided into three categories: full reference (FR-IQA), semi-reference (RR-IQA) and no reference (NR-IQA). Full reference image quality assessment can refer to the original undistorted image and obtain the image quality score of the distorted image based on the difference between the distorted image and the original image. Semi-reference image quality assessment uses part of the information of the original image as a reference to predict image quality. However, in practical applications, reference images are difficult to obtain, making the above two methods inapplicable. Therefore, no-reference image quality assessment without reference information is the most studied but also the most difficult to design objective image quality assessment method.

[0003] Image quality assessment (IQA) models require a lot of manpower for data annotation due to the task setting. However, due to the different subjective evaluation systems between observers and the different settings of the environment in each experiment, it is difficult to directly make up for the lack of data in practice when there is a problem of insufficient data. To solve this problem, some researchers believe that the NR-IQA task is an unsupervised problem. Some researchers evaluate image quality by selecting a statistical feature called the local pattern statistical index extracted from the binary pattern of the local image structure. Some researchers quantify image quality degradation by measuring the changes in structure, naturalness, and perceptual quality of distorted images compared to original natural images. But since the No Reference Image Quality Assessment (NR-IQA) is intended to simulate a subjective human system, the performance of unsupervised methods is far from satisfactory.

[0004] With the development of deep learning, many researchers fine-tune pre-trained models to solve the problem of insufficient data. For example, CNN-based NR-IQA methods directly use or fine-tune pre-trained CNN classification models as feature extractors to further predict image quality scores. Similarly, there are also methods that use Transformer as the backbone for feature extraction. However, this pipeline not only has high requirements on the correlation between upstream tasks and NR-IQA tasks, but also cannot fundamentally solve the problem of insufficient data. Summary of the invention

[0005] The purpose of this invention is to propose a semi-supervised no-reference image quality assessment method based on distillation learning and incremental learning, which decomposes the problem of insufficient data in image quality assessment into semi-supervised learning and catastrophic forgetting problems. A unified semi-supervised and incremental learning framework is proposed to solve the above problems. In order to simulate this problem, two datasets are used. Dataset D A Contains all labeled data, and dataset D B Contains all unlabeled data. The task to be solved by the present invention is to transform the dataset D A The knowledge is transferred to the dataset D B , that is, D A ->D B First, in the dataset D A The teacher model is trained on the dataset D. B Set pseudo labels, which are used to transform the dataset D A The manifold covers the dataset D B To ensure that the model is B and in D A Finally, a replay-based incremental learning algorithm is utilized during each retraining subtask to preserve analytical power by revisiting representative examples from previous subtasks to prevent catastrophic forgetting.

[0006] The present invention comprises the following steps:

[0007] 1) Use a labeled dataset D A Initialize the model;

[0008] 2) The initial model is copied into a teacher model and a student model;

[0009] 3) For the unlabeled dataset D B Block

[0010] 4) Using the knowledge distillation algorithm based on kernel ridge regression (KRR) for the teacher model to extract the unlabeled dataset Generate pseudo labels;

[0011] 5) Based on the teacher model A and Perform simple sample selection and select the K samples that are most accurately predicted by the teacher model as representative samples ε A and

[0012] 6) ε A and Mixing, training the student model;

[0013] 7) Repeat steps 4) to 6) until the student model is complete. B training;

[0014] 8) Given any image, input it into the model and the model outputs its predicted quality score.

[0015] Features and effects of the present invention:

[0016] The semi-supervised no-reference image quality assessment method based on distillation learning and incremental learning proposed in this paper solves the problem of insufficient data in NR-IQA. The problem of insufficient data is divided into the semi-supervised learning problem and the catastrophic forgetting problem. In order to solve the semi-supervised learning problem, a new knowledge distillation algorithm based on kernel ridge regression (KRR) is proposed, which is used to assign pseudo labels to unlabeled data. This enables the present invention to transform the manifold from the dataset D A Transfer to dataset D B To ensure that the proposed model is B To solve the catastrophic forgetting problem, this paper proposes a new replay-based incremental learning algorithm to prevent performance degradation. By revisiting representative examples from previous semi-supervised learning, the analytical ability of the model is maintained. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a framework diagram of the present invention.

[0018] Figure 2 Comparison of the amount of labeled data and performance between the present invention and the supervised method (HyperIQA) in the LIVE->CSIQ task.

[0019] Figure 3 Comparison of the visualization and prediction results of the present invention and the benchmark model on some images. DETAILED DESCRIPTION

[0020] The present invention proposes a semi-supervised no-reference image quality assessment method based on distillation learning and incremental learning, which is described in detail below with reference to the accompanying drawings.

[0021] The present invention decomposes the data shortage problem into semi-supervised learning and catastrophic forgetting problems. Each retraining subtask is regarded as a semi-supervised learning problem and solved by a knowledge distillation algorithm. At the same time, the performance degradation after each retraining is regarded as a catastrophic forgetting problem and is handled by an incremental learning algorithm. The process of the method of the present invention is as follows: Figure 1 As shown in Figure 2, the model consists of two stages. In the first stage, D A Used to initialize the student model and the teacher model, corresponding to algorithm steps 1) to 2). In the second stage, in D B The KRR-Distill module generates pseudo signatures using a knowledge distillation algorithm based on kernel ridge regression. To prevent catastrophic forgetting during incremental learning, the SDK module selects K most accurate samples for each block of data as representative samples for playback. The embodiment of the present invention specifically includes the following steps:

[0022] 1) Use a labeled dataset D A Initialize the model;

[0023] 2) The initial model is copied into a teacher model and a student model;

[0024] 3) For the unlabeled dataset D B Block

[0025] 4) Using the knowledge distillation algorithm based on kernel ridge regression (KRR) for the teacher model to extract the unlabeled dataset Generate pseudo labels;

[0026] 41) The loss function of KRR can be defined as:

[0027]

[0028] where and f i and f z Indicates D A The features of the i-th and z-th samples in a i and a z is the corresponding factor, λ is the balance factor, K is the radial basis function, and N is D A The number of .

[0029] 42) The definition of radial basis function K can be written as:

[0030] K(f i ,f z )=exp(-γ‖f i -f z ‖ 2 ),γ>0 (2)

[0031] Among them, γ defines the influence range of a single sample.

[0032] 43) Use the feature matrix as input and obtain unlabeled data D B 's pseudo-labels.

[0033]

[0034] Among them, y j Yes D B The pseudo label of the jth sample in .

[0035] 5) Based on the teacher model A and Perform simple sample selection and select the K samples that are most accurately predicted by the teacher model as representative samples ε A and

[0036] 51) Teacher model calculation block D A The distance between the sample prediction value and the true value in is:

[0037] Distance=||GT-Predition|| (4)

[0038] 52) Select the K samples with the smallest distance, that is, the K samples with the most accurate prediction as representative samples ε A , review incremental learning;

[0039] 53) Yes Repeat steps 51) and 52) to generate representative samples

[0040] 6) ε A and Mixing, training the student model;

[0041] The ViT-S architecture proposed by DeiT (Touvron H, Cord M, Jégou H. Deit iii: Revenge of the vit[C] / / Computer Vision-ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part XXIV. Cham: Springer Nature Switzerland, 2022: 516-533.) is adopted and pre-trained on ImageNet-1K for 400 epochs. During the training process, the Adam optimizer is used, and the loss function adopts a smooth L1 loss. In the initial stage, DA The dataset was trained for 9 epochs, including 3 warm-up epochs with a learning rate of 2×10 -4 , and decreases by 0.1 every 3 epochs. In the incremental phase, on the dataset D B The training was performed on blocks of 6 epochs with an initial learning rate of 2×10 -5 , and decreases by 0.1 every 2 periods. Dataset D B Divided into 3 blocks. For data set D A and D B , 100 samples are selected as representative old samples.

[0042] 7) Repeat steps 4) to 6) until the student model is complete. B training;

[0043] 8) Given any image, input it into the model and the model outputs its predicted quality score.

[0044] No-reference image quality assessment aims to simulate human assessment of image distortion. Therefore, it has a great demand for labeled data, but labeled data is far from enough in practice. To this end, the present invention proposes a unified semi-supervised and incremental learning framework to solve the above problems. When the training data is insufficient, semi-supervised learning is required to infer a large amount of unlabeled data. At the same time, multiple semi-supervised learning can easily lead to catastrophic forgetting problems, requiring incremental learning. More specifically, knowledge distillation is used to provide pseudo labels for unlabeled data to maintain analytical capabilities, thereby achieving semi-supervised learning. At the same time, by selecting representative examples in multiple semi-supervised learning processes, incremental learning is used to correct previous data, thereby ensuring that the model of the present invention does not degenerate. In summary, the method proposed in the present invention demonstrates its potential to solve image quality assessment problems in actual production.

[0045] Figure 2 This is a comparison of the amount of labeled data and performance of the present invention and the supervised method (HyperIQA) in the LIVE->CSIQ task. As shown in the above part, HyperIQA A Pre-train on D B The method of the present invention includes fine-tuning D A Pre-training is performed and D is directly trained with pseudo labels B Therefore, D A The performance of HyperIQA is maintained and exceeds that of HyperIQA in D B This shows that the method of the present invention achieves better performance when there is less labeled data.

[0046] Figure 3Comparison of the visualization and prediction results of the present invention and the baseline model on some images. The present invention's model pays more attention to the features related to image distortion after each distillation and increment, and the predicted image quality score is closer to the true value. The numbers under each row of images represent the predicted values ​​of the model, and the numbers in brackets represent the distance from the true value.

Claims

1. A semi-supervised no-reference image quality assessment method based on distillation learning and incremental learning, characterized by The following steps are involved: 1) Use a labeled dataset D A Initialize the model; 2) The initial model is copied into a teacher model and a student model; 3) For the unlabeled dataset D B Block 4) Use the knowledge distillation algorithm based on kernel ridge regression for the teacher model to extract the unlabeled dataset Generate pseudo labels; The knowledge distillation algorithm based on kernel ridge regression is used for unlabeled datasets. Generate pseudo labels. The specific steps are as follows: 41) The loss function of KRR is defined as: Among them, f i and f z Indicates D A The features of the i-th and z-th samples in a i and a z is the corresponding factor, λ is the balance factor, K is the radial basis function, and N is D A The number of 42) The definition of the radial basis function K is written as: K(f i ,f z )=exp(-γ||f i -f z || 2 ),γ>0 (2) Among them, γ defines the influence range of a single sample; 43) Use the feature matrix as input and obtain unlabeled data D B Pseudo labels of Among them, y j Yes D B The pseudo label of the jth sample in ; 5) Based on the teacher model A and Perform simple sample selection and select the K samples that are most accurately predicted by the teacher model as representative samples ε A and 6) ε A and Mixing, training the student model; 7) Repeat steps 4) to 6) until the student model is complete. B training; 8) Given any image, input it into the model and the model outputs its predicted quality score.

2. The semi-supervised no-reference image quality assessment method based on distillation learning and incremental learning as claimed in claim 1, characterized in that In step 5), the K samples most accurately predicted by the teacher model are selected, and the specific steps are as follows: 51) Teacher model calculation block D A The distance between the sample prediction value and the true value in is: Distance=||GT-Predition|| (4) 52) Select the K samples with the smallest distance, that is, the K samples with the most accurate prediction as representative samples ε A , review incremental learning; 53) Yes Repeat steps 51) and 52) to generate representative samples