A Semi-Supervised Blind Reference Image Quality Assessment Method Based on PU Learning
By adopting a semi-supervised method based on PU learning in the field of blind reference image quality evaluation, the diffusion model is used to generate label-free data and exclude outliers, which solves the problems of insufficient data sets and outlier damage, and significantly improves the performance of the image quality evaluation model.
Patent Information
- Application Number
- CN202311137352.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-05
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-09-05
AI Technical Summary
The field of blind reference image quality evaluation datasets are insufficient, and outliers in unlabeled datasets impair model performance.
The blind reference image quality evaluation semi-supervised method based on PU learning is adopted to generate label-free data through diffusion model to amplify the data set, and PU learning is used to exclude outliers, combined with teacher-student model training, and model performance is improved.
The blind image quality evaluation data set is effectively expanded, the semantic information interference is reduced, and the model's ability to exclude outliers is improved, thereby improving the performance of the image quality evaluation model.
Smart Images

Figure CN117173507B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, relates to the technical field of blind reference image quality assessment in computer vision, and particularly relates to a semi-supervised method for blind reference image quality assessment based on PU learning. Background Art
[0002] Image Quality Assessment (IQA) is an important computer vision task, which is a process of evaluating the perceptual quality of an image according to human subjective perception by using a computational model, aiming to objectively predict the Mean Opinion Scores (MOS) of the image. Currently, images are widely used in fields such as digital photography, medical imaging, video transmission, and visual recognition. However, due to limitations in image acquisition devices, transmission channels, or processing algorithms, images may be affected by various noises, distortions, or artifacts, resulting in a decline in image quality. Therefore, accurately evaluating image quality is crucial for ensuring the reliability of image applications and user experience.
[0003] According to the availability of a reference high-quality image, image quality assessment methods are classified into three categories: full-reference image quality assessment (Full-Reference IQA, FR-IQA) (Wang, Z., et al. “Image Quality Assessment: From Error Visibility to Structural Similarity.” IEEE Transactions on Image Processing, Apr. 2004, pp. 600-12), reduced-reference image quality assessment (Reduced-Reference IQA, RR-IQA) (Soundararajan, Rajiv, and Alan C. Bovik. “RRED Indices: Reduced Reference Entropic Differencing Framework for Image Quality Assessment.” 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2011) and blind reference image quality assessment (Blind IQA, BIOA) (Moorthy, A.K., and A.C. Bovik. “Blind Image Quality Assessment: From Natural Scene Statistics to Perceptual Quality.” IEEE Transactions on Image Processing, Apr. 2011, pp. 3350-64). Full-reference methods use a high-quality reference image to compare the quality of the image being evaluated. Reduced-reference methods use limited information about the original image to evaluate the quality of the distorted image, while blind reference methods do not require a reference image but rely on computational models to evaluate the quality of the image, which is more challenging than full-reference and reduced-reference image quality assessment and is more practical for many applications due to the lack of a general reference image.
[0004] Blind reference image quality assessment includes traditional methods (Saad, Michele A., Alan C. Bovik, and Christophe Charrier. "Blind image quality assessment: A natural scene statistics approach in the DCT domain." IEEE transactions on Image Processing 21.8 (2012): 3339 - 3352) and deep learning methods (LeCun, Yann, Yoshua Bengio, and Geoffrey Hinton. "Deep learning." nature 521.7553 (2015): 436 - 444). The natural scene statistics theory (Simoncelli, Eero P., and Bruno A. Olshausen. "Natural image statistics and neural representation." Annual review of neuroscience 24.1 (2001): 1193 - 1216) is a relatively popular traditional method. This theory holds that high-quality natural images follow certain statistical laws, while various types of distorted images will disrupt these laws, thus providing a reasonable method to approximately perceive visual quality. However, when faced with complex and mixed distortions, the performance of models based on traditional methods is often insufficient. With the rise of convolutional neural networks (Yamashita, Rikiya, et al. "Convolutional neural networks: an overview and application in radiology." Insights into imaging 9 (2018): 611 - 629), deep learning has developed rapidly, and methods based on deep learning have become the mainstream methods for image quality assessment. Deep learning methods automatically extract quality-related features of images in an end-to-end manner. Compared with traditional methods, these methods have stronger quality perception ability and stronger generalization ability. However, the high cost of annotating image quality assessment datasets has led to relatively small image quality assessment datasets, making data-driven deep learning models prone to overfitting. Therefore, it is difficult to fully exploit the performance of deep learning models relying solely on scarce labeled data.In this case, it is crucial to expand the image quality assessment dataset and utilize semi-supervised learning methods to extract potential information from unlabeled data. Existing methods usually adopt generative adversarial networks to generate unlabeled data, but the data generated by this method often fails to achieve the desired effect. Moreover, the generated data inevitably contains noise, which has a negative impact on model training. Therefore, it is particularly important to use more advanced methods to generate pictures and eliminate the generated noise. Summary of the Invention
[0005] Aiming at the problems of insufficient datasets in the field of blind reference image quality assessment and the damage of outliers in unlabeled datasets to model performance, the present invention provides a semi-supervised method for blind reference image quality assessment based on PU learning. First, a diffusion model is used to generate unlabeled data to expand the blind image quality assessment dataset. Furthermore, an unlabeled dataset with the same semantic information as the positive samples is generated and added to the negative samples to reduce the interference of semantic information and improve the ability of PU learning to exclude outliers. For the obtained pure unlabeled data and labeled data, an efficient semi-supervised learning method - the classroom-student architecture is used to train the network model to improve model performance.
[0006] The present invention includes the following steps:
[0007] 1) Scale the image quality score S corresponding to the data in the non-target dataset to [0 - 9] and round it to an integer y as its score level, and use this as a label to train the guided diffusion model (M 1 ) to generate unlabeled data U 1 ;
[0008] 2) Use the target dataset L to train the guided diffusion model (M 2 ) in the manner of step 1) to generate unlabeled data U 2 with semantics similar to the target dataset;
[0009] 3) Pre-train the teacher model T using the target dataset L;
[0010] 4) Sample a batch of labeled data B l from the target dataset L as positive samples, and use M 1 and M 2 to sample a batch of unlabeled data B u1 as negative samples respectively, and use this to train the PU learning model b(x);
[0011] 5) Sample n more batches of unlabeled data B 1 and M 2 in M u, obtain the confidence of the data through the PU - learning model \(b(x)\), regard the data with confidence lower than the specified threshold \(\lambda\) as outliers and delete them to obtain relatively pure unlabeled data \(B\) uυ ;
[0012] 6) Use the teacher model \(T\) to assign pseudo - labels \(y^*\) to the pure unlabeled data, and no data augmentation is performed on the data input to the teacher model;
[0013] 7) After data augmentation on the above - mentioned batch of labeled data \(B\) l and the pure unlabeled data \(B\) with pseudo - labels \(y^*\) uv , use them to supervise the training of the student model; among them, add the average confidence of the unlabeled data to its loss term;
[0014] 8) Repeat steps 4) - 7) until the maximum number of iterations is reached.
[0015] In step 1), training the guided diffusion model includes training the classifier and the diffusion model; training the classifier \(p\) φ (y|x t , t): Let the maximum value \(s\) of the average subjective scores of the non - target data set max and the minimum value \(s\) min , the picture quality score \(S\), and perform the following operations on \(S\) to obtain the quality - level label corresponding to the picture:
[0016]
[0017] The classifier \(p\) used to supervise the training of the guided diffusion model sampling φ (y|x t , t) is composed of a Unet encoder plus a fully - connected classification head;
[0018] Train the diffusion model \(\epsilon\) θ (x t , t) as follows: Gradually add prior random noise \(\epsilon\) 0 to the original picture \(x\) i to obtain noise samples \(x\) 1 ,...·x T-1 , x T , where \(x\) T is isotropic Gaussian noise, and use \((x\) i , \(\epsilon\) i ) as data to train the diffusion model \(\epsilon\) θ (x t , t) until the maximum number of iterations is reached; the diffusion model uses a complete Unet network, including an encoder and a decoder.
[0019] In step 3), the pre-training process is to preprocess the labeled data to obtain a preprocessed training set, denoted as where represents the i-th distorted image, and N is the number of images; the preprocessing is random cropping and flipping operations.
[0020] In step 4), the PU-learning model uses VGG19, loads the pre-trained model trained on the imagenet dataset, deletes the classification head, adds a binary classification fully connected layer as the new classification head, uses the focal loss as the loss function, uses the SGD optimizer, and randomly crops and flips the data before training the model.
[0021] In step 5), the sampling process is as follows: First, randomly sample an isotropic Gaussian noise X from the standard normal distribution T , and according to the diffusion model ∈ θ (x t , t), calculate the mean μ ← μ θ (x t , t) and variance ∑ ← ∑ θ (x t , t), then calculate the classifier gradient and multiply it by the gradient scaling factor a to get g, which is used to guide the sampling direction. Then, randomly sample in the distribution to obtain a denoised image, and the sampled image is obtained after 1000 steps of denoising; where α t : = 1 - β t , β t is a predefined constant, and fix μ θ (x t , t) as β t ; Pass the sampled image through the PU-learning model, and filter out those with a confidence threshold lower than 0.4.
[0022] In step 7), perform random cropping and flipping data augmentation on the above batch of labeled data B l , perform random flipping data augmentation on the pure unlabeled data B uυ , and the average confidence of the unlabeled data given by the PU-learning model is The loss function of the student model is defined as follows:
[0023]
[0024] where σ is a hyperparameter used to weigh the weights of labeled data and unlabeled data in the loss term.
[0025] First, the present invention slices according to the average subjective score of the image to obtain the corresponding quality level of the image as the label of the image, and trains the guided diffusion model to generate a large amount of unlabeled data to expand the BIQA dataset; secondly, applies PU-learning to exclude outliers in the unlabeled data, that is, assigns positive labels to the labeled data and negative labels to the unlabeled data, and in the unlabeled data, adds an unlabeled dataset generated by the diffusion model with semantic information similar to the labeled data to prevent the PU-learning model from simply distinguishing positive and negative samples according to semantic information and improve its ability to exclude outliers; thirdly, based on the labeled data and the pure unlabeled data, with the help of teacher-student model training, uses the pre-trained teacher model to generate pseudo-labels for the unlabeled data, and adaptively adjusts the loss weight of the unlabeled data according to the confidence of the data given by the PU-learning model. The present invention uses the diffusion model to generate pictures to expand the image quality assessment dataset, combines the diffusion model with PU learning to provide a novel method to screen out outliers in the generated data, and uses an efficient semi-supervised learning method - the classroom-student architecture to train the network model for the obtained pure unlabeled data and labeled data to improve the model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of the overall framework of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The following embodiments will describe the present invention in detail with reference to the accompanying drawings.
[0028] As Figure 1 , the embodiments of the present invention include the following steps:
[0029] 1) Use the non-target dataset to train the guided diffusion model (M 1 ) to generate unlabeled data U 1 . Training the guided diffusion model includes training the classifier p φ (y|x t , t) and the diffusion model ∈ θ (x t , t).
[0030] Training the classifier p φ (y|x t , t): Let the maximum value s max and the minimum value s min of the average subjective score of the dataset, the picture quality score S, and perform the following operations on S to obtain the corresponding quality level label of the picture:
[0031]
[0032] The classifier p used to supervise the training of the guided diffusion model sampling φ (y|x t , t), and the classifier is composed of a Unet encoder and a fully connected classification head. Training the diffusion model ∈ θ (x t , t) is as follows: Gradually add prior random noise ∈ 0 to the original image x i to obtain noise samples x 1 ,....x T-1 , x T , where x T is isotropic Gaussian noise, and use (x i , ε i ) as data to train the diffusion model ∈ θ (x t , t) until the maximum number of iterations is reached. The diffusion model uses a complete Unet network, including an encoder and a decoder.
[0033] 2) Use the target dataset L to train the guided diffusion model (M 2 ) in the above way to generate unlabeled data U 2 whose semantics are similar to the target dataset.
[0034] 3) Pre-train the teacher model T using the target dataset L. Before training, preprocess the labeled data to obtain the preprocessed training set, denoted as where represents the i-th distorted image, and N is the number of images; the preprocessing is random cropping and flipping operations.
[0035] 4) Sample a batch of labeled data B l from the target dataset L as positive samples, and use M 1 and M 2 to sample a batch of unlabeled data B u1 as negative samples respectively, and train the PU learning model b(x) with this. The PU-learning model b(x) uses the VGG19 model, initializes the VGG19 model with the model parameters pre-trained on imagenet, uses the focal loss as the loss function, uses the SGD optimizer, and randomly crops and flips the data before training the model.
[0036] 5) Sample n more batches of unlabeled data B 1 from M 2 and M u, the confidence of the data is obtained through the PU - learning model b(x), and the data with confidence lower than the specified threshold A is regarded as outliers and deleted to obtain relatively pure unlabeled data B uv . For each sampling, first, an isotropic Gaussian noise X is randomly sampled from the standard normal distribution T , according to the diffusion model ∈ θ (x t , t), calculate the mean μ←μ θ (x t , t) and variance ∑←∑ θ (x t , t), then calculate the classifier gradient and multiply it by the gradient scaling factor a to get g, which is used to guide the sampling direction. Then, a random sampling is performed in the distribution to obtain a denoised image. After 1000 steps of denoising, a sampled image is obtained. Among them, α t := 1 - β t , β t is a predefined constant, fix μ θ (x t , t) as β t . The sampled pictures are passed through the PU - learning model, and those with a confidence threshold lower than 0.4 are filtered out.
[0037] 6) Use the teacher model T to assign pseudo - labels y* to the pure unlabeled data. The data input to the teacher model is not data - augmented.
[0038] 7) Perform random cropping and flipping data augmentation on the above - mentioned batch of labeled data B l , and perform random flipping data augmentation on the pure unlabeled data B with pseudo - labels y* uυ to train the student model. Let the average confidence of the unlabeled data given by the PU - learning model be The loss function of the student model is defined as follows:
[0039]
[0040] Among them, σ is a weight hyperparameter used to balance the labeled data and unlabeled data in the loss term.
[0041] 8) Repeat steps 4) to 7) until the maximum number of iterations is satisfied.
[0042] The present invention conducts verification experiments on the datasets LIVE, CSIQ, TID2013, and KADID and compares with the state - of - the - art methods. The results are shown in Table 1:
[0043] Table 1
[0044]
[0045] In Table 1, SL is the performance of the model under full supervision, and DPIQA is the result using the method of the present invention under semi-supervised learning conditions. Comparing with the performance of existing state-of-the-art methods, the method proposed in the present invention achieves the best results on most datasets.
[0046] DIIVINE: Saad, M.A.; Bovik, A.C.; and Charrier, C. 2012. Blind image quality assessment: A natural scene statistics approach in the DCT domain. IEEE transactions on Image ing, 21(8): 3339 - 3352. "Blind image quality assessment: A natural scene statistics approach in the DCT domain".
[0047] BRISQUE: Mittal, A.; Moorthy, A.K.; and Bovik, A.C. 2012a. reference image quality assessment in the spatial domain. IEEE Transactions on image processing, 21(12): 4695 - 4708.
[0048] ILNIQE: Zhang, L.; Zhang, L.; and Bovik, A.C. 2015a. A enriched completely blind image quality evaluator. IEEE Transactions on Image Processing, 24(8): 2579–2591.
[0049] BIECON: Kim, J.; and Lee, S. 2016. Fully deep blind image quality predictor. IEEE Journal of selected topics in signal ing, 11(1): 206–220.
[0050] MEON: Ma, K.; Liu, W.; Zhang, K.; Duanmu, Z.; Wang, Z.; and Zuo, W. 2017. End-to-end blind image quality assessment using deep neural networks. IEEE Transactions on Image Processing, 27(3): 1202–1213.
[0051] WaDIQaM: Bosse, S.; Maniry, D.; M¨uller, K.-R.; Wiegand, T.; and Samek, W. 2017. Deep neural networks for no-reference and full-reference image quality assessment. IEEE Transactions on image processing, 27(1): 206–219.
[0052] DBCNN: Zhang, W.; Ma, K.; Yan, J.; Deng, D.; and Wang, Z. 2018. Blind image quality assessment using a deep bilinear convolutional neural network. IEEE Transactions on Circuits and Systems for Video Technology, 30(1): 36–47.
[0053] TIQA: You, J.; and Korhonen, J. 2021a. Transformer for image quality assessment. In 2021 IEEE international conference on image processing (ICIP), 1389–1393. IEEE.
[0054] MetaIQA: Zhu, H.; Li, L.; Wu, J.; Dong, W.; and Shi, G. 2020. MetaIQA: Deep meta-learning for no-reference image ity assessment. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 14143–14152. Zhu, X. J. 2005. Semi-supervised learning literature survey。
[0055] P2P-BM: Sohn, K.; Berthelot, D.; Carlini, N.; Zhang, Z.; Zhang, H.; Raffel, C. A.; Cubuk, E. D.; Kurakin, A.; and Li, C.-L. 2020. Fixmatch: Simplifying semi-supervisedlearning with consistency and confidence. Advances in neural informationprocessing systems, 33:596–608。
[0056] HyperIQA: Su, S.; Yan, Q.; Zhu, Y.; Zhang, C.; Ge, X.; Sun, J.; and Zhang, Y. 2020b. Blindly assess image quality in the wild guided by a self-adaptivehyper network. In Proceedings of the IEEE / CVF Conference on Computer Visionand Pattern Recognition, 3667–3676。
[0057] TReS: Golestaneh, S.A.; Dadsetan, S.; and Kitani, K.M. 2022a. No-reference image quality assessment via transformers, relative ranking, and self-consistency. In Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision, 1220–1230.
[0058] MUSIQ: Ke, J.; Wang, Q.; Wang, Y.; Milanfar, P.; and Yang, F. 2021. Musiq: Multi-scale image quality transformer. In Proceedings of the IEEE / CVF International Conference on Computer Vision, 5148–5157.
[0059] DACNN: Pan, Z.; Zhang, H.; Lei, J.; Fang, Y.; Shao, X.; Ling, N.; and Kwong, S. 2022. Dacnn: Blind image quality assessment via a distortion-aware convolutional neural network. IEEE Transactions on Circuits and Systems for Video Technology, 32(11): 7518–7531.
[0060] DEIQT: Qin, G.; Hu, R.; Liu, Y.; Zheng, X.; Liu, H.; Li, X.; and Zhang, Y. 2023b. Data-Efficient Image Quality Assessment with Attention-Panel Decoder. arXiv preprint arXiv:2304.04952.
Claims
1. A semi-supervised method for blind reference image quality assessment based on PU learning, characterized in that it includes the following steps: 1) Scale the image quality score a corresponding to the data in the non-target dataset to [0-9] and round it to an integer y as its score level, and use this as a label to train the guided diffusion model M 1 to generate unlabeled data U 1 ; Training the guided diffusion model includes training a classifier and a diffusion model; Train classifier p φ (y|x t , t): Let the maximum value s max and the minimum value s min of the average subjective scores of the non-target dataset. For the picture quality score s, perform the following operations to obtain the quality level label corresponding to the picture: Classifier p used to supervise the training of the guided diffusion model sampling φ (y|x t , t), the classifier is composed of a Unet encoder and a fully connected classification head; Training diffusion model ε θ (x t , t): Gradually add prior random noise ∈ 0 to the original image x i to obtain noise samples x 1 ,....x T-1 , x T , where x T is isotropic Gaussian noise, and use (x i , ε i ) as data to train the diffusion model ∈ θ (x t , t) until the maximum number of iterations is reached; the diffusion model uses a complete Unet network, including an encoder and a decoder; 2) Use the target dataset L to train the guided diffusion model M in the manner of step 1) 2 to generate unlabeled data U with semantics similar to the target dataset 2 ; 3) Pre-train the teacher model T using the target dataset L; 4) Sample a batch of labeled data B from the target dataset L l As positive samples, use M 1 and M 2 to sample a batch of unlabeled data B u1 as negative samples, and train the PU-learning model b(x) with this; The PU-learning model uses VGG19, loads the pre-trained model trained on the imagenet dataset, deletes the classification head, adds a binary classification fully connected layer as the new classification head, uses the focal loss as the loss function, uses the SGD optimizer, and randomly crops and flips the data before training the model; 5) Sample n batches of unlabeled data B in M 1 and M 2 respectively, obtain the confidence of the data through the PU-learning model b(x), regard the data with confidence lower than the specified threshold λ as outliers and delete them, and obtain relatively pure unlabeled data B u ; uv ; 6) Use the teacher model T to assign pseudo-labels y to the pure unlabeled data * , and no data augmentation is performed on the data input to the teacher model; 7) For the above-mentioned batch of labeled data B l and the clean unlabeled data B * with pseudo label y uv After data augmentation, it is used to supervise the training of the student model; among them, the average confidence of the unlabeled data is added to its loss term; 8) Repeat steps 4) to 7) until the maximum number of iterations is satisfied.
2. The semi-supervised method for blind reference image quality assessment based on PU learning according to claim 1, characterized in that In step 3), the pre-training process is to preprocess the labeled data to obtain a preprocessed training set, denoted as where represents the i-th distorted image, N is the number of images; the preprocessing is random cropping and flipping operations.
3. The semi-supervised method for blind reference image quality assessment based on PU learning according to claim 1, characterized in that In step 5), the sampling process is as follows: First, randomly sample an isotropic Gaussian noise X from a standard normal distribution T , and according to the diffusion model ∈ θ (x t , t), calculate the mean μ ← μ θ (x t , t) and variance ∑ ← ∑ θ (x t , t) of the denoised image. Then calculate the classifier gradient and multiply it by the gradient scaling factor a to get g, which is used to guide the sampling direction. Next, randomly sample in the distribution to obtain the denoised image, and the sampled image is obtained after 1000 steps of denoising; where β t is a predefined constant, fix μ θ (x t , t) as β t ; Pass the sampled image through the PU-learning model, and filter out those with a confidence threshold lower than 0.
4.
4. The semi-supervised method for blind reference image quality assessment based on PU learning according to claim 1, characterized in that In step 7), the above-mentioned batch of labeled data B l and the clean unlabeled data B * with pseudo-label y uv are subjected to data augmentation. Specifically, the above-mentioned batch of labeled data B l is subjected to random cropping and flipping data augmentation, and the clean unlabeled data B uv is subjected to random flipping data augmentation. The average confidence of the unlabeled data given by the PU-learning model is The loss function of the student model is defined as follows: wherein, σ is a weight hyperparameter used to balance the labeled data and the unlabeled data in the loss term.
Citation Information
Patent Citations
Three-dimensional grid generation method and device, electronic equipment and storage medium
CN115994992A
Image analysis method, device and system based on consistency semantic segmentation
CN116030461A