Data-efficient blind super-resolution image quality evaluation method and system based on model sorting
The pseudo-label data set is generated by a method based on model sorting, and a blind super-resolution image quality evaluation model is trained using semi-supervised learning method, which solves the problem of dependence on manual annotation data in the prior art, and improves model performance and prediction consistency.
Patent Information
- Application Number
- CN202510183646.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The existing blind image quality evaluation method requires a large amount of manual annotation data for supervision and training, which leads to high data acquisition costs and is difficult to achieve large-scale application.
A model-based sorting method is adopted to generate a pseudo-label dataset through maximum difference competition and pairwise sorting learning, and combined with the existing image quality evaluation dataset, the blind super-resolution image quality evaluation model is trained in a semi-supervised learning method.
It effectively reduces the dependence on manual labeled data in model training, improves the performance of blind image quality evaluation models, and makes its prediction results more consistent with subjective evaluation.
Smart Images

Figure CN120107210A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image quality assessment, and in particular to a blind super-resolution image quality assessment method and system. Background Art
[0002] Image quality assessment is an important research field in computer vision, which aims to quantify the perceived quality of images, thereby providing important performance indicators for other visual tasks (such as image super-resolution). Research in this field mainly focuses on two evaluation methods: subjective quality assessment and objective quality assessment. Subjective quality assessment can best reflect the perceived quality of images by humans, but it is limited by high labor and time costs, lacks real-time performance, and is difficult to apply on a large scale. As an alternative to subjective quality assessment, objective quality assessment aims to automatically evaluate the quality of images using computational models, which is efficient and has good scalability.
[0003] According to whether it relies on reference images, objective quality evaluation can be further divided into full-reference quality evaluation and no-reference quality evaluation, the latter of which is also called blind image quality evaluation. Full-reference quality evaluation quantifies the quality of an image by directly comparing the difference between a reference image and a test image. Relevant research has made significant progress in recent years. However, in practical applications, especially in image super-resolution tasks, the application scenarios of full-reference quality evaluation are actually very limited because the original reference image is often not available. In contrast, blind image quality evaluation does not rely on reference images, but only evaluates the quality of an image based on the information of the distorted image itself. In recent years, with the rapid development of deep learning, blind image quality evaluation methods based on deep learning have gradually become a research hotspot. Such methods usually require a large number of manually annotated image-quality score pairs for supervised training, but high-quality manually annotated data requires high costs, and it is often unrealistic to obtain a large amount of high-quality manually annotated data.
[0004] Therefore, how to reduce the dependence on manually labeled data in model training while designing a data-efficient blind super-resolution image quality assessment learning method has become a key issue that needs to be solved urgently. Summary of the invention
[0005] In view of the above-mentioned deficiencies in the prior art, an object of the present invention is to provide a data-efficient blind super-resolution image quality assessment method and system based on model sorting.
[0006] According to one aspect of the present invention, a data-efficient blind super-resolution image quality assessment learning method based on model sorting is provided, comprising:
[0007] Generate a super-resolution image set from the original low-resolution image set using multiple super-resolution methods;
[0008] A maximum difference competition method is used to select a subset that maximizes the difference between each pair of super-resolution methods from the original low-resolution image set;
[0009] Performing subjective testing on the super-resolution images corresponding to the subsets selected by each pair of super-resolution methods to obtain a global ranking of the multiple super-resolution methods;
[0010] Transferring the global rankings of the multiple super-resolution methods to instance-level pseudo labels of unlabeled super-resolution images to obtain a pseudo-label dataset;
[0011] A blind super-resolution image quality assessment model is trained in a semi-supervised learning manner by adopting a pairwise ranking learning method and combining the pseudo-label dataset with an existing image quality assessment dataset.
[0012] Preferably, the super-resolution image set is generated from the original low-resolution image set using multiple super-resolution methods, wherein:
[0013] The original low-resolution image set is an unlabeled, reference-free low-resolution image set from the real world.
[0014] The multiple super-resolution methods can be selected from mainstream image super-resolution methods.
[0015] Preferably, the adopting the maximum difference competition method to select a subset that maximizes the difference between each pair of super-resolution methods from the original low-resolution image set comprises:
[0016] For each pair of different super-resolution methods, we traverse each sample in the original low-resolution image set and iteratively select K samples that maximize the difference between the two methods.
[0017] Preferably, a subjective test is performed on the super-resolution images corresponding to the subsets selected by each pair of super-resolution methods to obtain a global ranking of the multiple super-resolution methods, including:
[0018] In each round of the subjective test, a pair of super-resolution images generated from the same original low-resolution image are shown, using a two-alternative forced choice (2AFC) method.
[0019] A count matrix is constructed based on the human preference data obtained from the subjective test, and maximum likelihood estimation is used under the Thurstone model to infer the global ranking of the multiple super-resolution methods.
[0020] Preferably, the global rankings of the multiple super-resolution methods are transferred to instance-level pseudo labels of unlabeled super-resolution images to obtain a pseudo-label dataset, wherein:
[0021] The unlabeled images include all remaining images in the super-resolution image set generated by the multiple super-resolution methods that are not selected by the maximum difference competition method;
[0022] The pseudo-label is the relative position of the super-resolution method used to generate the current unlabeled image in the global ranking vector.
[0023] Preferably, a pairwise ranking learning method is adopted, and the pseudo-label dataset is combined with an existing image quality assessment dataset to train a blind super-resolution image quality assessment model in a semi-supervised learning manner, including:
[0024] For every two samples x and y in the same batch of training data, calculate a 0-1 binary relative quality label and calculate the probability that the predicted value of x is greater than the predicted value of y under the Thurstone model;
[0025] Calculate the fidelity loss of every two samples based on the 0-1 binary relative quality label and the probability;
[0026] The fidelity loss of the pseudo-label dataset and the existing image quality assessment dataset is calculated respectively, and the total loss is obtained by adding them after setting the weight factor, and back propagation is performed to update the model parameters.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] The present invention can effectively alleviate the problem that blind image quality evaluation methods require a large amount of manually annotated data for supervised training;
[0029] The present invention can effectively improve the performance of the existing blind image quality evaluation model and obtain a prediction result that is more consistent with the subjective evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0031] Figure 1 A flow chart of a data-efficient blind super-resolution image quality assessment learning method based on model sorting according to an embodiment of the present invention; DETAILED DESCRIPTION
[0032] The following is a detailed description of the embodiments of the present invention: This embodiment is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and a specific operation process are given. It should be pointed out that for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.
[0033] like Figure 1 As shown, in one embodiment of the present invention, a data-efficient blind super-resolution image quality assessment learning method based on model sorting includes:
[0034] S1: Generate a super-resolution image set from the original low-resolution image set using multiple super-resolution methods to increase the resolution of the image, thereby improving the details and clarity of the image. Specifically, it processes unlabeled, reference-free low-resolution image sets from the real world. Super-resolution methods can be selected from mainstream image super-resolution methods.
[0041] In this embodiment, 12 super-resolution methods are selected, namely USRGAN, Real-ESRGAN, BSRGAN, DASR, FeMaSR, LDL, DiffBIR, RGT, ResShift, SinSR, StableSR, and SeeSR. For a given set of original low-resolution images, the above 12 super-resolution methods are used to process them respectively. Each method will perform feature extraction, feature mapping and high-resolution image reconstruction on the input low-resolution image according to its own algorithm and model, including forward propagation calculation, that is, mapping the low-resolution image to the high-resolution image space through a deep learning model. Finally, each method will generate a corresponding super-resolution image set. These image sets not only have improved resolution, but may also contain richer details and clearer edges.
[0035] S2: Use the maximum difference competition method to select a subset from the original low-resolution image set that maximizes the difference between each pair of super-resolution methods in step S1; that is, use the maximum difference competition method to select a subset from the original low-resolution image set that maximizes the difference between the two super-resolution methods. Specifically, for each pair of different super-resolution methods f 1 With f 2 , traverse each sample x in the low-resolution image set X, and iteratively solve the following expression to select the kth sample that maximizes the difference between the two methods:
[0043]
[0044] Among them, S is the set of k-1 samples selected according to the maximum difference competition method, D 1 is a method for estimating two super-resolution methods f 1 With f 2 The measure of the perceived distance between 2It is a measure used to estimate the semantic distance between an image sample x and a selected sample set S. λ is a balance factor between the two distances D1 and D2, and its value range is usually between 0 and 1. It is used to adjust the importance of perceptual difference and semantic diversity. Perceptual distance usually reflects the subjective feeling of the human visual system on image differences, which can be measured by structural similarity (SSIM), peak signal-to-noise ratio (PSNR) or more advanced perceptual quality assessment indicators. Semantic distance measures the similarity of image content.
[0036] S3: Perform subjective testing on the super-resolution images corresponding to the subsets selected by each pair of super-resolution methods in step S2 to obtain a global ranking of multiple super-resolution methods in S1; Specifically, based on the human preference data obtained from subjective tests, an N×N count matrix C is constructed, where N is the number of super-resolution methods selected, and C ij is the super-resolution method f i In with f j The number of times it is selected in the comparison. After that, the maximum likelihood estimation is used under the Thurstone model to infer the global ranking of the selected multiple super-resolution methods. Specifically, the global ranking vector that maximizes the log-likelihood function of the following count matrix C is solved:
[0047]
[0048] Where μ = [μ 1 ,μ 2 ,...,μ N ] is the global ranking vector of the multiple super-resolution methods, Φ is the standard normal distribution function, and satisfies Σ i μ i =0.
[0037] S4: Transfer the global rankings of multiple super-resolution methods obtained in S3 to instance-level pseudo labels of unlabeled super-resolution images to obtain a pseudo-label dataset; Specifically, the instance-level pseudo-label is the relative position of the super-resolution method that generates the current unlabeled image in the global ranking, with a value range of {1, 2, ..., N}, and each pseudo-label data sample is an image-pseudo-label data pair.
[0038] S5: Using the paired ranking learning method, and combining the pseudo-label dataset obtained in S4 with the existing image quality assessment dataset, a blind super-resolution image quality assessment model is trained in a semi-supervised learning manner. Specifically, for every two samples x and y in the same batch of training data, a relative quality label is calculated:
[0051]
[0052] where μ x ,μ y is the average opinion score of x and y.
[0053] For a blind image quality assessment model q with a predefined calculation structure w (·), under the Thurstone model, calculate the probability that the predicted value of x is higher than the predicted value of y:
[0054]
[0055] Where Φ(·) is the probability cumulative function of the standard normal distribution.
[0056] As a preferred embodiment, a blind image quality assessment model q having a predefined calculation structure w (·) LIQE was selected.
[0057] The model is trained using fidelity loss, specifically:
[0058]
[0059] For pseudo-labeled datasets, instance-level pseudo-labels can only be compared between super-resolution images generated from the same original low-resolution image. Specifically, the data loader loads all super-resolution images generated from the same original low-resolution image as a training batch each time, and the relative quality labels are calculated by the global ranking of the corresponding super-resolution methods.
[0060] Combine the pseudo-labeled dataset with the existing image quality assessment dataset for joint training. Specifically, for the image quality assessment dataset with real labels, calculate the additional fidelity loss term and obtain a semi-supervised learning loss function:
[0061] l=l t +αl p
[0062] Among them, l t is the fidelity loss of the image quality evaluation dataset with true labels, l p is the fidelity loss of the pseudo-labeled dataset, and α is the weight factor.
[0063] As a preferred embodiment, KonIQ-10k and PIPAL are selected as the existing image quality assessment data sets, and the weight factor α is selected as 0.1.
[0064] Afterwards, error backpropagation is performed and model parameters are updated.
[0065] Implementation effect:
[0066] In order to verify the effectiveness of the data-efficient blind super-resolution image quality assessment learning method based on model sorting provided in the above embodiments of the present invention, the average 2AFC score of the model on all test images can be calculated based on the results of subjective tests. The 2AFC score is defined as:
[0067] 2AFCscore=pq+(1-p)(1-q)
[0068] Among them, p is the vote rate and q∈{0,1} is the model preference.
[0069] The performance test results are shown in Table 1. The experiment also tested several baseline models for comparison. It can be seen that the embodiment of the present invention is significantly better than the baseline model in performance.
[0070] Table 1
[0071]
[0072] In order to verify the versatility of the data-efficient blind super-resolution image quality assessment learning method based on model sorting provided in the above embodiment of the present invention, it can be tested on multiple image quality assessment datasets, including: QADS, Ma17, PIPAL, KonIQ-10k, KADID-10k, SPAQ. The experiment mainly measures the model performance through two classic image quality assessment indicators, including PLCC and SRCC. The performance test results are shown in Table 2. The experiment also tested several baseline models for comparison. It can be seen that the embodiment of the present invention is better than most baseline models in performance.
[0073] Table 2
[0074]
[0075] The above-mentioned embodiment of the present invention provides a data-efficient blind super-resolution image quality assessment learning method based on model sorting, which can effectively alleviate the problem that the blind image quality assessment method requires a large amount of manually annotated data for supervised training, effectively improve the performance of the existing blind image quality assessment model, and obtain prediction results that are more consistent with subjective evaluation.
[0076] It should be noted that the steps in the method provided by the present invention can be implemented by using corresponding modules, devices, units, etc. in the system. Those skilled in the art can refer to the technical solution of the system to implement the step flow of the method, that is, the embodiments in the system can be understood as preferred examples for implementing the method, which will not be elaborated here.
[0077] Those skilled in the art know that, in addition to implementing the system and its various devices provided by the present invention in a purely computer-readable program code, the system and its various devices provided by the present invention can be made to implement the same functions in the form of logic gates, switches, application-specific integrated circuits, editable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices provided by the present invention can be considered as a hardware component, and the devices for implementing various functions contained therein can also be considered as structures within the hardware component; the devices for implementing various functions can also be considered as both software modules for implementing the method and structures within the hardware component.
[0078] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various modifications or variations within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the above embodiments and features in the embodiments can be combined with each other.
Claims
1. A data-efficient blind super-resolution image quality assessment learning method based on model sorting, characterized in that: include: S1. Obtain an original low-resolution image set, and generate a corresponding super-resolution image set using multiple super-resolution methods; S2. adopting the maximum difference competition method, selecting a subset that maximizes the difference between each pair of super-resolution methods from the original low-resolution image set; S3. Performing subjective testing on the super-resolution images corresponding to the selected subset and determining the global ranking of the multiple super-resolution methods; S4. transferring the global rankings of the multiple super-resolution methods to instance-level pseudo labels of unlabeled super-resolution images to construct a pseudo-label dataset; S5. Combining the pseudo-label dataset with the existing image quality assessment dataset, a blind super-resolution image quality assessment model is trained using paired ranking learning and semi-supervised learning methods.
2. According to claim 1, a data-efficient blind super-resolution image quality assessment learning method based on model sorting is characterized in that: The original low-resolution image set in step S1 is a label-free, reference-free low-resolution image set from the real world; the multiple super-resolution methods are selected from mainstream image super-resolution methods.
3. The data-efficient blind super-resolution image quality assessment learning method based on model sorting according to claim 1, characterized in that: The step 2 specifically includes: S2.1 obtaining the original low-resolution image set X in step S1, wherein the original low-resolution image set X includes a plurality of low-resolution image samples; S2.2 select at least two different super-resolution methods f1 and f2, wherein each super-resolution method is capable of converting a low-resolution image into a high-resolution image; S2.3 defines a perceptual distance D1, which is used to estimate the perceptual distance between the reconstruction results of two super-resolution methods f1 and f2 for the same low-resolution image sample; S2.4 defines a semantic distance D2, which is used to estimate the semantic distance between the low-resolution image sample and the selected sample set S; S2.5 sets a balance factor λ, where 0≤λ≤1, for adjusting the weights of the perceptual distance and the semantic distance in the calculation of the comprehensive score; S2.6 initialize an empty set S as the set of selected samples; S2.7 For each pair of different super-resolution methods f1, f2, traverse each sample x in the low-resolution image set X and perform the following operations: Calculate the perceptual distance D1(f1(x),f2(x)) between the images of sample x reconstructed by different super-resolution methods f1 and f2; Calculate the semantic distance D2(x,S) between sample x and all samples in set S; According to the perceptual distance D1, semantic distance D2 and balance factor λ, the comprehensive score is calculated, and the sample with the largest comprehensive score is selected. As the kth sample, add it to the set S. The formula is as follows: S2.8 repeat step S2.7 until the predetermined number of samples is reached; S2.9 outputs the set S as the subset that maximizes the difference between super-resolution methods.
4. The data-efficient blind super-resolution image quality assessment learning method based on model sorting according to claim 1, characterized in that: The step S3 specifically includes: S3.1 obtain image subsets generated by different super-resolution methods in step S2, wherein each subset corresponds to each pair of super-resolution methods selected in step S2; S3.2 conducts subjective testing on each subset, asking human observers to compare images generated by different super-resolution methods and select the images they consider to be of better quality; S3.3 Based on the results of the subjective test, construct an N×N count matrix C, where N is the number of super-resolution methods selected, and C ij is the number of times the super-resolution method f1 is selected in comparison with the super-resolution method f2; S3.4 Under the Thurstone model, a maximum likelihood estimation method is used to solve a global ranking vector that maximizes the log-likelihood function based on the count matrix C, wherein the global ranking vector represents the relative ranking of multiple super-resolution methods; S3.5 outputs the global ranking vector as the global ranking of the selected super-resolution method.
5. The data-efficient blind super-resolution image quality assessment learning method based on model sorting according to claim 1, characterized in that: The step S4 specifically includes: S4.1 obtain the global ranking of multiple super-resolution methods obtained in step S3; S4.2 provides a set of unlabeled super-resolution images; S4.3 For each unlabeled image, select a super-resolution method to process it and generate a high-resolution version of the image; S4.4 assigns an instance-level pseudo-label to the generated high-resolution image according to the position of the selected super-resolution method in the global ranking, wherein the pseudo-label represents the relative position of the super-resolution method used to generate the current image in the global ranking, and the value range is {1, 2, ..., N}, where N is the number of super-resolution methods in the global ranking; S4.5 combines each unlabeled image and its corresponding pseudo label into a data pair to form a pseudo-label dataset.
6. The data-efficient blind super-resolution image quality assessment learning method based on model sorting according to claim 1, characterized in that: The step S5 specifically includes: S5.1 For every two samples x and y in the same batch of training data, calculate a 0-1 binary relative quality label and calculate the probability that the predicted value of sample x is greater than the predicted value of sample y under the Thurstone model; S5.2 calculating the fidelity loss of every two samples based on the 0-1 binary relative quality label and the probability; S5.3 respectively calculates the fidelity loss of the pseudo-label dataset and the existing image quality assessment dataset, sets the weight factor and adds them together to obtain the total loss, performs back propagation and updates the model parameters.
7. A data-efficient blind super-resolution image quality assessment terminal based on model sorting, characterized in that The method comprises a memory, a processor and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the method according to any one of claims 1 to 6 when executing the program.
Citation Information
Patent Citations
Reference-free image quality evaluation method for image super-resolution reconstruction
CN108846800A
Super-resolution reconstructed image quality evaluation method
CN116363094A
Super-resolution image quality evaluation method and system in combination with sort learning
CN118115495A
System and Method for Learning-Based Image Super-Resolution
US20190304063A1
Automated method and system for generating models from data
US7480640B1