Super-resolution image evaluation method combining visual effect and reliability measurement
By combining Vision Transformer and ResNet networks to extract features and use reconstruction scale factors for adaptive fusion and spatial attention learning, the unifiedness and applicability of super-resolution image quality evaluation are solved, and quality prediction scores that are more in line with human eye feelings are generated.
Patent Information
- Application Number
- CN202410176402.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-08
- Publication Date
- 2025-08-08
AI Technical Summary
The existing super-resolution image quality evaluation methods are difficult to effectively combine visual effects and reliability, and subjective evaluation is easily disturbed by external factors, and objective evaluation lacks uniformity and applicability.
Super-resolution and low-resolution image features are extracted using Vision Transformer and ResNet networks, combined with reconstruction scale factors, visual effect and reliability chunking measurements are performed through adaptive fusion and spatial attention learning, and quality prediction scores are finally generated.
It is realized that without the need for original high-resolution images, the quality evaluation results are more in line with the subjective feelings of the human eye, enhance the flexibility and accuracy of the evaluation, take into account both global and local characteristics, and improve the uniformity and applicability of the evaluation.
Smart Images

Figure BDA0004702266010000041 
Figure BDA0004702266010000044 
Figure BDA0004702266010000045
Abstract
Description
Technical Field
[0001] The present invention relates to a super-resolution image evaluation method combining visual effects and reliability measurement, and belongs to the field of computer vision and intelligent information technology. Background Art
[0002] Image super-resolution reconstruction has been a hot research topic in computer vision in recent years. It aims to reconstruct low-resolution images using image processing techniques, thereby restoring more details and generating high-resolution images. With the rapid development of artificial intelligence (AI), deep learning-based image super-resolution reconstruction has achieved remarkable results and is widely used in various fields, such as security and surveillance, medical imaging, and remote sensing.
[0003] Research in image super-resolution technology has made significant progress, and effective quality assessment of reconstructed super-resolution images has become a pressing task. To measure the quality of reconstructed images and better compare the reconstruction performance of various super-resolution algorithms, it is necessary to adopt appropriate quality assessment methods for super-resolution images. Image quality assessment algorithms can be divided into subjective and objective quality assessments based on the evaluation subject. Subjective quality assessment scores the reconstructed image based on the human eye's subjective visual effects, but is easily affected by external factors and has certain limitations. Therefore, objective quality assessment methods that are consistent with subjective quality assessments have become a research hotspot in image quality assessment. Super-resolution image quality assessment must not only consider the visual effects of the super-resolution image, but also ensure consistency and correspondence with its corresponding low-resolution image. Furthermore, different reconstruction scale factors have a statistically significant impact on the subjective quality score of super-resolution reconstructed images. Therefore, this paper introduces the low-resolution image and reconstruction scale factor as prior knowledge to assist in super-resolution image quality assessment. A semi-reference evaluation method for super-resolution reconstructed image quality that combines visual effects and reliability metrics is proposed. This method aims to enhance the network's quality assessment performance and ensure that the prediction results are more consistent with the subjective evaluation results of the human eye. Summary of the Invention
[0004] This paper proposes a semi-referenced method for evaluating the quality of super-resolution reconstructed images using a combined visual effect and reliability metric. Using a low-resolution image as a reference and combining it with the reconstruction scale factor information, the method comprehensively evaluates the visual effect and reliability of the super-resolution image. It uses a residual neural network (ResNet) and a Vision Transformer (ViT) to simultaneously extract image features from both the super-resolution image and the low-resolution image, and fuses these features with the features represented by the reconstruction scale factor. Furthermore, it designs visual effect metric and reliability metric branches to measure the visual effect and reliability of the super-resolution image in blocks, and adaptively fuses them through learning spatial attention, making the predicted scores more consistent with the subjective visual effect of the human eye.
[0005] A super-resolution image evaluation method combining visual effect and reliability measurement includes the following steps:
[0006] (1) Input the super-resolution reconstructed image and its paired low-resolution image, and use the ViT network and ResNet network to extract features respectively;
[0007] (2) In the feature fusion stage, the weights of ViT features and ResNet features are learned for adaptive fusion, and then fused with the feature representation of the reconstructed scale factor to obtain more comprehensive image features; visual features and difference features are obtained through feature operations and used as inputs of the visual effect measurement branch and the reliability measurement branch respectively;
[0008] (3) Using visual features and difference features to measure the visual effect and reliability of the super-resolution image in blocks, and obtain a visual effect score map and a reliability score map respectively;
[0009] (4) Using visual features and difference features, the weights of each block are learned through spatial attention to obtain the visual effect score weight map and the reliability score weight map respectively;
[0010] (5) The final prediction score is obtained by combining the visual effect score map and the reliability score map with the attention weight and performing weighted summation.
[0011] Compared with the prior art, the present invention has the following beneficial effects:
[0012] 1. This paper proposes a semi-referenced super-resolution reconstruction image quality evaluation method, which does not require the original high-resolution image and is more applicable and flexible in real scenes;
[0013] 2. The present invention introduces low-resolution images and reconstruction scale factors as prior knowledge to assist in completing super-resolution image quality evaluation, making better use of existing information;
[0014] 3. The present invention simultaneously utilizes the ViT network and the ResNet network to extract features from the input image pair, taking into account both global and local feature information. In the feature fusion stage, the weights can be adaptively learned to further fuse with the features represented by the reconstructed scale factor, further enhancing the feature information related to the super-resolution image quality assessment task;
[0015] 4. The present invention evaluates the quality of super-resolution reconstructed images from the perspectives of visual effect and reliability, and obtains the final prediction results through two metric branches and a weight branch, thereby enhancing the quality evaluation performance of the network. The prediction results can better reflect people's subjective visual experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a block diagram of the overall network model of the present invention;
[0017] Figure 2 This is a structural diagram of the adaptive feature fusion module of the present invention;
[0018] Figure 3 This is a structural diagram of the block measurement module of the present invention;
[0019] Figure 4 This is a structural diagram of the weight generation module of the present invention;
[0020] Figure 5 This is a data consistency comparison chart between the present invention and other image quality assessment methods on the RealSRQ dataset. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0022] 1. Overall network framework
[0023] like Figure 1 As shown in Figure 1, the network mainly consists of a visual effect evaluation branch and a reliability evaluation branch. The entire network can be divided into six stages: feature extraction, feature operation, feature fusion, block measurement, fusion weight generation and final fusion stage. Given a super-resolution reconstructed image I SR and its paired low-resolution image I LR ,The goal of the super-resolution reconstruction image quality evaluation ,task is to generate a prediction score that is consistent with the ,subjective visual effect of the human eye.
[0024] 2. Feature extraction stage
[0025] Since Vision Transformer (ViT) can capture global image information, while convolutional neural networks can capture local image information, the ViT-B / 8 network and ResNet50 network pre-trained on ImageNet are selected to simultaneously extract features from super-resolution images and their corresponding low-resolution images, so as to better capture the features of various details in the super-resolution image itself. In the feature extraction stage, the network model is based on the low-resolution image I LR and super-resolution image I SR As input, the low-resolution image I LR After interpolation and amplification, the ResNet network and ViT network are used to extract features respectively. The process can be expressed as:
[0026]
[0027] Among them, f ResNet and f ViT They are ResNet network and ViT network, F ResNet_LR 、F ResNet_SR 、F ViT_LR and F ViT_SR It is the feature result extracted by ResNet network and ViT network.
[0028] For the ViT network, since the number of Transformer encoders used by ViT is 12, that is, there are 12 cascaded Transformer encoders for encoding, 12 different levels of encoder output features will be obtained in the middle layer. This paper selects the intermediate features of the 0th, 1st, 2nd, 3rd, and 4th layers (a total of 5 layers) as output and merges them in the channel dimension, that is, the output feature F of ViT is obtained. ViT_LR and F ViT_SR The dimension is shape = ((H×W), 5N), where N is the number of channels of the output feature in ViT, and N = 784. In order to better interact and fuse with the features extracted by the ResNet network, its shape is changed. Finally, the dimensionality reduction is performed through 1×1 convolution to obtain the output features of the final ViT network. For the ResNet network, we extract the multi-scale features of the four different stages of ResNet, then interpolate the four features to make H=W=28, and finally merge the channel dimensions, and then use 1×1 convolution to reduce the dimension to obtain the output features of the final ResNet network.
[0029] 3. Feature calculation stage
[0030] Since low-resolution images and super-resolution images have certain differences in feature space, after obtaining the features extracted by the ResNet network and the ViT network, in order to represent the difference between the super-resolution image and its paired low-resolution image, the following feature operation steps are used:
[0031]
[0032] F ResNet_SR and F ViT_SR As the input of the visual effect metric branch, F ResNet_Diff and F ViT_Diff Serves as input to the reliability measurement branch.
[0033] 4. Feature fusion stage
[0034] In order to better interactively fuse the extracted ViT features and ResNet features, the present invention proposes an adaptive fusion module, which can learn the importance of global and local visual features, learn the respective weights of ViT features and ResNet features, and then adaptively fuse them to obtain more comprehensive image features.
[0035] like Figure 2 As shown in Figure 2, the learned global and local features as well as the reconstruction scale factors captured by the ViT and ResNet feature extraction modules are used as input to the adaptive fusion module. and ResNet features First, input it into the adaptive fusion module and perform dimensionality increase operation, namely Then concatenate them together in the last dimension, which can be expressed as:
[0036]
[0037] Among them, Concat(·) is a connection function, which connects the extracted features. In order to learn the weights of ViT features and ResNet features and evaluate the importance of the two features, the adaptive fusion operation can be expressed as:
[0038]
[0039] Among them, Linear(·) represents linear transformation, which is implemented by the fully connected layer; ReLU(·) represents the ReLU activation function; after fusion
[0040] The reconstruction scale factor has a statistically significant effect on the subjective quality score of the super-resolution image, indicating that scale information can be used to guide the super-resolution reconstruction quality evaluation task. Therefore, the present invention inputs the vector represented by the reconstruction scale factor S into the adaptive fusion module. First, the reconstruction scale factor S is passed through two fully connected layers to obtain the feature F S , and change the shape so that Then, the feature information after adaptive fusion is merged in the channel dimension. The whole process can be expressed as:
[0041]
[0042] Among them, Conv3 represents 3×3 convolution, and the feature fusion stage outputs visual features and difference features
[0043] 5. Block Metrics
[0044] Since each pixel corresponds to a different block of the input image and contains rich information, the quality of the image depends on its different regions, so the information of the spatial dimension is indispensable. Figure 3 As shown in Figure 2, the visual effect and reliability of the super-resolution image are measured in blocks using visual features and difference features, and the visual effect score map S1 and the reliability score map S2 are obtained respectively, which can be expressed as:
[0045]
[0046] Among them, the visual effect score map and the reliability score map
[0047] 6. Fusion weight generation
[0048] While obtaining the visual effect score map and reliability score map of the super-resolution image, the attention weight map is obtained by learning the weights of each block through spatial attention. The weight generation module proposed in this invention is as follows: Figure 4 As shown, the visual effect score weight map w1 and the reliability score weight map w2 can be expressed as:
[0049]
[0050] Among them, Sigmoid(·) represents the Sigmoid activation function, the visual effect score weight map and the reliability score weight map
[0051] 7. Final fusion stage
[0052] The metric branch computes a visual quality score and reliability score for each block in the feature map, while the spatial attention branch generates attention weights for each corresponding score. The final fusion stage combines the attention weight maps w1 and w2, performing a weighted sum of the visual quality score map S1 and the reliability score map S2 to obtain the final prediction score. This weighted sum operation helps model the importance of regions to simulate the human visual system. It can be expressed as:
[0053]
[0054] Where * represents the Hadamard product, S PF is the final prediction score. PF and the true score S GT The mean squared error (MSE) loss between the two is used in the training process of the method of the present invention.
[0055] In order to verify the effectiveness of the super-resolution image evaluation method combining visual effects and reliability metrics described in the present invention, a detailed comparison will be conducted through experiments below.
[0056] Experimental environment: The operating system is the Linux distribution Ubuntu 20.04.5LTS, the deep learning framework Pytorch 1.10, and Python version 3.8. This paper uses three public datasets commonly used in the field of super-resolution reconstruction image quality assessment: QADS, Waterloo, and RealSRQ for training and testing, and compares the PLCC and SRCC results of the model (the larger the PLCC and SRCC, the better the network performance). This paper selects five mainstream image quality evaluation methods in recent years for comparison, specifically:
[0057] CNNIQA: Kang et al., “Kang L, Ye P, Li Y, et al. Convolutional neural networks for no-reference image quality assessment[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2014: 1733-1740.”
[0058] NRQM: Ma et al., “Ma C, Yang CY, Yang X, et al. Learning ano-reference quality metric for single-image super-resolution[J]. Computer Vision and Image Understanding, 2017, 158:1-16.”
[0059] DISQ: The method proposed by Gong et al., reference "Zhao T, Lin Y, Xu Y, et al. Learning-based quality assessment for image super-resolution[J]. IEEE Transactions on Multimedia, 2021, 24: 3570-3581."
[0060] C 2 MT: The method proposed by Li et al., Reference “Li H, Zhang K, Niu Z, et al.C 2 MT: ACredible and Class-Aware Multi-Task Transformer for SR-IQA[J]. IEEE SignalProcessing Letters, 2022, 29: 2662-2666."
[0061] SGH: Method proposed by Fu et al., reference paper "Fu J. Scale Guided Hypernetwork for Blind Super-Resolution Image Quality Assessment [EB / OL]. arXiv preprint arXiv:2306.02398, 2023."
[0062] The test results are shown in Table 1. The method proposed in this paper uses PLCC and SRCC as evaluation indicators on three data sets. Compared with the other five methods, the method proposed in this paper has a greater advantage. Figure 5 A data consistency comparison chart of the method of the present invention and other methods on the RealSRQ dataset is shown. It can be observed that the prediction score of the method of the present invention is more consistent with the subjective visual effect of the human eye.
[0063] Table 1 Comparison of results of super-resolution reconstruction image quality evaluation methods
[0064]
[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A super-resolution image evaluation method combining visual effect and reliability measurement, characterized by The following steps are involved: (1) Input super-resolution reconstructed image I SR and its paired low-resolution image I LR , the Vision Transformer (ViT) network and the residual neural network (ResNet) are used to extract features and obtain the feature F ResNet_SR 、F ViT_SR 、F ResNet_Diff and F ViT_Diff , where F ResNet_SR 、F ViT_SR As the input of the visual effect metric branch, F ResNet_Diff and F ViT_Diff As input to the reliability measurement branch; (2) In the feature fusion stage, the features extracted by the ViT network (F ViT_SR ,F ViT_Diff ) and the features extracted by ResNet network (F ResNet_SR ,F ResNet_Diff ) and fuse their respective weights to obtain a more comprehensive image feature F SR and F Diff , and then reconstruct the feature F represented by the scale factor S S Fusion is performed and visual features F are obtained through feature operation SR_S and differential feature F Diff_S ; (3) Using visual features F SR_S and differential feature F Diff_S The visual effect and reliability of the super-resolution image are measured in blocks, and the visual effect score map S1 and the reliability score map S2 are obtained respectively; (4) Using visual features F SR_S and differential feature F Diff_S , by learning the weights of each block through spatial attention, we obtain the visual effect score weight map w1 and the reliability score weight map w2 respectively; (5) Combine the score weight maps w1 and w2 to perform weighted summation on the visual effect score map S1 and the reliability score map S2 to obtain the final prediction score S PF .
2. The method according to claim 1, characterized in that In step (1), the low-resolution image is used as a reference, and the ViT network and the ResNet network are used to extract features from the super-resolution image and its corresponding low-resolution image, which can be expressed as: Among them, f ResNet and f ViT They are ResNet network and ViT network, F ResNet_LR 、F ResNet_SR 、F ViT_LR and F ViT_SR is its output feature result; After obtaining the features extracted by the ResNet network and the ViT network, in order to calculate the difference between them, the following feature processing steps are used: F ResNet_SR and F ViT_SR As the input of the visual effect metric branch, F ResNet_Diff and F ViT_Diff Serves as input to the reliability measurement branch.
3. The method according to claim 1, characterized in that In the step (2), the ViT feature and the ResNet feature are adaptively fused, and a reconstruction scale factor is introduced to further fuse the feature representation of the reconstruction scale factor with the image feature after adaptive fusion; Vi T Features and ResNet features First, input it into the adaptive fusion module and perform dimensionality increase operation, namely Then concatenate them together in the last dimension, which can be expressed as: Among them, Concat(·) is a connection function, which connects the extracted features. In order to evaluate the importance of the two features, the weights of ViT features and ResNet features are learned to perform adaptive fusion operations, which can be expressed as: Among them, Linear(·) represents linear transformation, which is implemented by the fully connected layer; ReLU(·) represents the ReLU activation function; after fusion The reconstructed scale factor S is passed through two fully connected layers to obtain the feature F S , and change the shape so that Then, the channel dimension is merged with the above adaptively fused image features. The whole process can be expressed as: Among them, Conv3 represents 3×3 convolution, and the feature fusion stage outputs visual features and difference features 4. The method according to claim 1, wherein In step (3), the visual feature F is used SR_S and differential feature F Diff_S The visual effect and reliability of the super-resolution image are measured in blocks, and the visual effect score map S1 and the reliability score map S2 are obtained respectively, which can be expressed as: Among them, the visual effect score map and the reliability score map 5. The method according to claim 1, characterized in that In step (4), the weights of each block are learned by spatial attention to obtain the attention weight map to obtain the visual effect score weight map w1 and the reliability score weight map w2, which can be expressed as: Among them, Sigmoid(·) represents the Sigmoid activation function, the visual effect score weight map and the reliability score weight map 6. The method according to claim 1, characterized in that In step (5), the final prediction score is obtained by weighted summing the visual effect score map S1 and the reliability score map S2 in combination with the attention weights w1, w2, which can be expressed as: Where * represents Hadamard product, S PF is the final prediction score.