Enhanced image quality evaluation method and system based on semantic guidance residual learning, and terminal
Through the method based on semantic-guided residual learning, using the pseudo-twin semantic extraction module and multi-scale enhanced feature extraction, the problem that the image quality evaluation method in the prior art is difficult to obtain distortion-free original images, and more accurate image quality evaluation is achieved.
Patent Information
- Application Number
- CN202510234261.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
AI Technical Summary
The existing image quality evaluation methods are difficult to obtain distortion-free original images, and ignore the connection between semantic information and the features of distorted areas, resulting in lower enhanced image quality evaluation performance.
Using a semantic-guided residual learning method, the enhanced image quality evaluation model is trained by obtaining the training image set, the pseudo-twin semantic extraction module is used to extract semantic information, combine spatial attention maps and channel attention maps, and multi-scale enhanced feature extraction is performed, and model parameters are optimized through the cross-covariance loss function, paying attention to semantic information and distortion information in the image.
It improves the accuracy and performance of image quality evaluation, can better capture semantic information and structural information in the image, and improves the effect of image quality evaluation.
Smart Images

Figure CN120163786A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of enhanced image quality assessment, and particularly to an enhanced image quality assessment method, system, terminal and computer-readable storage medium based on semantic-guided residual learning. Background Art
[0002] Images taken under adverse weather conditions, such as rainy days, foggy days, and cloudy days, usually show a significant quality decline. Rain streaks, fog, and image distortion caused by insufficient brightness will appear in the pictures. These distortions will lead to information loss, visual distortion, and blurring, hindering human visual perception and the performance of vision-related tasks. To solve these problems, image enhancement algorithms such as de-raining, de-fogging, and low-light enhancement have emerged. However, these algorithms may introduce additional distortions during the image enhancement process, such as background information loss and structural damage. Therefore, it is necessary to study an image quality evaluation method to evaluate the performance of related enhancement algorithms.
[0003] Currently, existing image quality evaluation methods are mainly divided into full-reference, semi-reference, and no-reference. On the one hand, since it is difficult to obtain distortion-free original images for existing methods, mainly no-reference methods are used, resulting in poor performance of the methods. On the other hand, existing methods only focus on semantic information or the features of the distorted area, ignoring the connection between the two, which are both very important for the quality of the image, resulting in poor performance of the current enhanced image evaluation methods.
[0004] Therefore, the prior art still needs to be improved and developed. Summary of the Invention
[0005] The main purpose of the present invention is to provide an enhanced image quality assessment method, system and terminal based on semantic-guided residual learning, aiming to solve the problem that it is difficult to obtain distortion-free original images in the prior art, and ignoring the connection between semantic information and the features of the distorted area, resulting in low performance of enhanced image quality assessment.
[0006] To achieve the above object, the present invention provides an enhanced image quality assessment method based on semantic-guided residual learning. The enhanced image quality assessment method based on semantic-guided residual learning includes the following steps:
[0007] Obtain a training image set, and train the created enhanced image quality assessment test model based on the training image set to obtain an enhanced image quality assessment model, where the training image set includes multiple groups of degraded images and enhanced images;
[0008] Obtain the enhanced image to be evaluated, input the enhanced image to be evaluated into the enhanced image quality assessment model, and output the enhanced image quality assessment result.
[0009] Optionally, in the enhanced image quality assessment method based on semantic-guided residual learning, the obtaining of the training image set, and the model training of the created enhanced image quality assessment test model based on the training image set to obtain an enhanced image quality assessment model specifically includes:
[0010] Obtain a training image set of the target object, create an enhanced image quality assessment test model, and input a group of degraded images and enhanced images in the training image set into the enhanced image quality assessment test model;
[0011] Obtain the semantic features of the enhanced image, generate a spatial attention map and a channel attention map according to the semantic features, and perform a combination process on the spatial attention map and the channel attention map to obtain a target feature map;
[0012] Obtain an edge residual map according to the degraded image and the enhanced image, obtain an enhanced feature map according to the edge residual map and the semantic features, and perform multi-scale enhancement on the enhanced feature map to obtain a multi-scale enhanced feature map;
[0013] Perform a splicing process on the multi-scale enhanced feature map and the target feature map to obtain a predicted enhanced image quality assessment result, and calculate a loss between the predicted enhanced image quality assessment result and the true quality score of the enhanced image to obtain an image quality loss value;
[0014] Extract features from the degraded image and the enhanced image respectively to obtain a first feature map and a second feature map, and calculate a loss between the first feature map and the second feature map to obtain a covariance loss value;
[0015] Obtain an overall loss value according to the image quality loss value and the covariance loss value, and correct the parameters of the enhanced image quality assessment test model according to the overall loss value;
[0016] Input the next group of degraded images and enhanced images in the training image set into the enhanced image quality assessment test model until the training situation of the enhanced image quality assessment test model meets a preset condition to obtain a trained enhanced image quality assessment model.
[0017] Optionally, in the enhanced image quality assessment method based on semantic-guided residual learning, the obtaining of the semantic features of the enhanced image, generating a spatial attention map and a channel attention map according to the semantic features, and performing a combination process on the spatial attention map and the channel attention map to obtain a target feature map specifically includes:
[0018] Feature extraction is performed on the enhanced image to obtain semantic features. Max pooling and average pooling are performed on the semantic features to obtain an enhanced image feature map. Concatenation processing is performed on the enhanced image feature map to obtain a preliminary spatial attention map, and distribution calculation of the spatial weights of the preliminary spatial attention map is performed to obtain a spatial attention map;
[0019] Max pooling convolution processing is performed on the semantic features to obtain a first feature map. Average pooling convolution processing is performed on the semantic features to obtain a second feature map. Convolution concatenation is performed on the first feature map and the second feature map to obtain a channel attention map;
[0020] The spatial attention map and the channel attention map are combined to obtain a target feature map.
[0021] Optionally, in the enhanced image quality assessment method based on semantic-guided residual learning, the edge residual map is obtained according to the degraded image and the enhanced image, the enhanced feature map is obtained according to the edge residual map and the semantic features, and multi-scale enhancement is performed on the enhanced feature map to obtain a multi-scale enhanced feature map, which specifically includes:
[0022] Edge information maps are respectively extracted from the degraded image and the enhanced image to obtain a degraded edge information map and an enhanced edge information map. The degraded edge information map is subtracted from the enhanced edge information map to obtain an edge residual map;
[0023] Max pooling and convolution processing are performed on the edge residual map to obtain an edge residual feature map. Convolution processing is performed on the semantic features to obtain a convolution feature map. Spatial weight distribution calculation is performed on the edge residual feature map and the convolution feature map to obtain an enhanced edge residual map;
[0024] The enhanced edge residual map and the convolution feature map are multiplied to obtain a target edge residual map, and the target edge residual map and the semantic features are fused to obtain a semantic feature map;
[0025] Multi-scale enhancement convolution processing is performed on the semantic feature map to obtain a first-scale feature map, a second-scale feature map, and a third-scale feature map. The first-scale feature map, the second-scale feature map, and the third-scale feature map are fused by convolution to obtain a multi-scale enhanced feature map.
[0026] Optionally, in the enhanced image quality assessment method based on semantic-guided residual learning, the step of splicing the multi-scale enhanced feature map with the target feature map to obtain a predicted enhanced image quality assessment result, and calculating a loss between the predicted enhanced image quality assessment result and the true quality score of the enhanced image to obtain an image quality loss value specifically includes:
[0027] Splice the fusion features at each stage in the multi-scale enhanced feature map to obtain a multi-scale aggregated feature, and splice the multi-scale aggregated feature with the target feature map to obtain a predicted enhanced image quality assessment result;
[0028] Calculate a loss between the predicted quality score of the predicted enhanced image quality assessment result and the true quality score of the enhanced image according to a loss function to obtain an image quality loss value.
[0029] Optionally, in the enhanced image quality assessment method based on semantic-guided residual learning, the step of respectively extracting features from the degraded image and the enhanced image to obtain a first feature map and a second feature map, and calculating a loss between the first feature map and the second feature map to obtain a covariance loss value specifically includes:
[0030] Respectively perform edge extraction and difference processing on the degraded image and the enhanced image to obtain a first feature map and a second feature map, and calculate the cross-covariance of the first feature map and the second feature map;
[0031] Calculate the loss values of the first feature map and the second feature map according to a first formula and the cross-covariance to obtain a covariance loss value.
[0032] Optionally, in the enhanced image quality assessment method based on semantic-guided residual learning, the first formula is:
[0033]
[0034] where Lcc is the covariance loss value, C is the number of channels of the feature map, diag is the diagonal element, is the transpose of the feature map corresponding to the degraded image at the i-th layer of the pseudo Siamese network, is the feature map corresponding to the enhanced image at the i-th layer of the pseudo Siamese network, and both i and j are the layers of the pseudo Siamese network.
[0035] Optionally, in the enhanced image quality assessment method based on semantic-guided residual learning, the enhanced image quality assessment system based on semantic-guided residual learning includes:
[0036] An evaluation model training module, configured to obtain a training image set, and perform model training on the created enhanced image quality evaluation test model based on the training image set to obtain an enhanced image quality evaluation model, wherein the training image set includes multiple groups of degraded images and enhanced images;
[0037] An image quality evaluation module, configured to obtain an enhanced image to be evaluated, input the enhanced image to be evaluated into the enhanced image quality evaluation model, and output an enhanced image quality evaluation result.
[0038] In addition, to achieve the above object, the present invention further provides a terminal, wherein the terminal includes: a memory, a processor, and an enhanced image quality evaluation program based on semantic-guided residual learning stored on the memory and executable on the processor. When the enhanced image quality evaluation program based on semantic-guided residual learning is executed by the processor, the steps of the enhanced image quality evaluation method based on semantic-guided residual learning as described above are implemented.
[0039] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores an enhanced image quality evaluation program based on semantic-guided residual learning. When the enhanced image quality evaluation program based on semantic-guided residual learning is executed by a processor, the steps of the enhanced image quality evaluation method based on semantic-guided residual learning as described above are implemented.
[0040] In the present invention, a training image set is obtained, and model training is performed on the created enhanced image quality evaluation test model based on the training image set to obtain an enhanced image quality evaluation model, wherein the training image set includes multiple groups of degraded images and enhanced images; an enhanced image to be evaluated is obtained, and the enhanced image to be evaluated is input into the enhanced image quality evaluation model, and an enhanced image quality evaluation result is output. By using the semantic information of the image to guide the learning of edge residual features, the present invention can pay attention to the semantic information and distortion information in the image while training the model, and combine the semantic information and distortion information to improve the performance of enhanced image quality evaluation. Description of the Drawings
[0041] Figure 1 is a flowchart of a preferred embodiment of the enhanced image quality evaluation method based on semantic-guided residual learning of the present invention;
[0042] Figure 2 is a schematic diagram of the overall architecture of the enhanced image quality evaluation method based on semantic-guided residual learning of the present invention;
[0043] Figure 3 is a schematic diagram of the processing architecture of the semantic optimization module in a preferred embodiment of the present invention;
[0044] Figure 4 It is a schematic diagram of the processing architecture of the multi-scale semantic-guided edge residual module in a preferred embodiment of the present invention;
[0045] Figure 5 It is a structural diagram of a preferred embodiment of the enhanced image quality assessment system based on semantic-guided residual learning of the present invention;
[0046] Figure 6 It is a schematic diagram of the operating environment of a preferred embodiment of the terminal of the present invention. Specific embodiments
[0047] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0048] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, then such directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If this specific posture changes, then such directional indications will also change accordingly.
[0049] In addition, if there are descriptions such as "first", "second", etc. involved in the embodiments of the present invention, then such descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first", "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those skilled in the art can implement it. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0050] The enhanced image quality assessment method based on semantic-guided residual learning described in a preferred embodiment of the present invention, as Figure 1 shown, the enhanced image quality assessment method based on semantic-guided residual learning includes the following steps:
[0051] Step S10: Obtain a training image set, and train the created enhanced image quality assessment test model based on the training image set to obtain an enhanced image quality assessment model, where the training image set includes multiple groups of degraded images and enhanced images.
[0052] The step S10 includes:
[0053] Step S11: Obtain the training image set of the target object, create an enhanced image quality assessment test model, and input a set of degraded images and enhanced images in the training image set into the enhanced image quality assessment test model;
[0054] Step S12: Obtain the semantic features of the enhanced image, generate a spatial attention map and a channel attention map according to the semantic features, respectively combine the spatial attention map and the channel attention map with the semantic features to obtain a target spatial attention map and a target channel attention map, and combine the target spatial attention map and the target channel attention map to obtain a target feature map;
[0055] Step S13: Obtain an edge residual map according to the degraded image and the enhanced image, obtain an enhanced feature map according to the edge residual map and the semantic features, and perform multi-scale enhancement on the enhanced feature map to obtain a multi-scale enhanced feature map;
[0056] Step S14: Perform splicing processing on the multi-scale enhanced feature map and the target feature map to obtain a predicted enhanced image quality assessment result, and calculate the loss between the predicted enhanced image quality assessment result and the true quality score of the enhanced image to obtain an image quality loss value;
[0057] Step S15: Respectively extract features from the degraded image and the enhanced image to obtain a first feature map and a second feature map, and calculate the loss between the first feature map and the second feature map to obtain a covariance loss value;
[0058] Step S16: Obtain an overall loss value according to the image quality loss value and the covariance loss value, and correct the parameters of the enhanced image quality assessment test model according to the overall loss value;
[0059] Step S17: Input the next set of degraded images and enhanced images in the training image set into the enhanced image quality assessment test model until the training situation of the enhanced image quality assessment test model meets the preset conditions to obtain a trained enhanced image quality assessment model.
[0060] Specifically, in the embodiments of the present invention, for the existing image quality assessment methods, it is difficult to obtain a distortion-free original image, and the relationship between semantic information and the characteristics of the distortion region is ignored, resulting in a low performance of enhanced image quality assessment. The present invention proposes a method based on SGRQA (Semantic-Guided Residual Learning for the Quality Assessment, enhanced image quality assessment with semantic-guided residual learning). By using the internal relationship between the degraded image (referring to an image with rain streaks, haze, or low light, which is an image with distortion and a degraded normal photo quality image, and can be called a degraded image or a low-quality image) and the enhanced image (referring to the image obtained by processing the degraded image through an image enhancement algorithm) for assessment. Among them, the encoder extracts semantic information, and the decoder uses this semantic information to guide residual learning, which can capture the crucial semantic information and structural information in image quality assessment, thereby achieving more accurate assessment. And the overall framework diagram corresponding to this method is as shown in Figure 2 shown, and it can be seen from Figure 2 that this framework includes a PSEM (Pseudo-Siamese Semantic Extraction Module), an SRM (Semantic Refinement Module), and an MSERM (Multi-Scale Semantic-Guided Edge Residual Module). Among them, the pseudo-siamese semantic extraction module (for example, the ResNet50 module) can extract features from the enhanced image and the degraded image, capture the semantic information related to the enhanced image and the degraded image using the cross-covariance matrix, and then guide the edge residual information learning in the subsequent modules. To better model the relationship between these two images and optimize the feature representation, a cross-covariance loss function is proposed. The semantic refinement module optimizes the semantic features in the channel and spatial dimensions, bridges the encoder and the decoder, and enables the decoder to perform precise semantic guidance in fine-grained tasks. The multi-scale semantic-guided edge residual module integrates the semantic information from the encoder, focuses on semantic features based on edge residuals, and extracts detailed context information, thereby improving the performance of image quality assessment.
[0061] Before training the corresponding enhanced image quality assessment test model, it is necessary to obtain a training image set of the target object. Among them, the training image set includes multiple groups of degraded images and enhanced images, and an enhanced image quality assessment test model is created, and a group of degraded images and enhanced images in the training image set are input into the enhanced image quality assessment test model. Subsequently, first, semantic information is extracted from the corresponding images through the pseudo-twin semantic extraction module. Since the enhanced image has many similarities with the corresponding degraded image in terms of semantics, content, and feature distribution, and these shared features, especially semantic information, have an important impact on image quality assessment. To effectively utilize these similar semantic information, a pseudo-twin semantic extraction module is introduced in the encoder to extract these features.
[0062] After extracting the semantic features X corresponding to the enhanced image, it is further enhanced through the semantic optimization module, and the corresponding processing process is as Figure 3 shown. In the embodiment of the present invention, the semantic optimization module acts as a bridge, which can seamlessly integrate the spatial features, channel features, and semantic features of the encoder, and provide strong semantic guidance for the decoder. By capturing global context information and enhancing local spatial and channel dependencies, the semantic optimization module ensures that the decoder receives a comprehensive semantic representation. The semantic branch of the semantic optimization module uses dilated convolution to effectively capture the global context and improve the receptive field of feature learning. The fusion of semantic information with spatial and channel enhanced features finally generates an output feature map rich in semantic content, enabling the decoder to perform precise semantic guidance in fine-grained tasks. The specific processing process is as follows: on the spatial branch, first, a feature map is generated by using max-pooling and average-pooling operations in the channel dimension to generate a spatial attention map. Specifically, max-pooling and average-pooling are performed on the semantic features to obtain an enhanced image feature map, and the enhanced image feature map is concatenated to obtain a preliminary spatial attention map. Secondly, a 5×5 convolution is applied, and then the spatial weight distribution is calculated through the Sigmoid activation function to effectively capture spatial dependency relationships and enhance feature representation. Specifically, the spatial weight of the preliminary spatial attention map is distributed and calculated to obtain a spatial attention map. The mathematical expressions corresponding to the above process are: X' s = Max(X) c Mean(X), X s = Sigmoid(Conv 5x5 (X' s ))), where X' s is the enhanced image feature map, X is the semantic feature, Max is the max-pooling operation, Mean is the average-pooling operation, c is the feature concatenation, X s is the spatial attention map, sigmoid is the activation function, and Conv 5x5 is a 5×5 convolution.
[0063] On the channel branch, average pooling and max pooling operations are used to extract the average value and the maximum value in the feature map. Then, these outputs go through a 1×1 convolution to reduce the number of parameters. After that, the obtained features are merged through element-wise addition and restored to the original channel dimension through an additional 1×1 convolution. Specifically, perform max pooling convolution on the semantic features to obtain a first feature map, perform average pooling convolution on the semantic features to obtain a second feature map, and perform convolutional splicing on the first feature map and the second feature map to obtain a channel attention map; the mathematical expression of the above process is: X max = Conv 1×1 (MaxPool(X)); X mean = Conv 1×1 (MeanPool(X)); X c = Conv 1×1 (X max + X mean ); where X max is the feature map after max pooling and convolution, that is, the first feature map, MaxPool is max pooling, X mean is the feature map after average pooling and convolution, that is, the second feature map, MeanPool is average pooling, Conv 1×1 is a 1×1 convolution, and X c is the channel attention map.
[0064] After that, it is necessary to obtain the target feature map through the spatial attention map and the channel attention map. Specifically, the outputs of the channel branch and the spatial branch are combined through element-wise multiplication to enhance the features, and the enhanced result is then multiplied element-wise with the original input. Finally, a feature map containing spatial attention and channel attention is obtained. The semantic information branch uses a 3×3 stride convolution to extract the semantic features of the image, and the stride convolution helps to increase the receptive field of feature learning; finally, the semantic features are fused with the feature map containing spatial attention and channel attention to obtain the target feature map. Specifically, the spatial attention map and the channel attention map are respectively combined with the semantic features to obtain a target spatial attention map and a target channel attention map, and the target spatial attention map and the target channel attention map are combined to obtain the target feature map. The mathematical expression corresponding to the above process is: X in = dilatConv 3×3,d=2 (X); where X' is the feature map containing spatial attention and channel attention, is element-wise multiplication, X sac is the feature map after the fusion of the spatial branch and the channel branch, X inis the convolutional feature map, dilatConv 3×3,d=2 is the strided convolution with a stride of 2 and a convolution kernel of 3, is the element-wise addition, X out is the target feature map.
[0065] After that, in the multi-scale semantic-guided edge residual module, the enhanced semantic information is used to guide the edge residual learning, and the multi-scale semantic-guided edge residual module uses the semantic information from the encoder to guide the edge residual learning and adopts a multi-scale strategy to enhance the feature extraction. As Figure 4 shown, the edge residual map highlights the differences between the enhanced image and the degraded image, capturing the enhanced regions and the affected background. To effectively integrate this information, both the edge residual map and the semantic features are processed through 1×1 convolutions; through the multi-scale max pooling operation, the edge residual map is adjusted to align with the corresponding feature map dimensions. Specifically, the edge information maps are extracted from the degraded image and the enhanced image respectively to obtain the degraded edge information map and the enhanced edge information map, and the degraded edge information map is subtracted from the enhanced edge information map to obtain the edge residual map; the edge residual map is subjected to max pooling and convolution processing to obtain the edge residual feature map. The mathematical expression of this process is: Q i = Conv 1×1 (maxpool(x res ))), where Q i is the edge residual feature map, Conv 1×1 is the 1×1 convolution, maxpool is the max average pooling, and x res is the edge residual map; then, the adjusted edge residual map (i.e., the edge residual feature map) is multiplied element-wise with the semantic-guided feature map to generate a semantic-rich edge residual map, which is closely related to the image quality. Subsequently, the original semantic information is further optimized through the feature map. Finally, the enhanced features are fused with the original feature map through the residual connection to ensure the retention of the basic information. Specifically, the semantic features are convolved to obtain the convolutional feature map, and the spatial weight distribution of the edge residual feature map and the convolutional feature map is calculated to obtain the enhanced edge residual map; the enhanced edge residual map is multiplied with the convolutional feature map to obtain the target edge residual map, and the target edge residual map is fused with the semantic features to obtain the semantic feature map. The mathematical expression corresponding to this process is: K i = Conv 1×1 (F i ), V i = Conv 1×1 (F i ); where, K i and Vi are the convolutional feature maps of the enhanced image, F i is the semantic feature, F i ' is the semantic feature map, is the transpose of the convolutional feature map, and sigmoid is the activation function.
[0066] Image degradation caused by rain, fog, and low-light conditions exhibits diverse and extensive characteristics. To effectively capture the features related to these degradations, we perform multi-scale enhancement on the difference feature map. This process starts with a 1×1 convolution, first reducing the channel dimension, and then performing three parallel depthwise separable dilated convolutions with kernel sizes of 3×3, 5×5, and 7×7 respectively. These operations can extract context information at multiple scales. Through these operations, the module can handle variations in different degradation modes. The outputs from the parallel convolutions are fused through a summation operation, and finally, a 1×1 convolution is applied to adjust the dimension of the feature map to ensure its effective integration into the subsequent processing stage; specifically, multi-scale enhancement convolution processing is performed on the semantic feature map to obtain the first-scale feature map, the second-scale feature map, and the third-scale feature map, and the first-scale feature map, the second-scale feature map, and the third-scale feature map are subjected to fusion convolution processing to obtain the multi-scale enhanced feature map; the mathematical expression corresponding to this process is: X1 = dsConv 3×3 (Conv 1×1 (F i ')); X2 = dsConv 5×5 (Conv 1×1 (F i ')); X3 = dsConv 7×7 (Conv 1×1 (F i ')); F i+1 = Conv 1×1 (X1 + X2 + X3); where X1, X2, and X3 are the first-scale feature map, the second-scale feature map, and the third-scale feature map respectively, Conv 1×1 is a 1×1 convolution, dsConv 3×3 is a 3×3 depthwise separable convolution, dsConv 5×5 is a 5×5 depthwise separable convolution, dsConv 7×7 is a 7×7 depthwise separable convolution, F i ' is the semantic feature map, F i+1 is the multi-scale enhanced feature map.
[0067] After that, the fused features at each stage in the multi-scale enhanced feature map are concatenated to obtain multi-scale aggregated features, and the multi-scale aggregated features are concatenated with the target feature map to obtain a predicted enhanced image quality assessment result; the loss function is used to calculate the loss between the predicted quality score of the predicted enhanced image quality assessment result and the true quality score of the enhanced image, and an image quality loss value is obtained; the corresponding calculation formula is: where L Q is the loss function, Q p (x i ) and Q g (x i ) are the predicted quality score and the true quality score of the i-th input image respectively, and N is the total number of images.
[0068] For the pseudo-twin semantic extraction module, in the embodiments of the present invention, a cross-covariance loss function is further introduced to guide the encoder to capture and optimize feature similarity, so as to achieve more robust quality assessment. ResNet50 is used as the basic architecture of the pseudo-twin network, and the input is a pair of images (x d , x e ), where x d is the degraded image and x e is the enhanced image; the feature map of the degraded image from the i-th layer of the pseudo-twin network is denoted as The feature map of the enhanced image from the i-th layer of the pseudo-twin network is denoted as After that, the similarity between the feature maps is calculated through the cross-covariance matrix, and the cross-covariance of the feature pair can be expressed as: where the cross-covariance loss aligns the feature pairs by encouraging the cross-covariance matrix to approximate the identity matrix, which reflects the idea that the features should contain the same semantic information. In the embodiments of the present invention, an adjustment is made, focusing on driving the diagonal elements of the covariance matrix to be 1 instead of forcing the consistency of the entire matrix. This adjustment promotes the similarity learning between image pairs by setting custom constraints in the loss function; specifically, according to the first formula and the cross-covariance, the loss value between the first feature map and the second feature map is calculated to obtain a covariance loss value, and the first formula is: where Lcc is the covariance loss value, C is the number of channels of the feature map, diag is the diagonal element, is the transpose of the feature map corresponding to the degraded image of the i-th layer of the pseudo-twin network, and both i and j are the layers of the pseudo-twin network.
[0069] After that, the overall loss value is obtained according to the image quality loss value and the covariance loss value, and the corresponding calculation formula is: L = L Q + α·L cc; where α is a weighted parameter for balancing the two parts, and in the embodiments of the present invention, the value of α is set to 0.2; and the parameters of the enhanced image quality assessment test model are corrected according to the overall loss value; subsequently, the next set of degraded images and enhanced images in the training image set are input into the enhanced image quality assessment test model, and the above process is repeated, which will not be elaborated here, until the training situation of the enhanced image quality assessment test model meets the preset conditions, and a trained enhanced image quality assessment model is obtained.
[0070] Further, after obtaining the enhanced image quality assessment model, in the embodiments of the present invention, in order to test the performance of the enhanced image quality assessment model, five image quality assessment data sets from three scenarios (for example, rain, fog, and low light) are tested, including the IVIPC (Subjective and Objective De-Raining Quality Assessment Towards Authentic Rain Image) data set, the SHRQ-R (Quality evaluation of image dehazing methods using synthetic hazy images) data set, the DHQ (Objective quality evaluation of dehazed images) data set, the LIEQ (Perceptual quality assessment of low-light image enhancement) data set, and the LEISD (No-reference quality assessment for low-light image enhancement) data set. Among them, the IVIPC data set contains 206 real images and 1236 de-rained images; the SHRQ-R data set is a regular image subset from the SHRQ database, containing 360 de-hazed images, which are generated from 45 synthetic haze images; the DHQ de-hazing contains 250 real haze images and 1750 de-hazed images; the LIEQ de-hazing contains 100 low-light images and 1000 enhanced images; the LEISD de-hazing includes 255 real low-light images and 2040 enhanced images. The performance of the target metrics is evaluated by using two criteria: the Spearman rank correlation coefficient (SRCC) and the Pearson linear correlation coefficient (PLCC). Higher values (up to 1) of these metrics indicate better prediction performance. For each database, 80% of the images are used for training, and the remaining 20% of the images are used to verify the performance of the model. The mean opinion score (MOS) corresponding to the enhanced images is also provided, and a higher score indicates better image quality.
[0071] To evaluate the performance of the method corresponding model of the proposed present invention, 10 advanced image quality assessment (IQA) metrics were compared, including five general models (IQA model, NIQMC model, HyperIQA model, VCRNet model, and TReS model respectively), one IQA metric for de-rained images (B-EFN), two IQA metrics for low-light enhanced images (NLIEE, MFRQA), and two IQA metrics for de-hazed images (FADE, TRG-DQA). The corresponding experimental results are shown in Table 1.
[0072] Table 1: Performance results of multiple models for enhanced image quality assessment on test data
[0073]
[0074] As shown in Table 1, the optimal results are marked in bold in Table 1. In the last column, the average SRCC and PLCC values of all databases are provided. The results in Table 1 demonstrate the superiority of the method proposed by the present invention on multiple models and datasets.
[0075] Step S20: Obtain the enhanced image to be evaluated and the corresponding target degraded image, input the enhanced image to be evaluated and the target degraded image into the enhanced image quality assessment model, and output the enhanced image quality assessment result.
[0076] Specifically, after obtaining the enhanced image quality assessment model, obtain the enhanced image to be evaluated of the target object and the corresponding target degraded image, and input the enhanced image to be evaluated and the target degraded image into the enhanced quality assessment model; then, perform quality assessment on the enhanced image to be evaluated and the target degraded image through the enhanced quality assessment model to obtain the corresponding enhanced image quality assessment result.
[0077] Furthermore, as Figure 5 shown, based on the above enhanced image quality assessment method based on semantic-guided residual learning, the present invention also correspondingly provides an enhanced image quality assessment system based on semantic-guided residual learning. Among them, the enhanced image quality assessment system based on semantic-guided residual learning includes:
[0078] An evaluation model training module 51, configured to obtain a training image set, and perform model training on the created enhanced image quality assessment test model based on the training image set to obtain an enhanced image quality assessment model, where the training image set includes multiple groups of degraded images and enhanced images;
[0079] The image quality evaluation module 52 is configured to obtain the enhanced image to be evaluated and the corresponding target degraded image, input the enhanced image to be evaluated and the target degraded image into the enhanced image quality evaluation model, and output the enhanced image quality evaluation result.
[0080] Further, as Figure 6 shown, based on the above enhanced image quality evaluation method based on semantic-guided residual learning, the present invention also correspondingly provides a terminal, which includes a processor 10, a memory 20, and a display 30. Figure 6 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0081] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as the hard disk or memory of the terminal. The memory 20 may also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk equipped on the terminal, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as the program code for installing the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, an enhanced image quality evaluation program 40 based on semantic-guided residual learning is stored on the memory 20, and the enhanced image quality evaluation program 40 based on semantic-guided residual learning can be executed by the processor 10, so as to implement the enhanced image quality evaluation method based on semantic-guided residual learning in the present application.
[0082] The processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 20 or process data, such as executing the enhanced image quality evaluation method based on semantic-guided residual learning, etc.
[0083] The display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. in some embodiments. The display 30 is used to display information on the terminal and a visual user interface. The components of the terminal communicate with each other through a system bus.
[0084] In one embodiment, when the processor 10 executes the enhanced image quality assessment program 40 based on semantic-guided residual learning in the memory 20, the following steps are implemented:
[0085] Obtain a training image set, and perform model training on the created enhanced image quality assessment test model based on the training image set to obtain an enhanced image quality assessment model, where the training image set includes multiple groups of degraded images and enhanced images;
[0086] Obtain the enhanced image to be evaluated and the corresponding target degraded image, input the enhanced image to be evaluated and the target degraded image into the enhanced image quality assessment model, and output the enhanced image quality assessment result.
[0087] Among them, the obtaining of the training image set, performing model training on the created enhanced image quality assessment test model based on the training image set to obtain an enhanced image quality assessment model specifically includes;
[0088] Obtain the training image set of the target object, create an enhanced image quality assessment test model, and input a group of degraded images and enhanced images in the training image set into the enhanced image quality assessment test model;
[0089] Obtain the semantic features of the enhanced image, generate a spatial attention map and a channel attention map according to the semantic features, respectively combine the spatial attention map and the channel attention map with the semantic features to obtain a target spatial attention map and a target channel attention map, and perform a combination process on the target spatial attention map and the target channel attention map to obtain a target feature map;
[0090] Obtain an edge residual map according to the degraded image and the enhanced image, obtain an enhanced feature map according to the edge residual map and the semantic features, and perform multi-scale enhancement on the enhanced feature map to obtain a multi-scale enhanced feature map;
[0091] Perform a splicing process on the multi-scale enhanced feature map and the target feature map to obtain a predicted enhanced image quality assessment result, and calculate the loss between the predicted enhanced image quality assessment result and the true quality score of the enhanced image to obtain an image quality loss value;
[0092] Perform feature processing on the degraded image and the enhanced image respectively to obtain a first feature map and a second feature map, and calculate the loss between the first feature map and the second feature map to obtain a covariance loss value;
[0093] Obtain an overall loss value according to the image quality loss value and the covariance loss value, and correct the parameters of the enhanced image quality assessment test model according to the overall loss value;
[0094] Input the next set of degraded images and enhanced images in the training image set into the enhanced image quality assessment test model until the training condition of the enhanced image quality assessment test model meets the preset condition, and obtain the trained enhanced image quality assessment model.
[0095] Among them, the obtaining of the semantic features of the enhanced image, generating a spatial attention map and a channel attention map according to the semantic features, respectively combining the spatial attention map and the channel attention map with the semantic features to obtain a target spatial attention map and a target channel attention map, and combining the target spatial attention map and the target channel attention map for processing to obtain a target feature map specifically includes:
[0096] Extract features from the enhanced image to obtain semantic features, perform max pooling and average pooling on the semantic features to obtain an enhanced image feature map, perform splicing processing on the enhanced image feature map to obtain a preliminary spatial attention map, and calculate the spatial weight distribution of the preliminary spatial attention map to obtain a spatial attention map;
[0097] Perform max pooling convolution processing on the semantic features to obtain a first feature map, perform average pooling convolution processing on the semantic features to obtain a second feature map, and perform convolution splicing on the first feature map and the second feature map to obtain a channel attention map;
[0098] Respectively combine the spatial attention map and the channel attention map with the semantic features to obtain a target spatial attention map and a target channel attention map, and combine the target spatial attention map and the target channel attention map for processing to obtain a target feature map.
[0099] Among them, the obtaining of an edge residual map according to the degraded image and the enhanced image, obtaining an enhanced feature map according to the edge residual map and the semantic features, and performing multi-scale enhancement on the enhanced feature map to obtain a multi-scale enhanced feature map specifically includes:
[0100] Extract edge information maps from the degraded image and the enhanced image respectively to obtain a degraded edge information map and an enhanced edge information map, and perform subtraction processing on the degraded edge information map and the enhanced edge information map to obtain an edge residual map;
[0101] Perform max pooling and convolution processing on the edge residual map to obtain an edge residual feature map, perform convolution processing on the semantic features to obtain a convolution feature map, and calculate the spatial weight distribution of the edge residual feature map and the convolution feature map to obtain an enhanced edge residual map;
[0102] Multiply the enhanced edge residual map with the convolutional feature map to obtain a target edge residual map, and fuse the target edge residual map with the semantic feature to obtain a semantic feature map;
[0103] Perform multi-scale enhanced convolution processing on the semantic feature map to obtain a first-scale feature map, a second-scale feature map, and a third-scale feature map, and perform fusion convolution processing on the first-scale feature map, the second-scale feature map, and the third-scale feature map to obtain a multi-scale enhanced feature map.
[0104] Among them, the process of splicing the multi-scale enhanced feature map with the target feature map to obtain a predicted enhanced image quality assessment result, and calculating the loss between the predicted enhanced image quality assessment result and the true quality score of the enhanced image to obtain an image quality loss value specifically includes:
[0105] Splice the fused features at each stage in the multi-scale enhanced feature map to obtain a multi-scale aggregated feature, and splice the multi-scale aggregated feature with the target feature map to obtain a predicted enhanced image quality assessment result;
[0106] Calculate the loss between the predicted quality score of the predicted enhanced image quality assessment result and the true quality score of the enhanced image according to the loss function to obtain an image quality loss value.
[0107] Among them, the process of respectively extracting features from the degraded image and the enhanced image to obtain a first feature map and a second feature map, and calculating the loss between the first feature map and the second feature map to obtain a covariance loss value specifically includes:
[0108] Respectively perform edge extraction and difference processing on the degraded image and the enhanced image to obtain a first feature map and a second feature map, and calculate the cross-covariance of the first feature map and the second feature map;
[0109] Calculate the loss values of the first feature map and the second feature map according to the first formula and the cross-covariance to obtain a covariance loss value.
[0110] Among them, the first formula is:
[0111]
[0112] Among them, Lcc is the covariance loss value, C is the number of channels of the feature map, diag is the diagonal element, is the transpose of the feature map corresponding to the degraded image of the i-th layer of the pseudo-twin network, is the feature map corresponding to the enhanced image of the i-th layer of the pseudo-twin network, and both i and j are the layers of the pseudo-twin network.
[0113] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an enhanced image quality assessment program based on semantic-guided residual learning. When the enhanced image quality assessment program based on semantic-guided residual learning is executed by a processor, the steps of the enhanced image quality assessment method based on semantic-guided residual learning as described above are implemented.
[0114] In summary, the present invention provides an enhanced image quality assessment method, system and terminal based on semantic-guided residual learning. The method includes: obtaining a training image set, training a created enhanced image quality assessment test model based on the training image set to obtain an enhanced image quality assessment model, wherein the training image set includes multiple groups of degraded images and enhanced images; obtaining an enhanced image to be evaluated and a corresponding target degraded image, inputting the enhanced image to be evaluated and the target degraded image into the enhanced image quality assessment model, and outputting an enhanced image quality assessment result. By using the semantic information of the image to guide the learning of edge residual features, the present invention can pay attention to the semantic information and distortion information in the image while training the model, and combine the semantic information and distortion information to improve the performance of the enhanced image quality assessment.
[0115] It should be noted that, in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal including the element.
[0116] Certainly, those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.
[0117] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. An enhanced image quality assessment method based on semantically guided residual learning, characterized in that: The enhanced image quality assessment method based on semantic guided residual learning includes: Acquire a training image set, and perform model training on the created enhanced image quality assessment test model based on the training image set to obtain an enhanced image quality assessment model, wherein the training image set includes multiple groups of degraded images and enhanced images; The enhanced image to be evaluated and the corresponding target degraded image are obtained, the enhanced image to be evaluated and the target degraded image are input into the enhanced image quality evaluation model, and the enhanced image quality evaluation result is output.
2. The enhanced image quality assessment method based on semantically guided residual learning according to claim 1, characterized in that: The acquiring of a training image set and performing model training on the created enhanced image quality assessment test model based on the training image set to obtain the enhanced image quality assessment model specifically includes: Acquire a training image set of the target object, create an enhanced image quality assessment test model, and input a group of degraded images and enhanced images in the training image set into the enhanced image quality assessment test model; Acquire semantic features of the enhanced image, generate a spatial attention map and a channel attention map according to the semantic features, respectively combine the spatial attention map and the channel attention map with the semantic features to obtain a target spatial attention map and a target channel attention map, and combine the target spatial attention map and the target channel attention map to obtain a target feature map; Obtaining an edge residual map according to the degraded image and the enhanced image, obtaining an enhanced feature map according to the edge residual map and the semantic features, and performing multi-scale enhancement on the enhanced feature map to obtain a multi-scale enhanced feature map; The multi-scale enhanced feature map is spliced with the target feature map to obtain a predicted enhanced image quality assessment result, and the predicted enhanced image quality assessment result is subjected to loss calculation with the true quality score of the enhanced image to obtain an image quality loss value; Performing feature extraction on the degraded image and the enhanced image respectively to obtain a first feature map and a second feature map, and performing loss calculation on the first feature map and the second feature map to obtain a covariance loss value; Obtaining an overall loss value according to the image quality loss value and the covariance loss value, and modifying parameters of the enhanced image quality assessment test model according to the overall loss value; A group of degraded images and enhanced images under the training image set are input into the enhanced image quality assessment test model until the training status of the enhanced image quality assessment test model meets the preset conditions, thereby obtaining a trained enhanced image quality assessment model.
3. The enhanced image quality assessment method based on semantically guided residual learning according to claim 2, characterized in that: The acquiring of the semantic features of the enhanced image, generating a spatial attention map and a channel attention map according to the semantic features, respectively combining the spatial attention map and the channel attention map with the semantic features to obtain a target spatial attention map and a target channel attention map, and combining the target spatial attention map and the target channel attention map to obtain a target feature map, specifically includes: Performing feature extraction on the enhanced image to obtain semantic features, performing maximum pooling and average pooling on the semantic features to obtain an enhanced image feature map, performing splicing processing on the enhanced image feature map to obtain a preliminary spatial attention map, and performing distribution calculation on the spatial weight of the preliminary spatial attention map to obtain a spatial attention map; Performing a maximum pooling convolution process on the semantic features to obtain a first feature map, performing an average pooling convolution process on the semantic features to obtain a second feature map, and performing convolution splicing on the first feature map and the second feature map to obtain a channel attention map; The spatial attention map and the channel attention map are respectively combined with the semantic features to obtain a target spatial attention map and a target channel attention map, and the target spatial attention map and the target channel attention map are combined to obtain a target feature map.
4. The enhanced image quality assessment method based on semantically guided residual learning according to claim 2 is characterized in that: The step of obtaining an edge residual map according to the degraded image and the enhanced image, obtaining an enhanced feature map according to the edge residual map and the semantic feature, and performing multi-scale enhancement on the enhanced feature map to obtain a multi-scale enhanced feature map specifically includes: Extracting edge information maps from the degraded image and the enhanced image respectively to obtain a degraded edge information map and an enhanced edge information map, and performing subtraction processing on the degraded edge information map and the enhanced edge information map to obtain an edge residual map; Performing maximum pooling and convolution processing on the edge residual map to obtain an edge residual feature map, performing convolution processing on the semantic feature to obtain a convolution feature map, and performing spatial weight distribution calculation on the edge residual feature map and the convolution feature map to obtain an enhanced edge residual map; The enhanced edge residual map is multiplied by the convolution feature map to obtain a target edge residual map, and the target edge residual map is fused with the semantic feature to obtain a semantic feature map; The semantic feature map is subjected to multi-scale enhanced convolution processing to obtain a first-scale feature map, a second-scale feature map, and a third-scale feature map, and the first-scale feature map, the second-scale feature map, and the third-scale feature map are subjected to fusion convolution processing to obtain a multi-scale enhanced feature map.
5. The enhanced image quality assessment method based on semantic guided residual learning according to claim 2, characterized in that: The step of performing splicing processing on the multi-scale enhanced feature map and the target feature map to obtain a predicted enhanced image quality assessment result, and performing loss calculation on the predicted enhanced image quality assessment result and the true quality score of the enhanced image to obtain an image quality loss value specifically includes: Splicing the fused features of each stage in the multi-scale enhanced feature map to obtain a multi-scale aggregated feature, and splicing the multi-scale aggregated feature with the target feature map to obtain a predicted enhanced image quality assessment result; A loss calculation is performed on the predicted quality score of the predicted enhanced image quality assessment result and the true quality score of the enhanced image according to the loss function to obtain an image quality loss value.
6. The enhanced image quality assessment method based on semantic guided residual learning according to claim 2, characterized in that: The extracting features of the degraded image and the enhanced image respectively to obtain a first feature map and a second feature map, and calculating the loss of the first feature map and the second feature map to obtain a covariance loss value specifically includes: Performing edge extraction and difference processing on the degraded image and the enhanced image respectively to obtain a first feature map and a second feature map, and calculating a cross covariance between the first feature map and the second feature map; The loss values of the first feature map and the second feature map are calculated according to the first formula and the cross covariance to obtain a covariance loss value.
7. The enhanced image quality assessment method based on semantically guided residual learning according to claim 6, characterized in that: The first formula is: Among them, Lcc is the covariance loss value, C is the number of channels of the feature map, and diag is the diagonal element. is the transpose of the feature map corresponding to the degraded image of the i-th layer of the pseudo-twin network, is the feature map corresponding to the enhanced image of the i-th layer of the pseudo-twin network, and i and j are the number of layers of the pseudo-twin network.
8. An enhanced image quality assessment system based on semantically guided residual learning, characterized in that: The enhanced image quality assessment system based on semantic guided residual learning includes: An evaluation model training module is used to obtain a training image set, and perform model training on the created enhanced image quality evaluation test model based on the training image set to obtain an enhanced image quality evaluation model, wherein the training image set includes multiple groups of degraded images and enhanced images; The image quality assessment module is used to obtain the enhanced image to be assessed, input the enhanced image to be assessed into the enhanced image quality assessment model, and output the enhanced image quality assessment result.
9. A terminal, characterized in that: The terminal includes: a memory, a processor, and an enhanced image quality assessment program based on semantically guided residual learning stored in the memory and executable on the processor. When the enhanced image quality assessment program based on semantically guided residual learning is executed by the processor, the steps of the enhanced image quality assessment method based on semantically guided residual learning as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an enhanced image quality assessment program based on semantically guided residual learning, and when the enhanced image quality assessment program based on semantically guided residual learning is executed by a processor, the steps of the enhanced image quality assessment method based on semantically guided residual learning as described in any one of claims 1-7 are implemented.